{
  "cells": [
    {
      "cell_type": "markdown",
      "id": "8a62f379",
      "metadata": {
        "id": "8a62f379"
      },
      "source": [
        "# Day 3: Assignment — Production-Ready RAG System\n",
        "\n",
        "## Overview\n",
        "\n",
        "Build a complete, documented RAG system ready for production handoff. This assignment extends your independent lab work with deeper evaluation (RAG triad), comprehensive error analysis, and a full RAG playbook.\n",
        "\n",
        "### Grading Summary\n",
        "| # | Deliverable | Points |\n",
        "|---|-------------|--------|\n",
        "| 1 | Knowledge Base Design | 15 |\n",
        "| 2 | RAG System Prompt | 25 |\n",
        "| 3 | RAG Outputs (15+ questions) | — |\n",
        "| 4 | Golden Q&A Set (15+ items) | — |\n",
        "| 5 | RAG Triad Metrics | 20 |\n",
        "| 6 | Error Analysis | 25 |\n",
        "| 7 | RAG Playbook | 15 |\n",
        "| | **Total** | **100** |"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "add75c48",
      "metadata": {
        "id": "add75c48"
      },
      "source": [
        "## Production-Ready RAG System Using Routed Hybrid Retrieval\n",
        "\n",
        "This code takes the routed hybrid retrieval approach. Difference made between the two codes is as below:\n",
        "| Aspect              | simple filtered dense retrieval| routed hybrid retrieval              |\n",
        "| ------------------- | --------- | ------------------ |\n",
        "| Routing             | ❌ none    | ✅ `route_filter()` |\n",
        "| Filtered retrieval  | ✅ yes     | ✅ yes              |\n",
        "| Dense embeddings    | ✅ yes     | ✅ yes              |\n",
        "| Sparse (TF-IDF)     | ❌ no      | ✅ yes              |\n",
        "| Hybrid fusion       | ❌ no      | ✅ yes              |\n",
        "| Score normalization | ❌ no      | ✅ yes              |\n",
        "| Metadata returned   | ❌ partial | ✅ full dict        |"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "c76e1107",
      "metadata": {
        "id": "c76e1107"
      },
      "source": [
        "---\n",
        "## Setup"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 1,
      "id": "1f5b22af",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "1f5b22af",
        "outputId": "7748a7c1-0c2b-42a5-84ca-42f6292cd551"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "\u001b[?25l     \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m0.0/53.1 kB\u001b[0m \u001b[31m?\u001b[0m eta \u001b[36m-:--:--\u001b[0m\r\u001b[2K     \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m53.1/53.1 kB\u001b[0m \u001b[31m2.4 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
            "\u001b[?25h\u001b[?25l   \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m0.0/724.7 kB\u001b[0m \u001b[31m?\u001b[0m eta \u001b[36m-:--:--\u001b[0m\r\u001b[2K   \u001b[91m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m\u001b[91m╸\u001b[0m \u001b[32m716.8/724.7 kB\u001b[0m \u001b[31m22.7 MB/s\u001b[0m eta \u001b[36m0:00:01\u001b[0m\r\u001b[2K   \u001b[90m━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━\u001b[0m \u001b[32m724.7/724.7 kB\u001b[0m \u001b[31m16.4 MB/s\u001b[0m eta \u001b[36m0:00:00\u001b[0m\n",
            "\u001b[?25h"
          ]
        }
      ],
      "source": [
        "!pip install -q -U google-genai"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 2,
      "id": "b0ff385a",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "b0ff385a",
        "outputId": "59d0e9cd-1c5a-45ed-ac23-d2592ec4b457"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "API key loaded: yes\n",
            "Generation model: gemini-2.5-flash-lite\n",
            "Embedding model:  gemini-embedding-001\n"
          ]
        }
      ],
      "source": [
        "# ── Imports ──────────────────────────────────────────────\n",
        "# ── Standard library ───────────────────────────────────────────────────────\n",
        "import os\n",
        "import time\n",
        "import json\n",
        "import random\n",
        "import re\n",
        "from datetime import datetime, timezone\n",
        "from typing import List, Optional, Literal\n",
        "\n",
        "# ── Third-party libraries ───────────────────────────────────────────────────\n",
        "import numpy as np\n",
        "import pandas as pd\n",
        "from pydantic import BaseModel, Field\n",
        "from sklearn.preprocessing import minmax_scale\n",
        "\n",
        "# ── Google GenAI SDK ─────────────────────────────────────────────────────────\n",
        "from google import genai\n",
        "from google.genai import types, errors\n",
        "\n",
        "\n",
        "\n",
        "# ── API Key ───────────────────────────────────────────────\n",
        "try:\n",
        "    from google.colab import userdata\n",
        "    API_KEY = userdata.get(\"GEMINI_API_KEY\")\n",
        "except Exception:\n",
        "    API_KEY = None\n",
        "\n",
        "if not API_KEY:\n",
        "    import getpass\n",
        "    API_KEY = getpass.getpass(\"Enter your Gemini API key: \")\n",
        "\n",
        "client = genai.Client(api_key=API_KEY)\n",
        "\n",
        "MODEL_ID = \"gemini-2.5-flash-lite\"\n",
        "EMBEDDING_MODEL = \"gemini-embedding-001\"\n",
        "\n",
        "print(f\"API key loaded: {'yes' if API_KEY else 'no'}\")\n",
        "print(f\"Generation model: {MODEL_ID}\")\n",
        "print(f\"Embedding model:  {EMBEDDING_MODEL}\")"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 3,
      "id": "238cfbb0",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "238cfbb0",
        "outputId": "beb92291-53c8-48e7-c519-da9d8beb7d3c"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "All infrastructure ready (includes hybrid retrieval, metadata filtering, and retrieval metrics).\n"
          ]
        }
      ],
      "source": [
        "# ── Infrastructure: API, Logging, and Core Functions ─────────────────────────\n",
        "\n",
        "PROMPT_LOG = []  # Track all API calls\n",
        "\n",
        "def _now():\n",
        "    return datetime.now(timezone.utc).isoformat().replace('+00:00', 'Z')\n",
        "\n",
        "def generate(prompt, temperature=0.7, max_tokens=1000, log=True, label=None):\n",
        "    \"\"\"Generate free-form text. Returns raw string.\"\"\"\n",
        "    t0 = time.time()\n",
        "    response = client.models.generate_content(\n",
        "        model=MODEL_ID,\n",
        "        contents=prompt,\n",
        "        config={\"temperature\": temperature, \"max_output_tokens\": max_tokens},\n",
        "    )\n",
        "    latency = time.time() - t0\n",
        "    text = response.text\n",
        "    if log:\n",
        "        PROMPT_LOG.append({\n",
        "            \"timestamp\": _now(), \"label\": label or \"generate\",\n",
        "            \"type\": \"free_form\",\n",
        "            \"prompt\": prompt[:300] + \"...\" if len(prompt) > 300 else prompt,\n",
        "            \"prompt_length\": len(prompt),\n",
        "            \"temperature\": temperature,\n",
        "            \"response\": text[:300] + \"...\" if len(text) > 300 else text,\n",
        "            \"response_length\": len(text),\n",
        "            \"latency_s\": round(latency, 2),\n",
        "        })\n",
        "    return text\n",
        "\n",
        "def generate_structured(\n",
        "    prompt,\n",
        "    response_model: type[BaseModel],\n",
        "    label=\"default\",\n",
        "    temperature=0.3,\n",
        "    max_tokens=512,\n",
        "    retries=6\n",
        "):\n",
        "    \"\"\"Generate structured JSON output with retry/backoff. Returns Pydantic model instance.\"\"\"\n",
        "    for attempt in range(retries):\n",
        "        try:\n",
        "            response = client.models.generate_content(\n",
        "                model=MODEL_ID,\n",
        "                contents=prompt,\n",
        "                config={\n",
        "                    \"temperature\": temperature,\n",
        "                    \"max_output_tokens\": max_tokens,\n",
        "                    \"response_mime_type\": \"application/json\",\n",
        "                },\n",
        "            )\n",
        "\n",
        "            json_str = (response.text or \"\").strip()\n",
        "\n",
        "            # Strip code fences if present\n",
        "            if json_str.startswith(\"```\"):\n",
        "                json_str = json_str.split(\"\\n\", 1)[1].strip()\n",
        "                if json_str.endswith(\"```\"):\n",
        "                    json_str = json_str[:-3].strip()\n",
        "\n",
        "            data = json.loads(json_str)\n",
        "            return response_model.model_validate(data)\n",
        "\n",
        "        except errors.ServerError as e:\n",
        "            # 503 UNAVAILABLE (transient)\n",
        "            wait = (2 ** attempt) + random.uniform(0, 0.5)\n",
        "            print(f\"[WARN] 503 from model (attempt {attempt+1}/{retries}). Sleeping {wait:.1f}s...\")\n",
        "            time.sleep(wait)\n",
        "\n",
        "        except json.JSONDecodeError:\n",
        "            # Model returned non-JSON; treat as transient and retry\n",
        "            wait = (2 ** attempt) + random.uniform(0, 0.5)\n",
        "            print(f\"[WARN] Non-JSON response (attempt {attempt+1}/{retries}). Sleeping {wait:.1f}s...\")\n",
        "            time.sleep(wait)\n",
        "\n",
        "    raise RuntimeError(\"generate_structured failed after retries (503 or invalid JSON).\")\n",
        "\n",
        "\n",
        "def embed_texts(texts, task_type=\"RETRIEVAL_DOCUMENT\"):\n",
        "    \"\"\"Embed one or more texts using Gemini Embeddings API.\"\"\"\n",
        "    if isinstance(texts, str):\n",
        "        texts = [texts]\n",
        "    response = client.models.embed_content(\n",
        "        model=EMBEDDING_MODEL,\n",
        "        contents=texts,\n",
        "        config=types.EmbedContentConfig(task_type=task_type),\n",
        "    )\n",
        "    return [np.array(e.values) for e in response.embeddings]\n",
        "\n",
        "def cosine_similarity(a, b):\n",
        "    \"\"\"Compute cosine similarity between two vectors or matrices.\"\"\"\n",
        "    a = a / (np.linalg.norm(a, axis=-1, keepdims=True) + 1e-9)\n",
        "    b = b / (np.linalg.norm(b, axis=-1, keepdims=True) + 1e-9)\n",
        "    return a @ b.T\n",
        "\n",
        "import re\n",
        "\n",
        "def chunk_sentences(text, max_words=80, overlap_words=20):\n",
        "    \"\"\"\n",
        "    Sentence-based chunking with word-overlap.\n",
        "    If a single sentence exceeds max_words, we split it by words to avoid empty/oversized chunks.\n",
        "    \"\"\"\n",
        "    if text is None:\n",
        "        return []\n",
        "    text = str(text).strip()\n",
        "    if not text:\n",
        "        return []\n",
        "\n",
        "    # Split into sentences (simple heuristic)\n",
        "    sentences = re.split(r'(?<=[.!?])\\s+', text)\n",
        "    sentences = [s.strip() for s in sentences if s and s.strip()]\n",
        "\n",
        "    chunks = []\n",
        "    current_words = []\n",
        "\n",
        "    def flush_current():\n",
        "        if current_words:\n",
        "            chunks.append(\" \".join(current_words))\n",
        "\n",
        "    for sent in sentences:\n",
        "        words = sent.split()\n",
        "\n",
        "        # If a single sentence is too long, split it into max_words blocks\n",
        "        while len(words) > max_words:\n",
        "            # First, flush whatever is currently building\n",
        "            flush_current()\n",
        "            current_words = []\n",
        "\n",
        "            # Take a max_words slice as a chunk\n",
        "            chunk_words = words[:max_words]\n",
        "            chunks.append(\" \".join(chunk_words))\n",
        "\n",
        "            # Prepare overlap for next chunk\n",
        "            words = words[max_words - overlap_words:] if overlap_words > 0 else words[max_words:]\n",
        "\n",
        "        # Normal case: add sentence if fits, else flush and start new with overlap\n",
        "        if len(current_words) + len(words) > max_words and current_words:\n",
        "            flush_current()\n",
        "            # overlap from previous chunk\n",
        "            current_words = current_words[-overlap_words:] if overlap_words > 0 else []\n",
        "\n",
        "        current_words.extend(words)\n",
        "\n",
        "    flush_current()\n",
        "    return chunks\n",
        "\n",
        "\n",
        "def search(query, chunks, chunk_embeddings, top_k=3):\n",
        "    \"\"\"\n",
        "    Retrieve top-k relevant chunks using cosine similarity.\n",
        "    Returns list of (index, score, text) tuples.\n",
        "    \"\"\"\n",
        "    query_embedding = embed_texts([query], task_type=\"RETRIEVAL_QUERY\")\n",
        "    similarities = cosine_similarity(query_embedding, chunk_embeddings)[0]\n",
        "    top_indices = np.argsort(-similarities)[:top_k]\n",
        "    results = [\n",
        "        (int(idx), float(similarities[idx]), chunks[idx])\n",
        "        for idx in top_indices\n",
        "    ]\n",
        "    return results\n",
        "\n",
        "def rag_query(\n",
        "    question,\n",
        "    chunks,\n",
        "    chunk_embeddings,\n",
        "    chunk_metadata,\n",
        "    tfidf_vectorizer,\n",
        "    tfidf_matrix,\n",
        "    top_k=3,\n",
        "    alpha=0.7,\n",
        "    system_prompt=\"\"\n",
        "):\n",
        "    \"\"\"\n",
        "    RAG query:\n",
        "    1) retrieve (filtered dense OR hybrid)\n",
        "    2) build context\n",
        "    3) generate answer\n",
        "    Returns (answer, retrieved_results)\n",
        "    retrieved_results: list of (chunk_text, score, metadata)\n",
        "    \"\"\"\n",
        "\n",
        "    # 1) Retrieve\n",
        "    flt = route_filter(question)\n",
        "    print(\n",
        "    f\"[RAG] Retriever = {'FILTERED (search_with_filter)' if flt else 'HYBRID (dense + TF-IDF)'}\"\n",
        "    )\n",
        "\n",
        "    if flt:\n",
        "        retrieved = search_with_filter(\n",
        "            question,\n",
        "            chunks,\n",
        "            chunk_embeddings,\n",
        "            chunk_metadata,\n",
        "            top_k=top_k,\n",
        "            doc_title_contains=flt[\"doc_title_contains\"],\n",
        "        )\n",
        "    else:\n",
        "        retrieved = hybrid_search(\n",
        "            question,\n",
        "            chunks,\n",
        "            chunk_embeddings,\n",
        "            chunk_metadata,\n",
        "            tfidf_vectorizer,\n",
        "            tfidf_matrix,\n",
        "            top_k=top_k,\n",
        "            alpha=alpha\n",
        "        )\n",
        "    print(\n",
        "    \"[RAG] Docs:\",\n",
        "    [meta.get(\"doc_title\", \"\") for (_, _, meta) in retrieved]\n",
        "    )\n",
        "\n",
        "    # 2) Build context string\n",
        "    context = \"\\n\\n\".join(\n",
        "        f\"[Chunk {i+1}] (score={score:.3f}, doc={meta.get('doc_title','')})\\n{text}\"\n",
        "        for i, (text, score, meta) in enumerate(retrieved)\n",
        "    )\n",
        "\n",
        "    # 3) Generate\n",
        "    rag_prompt = f\"\"\"{system_prompt}\n",
        "\n",
        "Use ONLY the context below. If the answer isn't in the context, say you don't know.\n",
        "\n",
        "Context:\n",
        "{context}\n",
        "\n",
        "Question: {question}\n",
        "\n",
        "Answer:\"\"\"\n",
        "\n",
        "    answer = generate(rag_prompt, label=\"rag_query\")\n",
        "    return answer, retrieved\n",
        "\n",
        "def route_filter(question: str):\n",
        "    q = question.lower()\n",
        "\n",
        "    if any(k in q for k in [\"vat\", \"upgrade\", \"downgrade\", \"billing\", \"price\", \"eur\", \"per user\"]):\n",
        "        return {\"doc_title_contains\": [\"Billing\", \"Pricing\", \"Refund\"]}\n",
        "\n",
        "    if \"refund\" in q and any(k in q for k in [\"policy\", \"timeframe\", \"window\", \"14\", \"days\", \"partial\"]):\n",
        "        return {\"doc_title_contains\": [\"Billing\", \"Pricing\", \"Refund\"]}\n",
        "\n",
        "    if any(k in q for k in [\"sso\", \"saml\", \"enterprise\", \"idp\", \"metadata url\", \"attribute mapping\"]):\n",
        "        return {\"doc_title_contains\": [\"SSO\"]}\n",
        "\n",
        "    if any(k in q for k in [\"csv\", \"api\", \"100,000\", \"100000\"]) or (\"export\" in q and any(k in q for k in [\"limit\", \"rows\", \"failed\", \"error\"])):\n",
        "        return {\"doc_title_contains\": [\"Exports\", \"Capabilities\"]}\n",
        "\n",
        "    if any(k in q for k in [\"browser\", \"internet explorer\", \"chrome\", \"firefox\", \"edge\", \"safari\"]):\n",
        "        return {\"doc_title_contains\": [\"Browser\", \"Capabilities\"]}\n",
        "\n",
        "    if any(k in q for k in [\"2fa\", \"two-factor\", \"authy\", \"google authenticator\", \"sms\", \"recovery\"]):\n",
        "        return {\"doc_title_contains\": [\"2FA\"]}\n",
        "\n",
        "    if any(k in q for k in [\"locked\", \"lockout\", \"failed login\"]) or (\"password\" in q and \"reset\" in q):\n",
        "            return {\"doc_title_contains\": [\"Lockout\", \"Password\", \"Account\"]}\n",
        "\n",
        "    return None\n",
        "\n",
        "\n",
        "\n",
        "# ── Hybrid Retrieval ────────\n",
        "from sklearn.feature_extraction.text import TfidfVectorizer\n",
        "\n",
        "def build_sparse_index(chunks):\n",
        "    \"\"\"Build TF-IDF sparse index for keyword retrieval.\"\"\"\n",
        "    vectorizer = TfidfVectorizer(stop_words='english')\n",
        "    matrix = vectorizer.fit_transform(chunks)\n",
        "    return vectorizer, matrix\n",
        "\n",
        "def hybrid_search(question, chunks, chunk_embeddings, chunk_metadata,\n",
        "                  tfidf_vectorizer, tfidf_matrix, top_k=3, alpha=0.7):\n",
        "    # Dense scores\n",
        "    q_emb = embed_texts(question, task_type=\"RETRIEVAL_QUERY\")[0]\n",
        "    q_emb = q_emb / np.linalg.norm(q_emb)\n",
        "    dense_scores = chunk_embeddings @ q_emb\n",
        "\n",
        "    # Sparse scores (TF-IDF)\n",
        "    q_tfidf = tfidf_vectorizer.transform([question])\n",
        "    sparse_scores = (tfidf_matrix @ q_tfidf.T).toarray().ravel()\n",
        "\n",
        "    # Normalize + fuse\n",
        "    dense_n = minmax_scale(dense_scores)\n",
        "    sparse_n = minmax_scale(sparse_scores)\n",
        "    fused = alpha * dense_n + (1 - alpha) * sparse_n\n",
        "\n",
        "    top_idx = np.argsort(-fused)[:top_k]\n",
        "    return [(chunks[i], float(fused[i]), chunk_metadata[i]) for i in top_idx]\n",
        "\n",
        "\n",
        "def search_with_filter(question, chunks, chunk_embeddings, chunk_metadata, top_k=3, doc_title_contains=None):\n",
        "    # pick candidate indices based on metadata\n",
        "    candidate_idx = list(range(len(chunks)))\n",
        "    if doc_title_contains:\n",
        "        candidate_idx = [\n",
        "            i for i, m in enumerate(chunk_metadata)\n",
        "            if any(key.lower() in m[\"doc_title\"].lower() for key in doc_title_contains)\n",
        "        ]\n",
        "        if not candidate_idx:  # fallback if filter too strict\n",
        "            candidate_idx = list(range(len(chunks)))\n",
        "\n",
        "    # dense retrieval over candidates\n",
        "    q_emb = embed_texts(question, task_type=\"RETRIEVAL_QUERY\")[0]\n",
        "    cand_emb = chunk_embeddings[candidate_idx]  # assumes np.array (n, d)\n",
        "    scores = cand_emb @ (q_emb / np.linalg.norm(q_emb))  # or cosine function\n",
        "\n",
        "    ranked = sorted(zip(candidate_idx, scores), key=lambda x: x[1], reverse=True)[:top_k]\n",
        "    return [(chunks[i], float(s), chunk_metadata[i]) for i, s in ranked]\n",
        "\n",
        "# ── Retrieval Metrics ─────────────────────────────────────\n",
        "\n",
        "def precision_recall_at_k(retrieved_indices, expected_indices, k):\n",
        "    \"\"\"Compute Precision@k and Recall@k.\"\"\"\n",
        "    retrieved_set = set(retrieved_indices[:k])\n",
        "    expected_set = set(expected_indices)\n",
        "    if not expected_set:\n",
        "        return None, None\n",
        "    hits = retrieved_set & expected_set\n",
        "    precision = len(hits) / k if k > 0 else 0\n",
        "    recall = len(hits) / len(expected_set)\n",
        "    return precision, recall\n",
        "\n",
        "print(\"All infrastructure ready (includes hybrid retrieval, metadata filtering, and retrieval metrics).\")"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 4,
      "id": "85001cfe",
      "metadata": {
        "id": "85001cfe"
      },
      "outputs": [],
      "source": [
        "class RAGTriadScore(BaseModel):\n",
        "    \"\"\"RAG triad evaluation for a single question.\"\"\"\n",
        "    context_relevance: int = Field(description=\"1-5: Are the retrieved chunks relevant to the question?\")\n",
        "    groundedness: int = Field(description=\"1-5: Is the answer supported by the retrieved chunks?\")\n",
        "    answer_relevance: int = Field(description=\"1-5: Does the answer address the question asked?\")\n",
        "    explanation: str = Field(description=\"Brief explanation of all three scores\")\n",
        "\n",
        "def evaluate_rag_triad(question, answer, retrieved_chunks, label=\"triad\"):\n",
        "    \"\"\"Score a single RAG output on all three triad dimensions.\"\"\"\n",
        "    chunks_text = \"\\n\\n\".join(f\"[Chunk {i+1}]: {c}\" for i, c in enumerate(retrieved_chunks))\n",
        "\n",
        "    eval_prompt = f\"\"\"\n",
        "You are evaluating a Retrieval-Augmented Generation (RAG) output.\n",
        "\n",
        "Return ONLY valid JSON (no markdown, no code fences, no extra keys) with this exact schema:\n",
        "{{\n",
        "  \"context_relevance\": 1-5,\n",
        "  \"groundedness\": 1-5,\n",
        "  \"answer_relevance\": 1-5,\n",
        "  \"explanation\": \"brief explanation covering all three scores\"\n",
        "}}\n",
        "\n",
        "Scoring guidance:\n",
        "- context_relevance: Do the retrieved chunks contain information needed to answer the question?\n",
        "- groundedness: Is every claim in the answer supported by the retrieved chunks? (5 = fully supported, 1 = major hallucination)\n",
        "- answer_relevance: Does the answer actually address what was asked?\n",
        "\n",
        "Question:\n",
        "{question}\n",
        "\n",
        "Retrieved Context:\n",
        "{chunks_text}\n",
        "\n",
        "Generated Answer:\n",
        "{answer}\n",
        "\"\"\"\n",
        "    return generate_structured(eval_prompt, RAGTriadScore, label=label)\n"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "d78f3fa0",
      "metadata": {
        "id": "d78f3fa0"
      },
      "source": [
        "---\n",
        "## Part 1: Knowledge Base Design (15 points)\n",
        "\n",
        "Design your knowledge base with 4-6 documents. You may reuse and expand documents from Lab 2 or create new ones.\n",
        "\n",
        "**Document your decisions:**\n",
        "- Domain and why you chose it\n",
        "- Chunk size, overlap, and rationale\n",
        "- Number of resulting chunks"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 5,
      "id": "c2548b2b",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "c2548b2b",
        "outputId": "87fa0b01-b415-435d-8363-b5678ae060db"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Documents: 6\n",
            "Total words: 1184\n"
          ]
        }
      ],
      "source": [
        "# ── Knowledge base (CloudBase internal documentation excerpts) ───────────────\n",
        "# 5 documents, ~1,200 total words. Each document starts with a title line.\n",
        "\n",
        "documents = [\n",
        "    \"\"\"Document 1 — CloudBase Account Access, Login Flow, and Password Reset\n",
        "\n",
        "CloudBase accounts are accessed through the web-based login page using a registered email address and password. Each user account is tied to a unique email address that serves as the primary identifier. To protect user privacy and security, CloudBase does not publicly confirm whether a specific email address is registered beyond the standard login and reset workflows.\n",
        "\n",
        "If a user forgets their password, they can initiate a self-service recovery by clicking the “Forgot Password” link on the login page. This action triggers CloudBase to send a password reset email to the registered email address. The email contains a secure reset link that allows the user to choose a new password. For security reasons, the reset link is valid for only 24 hours. If the link expires before use, the user must request a new reset email.\n",
        "\n",
        "Users are advised to check spam or junk folders if the reset email does not appear promptly. Mailbox filters or corporate email security tools may delay or block automated messages. Password resets are designed to minimize downtime and typically do not require administrator intervention. If repeated reset attempts fail or the user no longer has access to the registered email address, the recommended next step is to contact CloudBase support through the official support channels described in the billing or support documentation.\n",
        "\"\"\",\n",
        "\n",
        "    \"\"\"Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking\n",
        "\n",
        "CloudBase implements automatic account lockouts to protect users from unauthorized access and brute-force password attacks. An account is locked after five consecutive unsuccessful login attempts. These attempts may occur in a short period or across multiple sessions. Once the lockout threshold is reached, the account becomes temporarily inaccessible.\n",
        "\n",
        "The lockout duration is 30 minutes. During this time, all login attempts will fail, even if the correct password is provided. This policy ensures that attackers cannot continue guessing passwords once suspicious behavior is detected. After the 30-minute lockout period expires, the user may log in again normally without taking additional action.\n",
        "\n",
        "If access is required immediately, CloudBase support can assist with unlocking the account after verifying the user’s identity. This option is particularly useful for business-critical accounts or administrators. Users who experience frequent lockouts are encouraged to reset their passwords using the “Forgot Password” feature and confirm that they are entering the correct credentials.\n",
        "\n",
        "Administrators should consider enabling additional security measures, such as two-factor authentication (2FA), for users with elevated privileges or access to sensitive data. Lockout policies work best when combined with strong passwords, unique credentials, and user awareness about secure login practices.\n",
        "\"\"\",\n",
        "\n",
        "    \"\"\"Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options\n",
        "\n",
        "Two-factor authentication (2FA) provides an additional security layer beyond passwords. CloudBase supports 2FA to reduce the risk of unauthorized access, even if login credentials are compromised. Users can enable 2FA by navigating to Settings and selecting the Security section. Once enabled, 2FA is required during each login attempt.\n",
        "\n",
        "CloudBase supports two primary 2FA methods. The first is authenticator apps, such as Google Authenticator or Authy, which generate time-based one-time codes. The second method is SMS-based verification, where a one-time code is sent to the user’s registered phone number. Users may choose the method that best fits their security and accessibility needs.\n",
        "\n",
        "During the 2FA setup process, CloudBase provides recovery codes. Recovery codes are intended for use when the primary 2FA method is unavailable, such as when a phone is lost or replaced. These codes should be stored securely, as they can restore account access without a live 2FA prompt.\n",
        "\n",
        "For security reasons, CloudBase does not provide instructions on bypassing 2FA controls. If a user cannot access their authenticator device, phone number, or recovery codes, the recommended next step is to contact CloudBase support for identity-verified recovery assistance.\n",
        "\"\"\",\n",
        "\n",
        "    \"\"\"Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention\n",
        "\n",
        "CloudBase offers multiple subscription plans designed for different organization sizes and usage needs. The Free plan is intended for small teams or evaluation purposes and supports up to five users. It includes a total of 1 GB of shared storage across the account. The Team plan supports larger teams, allowing up to 100 users and providing 50 GB of total shared storage.\n",
        "\n",
        "When user or storage limits are exceeded, the organization must upgrade to a higher plan to continue adding users or storing additional data. Exceeding limits may restrict certain actions until an upgrade is completed. Plan limits should be reviewed carefully before onboarding large teams or importing significant amounts of data.\n",
        "\n",
        "Data retention policies vary by plan. The Free plan does not include backups for deleted data. Once data is deleted under the Free plan, it cannot be restored. The Team plan includes backups for deleted data with a retention period of 30 days. This allows organizations to recover accidentally deleted items within that window.\n",
        "\n",
        "Retention policies are especially important for teams that rely on historical data or compliance-related records. Administrators should ensure that their selected plan aligns with their data protection and recovery requirements.\n",
        "\"\"\",\n",
        "\n",
        "    \"\"\"Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes\n",
        "\n",
        "CloudBase pricing for the Team plan is set at EUR 12 per user per month. Pricing information listed in the Billing & Subscription Guide excludes VAT. VAT application depends on the customer’s billing address, tax status, and applicable regional regulations, but the base plan price itself is exclusive of VAT.\n",
        "\n",
        "Subscription upgrades and downgrades follow defined timing rules. Upgrades take effect immediately. When an upgrade occurs, CloudBase charges a prorated amount for the remainder of the current billing cycle, reflecting the higher plan level from the time of the change. Downgrades, by contrast, do not take effect immediately and instead apply at the start of the next billing cycle.\n",
        "\n",
        "CloudBase does not issue partial refunds for the current billing cycle when a downgrade occurs. Refund eligibility is limited to a specific timeframe. Full refunds are available within 14 days of an initial subscription purchase or a plan upgrade. After the 14-day window has passed, refunds are not provided. These policies help ensure billing transparency and predictable revenue management.\n",
        "\"\"\",\n",
        "\n",
        "    \"\"\"Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO\n",
        "\n",
        "CloudBase supports data export for reporting and operational use but enforces limits to maintain performance. CSV exports are limited to 100,000 rows per export. When datasets exceed this limit, CloudBase recommends using the API to retrieve data. The API is designed for large-scale data access, automation, and integration with external systems.\n",
        "\n",
        "CloudBase supports modern web browsers to ensure security and performance. Supported browsers include Chrome version 90 and newer, Firefox version 88 and newer, Edge version 90 and newer, and Safari version 15 and newer. Internet Explorer is not supported in any version. For mobile users, CloudBase provides a dedicated mobile application compatible with iOS 15+ and Android 12+.\n",
        "\n",
        "For enterprise customers, CloudBase supports Single Sign-On (SSO) using SAML 2.0. SSO is available on Enterprise plans. Setting up SSO requires an Identity Provider metadata URL, attribute mapping, and administrator approval. The configuration process typically takes 1–2 business days, depending on validation and setup complexity.\n",
        "\"\"\",\n",
        "]\n",
        "\n",
        "total_words = sum(len(doc.split()) for doc in documents)\n",
        "print(f\"Documents: {len(documents)}\")\n",
        "print(f\"Total words: {total_words}\")\n",
        "assert total_words >= 1000, f\"Need at least 1,000 words, have {total_words}\""
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 6,
      "id": "1286eaef",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "1286eaef",
        "outputId": "9106e4ee-8784-48c2-d90d-e07f014cee04"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Chunking: max_words=90, overlap=25\n",
            "Total chunks: 20\n",
            "Dense index:  20 embeddings\n",
            "Sparse index: 406 vocabulary terms\n",
            "\n",
            "Chunks per document:\n",
            "  Doc 0: 4 chunks\n",
            "  Doc 1: 3 chunks\n",
            "  Doc 2: 3 chunks\n",
            "  Doc 3: 4 chunks\n",
            "  Doc 4: 3 chunks\n",
            "  Doc 5: 3 chunks\n",
            "\n",
            "\n",
            "We chunked documents into ~90-word sentence-based chunks with a 25-word overlap.\n",
            "\n",
            "- Chunk size (90 words): The expanded knowledge base contains multi-sentence policy explanations followed by concrete constraints (e.g., limits, durations, prices). A ~90-word chunk is large enough to preserve the full meaning of a policy section (such as account lockouts, refunds, or exports) while remaining small enough to avoid mixing unrelated topics.\n",
            "- Overlap (25 words): Key conditions and exceptions often appear at sentence or paragraph boundaries (for example, a policy followed by timing rules or recovery options). A 25-word overlap ensures these dependencies are not split across chunks, improving retrieval robustness.\n",
            "- Trade-off validation: Smaller chunks (≤60 words) fragmented policy logic and increased partial matches. Larger chunks (≥120 words) reduced retrieval precision by combining multiple concepts (e.g., pricing and refunds). The chosen values balance semantic coherence with accurate, grounded retrieval.\n",
            "\n"
          ]
        }
      ],
      "source": [
        "# ── Chunking configuration ────────────────────────────────────────────────\n",
        "# We use sentence-based chunking to keep policy statements and numeric limits intact.\n",
        "\n",
        "CHUNK_MAX_WORDS = 90      # Target chunk size (words)\n",
        "CHUNK_OVERLAP_WORDS = 25  # Overlap to reduce boundary-splitting of key details\n",
        "\n",
        "# Document your rationale (used later in the playbook write-up):\n",
        "CHUNKING_RATIONALE = \"\"\"\n",
        "We chunked documents into ~90-word sentence-based chunks with a 25-word overlap.\n",
        "\n",
        "- Chunk size (90 words): The expanded knowledge base contains multi-sentence policy explanations followed by concrete constraints (e.g., limits, durations, prices). A ~90-word chunk is large enough to preserve the full meaning of a policy section (such as account lockouts, refunds, or exports) while remaining small enough to avoid mixing unrelated topics.\n",
        "- Overlap (25 words): Key conditions and exceptions often appear at sentence or paragraph boundaries (for example, a policy followed by timing rules or recovery options). A 25-word overlap ensures these dependencies are not split across chunks, improving retrieval robustness.\n",
        "- Trade-off validation: Smaller chunks (≤60 words) fragmented policy logic and increased partial matches. Larger chunks (≥120 words) reduced retrieval precision by combining multiple concepts (e.g., pricing and refunds). The chosen values balance semantic coherence with accurate, grounded retrieval.\n",
        "\"\"\"\n",
        "# ── Chunking the documents ────────────────────────────────────────────────\n",
        "all_chunks = []\n",
        "chunk_sources = []    # Track which document each chunk came from\n",
        "chunk_metadata = []   # Rich metadata for each chunk\n",
        "for doc_idx, doc in enumerate(documents):\n",
        "    doc_title = doc.strip().split('\\n')[0]  # First line as title\n",
        "    chunks = chunk_sentences(doc, max_words=CHUNK_MAX_WORDS, overlap_words=CHUNK_OVERLAP_WORDS)\n",
        "    for chunk_idx, chunk in enumerate(chunks):\n",
        "        all_chunks.append(chunk)\n",
        "        chunk_sources.append(doc_idx)\n",
        "        chunk_metadata.append({\n",
        "            \"chunk_id\": f\"DOC{doc_idx}::C{chunk_idx+1}\",\n",
        "            \"doc_index\": doc_idx,\n",
        "            \"doc_title\": doc_title,\n",
        "        })\n",
        "\n",
        "chunk_embeddings = np.vstack(embed_texts(all_chunks, task_type=\"RETRIEVAL_DOCUMENT\"))\n",
        "chunk_embeddings = chunk_embeddings / np.linalg.norm(\n",
        "    chunk_embeddings, axis=1, keepdims=True\n",
        ")\n",
        "\n",
        "\n",
        "# build sparse index for hybrid retrieval\n",
        "tfidf_vectorizer, tfidf_matrix = build_sparse_index(all_chunks)\n",
        "\n",
        "print(f\"Chunking: max_words={CHUNK_MAX_WORDS}, overlap={CHUNK_OVERLAP_WORDS}\")\n",
        "print(f\"Total chunks: {len(all_chunks)}\")\n",
        "print(f\"Dense index:  {len(chunk_embeddings)} embeddings\")\n",
        "print(f\"Sparse index: {tfidf_matrix.shape[1]} vocabulary terms\")\n",
        "print(f\"\\nChunks per document:\")\n",
        "for i in range(len(documents)):\n",
        "    count = chunk_sources.count(i)\n",
        "    print(f\"  Doc {i}: {count} chunks\")\n",
        "print(f\"\\n{CHUNKING_RATIONALE}\")\n",
        "\n"
      ]
    },
    {
      "cell_type": "code",
      "source": [
        "# ===== Inspect chunks (for debugging and finding Golden Set expected chunks) =====\n",
        "\n",
        "for i, text in enumerate(all_chunks):\n",
        "    meta = chunk_metadata[i]\n",
        "    print(\"=\" * 90)\n",
        "    print(f\"{i} | {meta['chunk_id']} | {meta['doc_title']}\")\n",
        "    print(\"-\" * 90)\n",
        "    print(text)\n",
        "    print()\n"
      ],
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "71vCFxtYJMt_",
        "outputId": "ecc65c54-3ebd-4a69-907a-3608ddb81fe0"
      },
      "id": "71vCFxtYJMt_",
      "execution_count": 7,
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "==========================================================================================\n",
            "0 | DOC0::C1 | Document 1 — CloudBase Account Access, Login Flow, and Password Reset\n",
            "------------------------------------------------------------------------------------------\n",
            "Document 1 — CloudBase Account Access, Login Flow, and Password Reset CloudBase accounts are accessed through the web-based login page using a registered email address and password. Each user account is tied to a unique email address that serves as the primary identifier. To protect user privacy and security, CloudBase does not publicly confirm whether a specific email address is registered beyond the standard login and reset workflows. If a user forgets their password, they can initiate a self-service recovery by clicking the “Forgot Password” link on the login page.\n",
            "\n",
            "==========================================================================================\n",
            "1 | DOC0::C2 | Document 1 — CloudBase Account Access, Login Flow, and Password Reset\n",
            "------------------------------------------------------------------------------------------\n",
            "and reset workflows. If a user forgets their password, they can initiate a self-service recovery by clicking the “Forgot Password” link on the login page. This action triggers CloudBase to send a password reset email to the registered email address. The email contains a secure reset link that allows the user to choose a new password. For security reasons, the reset link is valid for only 24 hours. If the link expires before use, the user must request a new reset email.\n",
            "\n",
            "==========================================================================================\n",
            "2 | DOC0::C3 | Document 1 — CloudBase Account Access, Login Flow, and Password Reset\n",
            "------------------------------------------------------------------------------------------\n",
            "security reasons, the reset link is valid for only 24 hours. If the link expires before use, the user must request a new reset email. Users are advised to check spam or junk folders if the reset email does not appear promptly. Mailbox filters or corporate email security tools may delay or block automated messages. Password resets are designed to minimize downtime and typically do not require administrator intervention.\n",
            "\n",
            "==========================================================================================\n",
            "3 | DOC0::C4 | Document 1 — CloudBase Account Access, Login Flow, and Password Reset\n",
            "------------------------------------------------------------------------------------------\n",
            "or corporate email security tools may delay or block automated messages. Password resets are designed to minimize downtime and typically do not require administrator intervention. If repeated reset attempts fail or the user no longer has access to the registered email address, the recommended next step is to contact CloudBase support through the official support channels described in the billing or support documentation.\n",
            "\n",
            "==========================================================================================\n",
            "4 | DOC1::C1 | Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking\n",
            "------------------------------------------------------------------------------------------\n",
            "Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking CloudBase implements automatic account lockouts to protect users from unauthorized access and brute-force password attacks. An account is locked after five consecutive unsuccessful login attempts. These attempts may occur in a short period or across multiple sessions. Once the lockout threshold is reached, the account becomes temporarily inaccessible. The lockout duration is 30 minutes. During this time, all login attempts will fail, even if the correct password is provided.\n",
            "\n",
            "==========================================================================================\n",
            "5 | DOC1::C2 | Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking\n",
            "------------------------------------------------------------------------------------------\n",
            "account becomes temporarily inaccessible. The lockout duration is 30 minutes. During this time, all login attempts will fail, even if the correct password is provided. This policy ensures that attackers cannot continue guessing passwords once suspicious behavior is detected. After the 30-minute lockout period expires, the user may log in again normally without taking additional action. If access is required immediately, CloudBase support can assist with unlocking the account after verifying the user’s identity. This option is particularly useful for business-critical accounts or administrators.\n",
            "\n",
            "==========================================================================================\n",
            "6 | DOC1::C3 | Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking\n",
            "------------------------------------------------------------------------------------------\n",
            "required immediately, CloudBase support can assist with unlocking the account after verifying the user’s identity. This option is particularly useful for business-critical accounts or administrators. Users who experience frequent lockouts are encouraged to reset their passwords using the “Forgot Password” feature and confirm that they are entering the correct credentials. Administrators should consider enabling additional security measures, such as two-factor authentication (2FA), for users with elevated privileges or access to sensitive data. Lockout policies work best when combined with strong passwords, unique credentials, and user awareness about secure login practices.\n",
            "\n",
            "==========================================================================================\n",
            "7 | DOC2::C1 | Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options\n",
            "------------------------------------------------------------------------------------------\n",
            "Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options Two-factor authentication (2FA) provides an additional security layer beyond passwords. CloudBase supports 2FA to reduce the risk of unauthorized access, even if login credentials are compromised. Users can enable 2FA by navigating to Settings and selecting the Security section. Once enabled, 2FA is required during each login attempt. CloudBase supports two primary 2FA methods. The first is authenticator apps, such as Google Authenticator or Authy, which generate time-based one-time codes.\n",
            "\n",
            "==========================================================================================\n",
            "8 | DOC2::C2 | Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options\n",
            "------------------------------------------------------------------------------------------\n",
            "each login attempt. CloudBase supports two primary 2FA methods. The first is authenticator apps, such as Google Authenticator or Authy, which generate time-based one-time codes. The second method is SMS-based verification, where a one-time code is sent to the user’s registered phone number. Users may choose the method that best fits their security and accessibility needs. During the 2FA setup process, CloudBase provides recovery codes. Recovery codes are intended for use when the primary 2FA method is unavailable, such as when a phone is lost or replaced.\n",
            "\n",
            "==========================================================================================\n",
            "9 | DOC2::C3 | Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options\n",
            "------------------------------------------------------------------------------------------\n",
            "provides recovery codes. Recovery codes are intended for use when the primary 2FA method is unavailable, such as when a phone is lost or replaced. These codes should be stored securely, as they can restore account access without a live 2FA prompt. For security reasons, CloudBase does not provide instructions on bypassing 2FA controls. If a user cannot access their authenticator device, phone number, or recovery codes, the recommended next step is to contact CloudBase support for identity-verified recovery assistance.\n",
            "\n",
            "==========================================================================================\n",
            "10 | DOC3::C1 | Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention\n",
            "------------------------------------------------------------------------------------------\n",
            "Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention CloudBase offers multiple subscription plans designed for different organization sizes and usage needs. The Free plan is intended for small teams or evaluation purposes and supports up to five users. It includes a total of 1 GB of shared storage across the account. The Team plan supports larger teams, allowing up to 100 users and providing 50 GB of total shared storage.\n",
            "\n",
            "==========================================================================================\n",
            "11 | DOC3::C2 | Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention\n",
            "------------------------------------------------------------------------------------------\n",
            "of shared storage across the account. The Team plan supports larger teams, allowing up to 100 users and providing 50 GB of total shared storage. When user or storage limits are exceeded, the organization must upgrade to a higher plan to continue adding users or storing additional data. Exceeding limits may restrict certain actions until an upgrade is completed. Plan limits should be reviewed carefully before onboarding large teams or importing significant amounts of data. Data retention policies vary by plan.\n",
            "\n",
            "==========================================================================================\n",
            "12 | DOC3::C3 | Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention\n",
            "------------------------------------------------------------------------------------------\n",
            "upgrade is completed. Plan limits should be reviewed carefully before onboarding large teams or importing significant amounts of data. Data retention policies vary by plan. The Free plan does not include backups for deleted data. Once data is deleted under the Free plan, it cannot be restored. The Team plan includes backups for deleted data with a retention period of 30 days. This allows organizations to recover accidentally deleted items within that window. Retention policies are especially important for teams that rely on historical data or compliance-related records.\n",
            "\n",
            "==========================================================================================\n",
            "13 | DOC3::C4 | Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention\n",
            "------------------------------------------------------------------------------------------\n",
            "allows organizations to recover accidentally deleted items within that window. Retention policies are especially important for teams that rely on historical data or compliance-related records. Administrators should ensure that their selected plan aligns with their data protection and recovery requirements.\n",
            "\n",
            "==========================================================================================\n",
            "14 | DOC4::C1 | Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes\n",
            "------------------------------------------------------------------------------------------\n",
            "Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes CloudBase pricing for the Team plan is set at EUR 12 per user per month. Pricing information listed in the Billing & Subscription Guide excludes VAT. VAT application depends on the customer’s billing address, tax status, and applicable regional regulations, but the base plan price itself is exclusive of VAT. Subscription upgrades and downgrades follow defined timing rules. Upgrades take effect immediately.\n",
            "\n",
            "==========================================================================================\n",
            "15 | DOC4::C2 | Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes\n",
            "------------------------------------------------------------------------------------------\n",
            "applicable regional regulations, but the base plan price itself is exclusive of VAT. Subscription upgrades and downgrades follow defined timing rules. Upgrades take effect immediately. When an upgrade occurs, CloudBase charges a prorated amount for the remainder of the current billing cycle, reflecting the higher plan level from the time of the change. Downgrades, by contrast, do not take effect immediately and instead apply at the start of the next billing cycle. CloudBase does not issue partial refunds for the current billing cycle when a downgrade occurs.\n",
            "\n",
            "==========================================================================================\n",
            "16 | DOC4::C3 | Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes\n",
            "------------------------------------------------------------------------------------------\n",
            "instead apply at the start of the next billing cycle. CloudBase does not issue partial refunds for the current billing cycle when a downgrade occurs. Refund eligibility is limited to a specific timeframe. Full refunds are available within 14 days of an initial subscription purchase or a plan upgrade. After the 14-day window has passed, refunds are not provided. These policies help ensure billing transparency and predictable revenue management.\n",
            "\n",
            "==========================================================================================\n",
            "17 | DOC5::C1 | Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO\n",
            "------------------------------------------------------------------------------------------\n",
            "Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO CloudBase supports data export for reporting and operational use but enforces limits to maintain performance. CSV exports are limited to 100,000 rows per export. When datasets exceed this limit, CloudBase recommends using the API to retrieve data. The API is designed for large-scale data access, automation, and integration with external systems. CloudBase supports modern web browsers to ensure security and performance.\n",
            "\n",
            "==========================================================================================\n",
            "18 | DOC5::C2 | Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO\n",
            "------------------------------------------------------------------------------------------\n",
            "data. The API is designed for large-scale data access, automation, and integration with external systems. CloudBase supports modern web browsers to ensure security and performance. Supported browsers include Chrome version 90 and newer, Firefox version 88 and newer, Edge version 90 and newer, and Safari version 15 and newer. Internet Explorer is not supported in any version. For mobile users, CloudBase provides a dedicated mobile application compatible with iOS 15+ and Android 12+. For enterprise customers, CloudBase supports Single Sign-On (SSO) using SAML 2.0. SSO is available on Enterprise plans.\n",
            "\n",
            "==========================================================================================\n",
            "19 | DOC5::C3 | Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO\n",
            "------------------------------------------------------------------------------------------\n",
            "application compatible with iOS 15+ and Android 12+. For enterprise customers, CloudBase supports Single Sign-On (SSO) using SAML 2.0. SSO is available on Enterprise plans. Setting up SSO requires an Identity Provider metadata URL, attribute mapping, and administrator approval. The configuration process typically takes 1–2 business days, depending on validation and setup complexity.\n",
            "\n"
          ]
        }
      ]
    },
    {
      "cell_type": "markdown",
      "id": "ad58bbfe",
      "metadata": {
        "id": "ad58bbfe"
      },
      "source": [
        "---\n",
        "## Part 2: RAG System Prompt (25 points)\n",
        "\n",
        "Your prompt must include ALL 8 components: role, task, grounding rules, scope, format, refusal, examples, conflict handling."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 8,
      "id": "7767d978",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "7767d978",
        "outputId": "1d71821a-3bcf-4552-9a90-12e11cb0a22e"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "System prompt:\n",
            "You are CloudBase Support RAG, a customer-support assistant for the CloudBase product.\n",
            "\n",
            "TASK\n",
            "Answer the user's question using ONLY the information in the provided Context sources.\n",
            "\n",
            "GROUNDING RULES (NON-NEGOTIABLE)\n",
            "- Use ONLY facts that appear in the Context (the text labeled [Source 0], [Source 1], etc.).\n",
            "- If the answer is not explicitly in the Context, reply exactly:\n",
            "  \"I don't have that information in the provided documents.\"\n",
            "- Do not guess, do not use outside knowledge, and do not invent product features, UI labels, prices, policies, timelines, or limits.\n",
            "- If sources conflict, prefer the most specific statement; if still ambiguous, say you don't have that information.\n",
            "\n",
            "SCOPE\n",
            "- In scope: CloudBase account security, 2FA, lockouts, plans/pricing/billing, refunds, exports/API limits, SSO, supported browsers/mobile, and basic troubleshooting as described in the Context.\n",
            "- Out of scope: Anything not in Context (e.g., phone numbers, unpublished features, legal advice, internal roadmaps).\n",
            "\n",
            "SAFETY / MISUSE\n",
            "- If the user requests hacking, bypassing security controls, or wrongdoing, refuse briefly and provide the safest official alternative that IS in the Context (e.g., recovery codes, support contact).\n",
            "\n",
            "FORMAT\n",
            "- Keep the answer under 90 words.\n",
            "- Use exact numbers/time windows/limits as written (e.g., “100,000 rows”, “14 days”, “30 minutes”).\n",
            "- Cite each key claim with the relevant source tag, e.g., [Source 0]. If multiple sources support a claim, cite multiple tags.\n",
            "\n",
            "EXAMPLE (GOOD)\n",
            "Q: What is the CSV export limit and what should I use instead?\n",
            "A: CSV exports are limited to 100,000 rows. For larger datasets, use the CloudBase API. [Source 0]\n",
            "\n",
            "Now answer the user question.\n",
            "\n",
            "Prompt length: 261 words\n"
          ]
        }
      ],
      "source": [
        "# ── Final production system prompt ───────────────────────────────────────\n",
        "# This prompt is designed for *grounded* customer support responses with citations.\n",
        "\n",
        "FINAL_SYSTEM_PROMPT = \"\"\"You are CloudBase Support RAG, a customer-support assistant for the CloudBase product.\n",
        "\n",
        "TASK\n",
        "Answer the user's question using ONLY the information in the provided Context sources.\n",
        "\n",
        "GROUNDING RULES (NON-NEGOTIABLE)\n",
        "- Use ONLY facts that appear in the Context (the text labeled [Source 0], [Source 1], etc.).\n",
        "- If the answer is not explicitly in the Context, reply exactly:\n",
        "  \"I don't have that information in the provided documents.\"\n",
        "- Do not guess, do not use outside knowledge, and do not invent product features, UI labels, prices, policies, timelines, or limits.\n",
        "- If sources conflict, prefer the most specific statement; if still ambiguous, say you don't have that information.\n",
        "\n",
        "SCOPE\n",
        "- In scope: CloudBase account security, 2FA, lockouts, plans/pricing/billing, refunds, exports/API limits, SSO, supported browsers/mobile, and basic troubleshooting as described in the Context.\n",
        "- Out of scope: Anything not in Context (e.g., phone numbers, unpublished features, legal advice, internal roadmaps).\n",
        "\n",
        "SAFETY / MISUSE\n",
        "- If the user requests hacking, bypassing security controls, or wrongdoing, refuse briefly and provide the safest official alternative that IS in the Context (e.g., recovery codes, support contact).\n",
        "\n",
        "FORMAT\n",
        "- Keep the answer under 90 words.\n",
        "- Use exact numbers/time windows/limits as written (e.g., “100,000 rows”, “14 days”, “30 minutes”).\n",
        "- Cite each key claim with the relevant source tag, e.g., [Source 0]. If multiple sources support a claim, cite multiple tags.\n",
        "\n",
        "EXAMPLE (GOOD)\n",
        "Q: What is the CSV export limit and what should I use instead?\n",
        "A: CSV exports are limited to 100,000 rows. For larger datasets, use the CloudBase API. [Source 0]\n",
        "\n",
        "Now answer the user question.\"\"\"\n",
        "\n",
        "print(\"System prompt:\")\n",
        "print(FINAL_SYSTEM_PROMPT)\n",
        "print(f\"\\nPrompt length: {len(FINAL_SYSTEM_PROMPT.split())} words\")\n"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "3208d190",
      "metadata": {
        "id": "3208d190"
      },
      "source": [
        "---\n",
        "## Part 3: Input Questions (15+)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 9,
      "id": "2916dfad",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "2916dfad",
        "outputId": "3a5ed047-5dee-4676-c073-d132ce27925f"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Questions: 16\n"
          ]
        }
      ],
      "source": [
        "# ── Questions to process (15+ required) ──────────────────────────────────\n",
        "QUESTIONS = [\n",
        "    \"How do I reset my CloudBase password?\",\n",
        "    \"How long is a password reset link valid?\",\n",
        "    \"What triggers an account lockout and how long does it last?\",\n",
        "    \"How do I enable two-factor authentication (2FA) and what methods are supported?\",\n",
        "    \"If I lose my phone, can I bypass 2FA?\",\n",
        "    \"What are the Free plan limits for users and storage?\",\n",
        "    \"What are the Team plan limits for users and storage?\",\n",
        "    \"What is the Team plan price per user per month, and does it include VAT?\",\n",
        "    \"How do upgrades and downgrades take effect? Are partial refunds offered?\",\n",
        "    \"What is the refund policy timeframe for an initial purchase or upgrade?\",\n",
        "    \"My CSV export failed\\u2014what is the CSV row limit and what should I use instead?\",\n",
        "    \"If a CSV export fails below the limit, what should I try before contacting support?\",\n",
        "    \"Does CloudBase support SSO and what is required to set it up?\",\n",
        "    \"Which browsers are supported? Is Internet Explorer supported?\",\n",
        "    \"Is there a CloudBase mobile app and what OS versions are supported?\",\n",
        "    \"What is the phone number for CloudBase support?\"\n",
        "]\n",
        "\n",
        "assert len(QUESTIONS) >= 15, f\"Need at least 15 questions, have {len(QUESTIONS)}\"\n",
        "print(f\"Questions: {len(QUESTIONS)}\")\n"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "340997f8",
      "metadata": {
        "id": "340997f8"
      },
      "source": [
        "---\n",
        "## Part 4: Run RAG Pipeline"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 10,
      "id": "50c764cc",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "50c764cc",
        "outputId": "e8e7429e-2004-4b29-c53e-ebf595fe3c86"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 1 — CloudBase Account Access, Login Flow, and Password Reset', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 1 — CloudBase Account Access, Login Flow, and Password Reset', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking', 'Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking', 'Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options']\n",
            "[RAG] Retriever = HYBRID (dense + TF-IDF)\n",
            "[RAG] Docs: ['Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention', 'Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention', 'Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention']\n",
            "[RAG] Retriever = HYBRID (dense + TF-IDF)\n",
            "[RAG] Docs: ['Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention', 'Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention', 'Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = HYBRID (dense + TF-IDF)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = HYBRID (dense + TF-IDF)\n",
            "[RAG] Docs: ['Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset']\n",
            "Processed 16 questions\n",
            "\n",
            "Q1: How do I reset my CloudBase password?\n",
            "A: To reset your CloudBase password, click the \"Forgot Password\" link on the login page. CloudBase will...\n",
            "\n",
            "Q2: How long is a password reset link valid?\n",
            "A: A password reset link is valid for 24 hours. If the link expires, you must request a new reset email...\n",
            "\n",
            "Q3: What triggers an account lockout and how long does it last?\n",
            "A: An account is locked after five consecutive unsuccessful login attempts [Source 2]. The lockout dura...\n",
            "\n"
          ]
        }
      ],
      "source": [
        "# Run RAG on all questions\n",
        "rag_outputs = []\n",
        "for q in QUESTIONS:\n",
        "    answer, retrieved = rag_query(\n",
        "        q,\n",
        "        all_chunks,\n",
        "        chunk_embeddings,\n",
        "        chunk_metadata,\n",
        "        tfidf_vectorizer,\n",
        "        tfidf_matrix,\n",
        "        top_k=3,\n",
        "        alpha=0.7,\n",
        "        system_prompt=FINAL_SYSTEM_PROMPT\n",
        "    )\n",
        "\n",
        "    rag_outputs.append({\n",
        "        \"question\": q,\n",
        "        \"answer\": answer,\n",
        "        \"retrieved_chunks\": [text for (text, _, _) in retrieved],\n",
        "        \"retrieval_scores\": [score for (_, score, _) in retrieved],\n",
        "        \"retrieved_docs\": [meta.get(\"doc_title\", \"\") for (_, _, meta) in retrieved],\n",
        "    })\n",
        "\n",
        "\n",
        "print(f\"Processed {len(rag_outputs)} questions\\n\")\n",
        "for i, out in enumerate(rag_outputs[:3]):\n",
        "    print(f\"Q{i+1}: {out['question']}\")\n",
        "    print(f\"A: {out['answer'][:100]}...\\n\")"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 11,
      "id": "0f01ae1f",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "0f01ae1f",
        "outputId": "875b6e5f-132b-4b2b-f7ba-d2939e8298be"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Exported to day3_assignment_rag_outputs.json\n"
          ]
        }
      ],
      "source": [
        "# Export RAG outputs\n",
        "with open(\"day3_assignment_rag_outputs.json\", \"w\") as f:\n",
        "    json.dump(rag_outputs, f, indent=2, default=str)\n",
        "print(\"Exported to day3_assignment_rag_outputs.json\")"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "0019c51b",
      "metadata": {
        "id": "0019c51b"
      },
      "source": [
        "---\n",
        "## Part 5: Golden Q&A Set (15+ items)\n",
        "\n",
        "Create your golden set with expected keywords. Include easy, medium, hard, and refusal questions."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 12,
      "id": "1bec81e7",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "1bec81e7",
        "outputId": "3fdd5322-9dd0-4487-de6b-54a0d5f45e9f"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Golden set items: 16\n"
          ]
        }
      ],
      "source": [
        "# ── Golden test set without expected chunks (16 items) ───────────────────────────────────────────\n",
        "# If this is used, the next block should not be used.\n",
        "# expected_chunks is left empty (optional) to avoid hard-coding chunk indices.\n",
        "# Each item includes expected keywords used for a simple evaluation metric.\n",
        "\n",
        "\n",
        "GOLDEN_SET = [\n",
        "    {\n",
        "        \"id\": \"Q01\",\n",
        "        \"question\": \"How do I reset my CloudBase password?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"Forgot Password\",\n",
        "            \"reset link\",\n",
        "            \"24 hours\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q02\",\n",
        "        \"question\": \"How can I enable two-factor authentication (2FA), and what methods are supported?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"Settings\",\n",
        "            \"Security\",\n",
        "            \"authenticator\",\n",
        "            \"SMS\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q03\",\n",
        "        \"question\": \"My account is locked-what triggered it and how long does it last?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"5\",\n",
        "            \"30 minutes\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q04\",\n",
        "        \"question\": \"What are the Free and Team plan limits for users and storage?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"Free\",\n",
        "            \"5 users\",\n",
        "            \"1 GB\",\n",
        "            \"Team\",\n",
        "            \"100 users\",\n",
        "            \"50 GB\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q05\",\n",
        "        \"question\": \"What is the Team plan price per user per month, and does it include VAT?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"EUR 12\",\n",
        "            \"exclude VAT\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q06\",\n",
        "        \"question\": \"Which browsers are supported, and is Internet Explorer supported?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"Chrome\",\n",
        "            \"Firefox\",\n",
        "            \"Edge\",\n",
        "            \"Safari\",\n",
        "            \"Internet Explorer\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q07\",\n",
        "        \"question\": \"How do upgrades and downgrades take effect, and are partial refunds offered?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"immediately\",\n",
        "            \"next billing cycle\",\n",
        "            \"No partial refunds\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"medium\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q08\",\n",
        "        \"question\": \"What is the refund policy timeframe for an initial purchase or upgrade?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"14 days\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"medium\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q09\",\n",
        "        \"question\": \"CSV export fails for a huge dataset. What iss the limit and what should I use instead?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"100,000\",\n",
        "            \"API\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"medium\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q10\",\n",
        "        \"question\": \"Does CloudBase support SSO, and what is required to set it up?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"SAML 2.0\",\n",
        "            \"metadata URL\",\n",
        "            \"attribute mapping\",\n",
        "            \"admin approval\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"medium\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q11\",\n",
        "        \"question\": \"If I lose my phone, can I bypass 2FA?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"recovery codes\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"hard\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q12\",\n",
        "        \"question\": \"If a CSV export fails below the limit, what should I try before contacting support?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"retry\",\n",
        "            \"connectivity\",\n",
        "            \"row count\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"hard\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q13\",\n",
        "        \"question\": \"Is there a CloudBase mobile app and what OS versions are supported?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"iOS 15\",\n",
        "            \"Android 12\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q14\",\n",
        "        \"question\": \"How long is the password reset link valid?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"24 hours\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q15\",\n",
        "        \"question\": \"What is the billing contact email for billing inquiries?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"billing@cloudbase.io\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q16\",\n",
        "        \"question\": \"What is the phone number for CloudBase support?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"I don't have that information in the provided documents\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"refusal\"\n",
        "    }\n",
        "]\n",
        "\n",
        "assert len(GOLDEN_SET) >= 15, f\"Need at least 15 golden questions, have {len(GOLDEN_SET)}\"\n",
        "print(f\"Golden set items: {len(GOLDEN_SET)}\")\n"
      ]
    },
    {
      "cell_type": "code",
      "source": [
        "# Run this if you want expected chunks ── Golden test set (16 items) ───────────────────────────────────────────\n",
        "# Each item includes expected keywords used for a simple evaluation metric.\n",
        "# expected_chunks filled using the chunk indices you provided.\n",
        "\n",
        "GOLDEN_SET = [\n",
        "    {\n",
        "        \"id\": \"Q01\",\n",
        "        \"question\": \"How do I reset my CloudBase password?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"Forgot Password\",\n",
        "            \"reset link\",\n",
        "            \"24 hours\"\n",
        "        ],\n",
        "        \"expected_chunks\": [0, 1, 2],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q02\",\n",
        "        \"question\": \"How can I enable two-factor authentication (2FA), and what methods are supported?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"Settings\",\n",
        "            \"Security\",\n",
        "            \"authenticator\",\n",
        "            \"SMS\"\n",
        "        ],\n",
        "        \"expected_chunks\": [7, 8],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q03\",\n",
        "        \"question\": \"My account is locked-what triggered it and how long does it last?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"5\",\n",
        "            \"30 minutes\"\n",
        "        ],\n",
        "        \"expected_chunks\": [4, 5],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q04\",\n",
        "        \"question\": \"What are the Free and Team plan limits for users and storage?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"Free\",\n",
        "            \"5 users\",\n",
        "            \"1 GB\",\n",
        "            \"Team\",\n",
        "            \"100 users\",\n",
        "            \"50 GB\"\n",
        "        ],\n",
        "        \"expected_chunks\": [10, 11],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q05\",\n",
        "        \"question\": \"What is the Team plan price per user per month, and does it include VAT?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"EUR 12\",\n",
        "            \"exclude VAT\"\n",
        "        ],\n",
        "        \"expected_chunks\": [14],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q06\",\n",
        "        \"question\": \"Which browsers are supported, and is Internet Explorer supported?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"Chrome\",\n",
        "            \"Firefox\",\n",
        "            \"Edge\",\n",
        "            \"Safari\",\n",
        "            \"Internet Explorer\"\n",
        "        ],\n",
        "        \"expected_chunks\": [18],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q07\",\n",
        "        \"question\": \"How do upgrades and downgrades take effect, and are partial refunds offered?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"immediately\",\n",
        "            \"next billing cycle\",\n",
        "            \"No partial refunds\"\n",
        "        ],\n",
        "        \"expected_chunks\": [14, 15],\n",
        "        \"difficulty\": \"medium\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q08\",\n",
        "        \"question\": \"What is the refund policy timeframe for an initial purchase or upgrade?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"14 days\"\n",
        "        ],\n",
        "        \"expected_chunks\": [16],\n",
        "        \"difficulty\": \"medium\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q09\",\n",
        "        \"question\": \"CSV export fails for a huge dataset. What iss the limit and what should I use instead?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"100,000\",\n",
        "            \"API\"\n",
        "        ],\n",
        "        \"expected_chunks\": [17],\n",
        "        \"difficulty\": \"medium\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q10\",\n",
        "        \"question\": \"Does CloudBase support SSO, and what is required to set it up?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"SAML 2.0\",\n",
        "            \"metadata URL\",\n",
        "            \"attribute mapping\",\n",
        "            \"admin approval\"\n",
        "        ],\n",
        "        \"expected_chunks\": [18, 19],\n",
        "        \"difficulty\": \"medium\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q11\",\n",
        "        \"question\": \"If I lose my phone, can I bypass 2FA?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"recovery codes\"\n",
        "        ],\n",
        "        \"expected_chunks\": [8, 9],\n",
        "        \"difficulty\": \"hard\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q12\",\n",
        "        \"question\": \"If a CSV export fails below the limit, what should I try before contacting support?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"retry\",\n",
        "            \"connectivity\",\n",
        "            \"row count\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"hard\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q13\",\n",
        "        \"question\": \"Is there a CloudBase mobile app and what OS versions are supported?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"iOS 15\",\n",
        "            \"Android 12\"\n",
        "        ],\n",
        "        \"expected_chunks\": [18],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q14\",\n",
        "        \"question\": \"How long is the password reset link valid?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"24 hours\"\n",
        "        ],\n",
        "        \"expected_chunks\": [1, 2],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q15\",\n",
        "        \"question\": \"What is the billing contact email for billing inquiries?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"billing@cloudbase.io\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"easy\"\n",
        "    },\n",
        "    {\n",
        "        \"id\": \"Q16\",\n",
        "        \"question\": \"What is the phone number for CloudBase support?\",\n",
        "        \"expected_keywords\": [\n",
        "            \"I don't have that information in the provided documents\"\n",
        "        ],\n",
        "        \"expected_chunks\": [],\n",
        "        \"difficulty\": \"refusal\"\n",
        "    }\n",
        "]\n",
        "\n",
        "assert len(GOLDEN_SET) >= 15, f\"Need at least 15 golden questions, have {len(GOLDEN_SET)}\"\n",
        "print(f\"Golden set items: {len(GOLDEN_SET)}\")\n"
      ],
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "6hWWTRoNLxJu",
        "outputId": "efbb2df4-1d75-4b80-8ec4-dc449376ca5e"
      },
      "id": "6hWWTRoNLxJu",
      "execution_count": 13,
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Golden set items: 16\n"
          ]
        }
      ]
    },
    {
      "cell_type": "markdown",
      "id": "dedc787c",
      "metadata": {
        "id": "dedc787c"
      },
      "source": [
        "---\n",
        "## Part 6: RAG Triad Metrics (20 points)\n",
        "\n",
        "Evaluate EVERY golden set question on all three RAG triad dimensions."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 14,
      "id": "e984caf7",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/",
          "height": 955
        },
        "id": "e984caf7",
        "outputId": "8e8167d1-053e-45d1-f543-ae091c707550"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 1 — CloudBase Account Access, Login Flow, and Password Reset', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking', 'Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking', 'Document 2 — Account Lockout Policy, Failed Login Protection, and Unlocking']\n",
            "[RAG] Retriever = HYBRID (dense + TF-IDF)\n",
            "[RAG] Docs: ['Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention', 'Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention', 'Document 4 — Subscription Plans, User Limits, Storage Limits, and Data Retention']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = HYBRID (dense + TF-IDF)\n",
            "[RAG] Docs: ['Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO', 'Document 6 — Exports, API Usage, Browser Support, Mobile Access, and SSO']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 1 — CloudBase Account Access, Login Flow, and Password Reset', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset']\n",
            "[RAG] Retriever = FILTERED (search_with_filter)\n",
            "[RAG] Docs: ['Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes', 'Document 5 — Pricing, VAT Treatment, Refund Policy, and Subscription Changes']\n",
            "[RAG] Retriever = HYBRID (dense + TF-IDF)\n",
            "[RAG] Docs: ['Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 3 — Two-Factor Authentication (2FA), Setup Process, and Recovery Options', 'Document 1 — CloudBase Account Access, Login Flow, and Password Reset']\n"
          ]
        },
        {
          "output_type": "execute_result",
          "data": {
            "text/plain": [
              "    id                                           question difficulty  \\\n",
              "0  Q01              How do I reset my CloudBase password?       easy   \n",
              "1  Q02  How can I enable two-factor authentication (2F...       easy   \n",
              "2  Q03  My account is locked-what triggered it and how...       easy   \n",
              "3  Q04  What are the Free and Team plan limits for use...       easy   \n",
              "4  Q05  What is the Team plan price per user per month...       easy   \n",
              "\n",
              "                                              answer  keyword_match  \\\n",
              "0  To reset your CloudBase password, click the “F...           True   \n",
              "1  You can enable two-factor authentication (2FA)...           True   \n",
              "2  Your account is locked after five consecutive ...           True   \n",
              "3  The Free plan supports up to five users and in...           True   \n",
              "4  The Team plan is EUR 12 per user per month [So...           True   \n",
              "\n",
              "   context_relevance  groundedness  answer_relevance  precision_at_3  \\\n",
              "0                  5             5                 5             0.0   \n",
              "1                  5             5                 5             0.0   \n",
              "2                  5             5                 5             0.0   \n",
              "3                  5             5                 5             0.0   \n",
              "4                  5             5                 5             0.0   \n",
              "\n",
              "   recall_at_3                                        explanation  \n",
              "0          0.0  The retrieved context fully explains how to re...  \n",
              "1          0.0  The retrieved context fully explains how to en...  \n",
              "2          0.0  The retrieved context clearly states that an a...  \n",
              "3          0.0  The retrieved context directly provides the us...  \n",
              "4          0.0  The retrieved context directly answers both pa...  "
            ],
            "text/html": [
              "\n",
              "  <div id=\"df-81a037ab-bd7f-4cfd-bde0-196973359f95\" class=\"colab-df-container\">\n",
              "    <div>\n",
              "<style scoped>\n",
              "    .dataframe tbody tr th:only-of-type {\n",
              "        vertical-align: middle;\n",
              "    }\n",
              "\n",
              "    .dataframe tbody tr th {\n",
              "        vertical-align: top;\n",
              "    }\n",
              "\n",
              "    .dataframe thead th {\n",
              "        text-align: right;\n",
              "    }\n",
              "</style>\n",
              "<table border=\"1\" class=\"dataframe\">\n",
              "  <thead>\n",
              "    <tr style=\"text-align: right;\">\n",
              "      <th></th>\n",
              "      <th>id</th>\n",
              "      <th>question</th>\n",
              "      <th>difficulty</th>\n",
              "      <th>answer</th>\n",
              "      <th>keyword_match</th>\n",
              "      <th>context_relevance</th>\n",
              "      <th>groundedness</th>\n",
              "      <th>answer_relevance</th>\n",
              "      <th>precision_at_3</th>\n",
              "      <th>recall_at_3</th>\n",
              "      <th>explanation</th>\n",
              "    </tr>\n",
              "  </thead>\n",
              "  <tbody>\n",
              "    <tr>\n",
              "      <th>0</th>\n",
              "      <td>Q01</td>\n",
              "      <td>How do I reset my CloudBase password?</td>\n",
              "      <td>easy</td>\n",
              "      <td>To reset your CloudBase password, click the “F...</td>\n",
              "      <td>True</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>0.0</td>\n",
              "      <td>0.0</td>\n",
              "      <td>The retrieved context fully explains how to re...</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>1</th>\n",
              "      <td>Q02</td>\n",
              "      <td>How can I enable two-factor authentication (2F...</td>\n",
              "      <td>easy</td>\n",
              "      <td>You can enable two-factor authentication (2FA)...</td>\n",
              "      <td>True</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>0.0</td>\n",
              "      <td>0.0</td>\n",
              "      <td>The retrieved context fully explains how to en...</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>2</th>\n",
              "      <td>Q03</td>\n",
              "      <td>My account is locked-what triggered it and how...</td>\n",
              "      <td>easy</td>\n",
              "      <td>Your account is locked after five consecutive ...</td>\n",
              "      <td>True</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>0.0</td>\n",
              "      <td>0.0</td>\n",
              "      <td>The retrieved context clearly states that an a...</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>3</th>\n",
              "      <td>Q04</td>\n",
              "      <td>What are the Free and Team plan limits for use...</td>\n",
              "      <td>easy</td>\n",
              "      <td>The Free plan supports up to five users and in...</td>\n",
              "      <td>True</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>0.0</td>\n",
              "      <td>0.0</td>\n",
              "      <td>The retrieved context directly provides the us...</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>4</th>\n",
              "      <td>Q05</td>\n",
              "      <td>What is the Team plan price per user per month...</td>\n",
              "      <td>easy</td>\n",
              "      <td>The Team plan is EUR 12 per user per month [So...</td>\n",
              "      <td>True</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>5</td>\n",
              "      <td>0.0</td>\n",
              "      <td>0.0</td>\n",
              "      <td>The retrieved context directly answers both pa...</td>\n",
              "    </tr>\n",
              "  </tbody>\n",
              "</table>\n",
              "</div>\n",
              "    <div class=\"colab-df-buttons\">\n",
              "\n",
              "  <div class=\"colab-df-container\">\n",
              "    <button class=\"colab-df-convert\" onclick=\"convertToInteractive('df-81a037ab-bd7f-4cfd-bde0-196973359f95')\"\n",
              "            title=\"Convert this dataframe to an interactive table.\"\n",
              "            style=\"display:none;\">\n",
              "\n",
              "  <svg xmlns=\"http://www.w3.org/2000/svg\" height=\"24px\" viewBox=\"0 -960 960 960\">\n",
              "    <path d=\"M120-120v-720h720v720H120Zm60-500h600v-160H180v160Zm220 220h160v-160H400v160Zm0 220h160v-160H400v160ZM180-400h160v-160H180v160Zm440 0h160v-160H620v160ZM180-180h160v-160H180v160Zm440 0h160v-160H620v160Z\"/>\n",
              "  </svg>\n",
              "    </button>\n",
              "\n",
              "  <style>\n",
              "    .colab-df-container {\n",
              "      display:flex;\n",
              "      gap: 12px;\n",
              "    }\n",
              "\n",
              "    .colab-df-convert {\n",
              "      background-color: #E8F0FE;\n",
              "      border: none;\n",
              "      border-radius: 50%;\n",
              "      cursor: pointer;\n",
              "      display: none;\n",
              "      fill: #1967D2;\n",
              "      height: 32px;\n",
              "      padding: 0 0 0 0;\n",
              "      width: 32px;\n",
              "    }\n",
              "\n",
              "    .colab-df-convert:hover {\n",
              "      background-color: #E2EBFA;\n",
              "      box-shadow: 0px 1px 2px rgba(60, 64, 67, 0.3), 0px 1px 3px 1px rgba(60, 64, 67, 0.15);\n",
              "      fill: #174EA6;\n",
              "    }\n",
              "\n",
              "    .colab-df-buttons div {\n",
              "      margin-bottom: 4px;\n",
              "    }\n",
              "\n",
              "    [theme=dark] .colab-df-convert {\n",
              "      background-color: #3B4455;\n",
              "      fill: #D2E3FC;\n",
              "    }\n",
              "\n",
              "    [theme=dark] .colab-df-convert:hover {\n",
              "      background-color: #434B5C;\n",
              "      box-shadow: 0px 1px 3px 1px rgba(0, 0, 0, 0.15);\n",
              "      filter: drop-shadow(0px 1px 2px rgba(0, 0, 0, 0.3));\n",
              "      fill: #FFFFFF;\n",
              "    }\n",
              "  </style>\n",
              "\n",
              "    <script>\n",
              "      const buttonEl =\n",
              "        document.querySelector('#df-81a037ab-bd7f-4cfd-bde0-196973359f95 button.colab-df-convert');\n",
              "      buttonEl.style.display =\n",
              "        google.colab.kernel.accessAllowed ? 'block' : 'none';\n",
              "\n",
              "      async function convertToInteractive(key) {\n",
              "        const element = document.querySelector('#df-81a037ab-bd7f-4cfd-bde0-196973359f95');\n",
              "        const dataTable =\n",
              "          await google.colab.kernel.invokeFunction('convertToInteractive',\n",
              "                                                    [key], {});\n",
              "        if (!dataTable) return;\n",
              "\n",
              "        const docLinkHtml = 'Like what you see? Visit the ' +\n",
              "          '<a target=\"_blank\" href=https://colab.research.google.com/notebooks/data_table.ipynb>data table notebook</a>'\n",
              "          + ' to learn more about interactive tables.';\n",
              "        element.innerHTML = '';\n",
              "        dataTable['output_type'] = 'display_data';\n",
              "        await google.colab.output.renderOutput(dataTable, element);\n",
              "        const docLink = document.createElement('div');\n",
              "        docLink.innerHTML = docLinkHtml;\n",
              "        element.appendChild(docLink);\n",
              "      }\n",
              "    </script>\n",
              "  </div>\n",
              "\n",
              "\n",
              "    </div>\n",
              "  </div>\n"
            ],
            "application/vnd.google.colaboratory.intrinsic+json": {
              "type": "dataframe",
              "variable_name": "triad_df",
              "summary": "{\n  \"name\": \"triad_df\",\n  \"rows\": 16,\n  \"fields\": [\n    {\n      \"column\": \"id\",\n      \"properties\": {\n        \"dtype\": \"string\",\n        \"num_unique_values\": 16,\n        \"samples\": [\n          \"Q01\",\n          \"Q02\",\n          \"Q06\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"question\",\n      \"properties\": {\n        \"dtype\": \"string\",\n        \"num_unique_values\": 16,\n        \"samples\": [\n          \"How do I reset my CloudBase password?\",\n          \"How can I enable two-factor authentication (2FA), and what methods are supported?\",\n          \"Which browsers are supported, and is Internet Explorer supported?\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"difficulty\",\n      \"properties\": {\n        \"dtype\": \"category\",\n        \"num_unique_values\": 4,\n        \"samples\": [\n          \"medium\",\n          \"refusal\",\n          \"easy\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"answer\",\n      \"properties\": {\n        \"dtype\": \"string\",\n        \"num_unique_values\": 14,\n        \"samples\": [\n          \"Yes, CloudBase supports Single Sign-On (SSO) using SAML 2.0 for enterprise customers [Source 1]. To set it up, you will need an Identity Provider metadata URL, attribute mapping, and administrator approval [Source 1]. The configuration process usually takes 1\\u20132 business days [Source 1]. SSO is available on Enterprise plans [Source 1].\",\n          \"I don't have that information in the provided documents.\",\n          \"To reset your CloudBase password, click the \\u201cForgot Password\\u201d link on the login page. CloudBase will send a password reset email to your registered email address with a secure reset link valid for 24 hours. If the link expires, request a new one. [Source 0] If you no longer have access to your registered email, contact CloudBase support. [Source 2]\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"keyword_match\",\n      \"properties\": {\n        \"dtype\": \"boolean\",\n        \"num_unique_values\": 2,\n        \"samples\": [\n          false,\n          true\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"context_relevance\",\n      \"properties\": {\n        \"dtype\": \"number\",\n        \"std\": 1,\n        \"min\": 1,\n        \"max\": 5,\n        \"num_unique_values\": 3,\n        \"samples\": [\n          5,\n          1\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"groundedness\",\n      \"properties\": {\n        \"dtype\": \"number\",\n        \"std\": 1,\n        \"min\": 1,\n        \"max\": 5,\n        \"num_unique_values\": 2,\n        \"samples\": [\n          1,\n          5\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"answer_relevance\",\n      \"properties\": {\n        \"dtype\": \"number\",\n        \"std\": 1,\n        \"min\": 1,\n        \"max\": 5,\n        \"num_unique_values\": 2,\n        \"samples\": [\n          1,\n          5\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"precision_at_3\",\n      \"properties\": {\n        \"dtype\": \"number\",\n        \"std\": 0.0,\n        \"min\": 0.0,\n        \"max\": 0.0,\n        \"num_unique_values\": 1,\n        \"samples\": [\n          0.0\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"recall_at_3\",\n      \"properties\": {\n        \"dtype\": \"number\",\n        \"std\": 0.0,\n        \"min\": 0.0,\n        \"max\": 0.0,\n        \"num_unique_values\": 1,\n        \"samples\": [\n          0.0\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"explanation\",\n      \"properties\": {\n        \"dtype\": \"string\",\n        \"num_unique_values\": 16,\n        \"samples\": [\n          \"The retrieved context fully explains how to reset a CloudBase password, including the self-service option via the 'Forgot Password' link, the email reset process, the link's validity, and what to do if the email is inaccessible. The answer accurately reflects this information, directly addressing the question and citing the relevant sources. All claims in the answer are supported by the context.\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    }\n  ]\n}"
            }
          },
          "metadata": {},
          "execution_count": 14
        }
      ],
      "source": [
        "# Run RAG + triad evaluation + retrieval metrics on golden set\n",
        "triad_results = []\n",
        "\n",
        "for qa in GOLDEN_SET:\n",
        "    # Run RAG (this will use hybrid_search or search_with_filter internally)\n",
        "    answer, retrieved = rag_query(\n",
        "        qa[\"question\"],\n",
        "        all_chunks,\n",
        "        chunk_embeddings,\n",
        "        chunk_metadata,\n",
        "        tfidf_vectorizer,\n",
        "        tfidf_matrix,\n",
        "        top_k=3,\n",
        "        alpha=0.7,\n",
        "        system_prompt=FINAL_SYSTEM_PROMPT\n",
        "    )\n",
        "\n",
        "    # retrieved is: (chunk_text, score, metadata)\n",
        "    retrieved_texts = [text for (text, _, _) in retrieved]\n",
        "    retrieved_indices = [idx for idx, _, _ in retrieved]\n",
        "    # Keyword check\n",
        "    answer_lower = answer.lower()\n",
        "    keyword_hit = any(kw.lower() in answer_lower for kw in qa[\"expected_keywords\"])\n",
        "\n",
        "    # RAG Triad evaluation\n",
        "    triad = evaluate_rag_triad(\n",
        "        qa[\"question\"],\n",
        "        answer,\n",
        "        retrieved_texts,\n",
        "        label=f\"triad_{qa['id']}\"\n",
        "    )\n",
        "\n",
        "    # expected_chunks is empty in your notebook → metrics will remain None\n",
        "    #p_at_k, r_at_k = None, None\n",
        "    #if qa.get(\"expected_chunks\"):\n",
        "        # Only compute if you later fill expected_chunks with real indices\n",
        "        # (Right now it's empty by design.)\n",
        "        #pass\n",
        "        # Precision@k / Recall@k (if expected_chunks provided)\n",
        "    p_at_k, r_at_k = None, None\n",
        "    if qa.get(\"expected_chunks\"):\n",
        "        p_at_k, r_at_k = precision_recall_at_k(retrieved_indices, qa[\"expected_chunks\"], k=3)\n",
        "\n",
        "    triad_results.append({\n",
        "        \"id\": qa[\"id\"],\n",
        "        \"question\": qa[\"question\"],\n",
        "        \"difficulty\": qa[\"difficulty\"],\n",
        "        \"answer\": answer,\n",
        "        \"keyword_match\": keyword_hit,\n",
        "        \"context_relevance\": triad.context_relevance,\n",
        "        \"groundedness\": triad.groundedness,\n",
        "        \"answer_relevance\": triad.answer_relevance,\n",
        "        \"precision_at_3\": p_at_k,\n",
        "        \"recall_at_3\": r_at_k,\n",
        "        \"explanation\": triad.explanation,\n",
        "    })\n",
        "\n",
        "triad_df = pd.DataFrame(triad_results)\n",
        "triad_df.head()\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 15,
      "id": "9d801247",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "9d801247",
        "outputId": "02729005-971b-4d78-b473-a0595294f9c2"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "======================================================================\n",
            "RAG TRIAD + RETRIEVAL EVALUATION RESULTS\n",
            "======================================================================\n",
            "\n",
            "Overall Metrics:\n",
            "  Keyword Accuracy:    88% (14/16)\n",
            "  Context Relevance:   4.4/5 ✓\n",
            "  Groundedness:        4.8/5 ✓\n",
            "  Answer Relevance:    4.2/5 ✓\n",
            "  Avg Precision@3:     0.00\n",
            "  Avg Recall@3:        0.00\n",
            "\n",
            "Per-Question Scores:\n",
            " id difficulty  keyword_match  context_relevance  groundedness  answer_relevance  precision_at_3  recall_at_3\n",
            "Q01       easy           True                  5             5                 5             0.0          0.0\n",
            "Q02       easy           True                  5             5                 5             0.0          0.0\n",
            "Q03       easy           True                  5             5                 5             0.0          0.0\n",
            "Q04       easy           True                  5             5                 5             0.0          0.0\n",
            "Q05       easy           True                  5             5                 5             0.0          0.0\n",
            "Q06       easy           True                  5             5                 5             0.0          0.0\n",
            "Q07     medium           True                  5             5                 5             0.0          0.0\n",
            "Q08     medium           True                  5             5                 5             0.0          0.0\n",
            "Q09     medium           True                  5             5                 5             0.0          0.0\n",
            "Q10     medium           True                  5             5                 5             0.0          0.0\n",
            "Q11       hard           True                  5             5                 5             0.0          0.0\n",
            "Q12       hard          False                  1             1                 1             NaN          NaN\n",
            "Q13       easy           True                  5             5                 5             0.0          0.0\n",
            "Q14       easy           True                  5             5                 5             0.0          0.0\n",
            "Q15       easy          False                  1             5                 1             NaN          NaN\n",
            "Q16    refusal           True                  3             5                 1             NaN          NaN\n",
            "\n",
            "Questions below target (< 4.0 on any dimension):\n",
            "  Q12: Context=1, Ground=1, Relevance=1 — The context does not contain information about what to do if a CSV export fails \n",
            "  Q15: Context=1, Relevance=1 — The context provided does not contain any information about a billing contact em\n",
            "  Q16: Context=3, Relevance=1 — The context mentions contacting CloudBase support for assistance in multiple sce\n"
          ]
        }
      ],
      "source": [
        "# Display RAG Triad + Retrieval Metrics\n",
        "print(\"=\" * 70)\n",
        "print(\"RAG TRIAD + RETRIEVAL EVALUATION RESULTS\")\n",
        "print(\"=\" * 70)\n",
        "\n",
        "keyword_acc = triad_df[\"keyword_match\"].mean()\n",
        "avg_context = triad_df[\"context_relevance\"].mean()\n",
        "avg_ground = triad_df[\"groundedness\"].mean()\n",
        "avg_relevance = triad_df[\"answer_relevance\"].mean()\n",
        "\n",
        "# Precision/Recall (only for non-refusal questions)\n",
        "pr_df = triad_df.dropna(subset=[\"precision_at_3\"])\n",
        "avg_precision = pr_df[\"precision_at_3\"].mean() if len(pr_df) > 0 else 0\n",
        "avg_recall = pr_df[\"recall_at_3\"].mean() if len(pr_df) > 0 else 0\n",
        "\n",
        "print(f\"\\nOverall Metrics:\")\n",
        "print(f\"  Keyword Accuracy:    {keyword_acc:.0%} ({triad_df['keyword_match'].sum()}/{len(triad_df)})\")\n",
        "print(f\"  Context Relevance:   {avg_context:.1f}/5 {'✓' if avg_context >= 4.0 else '✗ (target: ≥ 4.0)'}\")\n",
        "print(f\"  Groundedness:        {avg_ground:.1f}/5 {'✓' if avg_ground >= 4.0 else '✗ (target: ≥ 4.0)'}\")\n",
        "print(f\"  Answer Relevance:    {avg_relevance:.1f}/5 {'✓' if avg_relevance >= 4.0 else '✗ (target: ≥ 4.0)'}\")\n",
        "if len(pr_df) > 0:\n",
        "    print(f\"  Avg Precision@3:     {avg_precision:.2f}\")\n",
        "    print(f\"  Avg Recall@3:        {avg_recall:.2f}\")\n",
        "\n",
        "print(f\"\\nPer-Question Scores:\")\n",
        "display_cols = [\"id\", \"difficulty\", \"keyword_match\", \"context_relevance\",\n",
        "                \"groundedness\", \"answer_relevance\"]\n",
        "if len(pr_df) > 0:\n",
        "    display_cols.extend([\"precision_at_3\", \"recall_at_3\"])\n",
        "print(triad_df[display_cols].to_string(index=False))\n",
        "\n",
        "# Flag questions below target\n",
        "below_target = triad_df[\n",
        "    (triad_df[\"context_relevance\"] < 4) |\n",
        "    (triad_df[\"groundedness\"] < 4) |\n",
        "    (triad_df[\"answer_relevance\"] < 4)\n",
        "]\n",
        "if len(below_target) > 0:\n",
        "    print(f\"\\nQuestions below target (< 4.0 on any dimension):\")\n",
        "    for _, row in below_target.iterrows():\n",
        "        dims = []\n",
        "        if row[\"context_relevance\"] < 4: dims.append(f\"Context={row['context_relevance']}\")\n",
        "        if row[\"groundedness\"] < 4: dims.append(f\"Ground={row['groundedness']}\")\n",
        "        if row[\"answer_relevance\"] < 4: dims.append(f\"Relevance={row['answer_relevance']}\")\n",
        "        print(f\"  {row['id']}: {', '.join(dims)} — {row['explanation'][:80]}\")"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "e5b13cd9",
      "metadata": {
        "id": "e5b13cd9"
      },
      "source": [
        "## Part 7: Error Analysis (25 points)\n",
        "\n",
        "Write your error analysis below. Cover all three sections.\n",
        "\n",
        "### A. Retrieval Failures\n",
        "\n",
        "**Pattern 1: Missing knowledge coverage leads to “relevant-but-insufficient” retrieval**\n",
        "\n",
        "* **What went wrong:** In some cases, retrieval returns highly relevant chunks (e.g., billing or account management documentation), but **the specific fact requested by the question is not present anywhere in the indexed knowledge base**, so neither filtered dense retrieval nor hybrid retrieval can surface it.\n",
        "* **Example:** **Q15** (“What is the billing contact email for billing inquiries?”) — retrieved chunks discuss billing policies and subscription handling, but no billing email address is included in any document; therefore the system cannot retrieve `billing@cloudbase.io`.\n",
        "* **Root cause:** A **coverage gap** between the golden set and the knowledge base: the golden set expects a billing contact email that is not present in the KB, making retrieval failure unavoidable regardless of retrieval strategy.\n",
        "\n",
        "**Pattern 2: Retrieval surfaces policy constraints but not operational troubleshooting guidance**\n",
        "\n",
        "* **What went wrong:** Retrieval correctly identifies and ranks chunks describing system constraints and alternatives (e.g., CSV export limits and API usage), but the question asks for **procedural troubleshooting steps below the limit**, which are not described in the documentation.\n",
        "* **Example:** **Q12** (“If a CSV export fails below the limit, what should I try before contacting support?”) — hybrid retrieval retrieves the CSV limit and API guidance, but cannot provide troubleshooting steps because they are not documented.\n",
        "* **Root cause:** The knowledge base focuses on **policy-level constraints and recommended alternatives**, but lacks **operational diagnostic content** (retry steps, connectivity checks, session validation), limiting retrieval completeness.\n",
        "\n",
        "---\n",
        "\n",
        "### B. Generation Failures\n",
        "\n",
        "**Pattern 1: Correct refusal conflicts with keyword-based scoring when golden expects missing information**\n",
        "\n",
        "* **What went wrong:** The model correctly refuses to hallucinate answers when required information is absent from retrieved context, but keyword-based evaluation fails because expected keywords cannot appear in a refusal answer.\n",
        "* **Example:** **Q15** — the model correctly states that the billing email is not available in the documents, but `keyword_match` is false because the golden set expects `billing@cloudbase.io`, which does not exist in the KB.\n",
        "* **Root cause:** The generation strictly follows the “use only provided context” constraint, while the golden set assumes the presence of information not covered by the knowledge base, creating an unavoidable evaluation mismatch.\n",
        "\n",
        "**Pattern 2: Triad evaluation instability for justified refusal cases**\n",
        "\n",
        "* **What went wrong:** In some refusal scenarios, the triad evaluator assigns low groundedness and answer relevance scores despite its own explanation indicating that refusal is appropriate due to missing information.\n",
        "* **Example:** **Q12** — the triad explanation notes that troubleshooting steps are absent from the documents, yet groundedness and answer relevance scores are very low, even though the refusal is correct and constraint-compliant.\n",
        "* **Root cause:** The triad evaluation rubric does not consistently treat “correct refusal due to missing information” as a valid, grounded response, leading to scoring inconsistencies.\n",
        "\n",
        "---\n",
        "\n",
        "### C. Improvements & Next Steps\n",
        "\n",
        "* **Retrieval:** Close knowledge base coverage gaps that the golden set expects:\n",
        "\n",
        "  * Add a concise “Billing & Support Contacts” section explicitly listing the billing contact email and supported contact channels. This would directly resolve Q15.\n",
        "  * Add a short “CSV export troubleshooting below 100,000 rows” checklist (e.g., verify row count, retry export, check session/connectivity, then contact support). This would directly resolve Q12.\n",
        "\n",
        "* **Prompting:** Preserve strict grounding while making refusal behavior more evaluation-robust:\n",
        "\n",
        "  * When refusing due to missing information, instruct the model to clearly state unavailability and, where supported by context, mention safe next steps (e.g., “contact support”) without inventing new details.\n",
        "\n",
        "* **Evaluation:** Align triad scoring with constraint-aware correctness:\n",
        "\n",
        "  * Update the triad evaluator prompt to explicitly score correct refusals caused by missing knowledge as **high groundedness** and **high answer relevance under constraints**, reducing penalties for answers that correctly avoid hallucination.\n"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "47930c55",
      "metadata": {
        "id": "47930c55"
      },
      "source": [
        "Here is the **updated redraft**, fully aligned with the **current routed hybrid retrieval code and observed results**, while **preserving your exact structure, headings, and formatting**. I have only adjusted the content where the old “filtered dense only” assumption leaked in.\n",
        "\n",
        "---\n",
        "\n",
        "## Part 8: RAG Playbook (15 points)\n",
        "\n",
        "Complete ALL sections of the playbook template below.\n",
        "\n",
        "---\n",
        "\n",
        "# RAG PLAYBOOK: CloudBase Support RAG\n",
        "\n",
        "**Version:** 1.0\n",
        "**Author:** Ravi\n",
        "**Date:** 2026-02-09\n",
        "**Status:** Production Ready (with known coverage gaps)\n",
        "\n",
        "## 1. Purpose\n",
        "\n",
        "CloudBase Support RAG is designed to answer common customer-support questions related to account access and security, subscription plans and billing rules, data exports and API usage, SSO configuration, and supported browsers. The system prioritizes **grounded, citation-backed answers** and explicitly refuses to answer when required information is not present in the knowledge base. The retrieval layer uses a **routed hybrid strategy**, combining filtered dense retrieval for document-specific queries and hybrid dense + lexical retrieval for general policy questions. The current deployment is suitable for internal support agents or as a first-line self-service assistant for end users.\n",
        "\n",
        "## 2. Knowledge Base\n",
        "\n",
        "| Property         | Value                                                                                                                       |\n",
        "| ---------------- | --------------------------------------------------------------------------------------------------------------------------- |\n",
        "| Documents        | 6 internal documents (Account Access, Lockouts, 2FA & Recovery, Plans & Retention, Billing & Refunds, Exports/Browsers/SSO) |\n",
        "| Total words      | ~1,300+                                                                                                                     |\n",
        "| Domain           | CloudBase product support policies, limits, and procedures                                                                  |\n",
        "| Update frequency | Quarterly, or immediately after policy/pricing changes                                                                      |\n",
        "\n",
        "**Observed from results:**\n",
        "The KB provides strong coverage of policy rules and limits (lockouts, refunds, exports, pricing, SSO). However, it **does not include billing contact details or procedural troubleshooting steps for CSV exports below the limit**, which directly caused correct refusals and evaluation misses.\n",
        "\n",
        "## 3. Chunking Strategy\n",
        "\n",
        "| Parameter    | Value                   | Rationale                                                                                     |\n",
        "| ------------ | ----------------------- | --------------------------------------------------------------------------------------------- |\n",
        "| Strategy     | Sentence-based chunking | Preserves readable policy statements and avoids mid-sentence splits                           |\n",
        "| Chunk size   | 90 words                | Large enough to keep policies and numeric limits together, small enough to avoid topic mixing |\n",
        "| Overlap      | 25 words                | Ensures limits and conditions near boundaries are retained                                    |\n",
        "| Total chunks | ~20–22                  | Increased due to expanded KB                                                                  |\n",
        "\n",
        "**Observed from results:**\n",
        "Chunking performed reliably across both filtered and hybrid retrieval modes. Errors were not caused by chunk boundaries, but by **knowledge base coverage gaps**, confirming the chunking configuration is appropriate.\n",
        "\n",
        "## 4. System Prompt\n",
        "\n",
        "You are CloudBase Support RAG, a customer-support assistant for the CloudBase product.\n",
        "\n",
        "TASK\n",
        "Answer the user's question using ONLY the information in the provided Context sources.\n",
        "\n",
        "GROUNDING RULES (NON-NEGOTIABLE)\n",
        "\n",
        "* Use ONLY facts that appear in the Context (the text labeled [Chunk 1], [Chunk 2], etc.).\n",
        "* If the answer is not explicitly in the Context, reply exactly:\n",
        "  **\"I don't have that information in the provided documents.\"**\n",
        "* Do not guess, do not use outside knowledge, and do not invent product features, UI labels, prices, policies, timelines, or limits.\n",
        "* If sources conflict, prefer the most specific statement; if still ambiguous, say you don't have that information.\n",
        "\n",
        "SCOPE\n",
        "\n",
        "* In scope: Account access, password resets, lockouts, 2FA, plans/pricing/billing rules, refunds, exports/API limits, SSO, supported browsers/mobile access.\n",
        "* Out of scope: Missing contact details, unpublished features, legal advice, internal roadmaps.\n",
        "\n",
        "SAFETY / MISUSE\n",
        "\n",
        "* If the user requests hacking, bypassing security controls, or wrongdoing, refuse briefly.\n",
        "* If the Context includes a safe recovery path (e.g., recovery codes, contact support), mention it without inventing details.\n",
        "\n",
        "FORMAT\n",
        "\n",
        "* Keep the answer under 90 words.\n",
        "* Use exact numbers/time windows/limits as written (e.g., “100,000 rows”, “14 days”, “30 minutes”).\n",
        "* Cite each key claim with the relevant chunk tag (e.g., [Chunk 2]). Multiple tags are allowed.\n",
        "\n",
        "**Observed from results:**\n",
        "The prompt consistently prevented hallucination across both retrieval modes. Correct refusals occurred when information was missing (e.g., Q12, Q15), demonstrating strong grounding and constraint adherence.\n",
        "\n",
        "## 5. Model Settings\n",
        "\n",
        "| Setting          | Value                 | Rationale                                           |\n",
        "| ---------------- | --------------------- | --------------------------------------------------- |\n",
        "| Generation model | gemini-2.5-flash-lite | Fast and cost-effective for short, factual answers  |\n",
        "| Embedding model  | gemini-embedding-001  | Optimized for semantic document and query retrieval |\n",
        "| Temperature      | 0.3                   | Low creativity to reduce hallucinations             |\n",
        "| top_k            | 3                     | Balances sufficient context with low noise          |\n",
        "\n",
        "## 6. RAG Triad Performance\n",
        "\n",
        "| Metric            | Observed Result                                      | Target |\n",
        "| ----------------- | ---------------------------------------------------- | ------ |\n",
        "| Context Relevance | High for most questions; lower on missing-info cases | ≥ 4.0  |\n",
        "| Groundedness      | High overall; refusal cases sometimes under-scored   | ≥ 4.0  |\n",
        "| Answer Relevance  | High for covered policies; low when KB lacks content | ≥ 4.0  |\n",
        "| Keyword Accuracy  | 14/16 (87.5%)                                        | ≥ 80%  |\n",
        "\n",
        "**Interpretation:**\n",
        "The system meets accuracy targets. Misses are driven by **knowledge base coverage gaps**, not by retrieval strategy or generation behavior.\n",
        "\n",
        "## 7. Known Limitations\n",
        "\n",
        "1. **Knowledge base coverage gaps**\n",
        "\n",
        "   * The system can only answer questions whose facts are explicitly present in the documents. Missing details such as billing contact email addresses or step-by-step CSV export troubleshooting below the limit result in correct refusals but lower keyword and triad scores.\n",
        "\n",
        "2. **Refusal handling vs. evaluation metrics**\n",
        "\n",
        "   * Correct refusals (“I don't have that information in the provided documents.”) may still score poorly in keyword-based accuracy or triad metrics when the golden set expects information that is not available in the KB.\n",
        "\n",
        "3. **Evaluation instability under load**\n",
        "\n",
        "   * Batch evaluation (especially triad scoring) is sensitive to API availability and may require retries or backoff, increasing runtime and introducing minor variability in scoring explanations.\n",
        "\n",
        "---\n",
        "\n",
        "## 8. Deployment Considerations\n",
        "\n",
        "* **Human review triggers:**\n",
        "  Trigger human review when the system refuses an answer, when groundedness or answer relevance scores fall below 3, or when questions indicate gaps in KB coverage (e.g., contact details or operational diagnostics).\n",
        "\n",
        "* **Confidence threshold:**\n",
        "  Automatically surface answers only when groundedness ≥ 4 and answer relevance ≥ 4. Below this threshold, flag responses or route to human support.\n",
        "\n",
        "* **Monitoring:**\n",
        "  Monitor refusal rates, keyword-match accuracy, triad scores, retrieval routing behavior (filtered vs hybrid), and citation compliance to detect KB drift or prompt regressions.\n",
        "\n",
        "* **Update process:**\n",
        "  Any KB or prompt update should be followed by a full golden-set re-evaluation and comparison against previous triad and keyword metrics before deployment.\n",
        "\n",
        "---\n",
        "\n",
        "## 9. Operations\n",
        "\n",
        "* **Knowledge base refresh:**\n",
        "  Review and update the KB quarterly or immediately after changes to pricing, limits, billing rules, or security policies.\n",
        "\n",
        "* **Prompt versioning:**\n",
        "  Version system prompts explicitly (e.g., v1.0, v1.1) and store them alongside evaluation results to enable regression tracking and rollback if performance degrades.\n",
        "\n",
        "* **Incident response:**\n",
        "  If hallucination or incorrect answers are detected in production, disable automated responses for the affected topic, add or correct KB content, re-run evaluation, and redeploy only after metrics return to target thresholds.\n",
        "\n",
        "---\n",
        "\n",
        "## 10. Version History\n",
        "\n",
        "| Version | Changes                                                                                                     | Context Rel. | Groundedness | Answer Rel. |\n",
        "| ------- | ----------------------------------------------------------------------------------------------------------- | ------------ | ------------ | ----------- |\n",
        "| 1.0     | Initial end-to-end RAG pipeline with routed hybrid retrieval, strict grounding prompt, and triad evaluation | ~4.4         | ~4.8         | ~4.2        |\n",
        "\n",
        "---\n"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "5ad3f66a",
      "metadata": {
        "id": "5ad3f66a"
      },
      "source": [
        "---\n",
        "## Export & Submission"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 16,
      "id": "77f9b97c",
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "77f9b97c",
        "outputId": "18eef330-3d9b-449f-e064-6ac782374533"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "✓ Exported day3_assignment_golden_set.json\n",
            "✓ Exported day3_assignment_triad_results.json\n",
            "✓ Exported 32 API calls to day3_assignment_prompt_log.csv\n",
            "\n",
            "==================================================\n",
            "SUBMISSION CHECKLIST\n",
            "==================================================\n",
            "  [✓] 4-6 documents in knowledge base\n",
            "  [✓] 1,000+ total words\n",
            "  [✓] Chunking strategy applied\n",
            "  [✓] System prompt completed (no TODOs)\n",
            "  [✓] 15+ questions processed\n",
            "  [✓] 15+ golden set items\n",
            "  [✓] RAG triad metrics computed\n",
            "  [ ] Error analysis completed (check markdown cells)\n",
            "  [ ] RAG playbook completed (check markdown cells)\n"
          ]
        }
      ],
      "source": [
        "# ── Export all files ──────────────────────────────────────────────────────────\n",
        "\n",
        "# 1. Golden set\n",
        "with open(\"day3_assignment_golden_set.json\", \"w\") as f:\n",
        "    json.dump(GOLDEN_SET, f, indent=2)\n",
        "print(\"✓ Exported day3_assignment_golden_set.json\")\n",
        "\n",
        "# 2. Triad results\n",
        "triad_export = triad_df.to_dict(orient=\"records\")\n",
        "with open(\"day3_assignment_triad_results.json\", \"w\") as f:\n",
        "    json.dump(triad_export, f, indent=2, default=str)\n",
        "print(\"✓ Exported day3_assignment_triad_results.json\")\n",
        "\n",
        "# 3. Prompt log\n",
        "if PROMPT_LOG:\n",
        "    log_df = pd.DataFrame(PROMPT_LOG)\n",
        "    log_df.to_csv(\"day3_assignment_prompt_log.csv\", index=False)\n",
        "    print(f\"✓ Exported {len(PROMPT_LOG)} API calls to day3_assignment_prompt_log.csv\")\n",
        "\n",
        "# Submission checklist\n",
        "print(\"\\n\" + \"=\" * 50)\n",
        "print(\"SUBMISSION CHECKLIST\")\n",
        "print(\"=\" * 50)\n",
        "print(f\"  [{'✓' if len(documents) >= 4 else '✗'}] 4-6 documents in knowledge base\")\n",
        "print(f\"  [{'✓' if total_words >= 1000 else '✗'}] 1,000+ total words\")\n",
        "print(f\"  [{'✓' if len(all_chunks) > 0 else '✗'}] Chunking strategy applied\")\n",
        "print(f\"  [{'✓' if 'TODO' not in FINAL_SYSTEM_PROMPT else '✗'}] System prompt completed (no TODOs)\")\n",
        "print(f\"  [{'✓' if len(QUESTIONS) >= 15 else '✗'}] 15+ questions processed\")\n",
        "print(f\"  [{'✓' if len(GOLDEN_SET) >= 15 else '✗'}] 15+ golden set items\")\n",
        "print(f\"  [{'✓' if len(triad_results) > 0 else '✗'}] RAG triad metrics computed\")\n",
        "print(f\"  [ ] Error analysis completed (check markdown cells)\")\n",
        "print(f\"  [ ] RAG playbook completed (check markdown cells)\")"
      ]
    },
    {
      "cell_type": "markdown",
      "id": "ef7f4556",
      "metadata": {
        "id": "ef7f4556"
      },
      "source": [
        "---\n",
        "## Conclusion\n",
        "\n",
        "Congratulations on completing your first production-ready RAG system!\n",
        "\n",
        "In Day 4, you'll learn about **GenAI Agents** — systems that can call tools, reason over multiple steps, and use your RAG pipeline as one component of a larger autonomous workflow.\n",
        "\n",
        "Your Day 3 RAG system becomes a building block: the agent will be able to *decide when to search your knowledge base*, formulate the right query, and integrate the retrieved answer into a multi-step plan."
      ]
    }
  ],
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.13.6"
    },
    "colab": {
      "provenance": []
    }
  },
  "nbformat": 4,
  "nbformat_minor": 5
}