{
 "nbformat": 4,
 "nbformat_minor": 5,
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python",
   "version": "3.11.0"
  }
 },
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Day 3: Assignment — Production-Ready RAG System\n",
    "\n",
    "## Overview\n",
    "\n",
    "Build a complete, documented RAG system ready for production handoff. This assignment extends your independent lab work with deeper evaluation (RAG triad), comprehensive error analysis, and a full RAG playbook.\n",
    "\n",
    "### Grading Summary\n",
    "| # | Deliverable | Points |\n",
    "|---|-------------|--------|\n",
    "| 1 | Knowledge Base Design | 15 |\n",
    "| 2 | RAG System Prompt | 25 |\n",
    "| 3 | RAG Outputs (15+ questions) | — |\n",
    "| 4 | Golden Q&A Set (15+ items) | — |\n",
    "| 5 | RAG Triad Metrics | 20 |\n",
    "| 6 | Error Analysis | 25 |\n",
    "| 7 | RAG Playbook | 15 |\n",
    "| | **Total** | **100** |"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Setup"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "!pip install -q -U google-genai"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "import os\nimport json\nimport re\nfrom typing import Optional, List, Tuple\nfrom datetime import datetime, timezone\nimport numpy as np\nimport pandas as pd\nfrom pydantic import BaseModel, Field\n\n# Google GenAI\nfrom google import genai\nfrom google.genai import types\n\ntry:\n    from google.colab import userdata\n    GEMINI_API_KEY = userdata.get(\"GEMINI_API_KEY\")\nexcept ImportError:\n    GEMINI_API_KEY = os.environ.get(\"GEMINI_API_KEY\")\n    if not GEMINI_API_KEY:\n        raise ValueError(\"GEMINI_API_KEY not found. Please set it in Colab Secrets or environment.\")\n\nclient = genai.Client(api_key=GEMINI_API_KEY)\n\nMODEL_ID = \"gemini-2.5-flash-lite\"\nEMBEDDING_MODEL = \"gemini-embedding-001\"\n\nprint(f\"Using model: {MODEL_ID}\")\nprint(f\"Using embedding: {EMBEDDING_MODEL}\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Infrastructure: API, Logging, and Core Functions ─────────────────────────\n\nPROMPT_LOG = []  # Track all API calls\n\ndef _now():\n    return datetime.now(timezone.utc).isoformat().replace('+00:00', 'Z')\n\ndef generate(prompt, label=\"default\", temperature=0.3):\n    \"\"\"Generate text using the model.\"\"\"\n    response = client.models.generate_content(\n        model=MODEL_ID,\n        contents=prompt,\n        config={\"temperature\": temperature, \"max_output_tokens\": 1000},\n    )\n    text = response.text\n    PROMPT_LOG.append({\n        \"timestamp\": _now(),\n        \"type\": \"generate\",\n        \"label\": label,\n        \"prompt_length\": len(prompt),\n        \"response_length\": len(text),\n    })\n    return text\n\ndef generate_structured(prompt, response_model: BaseModel, label=\"default\", temperature=0.3):\n    \"\"\"Generate structured JSON response.\"\"\"\n    response = client.models.generate_content(\n        model=MODEL_ID,\n        contents=prompt,\n        config={\n            \"response_mime_type\": \"application/json\",\n            \"response_schema\": response_model,\n            \"temperature\": temperature,\n        },\n    )\n    json_str = response.text\n    PROMPT_LOG.append({\n        \"timestamp\": _now(),\n        \"type\": \"generate_structured\",\n        \"label\": label,\n        \"prompt_length\": len(prompt),\n        \"response_length\": len(json_str),\n    })\n    return response_model.model_validate_json(json_str)\n\ndef embed_texts(texts, task_type=\"RETRIEVAL_DOCUMENT\"):\n    \"\"\"Embed one or more texts using Gemini Embeddings API.\"\"\"\n    if isinstance(texts, str):\n        texts = [texts]\n    response = client.models.embed_content(\n        model=EMBEDDING_MODEL,\n        contents=texts,\n        config=types.EmbedContentConfig(task_type=task_type),\n    )\n    return [np.array(e.values) for e in response.embeddings]\n\ndef cosine_similarity(a, b):\n    \"\"\"Compute cosine similarity between two vectors.\"\"\"\n    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))\n\ndef chunk_sentences(text, max_words=80, overlap_words=20):\n    \"\"\"\n    Split text into chunks at sentence boundaries.\n    max_words: target chunk size\n    overlap_words: overlap between chunks\n    \"\"\"\n    sentences = re.split(r'(?<=[.!?])\\s+', text.strip())\n    sentences = [s.strip() for s in sentences if s.strip()]\n    \n    chunks = []\n    current_chunk = []\n    current_words = 0\n    \n    for sent in sentences:\n        sent_words = len(sent.split())\n        if current_words + sent_words > max_words and current_chunk:\n            chunks.append(' '.join(current_chunk))\n            # Overlap: keep last few sentences\n            overlap_sents = []\n            overlap_count = 0\n            for s in reversed(current_chunk):\n                overlap_count += len(s.split())\n                if overlap_count > overlap_words:\n                    break\n                overlap_sents.insert(0, s)\n            current_chunk = overlap_sents\n            current_words = sum(len(s.split()) for s in current_chunk)\n        \n        current_chunk.append(sent)\n        current_words += sent_words\n    \n    if current_chunk:\n        chunks.append(' '.join(current_chunk))\n    \n    return chunks\n\ndef search(query, chunks, chunk_embeddings, top_k=3):\n    \"\"\"\n    Retrieve top-k relevant chunks using cosine similarity.\n    Returns list of (index, score, text) tuples.\n    \"\"\"\n    query_emb = embed_texts(query, task_type=\"RETRIEVAL_QUERY\")[0]\n    scores = []\n    for i, doc_emb in enumerate(chunk_embeddings):\n        sim = cosine_similarity(query_emb, doc_emb)\n        scores.append((i, sim, chunks[i]))\n    scores.sort(key=lambda x: x[1], reverse=True)\n    return scores[:top_k]\n\ndef rag_query(query, chunks, chunk_embeddings, top_k=3, system_prompt=\"\"):\n    \"\"\"\n    Execute a RAG query: retrieve context, then generate answer.\n    Returns (answer, retrieved_results).\n    \"\"\"\n    # Retrieve\n    retrieved = search(query, chunks, chunk_embeddings, top_k=top_k)\n    context = \"\\n\\n\".join(f\"[Source {i}]: {text}\" for i, (_, _, text) in enumerate(retrieved))\n    \n    # Generate\n    rag_prompt = f\"\"\"{system_prompt}\n\nContext:\n{context}\n\nQuestion: {query}\n\nAnswer:\"\"\"\n    answer = generate(rag_prompt, label=\"rag_query\")\n    return answer, retrieved\n\n# ── Hybrid Retrieval (optional, for bonus points) ────────\nfrom sklearn.feature_extraction.text import TfidfVectorizer\n\ndef build_sparse_index(chunks):\n    \"\"\"Build TF-IDF sparse index for keyword retrieval.\"\"\"\n    vectorizer = TfidfVectorizer(stop_words='english')\n    matrix = vectorizer.fit_transform(chunks)\n    return vectorizer, matrix\n\ndef hybrid_search(query, doc_embeddings, chunks, vectorizer, tfidf_matrix,\n                  top_k=5, semantic_weight=0.6, keyword_weight=0.4):\n    \"\"\"Combine semantic and keyword search with weighted scoring.\"\"\"\n    from sklearn.metrics.pairwise import cosine_similarity as sklearn_cosine\n    # Semantic scores\n    query_emb = embed_texts(query, task_type=\"RETRIEVAL_QUERY\")[0]\n    sem_scores = [cosine_similarity(query_emb, emb) for emb in doc_embeddings]\n    # Sparse scores\n    query_vec = vectorizer.transform([query])\n    sparse_scores = sklearn_cosine(query_vec, tfidf_matrix).ravel()\n    # Combine\n    combined = []\n    for i in range(len(chunks)):\n        combo = semantic_weight * sem_scores[i] + keyword_weight * float(sparse_scores[i])\n        combined.append((i, combo, chunks[i]))\n    combined.sort(key=lambda x: x[1], reverse=True)\n    return combined[:top_k]\n\ndef search_with_filter(query, doc_embeddings, chunks, chunk_sources, top_k=3,\n                       filter_doc=None):\n    \"\"\"Semantic search with optional document source filter.\"\"\"\n    query_emb = embed_texts(query, task_type=\"RETRIEVAL_QUERY\")[0]\n    scores = []\n    for i, doc_emb in enumerate(doc_embeddings):\n        if filter_doc is not None and chunk_sources[i] != filter_doc:\n            continue\n        sim = cosine_similarity(query_emb, doc_emb)\n        scores.append((i, sim, chunks[i]))\n    scores.sort(key=lambda x: x[1], reverse=True)\n    return scores[:top_k]\n\n# ── Retrieval Metrics ─────────────────────────────────────\n\ndef precision_recall_at_k(retrieved_indices, expected_indices, k):\n    \"\"\"Compute Precision@k and Recall@k.\"\"\"\n    retrieved_set = set(retrieved_indices[:k])\n    expected_set = set(expected_indices)\n    if not expected_set:\n        return None, None\n    hits = retrieved_set & expected_set\n    precision = len(hits) / k if k > 0 else 0\n    recall = len(hits) / len(expected_set)\n    return precision, recall\n\nprint(\"All infrastructure ready (includes hybrid retrieval, metadata filtering, and retrieval metrics).\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── RAG Triad Evaluation ────────────────────────────────────────────────────\n",
    "\n",
    "class RAGTriadScore(BaseModel):\n",
    "    \"\"\"RAG triad evaluation for a single question.\"\"\"\n",
    "    context_relevance: int = Field(description=\"1-5: Are the retrieved chunks relevant to the question?\")\n",
    "    groundedness: int = Field(description=\"1-5: Is the answer supported by the retrieved chunks?\")\n",
    "    answer_relevance: int = Field(description=\"1-5: Does the answer address the question asked?\")\n",
    "    explanation: str = Field(description=\"Brief explanation of all three scores\")\n",
    "\n",
    "def evaluate_rag_triad(question, answer, retrieved_chunks, label=\"triad\"):\n",
    "    \"\"\"Score a single RAG output on all three triad dimensions.\"\"\"\n",
    "    chunks_text = \"\\n\\n\".join(f\"[Chunk {i+1}]: {c}\" for i, c in enumerate(retrieved_chunks))\n",
    "    eval_prompt = f\"\"\"Evaluate this RAG system output on three dimensions.\n",
    "\n",
    "Question: {question}\n",
    "\n",
    "Retrieved Context:\n",
    "{chunks_text}\n",
    "\n",
    "Generated Answer: {answer}\n",
    "\n",
    "Score each dimension 1-5:\n",
    "1. Context Relevance: Are the retrieved chunks relevant to answering this question?\n",
    "2. Groundedness: Is the generated answer fully supported by the retrieved chunks? (5=fully grounded, 1=hallucinated)\n",
    "3. Answer Relevance: Does the answer actually address what was asked?\"\"\"\n",
    "    return generate_structured(eval_prompt, RAGTriadScore, label=label)\n",
    "\n",
    "print(\"RAG Triad evaluation function loaded.\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 1: Knowledge Base Design (15 points)\n",
    "\n",
    "Design your knowledge base with 4-6 documents. You may reuse and expand documents from Lab 2 or create new ones.\n",
    "\n",
    "**Document your decisions:**\n",
    "- Domain and why you chose it\n",
    "- Chunk size, overlap, and rationale\n",
    "- Number of resulting chunks"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── TODO: Define your knowledge base ───────────────────────────────────────\n",
    "# Include 4-6 documents with at least 1,000 total words\n",
    "\n",
    "documents = [\n",
    "    \"\"\"TODO: Document 1 — [Title]\n",
    "    \n",
    "    [Your document content here, 150-300 words]\n",
    "    \"\"\",\n",
    "    \"\"\"TODO: Document 2 — [Title]\n",
    "    \n",
    "    [Your document content here, 150-300 words]\n",
    "    \"\"\",\n",
    "    \"\"\"TODO: Document 3 — [Title]\n",
    "    \n",
    "    [Your document content here, 150-300 words]\n",
    "    \"\"\",\n",
    "    \"\"\"TODO: Document 4 — [Title]\n",
    "    \n",
    "    [Your document content here, 150-300 words]\n",
    "    \"\"\",\n",
    "    # Add more documents as needed (5-6 recommended)\n",
    "]\n",
    "\n",
    "total_words = sum(len(doc.split()) for doc in documents)\n",
    "print(f\"Documents: {len(documents)}\")\n",
    "print(f\"Total words: {total_words}\")\n",
    "assert total_words >= 1000, f\"Need at least 1,000 words, have {total_words}\""
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── TODO: Configure and apply chunking ───────────────────\nCHUNK_MAX_WORDS = 80     # TODO: Adjust based on your content\nCHUNK_OVERLAP_WORDS = 20  # TODO: Adjust overlap\n\n# Document your rationale:\nCHUNKING_RATIONALE = \"\"\"\nTODO: Explain your chunking decisions:\n- Why this chunk size?\n- Why this overlap?\n- Did you try other values? What happened?\n\"\"\"\n\nall_chunks = []\nchunk_sources = []    # Track which document each chunk came from\nchunk_metadata = []   # Rich metadata for each chunk\nfor doc_idx, doc in enumerate(documents):\n    doc_title = doc.strip().split('\\n')[0]  # First line as title\n    chunks = chunk_sentences(doc, max_words=CHUNK_MAX_WORDS, overlap_words=CHUNK_OVERLAP_WORDS)\n    for chunk_idx, chunk in enumerate(chunks):\n        all_chunks.append(chunk)\n        chunk_sources.append(doc_idx)\n        chunk_metadata.append({\n            \"chunk_id\": f\"DOC{doc_idx}::C{chunk_idx+1}\",\n            \"doc_index\": doc_idx,\n            \"doc_title\": doc_title,\n        })\n\nchunk_embeddings = embed_texts(all_chunks, task_type=\"RETRIEVAL_DOCUMENT\")\n\n# Optional: build sparse index for hybrid retrieval\ntfidf_vectorizer, tfidf_matrix = build_sparse_index(all_chunks)\n\nprint(f\"Chunking: max_words={CHUNK_MAX_WORDS}, overlap={CHUNK_OVERLAP_WORDS}\")\nprint(f\"Total chunks: {len(all_chunks)}\")\nprint(f\"Dense index:  {len(chunk_embeddings)} embeddings\")\nprint(f\"Sparse index: {tfidf_matrix.shape[1]} vocabulary terms\")\nprint(f\"\\nChunks per document:\")\nfor i in range(len(documents)):\n    count = chunk_sources.count(i)\n    print(f\"  Doc {i}: {count} chunks\")\nprint(f\"\\n{CHUNKING_RATIONALE}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 2: RAG System Prompt (25 points)\n",
    "\n",
    "Your prompt must include ALL 8 components: role, task, grounding rules, scope, format, refusal, examples, conflict handling."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── TODO: Write your complete system prompt ─────────────────────────────────\n",
    "\n",
    "FINAL_SYSTEM_PROMPT = \"\"\"TODO: Write your production-quality system prompt.\n",
    "\n",
    "Include ALL of these components:\n",
    "\n",
    "1. ROLE: Who is this assistant?\n",
    "2. TASK: What does it do?\n",
    "3. GROUNDING RULES: How must it use the context?\n",
    "4. SCOPE: What topics are in/out of scope?\n",
    "5. FORMAT: Answer length, citation format, structure\n",
    "6. REFUSAL: What to say when the answer is not in context\n",
    "7. EXAMPLES: At least one example of a good answer\n",
    "8. CONFLICT HANDLING: What to do when chunks disagree\n",
    "\n",
    "\"\"\"\n",
    "\n",
    "print(\"System prompt:\")\n",
    "print(FINAL_SYSTEM_PROMPT)\n",
    "print(f\"\\nPrompt length: {len(FINAL_SYSTEM_PROMPT.split())} words\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 3: Input Questions (15+)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── TODO: Define 15+ questions to process ──────────────────────────────────\n",
    "QUESTIONS = [\n",
    "    \"TODO: Question 1\",\n",
    "    \"TODO: Question 2\",\n",
    "    \"TODO: Question 3\",\n",
    "    \"TODO: Question 4\",\n",
    "    \"TODO: Question 5\",\n",
    "    \"TODO: Question 6\",\n",
    "    \"TODO: Question 7\",\n",
    "    \"TODO: Question 8\",\n",
    "    \"TODO: Question 9\",\n",
    "    \"TODO: Question 10\",\n",
    "    \"TODO: Question 11\",\n",
    "    \"TODO: Question 12\",\n",
    "    \"TODO: Question 13\",\n",
    "    \"TODO: Question 14\",\n",
    "    \"TODO: Question 15\",\n",
    "    # Add more as needed\n",
    "]\n",
    "\n",
    "assert len(QUESTIONS) >= 15, f\"Need at least 15 questions, have {len(QUESTIONS)}\"\n",
    "print(f\"Questions: {len(QUESTIONS)}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 4: Run RAG Pipeline"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Run RAG on all questions\n",
    "rag_outputs = []\n",
    "for q in QUESTIONS:\n",
    "    answer, retrieved = rag_query(\n",
    "        q, all_chunks, chunk_embeddings,\n",
    "        top_k=3, system_prompt=FINAL_SYSTEM_PROMPT\n",
    "    )\n",
    "    rag_outputs.append({\n",
    "        \"question\": q,\n",
    "        \"answer\": answer,\n",
    "        \"retrieved_chunks\": [doc for _, _, doc in retrieved],\n",
    "        \"retrieval_scores\": [score for _, score, _ in retrieved],\n",
    "    })\n",
    "\n",
    "print(f\"Processed {len(rag_outputs)} questions\\n\")\n",
    "for i, out in enumerate(rag_outputs[:3]):\n",
    "    print(f\"Q{i+1}: {out['question']}\")\n",
    "    print(f\"A: {out['answer'][:100]}...\\n\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Export RAG outputs\n",
    "with open(\"day3_assignment_rag_outputs.json\", \"w\") as f:\n",
    "    json.dump(rag_outputs, f, indent=2, default=str)\n",
    "print(\"Exported to day3_assignment_rag_outputs.json\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 5: Golden Q&A Set (15+ items)\n",
    "\n",
    "Create your golden set with expected keywords. Include easy, medium, hard, and refusal questions."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── TODO: Create your golden test set ────────────────────\nGOLDEN_SET = [\n    # Easy (6-7 items)\n    {\"id\": \"Q01\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"easy\"},\n    {\"id\": \"Q02\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"easy\"},\n    {\"id\": \"Q03\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"easy\"},\n    {\"id\": \"Q04\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"easy\"},\n    {\"id\": \"Q05\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"easy\"},\n    {\"id\": \"Q06\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"easy\"},\n    # Medium (4-5 items)\n    {\"id\": \"Q07\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"medium\"},\n    {\"id\": \"Q08\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"medium\"},\n    {\"id\": \"Q09\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"medium\"},\n    {\"id\": \"Q10\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"medium\"},\n    # Hard / Edge (2-3 items)\n    {\"id\": \"Q11\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"hard\"},\n    {\"id\": \"Q12\", \"question\": \"TODO\", \"expected_keywords\": [\"TODO\"], \"expected_chunks\": [], \"difficulty\": \"hard\"},\n    # Refusal (2-3 items)\n    {\"id\": \"Q13\", \"question\": \"TODO\", \"expected_keywords\": [\"don't have\", \"not in\", \"no information\", \"outside\"], \"expected_chunks\": [], \"difficulty\": \"refusal\"},\n    {\"id\": \"Q14\", \"question\": \"TODO\", \"expected_keywords\": [\"don't have\", \"not in\", \"no information\", \"outside\"], \"expected_chunks\": [], \"difficulty\": \"refusal\"},\n    {\"id\": \"Q15\", \"question\": \"TODO\", \"expected_keywords\": [\"don't have\", \"not in\", \"no information\", \"outside\"], \"expected_chunks\": [], \"difficulty\": \"refusal\"},\n]\n\nassert len(GOLDEN_SET) >= 15, f\"Need at least 15 items, have {len(GOLDEN_SET)}\"\nprint(f\"Golden set: {len(GOLDEN_SET)} questions\")\nprint(\"NOTE: Fill in expected_chunks with the chunk indices you expect to be retrieved.\")\nprint(\"      Leave empty ([]) for refusal questions.\\n\")\nfor q in GOLDEN_SET:\n    print(f\"  [{q['difficulty']:7s}] {q['id']}: {q['question']}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 6: RAG Triad Metrics (20 points)\n",
    "\n",
    "Evaluate EVERY golden set question on all three RAG triad dimensions."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# Run RAG + triad evaluation + retrieval metrics on golden set\ntriad_results = []\n\nfor qa in GOLDEN_SET:\n    # Run RAG\n    answer, retrieved = rag_query(\n        qa[\"question\"], all_chunks, chunk_embeddings,\n        top_k=3, system_prompt=FINAL_SYSTEM_PROMPT\n    )\n    retrieved_texts = [doc for _, _, doc in retrieved]\n    retrieved_indices = [idx for idx, _, _ in retrieved]\n    \n    # Keyword check\n    answer_lower = answer.lower()\n    keyword_hit = any(kw.lower() in answer_lower for kw in qa[\"expected_keywords\"])\n    \n    # RAG Triad evaluation\n    triad = evaluate_rag_triad(qa[\"question\"], answer, retrieved_texts, label=f\"triad_{qa['id']}\")\n    \n    # Precision@k / Recall@k (if expected_chunks provided)\n    p_at_k, r_at_k = None, None\n    if qa.get(\"expected_chunks\"):\n        p_at_k, r_at_k = precision_recall_at_k(retrieved_indices, qa[\"expected_chunks\"], k=3)\n    \n    triad_results.append({\n        \"id\": qa[\"id\"],\n        \"question\": qa[\"question\"],\n        \"difficulty\": qa[\"difficulty\"],\n        \"answer\": answer,\n        \"keyword_match\": keyword_hit,\n        \"context_relevance\": triad.context_relevance,\n        \"groundedness\": triad.groundedness,\n        \"answer_relevance\": triad.answer_relevance,\n        \"precision_at_3\": p_at_k,\n        \"recall_at_3\": r_at_k,\n        \"explanation\": triad.explanation,\n    })\n\ntriad_df = pd.DataFrame(triad_results)"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# Display RAG Triad + Retrieval Metrics\nprint(\"=\" * 70)\nprint(\"RAG TRIAD + RETRIEVAL EVALUATION RESULTS\")\nprint(\"=\" * 70)\n\nkeyword_acc = triad_df[\"keyword_match\"].mean()\navg_context = triad_df[\"context_relevance\"].mean()\navg_ground = triad_df[\"groundedness\"].mean()\navg_relevance = triad_df[\"answer_relevance\"].mean()\n\n# Precision/Recall (only for non-refusal questions)\npr_df = triad_df.dropna(subset=[\"precision_at_3\"])\navg_precision = pr_df[\"precision_at_3\"].mean() if len(pr_df) > 0 else 0\navg_recall = pr_df[\"recall_at_3\"].mean() if len(pr_df) > 0 else 0\n\nprint(f\"\\nOverall Metrics:\")\nprint(f\"  Keyword Accuracy:    {keyword_acc:.0%} ({triad_df['keyword_match'].sum()}/{len(triad_df)})\")\nprint(f\"  Context Relevance:   {avg_context:.1f}/5 {'✓' if avg_context >= 4.0 else '✗ (target: ≥ 4.0)'}\")\nprint(f\"  Groundedness:        {avg_ground:.1f}/5 {'✓' if avg_ground >= 4.0 else '✗ (target: ≥ 4.0)'}\")\nprint(f\"  Answer Relevance:    {avg_relevance:.1f}/5 {'✓' if avg_relevance >= 4.0 else '✗ (target: ≥ 4.0)'}\")\nif len(pr_df) > 0:\n    print(f\"  Avg Precision@3:     {avg_precision:.2f}\")\n    print(f\"  Avg Recall@3:        {avg_recall:.2f}\")\n\nprint(f\"\\nPer-Question Scores:\")\ndisplay_cols = [\"id\", \"difficulty\", \"keyword_match\", \"context_relevance\",\n                \"groundedness\", \"answer_relevance\"]\nif len(pr_df) > 0:\n    display_cols.extend([\"precision_at_3\", \"recall_at_3\"])\nprint(triad_df[display_cols].to_string(index=False))\n\n# Flag questions below target\nbelow_target = triad_df[\n    (triad_df[\"context_relevance\"] < 4) |\n    (triad_df[\"groundedness\"] < 4) |\n    (triad_df[\"answer_relevance\"] < 4)\n]\nif len(below_target) > 0:\n    print(f\"\\nQuestions below target (< 4.0 on any dimension):\")\n    for _, row in below_target.iterrows():\n        dims = []\n        if row[\"context_relevance\"] < 4: dims.append(f\"Context={row['context_relevance']}\")\n        if row[\"groundedness\"] < 4: dims.append(f\"Ground={row['groundedness']}\")\n        if row[\"answer_relevance\"] < 4: dims.append(f\"Relevance={row['answer_relevance']}\")\n        print(f\"  {row['id']}: {', '.join(dims)} — {row['explanation'][:80]}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 7: Error Analysis (25 points)\n",
    "\n",
    "Write your error analysis below. Cover all three sections.\n",
    "\n",
    "### A. Retrieval Failures\n",
    "\n",
    "**Pattern 1:** [Name]\n",
    "- **What went wrong:** [Description — wrong chunks retrieved or correct chunks ranked too low]\n",
    "- **Example:** [Question ID, what was retrieved vs what should have been]\n",
    "- **Root cause:** [Chunk size issue? Embedding limitation? Query phrasing?]\n",
    "\n",
    "**Pattern 2:** [Name]\n",
    "- **What went wrong:** [Description]\n",
    "- **Example:** [Specific example]\n",
    "- **Root cause:** [Analysis]\n",
    "\n",
    "### B. Generation Failures\n",
    "\n",
    "**Pattern 1:** [Name — e.g., hallucination, over-hedging, wrong scope]\n",
    "- **What happened:** [Description]\n",
    "- **Example:** [Question, context provided, and incorrect answer]\n",
    "- **Prompt fix attempted:** [What you changed and whether it helped]\n",
    "\n",
    "**Pattern 2:** [Name]\n",
    "- **What happened:** [Description]\n",
    "- **Example:** [Specific example]\n",
    "- **Prompt fix attempted:** [Change and result]\n",
    "\n",
    "### C. Refusal Testing\n",
    "\n",
    "- **Correct refusals:** [How many out-of-scope questions were correctly refused?]\n",
    "- **False refusals:** [Did the system refuse when it should have answered? Examples?]\n",
    "- **Missed refusals:** [Did the system answer when it should have refused? Examples?]\n",
    "- **Production improvements:** [What would you change for production refusal handling?]"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 8: RAG Playbook (15 points)\n",
    "\n",
    "Complete ALL sections of the playbook template below.\n",
    "\n",
    "---\n",
    "\n",
    "# RAG PLAYBOOK: [Your System Name]\n",
    "\n",
    "**Version:** 1.0\n",
    "**Author:** [Your Name]\n",
    "**Date:** [Date]\n",
    "**Status:** Production Ready\n",
    "\n",
    "## 1. Purpose\n",
    "[TODO: What does this RAG system do? What problem does it solve? Who uses it?]\n",
    "\n",
    "## 2. Knowledge Base\n",
    "| Property | Value |\n",
    "|----------|-------|\n",
    "| Documents | [TODO] |\n",
    "| Total words | [TODO] |\n",
    "| Domain | [TODO] |\n",
    "| Update frequency | [TODO] |\n",
    "\n",
    "## 3. Chunking Strategy\n",
    "| Parameter | Value | Rationale |\n",
    "|-----------|-------|----------|\n",
    "| Strategy | [TODO] | [TODO] |\n",
    "| Chunk size | [TODO] | [TODO] |\n",
    "| Overlap | [TODO] | [TODO] |\n",
    "| Total chunks | [TODO] | — |\n",
    "\n",
    "## 4. System Prompt\n",
    "[TODO: Paste your complete final system prompt]\n",
    "\n",
    "## 5. Model Settings\n",
    "| Setting | Value | Rationale |\n",
    "|---------|-------|----------|\n",
    "| Generation model | gemini-2.5-flash-lite | [TODO] |\n",
    "| Embedding model | gemini-embedding-001 | [TODO] |\n",
    "| Temperature | [TODO] | [TODO] |\n",
    "| top_k | [TODO] | [TODO] |\n",
    "\n",
    "## 6. RAG Triad Performance\n",
    "| Metric | Score | Target |\n",
    "|--------|-------|--------|\n",
    "| Context Relevance | [TODO] | ≥ 4.0 |\n",
    "| Groundedness | [TODO] | ≥ 4.0 |\n",
    "| Answer Relevance | [TODO] | ≥ 4.0 |\n",
    "| Keyword Accuracy | [TODO] | ≥ 80% |\n",
    "\n",
    "## 7. Known Limitations\n",
    "1. [TODO: Limitation 1]\n",
    "2. [TODO: Limitation 2]\n",
    "3. [TODO: Limitation 3]\n",
    "\n",
    "## 8. Deployment Considerations\n",
    "- **Human review triggers:** [TODO]\n",
    "- **Confidence threshold:** [TODO]\n",
    "- **Monitoring:** [TODO]\n",
    "- **Update process:** [TODO]\n",
    "\n",
    "## 9. Operations\n",
    "- **Knowledge base refresh:** [TODO]\n",
    "- **Prompt versioning:** [TODO]\n",
    "- **Incident response:** [TODO]\n",
    "\n",
    "## 10. Version History\n",
    "| Version | Changes | Context Rel. | Groundedness | Answer Rel. |\n",
    "|---------|---------|-------------|-------------|-------------|\n",
    "| 1.0 | Initial | [TODO] | [TODO] | [TODO] |\n",
    "\n",
    "## 11. Handoff Checklist\n",
    "- [ ] Knowledge base documented\n",
    "- [ ] System prompt reviewed by team\n",
    "- [ ] Test set of 15+ questions provided\n",
    "- [ ] All RAG triad metrics meet targets\n",
    "- [ ] Limitations acknowledged\n",
    "- [ ] Monitoring plan in place"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Export & Submission"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Export all files ──────────────────────────────────────────────────────────\n",
    "\n",
    "# 1. Golden set\n",
    "with open(\"day3_assignment_golden_set.json\", \"w\") as f:\n",
    "    json.dump(GOLDEN_SET, f, indent=2)\n",
    "print(\"✓ Exported day3_assignment_golden_set.json\")\n",
    "\n",
    "# 2. Triad results\n",
    "triad_export = triad_df.to_dict(orient=\"records\")\n",
    "with open(\"day3_assignment_triad_results.json\", \"w\") as f:\n",
    "    json.dump(triad_export, f, indent=2, default=str)\n",
    "print(\"✓ Exported day3_assignment_triad_results.json\")\n",
    "\n",
    "# 3. Prompt log\n",
    "if PROMPT_LOG:\n",
    "    log_df = pd.DataFrame(PROMPT_LOG)\n",
    "    log_df.to_csv(\"day3_assignment_prompt_log.csv\", index=False)\n",
    "    print(f\"✓ Exported {len(PROMPT_LOG)} API calls to day3_assignment_prompt_log.csv\")\n",
    "\n",
    "# Submission checklist\n",
    "print(\"\\n\" + \"=\" * 50)\n",
    "print(\"SUBMISSION CHECKLIST\")\n",
    "print(\"=\" * 50)\n",
    "print(f\"  [{'✓' if len(documents) >= 4 else '✗'}] 4-6 documents in knowledge base\")\n",
    "print(f\"  [{'✓' if total_words >= 1000 else '✗'}] 1,000+ total words\")\n",
    "print(f\"  [{'✓' if len(all_chunks) > 0 else '✗'}] Chunking strategy applied\")\n",
    "print(f\"  [{'✓' if 'TODO' not in FINAL_SYSTEM_PROMPT else '✗'}] System prompt completed (no TODOs)\")\n",
    "print(f\"  [{'✓' if len(QUESTIONS) >= 15 else '✗'}] 15+ questions processed\")\n",
    "print(f\"  [{'✓' if len(GOLDEN_SET) >= 15 else '✗'}] 15+ golden set items\")\n",
    "print(f\"  [{'✓' if len(triad_results) > 0 else '✗'}] RAG triad metrics computed\")\n",
    "print(f\"  [ ] Error analysis completed (check markdown cells)\")\n",
    "print(f\"  [ ] RAG playbook completed (check markdown cells)\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Conclusion\n",
    "\n",
    "Congratulations on completing your first production-ready RAG system!\n",
    "\n",
    "In Day 4, you'll learn about **GenAI Agents** — systems that can call tools, reason over multiple steps, and use your RAG pipeline as one component of a larger autonomous workflow.\n",
    "\n",
    "Your Day 3 RAG system becomes a building block: the agent will be able to *decide when to search your knowledge base*, formulate the right query, and integrate the retrieved answer into a multi-step plan."
   ]
  }
 ]
}