{
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Day 2: Assignment — Production-Ready Extraction Pipeline\n",
    "\n",
    "## Overview\n",
    "\n",
    "Build and document a **production-ready extraction pipeline** using Structured Outputs.\n",
    "\n",
    "## Deliverables\n",
    "\n",
    "Submit ONE notebook (or PDF export) containing:\n",
    "\n",
    "1. **Final Pydantic schema** with field descriptions\n",
    "2. **Final extraction prompt** (v2 or v3)\n",
    "3. **Extracted outputs** for at least 12 items\n",
    "4. **Golden set** (8+ items) with ground truth labels\n",
    "5. **Metrics** (accuracy for classification and urgency fields)\n",
    "6. **Error analysis** (½–1 page)\n",
    "7. **Prompt playbook** documentation\n",
    "\n",
    "## Grading Criteria\n",
    "\n",
    "| Criterion | Weight | Description |\n",
    "|-----------|--------|-------------|\n",
    "| Schema Design | 15% | Appropriate fields, types, constraints |\n",
    "| Prompt Quality | 25% | Effective rules, examples, structure |\n",
    "| Accuracy | 20% | Performance on golden set |\n",
    "| Analysis | 25% | Error patterns, iteration insights |\n",
    "| Documentation | 15% | Complete playbook, clear explanations |"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Setup"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "!pip install -q -U google-genai"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "import os\nimport time\nimport json\nimport pandas as pd\nfrom datetime import datetime, timezone\nfrom typing import List, Optional, Literal\nfrom pydantic import BaseModel, Field\nfrom google import genai\n\n# Configure API — reads the GEMINI_API_KEY you set up on Day 1\ntry:\n    from google.colab import userdata\n    API_KEY = userdata.get(\"GEMINI_API_KEY\")\nexcept:\n    API_KEY = None\n\nif not API_KEY:\n    import getpass\n    API_KEY = getpass.getpass(\"Enter your Gemini API key: \")\n\nclient = genai.Client(api_key=API_KEY)\nMODEL_ID = \"gemini-2.5-flash-lite\"\n\nprint(f\"✓ API key loaded\")\nprint(f\"✓ Model: {MODEL_ID}\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Infrastructure\n",
    "PROMPT_LOG = []\n",
    "\n",
    "def _now():\n",
    "    return datetime.now(timezone.utc).isoformat().replace('+00:00', 'Z')\n",
    "\n",
    "def generate_structured(prompt, schema_model, temperature=0.2, log=True, label=None):\n",
    "    \"\"\"Generate structured output using Pydantic schema.\"\"\"\n",
    "    start_time = time.time()\n",
    "    \n",
    "    response = client.models.generate_content(\n",
    "        model=MODEL_ID,\n",
    "        contents=prompt,\n",
    "        config={\n",
    "            \"temperature\": temperature,\n",
    "            \"response_mime_type\": \"application/json\",\n",
    "            \"response_json_schema\": schema_model.model_json_schema(),\n",
    "        },\n",
    "    )\n",
    "    \n",
    "    latency = time.time() - start_time\n",
    "    raw_text = response.text or \"\"\n",
    "    result = schema_model.model_validate_json(raw_text)\n",
    "    \n",
    "    if log:\n",
    "        PROMPT_LOG.append({\n",
    "            \"timestamp\": _now(),\n",
    "            \"label\": label,\n",
    "            \"schema\": schema_model.__name__,\n",
    "            \"prompt_length\": len(prompt),\n",
    "            \"response_length\": len(raw_text),\n",
    "            \"latency_s\": round(latency, 3)\n",
    "        })\n",
    "    \n",
    "    return result\n",
    "\n",
    "def evaluate(predictions, golden, field):\n",
    "    \"\"\"Calculate accuracy for a field.\"\"\"\n",
    "    correct, total, errors = 0, 0, []\n",
    "    for id, expected in golden.items():\n",
    "        if id in predictions:\n",
    "            total += 1\n",
    "            pred_val = getattr(predictions[id], field)\n",
    "            if pred_val == expected[field]:\n",
    "                correct += 1\n",
    "            else:\n",
    "                errors.append({\"id\": id, \"expected\": expected[field], \"got\": pred_val})\n",
    "    return {\"correct\": correct, \"total\": total, \n",
    "            \"accuracy\": correct/total if total else 0, \"errors\": errors}\n",
    "\n",
    "print(\"✓ Infrastructure ready\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 1: Final Schema (15 points)\n",
    "\n",
    "Paste your final, refined Pydantic schema below."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# TODO: Paste your final schema here\n",
    "\n",
    "class FinalExtraction(BaseModel):\n",
    "    \"\"\"Your schema description here.\"\"\"\n",
    "    id: str = Field(description=\"Unique identifier\")\n",
    "    \n",
    "    # TODO: Add your fields\n",
    "    # category: Literal[\"...\", \"...\", \"...\"] = Field(description=\"...\")\n",
    "    # urgency: Literal[\"low\", \"medium\", \"high\"] = Field(description=\"...\")\n",
    "    # summary: str = Field(description=\"...\")\n",
    "    # next_step: str = Field(description=\"...\")\n",
    "    # missing_info: List[str] = Field(description=\"...\")\n",
    "    pass\n",
    "\n",
    "\n",
    "class FinalBatch(BaseModel):\n",
    "    \"\"\"Batch of extractions.\"\"\"\n",
    "    items: List[FinalExtraction]\n",
    "\n",
    "\n",
    "# Show schema\n",
    "print(\"Schema fields:\")\n",
    "for name, field in FinalExtraction.model_fields.items():\n",
    "    print(f\"  {name}: {field.annotation}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 2: Final Prompt (25 points)\n",
    "\n",
    "Paste your final, refined extraction prompt."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# TODO: Paste your final prompt here\n",
    "\n",
    "FINAL_PROMPT = \"\"\"\n",
    "# YOUR FINAL EXTRACTION PROMPT\n",
    "\n",
    "# Include:\n",
    "# - Role/context\n",
    "# - Clear instructions\n",
    "# - Category definitions/rules\n",
    "# - Urgency definitions/rules\n",
    "# - Any few-shot examples\n",
    "# - Constraints\n",
    "\n",
    "Items:\n",
    "{items}\n",
    "\"\"\"\n",
    "\n",
    "def format_items(items):\n",
    "    return \"\\n\".join([f\"{it['id']}: {it['text']}\" for it in items])\n",
    "\n",
    "def run_extraction(items, label=\"extraction\"):\n",
    "    prompt = FINAL_PROMPT.format(items=format_items(items))\n",
    "    return generate_structured(prompt, FinalBatch, label=label)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 3: Input Data (12+ items)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# TODO: Provide your input data (at least 12 items)\n",
    "\n",
    "inputs = [\n",
    "    # {\"id\": \"X1\", \"text\": \"...\"},\n",
    "    # {\"id\": \"X2\", \"text\": \"...\"},\n",
    "    # ... at least 12 items\n",
    "]\n",
    "\n",
    "print(f\"Input items: {len(inputs)}\")\n",
    "assert len(inputs) >= 12, \"Need at least 12 input items!\""
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 4: Run Extraction"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Run extraction\n",
    "# result = run_extraction(inputs, label=\"final_extraction\")\n",
    "# print(f\"Extracted: {len(result.items)} items\")\n",
    "\n",
    "# View results\n",
    "# df = pd.DataFrame([item.model_dump() for item in result.items])\n",
    "# df"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Save extracted results\n",
    "# with open(\"day2_assignment_extracted.json\", \"w\") as f:\n",
    "#     json.dump([item.model_dump() for item in result.items], f, indent=2, ensure_ascii=False)\n",
    "# print(\"✓ Saved: day2_assignment_extracted.json\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 5: Golden Set (8+ items)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# TODO: Define your golden set (at least 8 items with ground truth)\n",
    "\n",
    "GOLDEN = {\n",
    "    # \"X1\": {\"category\": \"...\", \"urgency\": \"...\"},\n",
    "    # \"X2\": {\"category\": \"...\", \"urgency\": \"...\"},\n",
    "    # ... at least 8 items\n",
    "}\n",
    "\n",
    "print(f\"Golden set size: {len(GOLDEN)}\")\n",
    "assert len(GOLDEN) >= 8, \"Need at least 8 labeled items!\""
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 6: Compute Metrics (20 points)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Create predictions dict\n",
    "# pred = {item.id: item for item in result.items}\n",
    "\n",
    "# Evaluate - update field names to match your schema!\n",
    "# cat_eval = evaluate(pred, GOLDEN, \"category\")\n",
    "# urg_eval = evaluate(pred, GOLDEN, \"urgency\")\n",
    "\n",
    "# print(\"=\"*50)\n",
    "# print(\"FINAL METRICS\")\n",
    "# print(\"=\"*50)\n",
    "# print(f\"Category Accuracy: {cat_eval['correct']}/{cat_eval['total']} = {cat_eval['accuracy']:.1%}\")\n",
    "# print(f\"Urgency Accuracy:  {urg_eval['correct']}/{urg_eval['total']} = {urg_eval['accuracy']:.1%}\")\n",
    "\n",
    "# Show errors\n",
    "# if cat_eval['errors']:\n",
    "#     print(\"\\nCategory Errors:\")\n",
    "#     for err in cat_eval['errors']:\n",
    "#         print(f\"  {err['id']}: expected '{err['expected']}', got '{err['got']}'\")\n",
    "\n",
    "# if urg_eval['errors']:\n",
    "#     print(\"\\nUrgency Errors:\")\n",
    "#     for err in urg_eval['errors']:\n",
    "#         print(f\"  {err['id']}: expected '{err['expected']}', got '{err['got']}'\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 7: Error Analysis (25 points)\n",
    "\n",
    "Write ½–1 page covering:"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Error Analysis\n",
    "\n",
    "#### 1. Three Common Error Patterns\n",
    "\n",
    "**Pattern 1: [Name]**\n",
    "- Description: [What happened]\n",
    "- Example: [Specific case]\n",
    "- Why it happens: [Root cause]\n",
    "\n",
    "**Pattern 2: [Name]**\n",
    "- Description: [What happened]\n",
    "- Example: [Specific case]\n",
    "- Why it happens: [Root cause]\n",
    "\n",
    "**Pattern 3: [Name]**\n",
    "- Description: [What happened]\n",
    "- Example: [Specific case]\n",
    "- Why it happens: [Root cause]\n",
    "\n",
    "#### 2. Two Prompt Changes That Helped\n",
    "\n",
    "**Change 1: [Description]**\n",
    "- What I changed: [Specific modification]\n",
    "- Impact: [How accuracy improved]\n",
    "\n",
    "**Change 2: [Description]**\n",
    "- What I changed: [Specific modification]\n",
    "- Impact: [How accuracy improved]\n",
    "\n",
    "#### 3. Remaining Risk + Mitigation\n",
    "\n",
    "**Risk:** [Describe a remaining edge case or failure mode]\n",
    "\n",
    "**Mitigation strategy:** [How you would handle this in production]\n",
    "- Example: \"Require human review when confidence is low\" or \"Add fallback rules for ambiguous cases\""
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 8: Prompt Playbook (15 points)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": "---\n\n# 📘 PROMPT PLAYBOOK\n\n## Extraction Pipeline: [YOUR PIPELINE NAME]\n\n**Version:** [e.g., 2.0]  \n**Author:** [Your name]  \n**Date:** [Date]  \n**Status:** Production Ready / Testing\n\n---\n\n### Purpose\n\n[What does this pipeline do? What business problem does it solve?]\n\n---\n\n### Schema\n\n```python\n# Paste your final schema here\nclass FinalExtraction(BaseModel):\n    ...\n```\n\n| Field | Type | Description |\n|-------|------|-------------|\n| id | str | ... |\n| category | Literal[...] | ... |\n| ... | ... | ... |\n\n---\n\n### The Prompt\n\n```\n[Paste your final prompt here]\n```\n\n---\n\n### Recommended Settings\n\n| Setting | Value | Rationale |\n|---------|-------|----------|\n| Model | gemini-2.5-flash-lite | [Why] |\n| Temperature | 0.2 | [Why - lower for structured output] |\n| Structured Output | Yes | Guarantees valid JSON |\n\n---\n\n### Performance Metrics\n\n| Metric | Value |\n|--------|-------|\n| Category Accuracy | [X]% |\n| Urgency Accuracy | [X]% |\n| Golden Set Size | [X] items |\n| Average Latency | [X] seconds |\n\n---\n\n### Category Decision Rules\n\n| Category | When to Use | Example |\n|----------|-------------|---------||\n| [Cat 1] | [Rule] | [Example] |\n| [Cat 2] | [Rule] | [Example] |\n| ... | ... | ... |\n\n---\n\n### Urgency Decision Rules\n\n| Level | When to Use | Example |\n|-------|-------------|---------||\n| high | [Rule] | [Example] |\n| medium | [Rule] | [Example] |\n| low | [Rule] | [Example] |\n\n---\n\n### Known Limitations\n\n1. [Limitation 1]\n2. [Limitation 2]\n3. [Limitation 3]\n\n---\n\n### Version History\n\n| Version | Date | Changes | Category Acc | Urgency Acc |\n|---------|------|---------|--------------|-------------|\n| 1.0 | [date] | Initial | [X]% | [X]% |\n| 2.0 | [date] | [Changes] | [X]% | [X]% |\n\n---\n\n### Production Deployment Notes\n\n- **Human review required when:** [conditions]\n- **Batch size recommendation:** [X] items per API call\n- **Rate limiting:** [considerations]\n- **Monitoring:** [what to track]\n\n---"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Export and Submission"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Export prompt log\n",
    "if PROMPT_LOG:\n",
    "    df_log = pd.DataFrame(PROMPT_LOG)\n",
    "    df_log.to_csv(\"day2_assignment_prompt_log.csv\", index=False)\n",
    "    print(\"✓ Saved: day2_assignment_prompt_log.csv\")\n",
    "    print(df_log)\n",
    "\n",
    "# Summary\n",
    "print(\"\\n\" + \"=\"*50)\n",
    "print(\"SUBMISSION CHECKLIST\")\n",
    "print(\"=\"*50)\n",
    "print(f\"☐ Schema defined with Literal types\")\n",
    "print(f\"☐ Final prompt with rules/examples\")\n",
    "print(f\"☐ Input items: {len(inputs) if 'inputs' in dir() else 0} (need 12+)\")\n",
    "print(f\"☐ Golden set: {len(GOLDEN) if 'GOLDEN' in dir() else 0} (need 8+)\")\n",
    "print(f\"☐ Metrics computed\")\n",
    "print(f\"☐ Error analysis written\")\n",
    "print(f\"☐ Prompt playbook completed\")\n",
    "print(\"\\nFiles to submit:\")\n",
    "print(\"  - This notebook (.ipynb or PDF)\")\n",
    "print(\"  - day2_assignment_extracted.json\")\n",
    "print(\"  - day2_assignment_prompt_log.csv\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "**Congratulations on completing Day 2!**\n",
    "\n",
    "You've learned how to:\n",
    "- Use Pydantic schemas for guaranteed-valid structured outputs\n",
    "- Build and iterate on extraction prompts\n",
    "- Evaluate against golden sets\n",
    "- Document production-ready prompt pipelines\n",
    "\n",
    "Tomorrow: **Retrieval-Augmented Generation (RAG)** — grounding LLM responses in your own data!"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python",
   "version": "3.10.0"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 4
}