{
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Day 2: Independent Lab — Build Your Own Extractor\n",
    "\n",
    "## Overview\n",
    "\n",
    "In this lab, you'll build a complete **extraction + evaluation pipeline** for a business domain of your choice.\n",
    "\n",
    "## Deliverables\n",
    "\n",
    "1. **Pydantic schema** defining your extraction structure\n",
    "2. **Extractor prompt v1** (baseline)\n",
    "3. **Extractor prompt v2** (improved based on evaluation)\n",
    "4. **Golden set** of 8+ labeled items\n",
    "5. **Metrics** comparing v1 vs v2 accuracy\n",
    "6. **Error analysis** documenting what you learned\n",
    "\n",
    "## Choose Your Track\n",
    "\n",
    "- **Track A:** Customer Support Triage (tickets → category, urgency, action)\n",
    "- **Track B:** Meeting Notes Extraction (notes → attendees, actions, decisions)\n",
    "- **Track C:** Product Review Analysis (reviews → sentiment, features, intent)\n",
    "- **Track D:** Job Posting Parser (postings → requirements, benefits, red flags)\n",
    "\n",
    "## Time Estimate: 60-90 minutes"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Setup"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "!pip install -q -U google-genai"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "import os\nimport time\nimport json\nimport pandas as pd\nfrom datetime import datetime, timezone\nfrom typing import List, Optional, Literal\nfrom pydantic import BaseModel, Field\nfrom google import genai\nfrom google.genai import types\n\n# Configure API — reads the GEMINI_API_KEY you set up on Day 1\ntry:\n    from google.colab import userdata\n    API_KEY = userdata.get(\"GEMINI_API_KEY\")\nexcept:\n    API_KEY = None\n\nif not API_KEY:\n    import getpass\n    API_KEY = getpass.getpass(\"Enter your Gemini API key: \")\n\nclient = genai.Client(api_key=API_KEY)\nMODEL_ID = \"gemini-2.5-flash-lite\"\n\nprint(f\"✓ API key loaded\")\nprint(f\"✓ Model: {MODEL_ID}\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Infrastructure\n",
    "PROMPT_LOG = []\n",
    "\n",
    "def _now():\n",
    "    return datetime.now(timezone.utc).isoformat().replace('+00:00', 'Z')\n",
    "\n",
    "def generate_structured(prompt, schema_model, temperature=0.2, log=True, label=None):\n",
    "    \"\"\"\n",
    "    Generate structured output using Pydantic schema.\n",
    "    The API guarantees the output matches your schema.\n",
    "    \"\"\"\n",
    "    start_time = time.time()\n",
    "    \n",
    "    response = client.models.generate_content(\n",
    "        model=MODEL_ID,\n",
    "        contents=prompt,\n",
    "        config={\n",
    "            \"temperature\": temperature,\n",
    "            \"response_mime_type\": \"application/json\",\n",
    "            \"response_json_schema\": schema_model.model_json_schema(),\n",
    "        },\n",
    "    )\n",
    "    \n",
    "    latency = time.time() - start_time\n",
    "    raw_text = response.text or \"\"\n",
    "    result = schema_model.model_validate_json(raw_text)\n",
    "    \n",
    "    if log:\n",
    "        PROMPT_LOG.append({\n",
    "            \"timestamp\": _now(),\n",
    "            \"label\": label,\n",
    "            \"schema\": schema_model.__name__,\n",
    "            \"prompt\": prompt[:300] + \"...\" if len(prompt) > 300 else prompt,\n",
    "            \"prompt_length\": len(prompt),\n",
    "            \"response\": raw_text[:300] + \"...\" if len(raw_text) > 300 else raw_text,\n",
    "            \"response_length\": len(raw_text),\n",
    "            \"latency_s\": round(latency, 3)\n",
    "        })\n",
    "    \n",
    "    return result\n",
    "\n",
    "def batch(items, batch_size=6):\n",
    "    \"\"\"Split items into batches.\"\"\"\n",
    "    for i in range(0, len(items), batch_size):\n",
    "        yield items[i:i+batch_size]\n",
    "\n",
    "def evaluate(predictions, golden, field):\n",
    "    \"\"\"Calculate accuracy for a field against golden set.\"\"\"\n",
    "    correct = 0\n",
    "    total = 0\n",
    "    errors = []\n",
    "    \n",
    "    for id, expected in golden.items():\n",
    "        if id in predictions:\n",
    "            total += 1\n",
    "            pred_value = getattr(predictions[id], field)\n",
    "            if pred_value == expected[field]:\n",
    "                correct += 1\n",
    "            else:\n",
    "                errors.append({\"id\": id, \"expected\": expected[field], \"got\": pred_value})\n",
    "    \n",
    "    return {\n",
    "        \"correct\": correct,\n",
    "        \"total\": total,\n",
    "        \"accuracy\": correct / total if total > 0 else 0,\n",
    "        \"errors\": errors\n",
    "    }\n",
    "\n",
    "def compare_versions(v1_preds, v2_preds, golden, fields):\n",
    "    \"\"\"Compare two versions across multiple fields.\"\"\"\n",
    "    results = []\n",
    "    for field in fields:\n",
    "        v1_eval = evaluate(v1_preds, golden, field)\n",
    "        v2_eval = evaluate(v2_preds, golden, field)\n",
    "        results.append({\n",
    "            \"field\": field,\n",
    "            \"v1_accuracy\": f\"{v1_eval['accuracy']:.1%}\",\n",
    "            \"v2_accuracy\": f\"{v2_eval['accuracy']:.1%}\",\n",
    "            \"improvement\": f\"{v2_eval['accuracy'] - v1_eval['accuracy']:+.1%}\"\n",
    "        })\n",
    "    return pd.DataFrame(results)\n",
    "\n",
    "print(\"✓ Infrastructure ready\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Step 1: Choose Your Track and Define Inputs\n",
    "\n",
    "Select your track and either use the provided sample data or create your own."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Mark your track\n",
    "CHOSEN_TRACK = \"\"  # Fill in: \"A\", \"B\", \"C\", or \"D\"\n",
    "\n",
    "print(f\"Selected track: {CHOSEN_TRACK}\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# Sample data for Track A: Customer Support\nTRACK_A_DATA = [\n    {\"id\": \"T1\", \"text\": \"App crashes when uploading photos. Tried reinstalling, still broken.\"},\n    {\"id\": \"T2\", \"text\": \"Can you add dark mode? Would really help with eye strain at night.\"},\n    {\"id\": \"T3\", \"text\": \"What's your refund policy for annual subscriptions?\"},\n    {\"id\": \"T4\", \"text\": \"Your service has been down for 3 hours! I'm losing business!\"},\n    {\"id\": \"T5\", \"text\": \"How do I export my data to CSV format?\"},\n    {\"id\": \"T6\", \"text\": \"I was promised a discount but charged full price.\"},\n    {\"id\": \"T7\", \"text\": \"The mobile app is much slower than the website.\"},\n    {\"id\": \"T8\", \"text\": \"Can I upgrade from Basic to Pro mid-subscription?\"},\n    {\"id\": \"T9\", \"text\": \"Dashboard shows wrong numbers. Revenue chart is definitely broken.\"},\n    {\"id\": \"T10\", \"text\": \"Love the product! Any plans for an API?\"},\n    {\"id\": \"T11\", \"text\": \"My account was charged twice this month!\"},\n    {\"id\": \"T12\", \"text\": \"Integration with Slack stopped working after your update.\"},\n]\n\n# Sample data for Track B: Meeting Notes\nTRACK_B_DATA = [\n    {\"id\": \"M1\", \"text\": \"Product sync - Jan 15. Attendees: Sarah (PM), Mike (Eng), Lisa (Design). Decided to prioritize mobile app over API. Lisa will deliver mockups by Jan 22. Next meeting: Jan 22, 2pm.\"},\n    {\"id\": \"M2\", \"text\": \"Quick sync with John and Amy about the customer complaint. Need to investigate database performance. John will look into it today. No decisions yet.\"},\n    {\"id\": \"M3\", \"text\": \"Q1 Planning. Team agreed on 3 OKRs. Budget approved for contractor hire. Sarah to draft job posting by Friday. Review in 2 weeks.\"},\n    {\"id\": \"M4\", \"text\": \"Design review with Lisa, Tom, and the CEO. Approved new logo concept. Tom to prepare brand guidelines. Launch planned for March.\"},\n    {\"id\": \"M5\", \"text\": \"1:1 with Alex. Discussed career growth and upcoming projects. Alex interested in leading the analytics initiative.\"},\n    {\"id\": \"M6\", \"text\": \"Incident postmortem. Root cause: misconfigured cache. Action items: add monitoring (DevOps, EOW), update runbook (SRE, next week), schedule training.\"},\n    {\"id\": \"M7\", \"text\": \"Sales pipeline review. 3 deals likely to close this quarter. Need demo environment fixed ASAP. Marketing to send case studies.\"},\n    {\"id\": \"M8\", \"text\": \"Brainstorm session for new features. No concrete decisions. Will reconvene after customer interviews.\"},\n    {\"id\": \"M9\", \"text\": \"Sprint retrospective with full dev team. Wins: shipped auth module on time. Struggles: too many context switches. Action: Mike to limit WIP to 2 items per person next sprint.\"},\n    {\"id\": \"M10\", \"text\": \"Client onboarding call with Acme Corp. Attendees: Rachel (CS), Dave (Solutions Eng), client reps. Agreed on 4-week rollout plan. Rachel to send onboarding checklist by Monday.\"},\n    {\"id\": \"M11\", \"text\": \"Weekly leadership sync — Feb 3. Headcount freeze confirmed through Q2. Decision: pause backfill for the open analyst role. CFO to share updated budget next week.\"},\n    {\"id\": \"M12\", \"text\": \"Security review meeting. Attendees: InfoSec team, two external auditors. Found 2 medium-severity vulnerabilities in payment flow. DevOps to patch by EOD Friday. Follow-up audit scheduled March 10.\"},\n]\n\n# Sample data for Track C: Product Review Analysis\nTRACK_C_DATA = [\n    {\"id\": \"R1\", \"text\": \"Absolutely love this laptop! The screen is gorgeous and performance is blazing fast. Only complaint: the fan gets a bit loud under heavy load. Would definitely buy again.\"},\n    {\"id\": \"R2\", \"text\": \"Terrible experience. Battery died after 3 months and customer support was unhelpful. Returning it.\"},\n    {\"id\": \"R3\", \"text\": \"It's fine for the price. Camera is decent, build quality is okay. Nothing special but gets the job done.\"},\n    {\"id\": \"R4\", \"text\": \"The noise cancellation on these headphones is incredible, best I've ever used. But the ear cups get uncomfortable after about 2 hours. Sound quality is top-notch though.\"},\n    {\"id\": \"R5\", \"text\": \"DO NOT BUY. Arrived damaged, replacement also had issues. The software is buggy and crashes constantly. Complete waste of money.\"},\n    {\"id\": \"R6\", \"text\": \"Great value for a budget tablet. Kids love it for games and videos. Screen could be sharper but at this price point I'm not complaining.\"},\n    {\"id\": \"R7\", \"text\": \"Upgraded from the previous model. The new camera system is a real step up and the battery easily lasts all day. Face unlock is lightning fast. Best phone I've owned.\"},\n    {\"id\": \"R8\", \"text\": \"Mixed feelings. The design is beautiful and it looks premium on my desk. But the keyboard is mushy and the trackpad is too small. Productivity suffers.\"},\n    {\"id\": \"R9\", \"text\": \"Bought this for my home office. Setup was painless. Print quality is sharp for documents but photo printing is mediocre. Decent for the price.\"},\n    {\"id\": \"R10\", \"text\": \"Six months in and I regret this purchase. The smart features barely work, the app is a mess, and it's slower than my old dumb TV. Wasted $800.\"},\n    {\"id\": \"R11\", \"text\": \"Perfect for running! Tracks GPS accurately, heart rate monitor is reliable, and it survived several rainstorms. The companion app needs work though.\"},\n    {\"id\": \"R12\", \"text\": \"Looks exactly like the pictures. Fabric quality is surprisingly good for the price. Runs a bit small — order one size up. Would recommend to friends.\"},\n]\n\n# Sample data for Track D: Job Posting Parser\nTRACK_D_DATA = [\n    {\"id\": \"J1\", \"text\": \"Senior Software Engineer - Remote. 5+ years experience in Python and cloud infrastructure. Competitive salary $150-180k. Must be comfortable with on-call rotations. Unlimited PTO.\"},\n    {\"id\": \"J2\", \"text\": \"Marketing Coordinator (Entry Level) - NYC Office. 0-2 years experience. Social media management, content creation. $45-55k. Great culture and mentorship program!\"},\n    {\"id\": \"J3\", \"text\": \"VP of Engineering. Lead a team of 50+ engineers. 15+ years experience required. Must relocate to SF. Equity package available. We work hard and play hard.\"},\n    {\"id\": \"J4\", \"text\": \"Data Analyst - Hybrid (3 days in office). SQL, Python, Tableau required. 2-4 years experience. $75-95k + bonus. Family-friendly workplace with flexible hours.\"},\n    {\"id\": \"J5\", \"text\": \"Full Stack Developer - URGENT HIRE. Must start immediately. Competitive pay DOE. Fast-paced startup environment. Wear many hats. Rockstar developers only.\"},\n    {\"id\": \"J6\", \"text\": \"Product Manager - Remote OK. 3-5 years PM experience, preferably B2B SaaS. Strong analytical skills. $120-150k. Transparent salary bands, 4-day work week trial.\"},\n    {\"id\": \"J7\", \"text\": \"Junior UX Designer - Onsite, Austin TX. Portfolio required. Figma proficiency a must. $55-70k. Collaborative team, regular design critiques, conference budget.\"},\n    {\"id\": \"J8\", \"text\": \"DevOps Lead - Hybrid. 8+ years experience. Kubernetes, Terraform, AWS. Manage team of 4. Must be available 24/7 for critical incidents. Salary not listed.\"},\n    {\"id\": \"J9\", \"text\": \"Customer Success Manager - Remote. 2+ years in SaaS CSM role. Manage portfolio of enterprise accounts. $80-100k + commission. Clear promotion path to Senior CSM.\"},\n    {\"id\": \"J10\", \"text\": \"Machine Learning Engineer - Onsite, Seattle. PhD preferred. PyTorch, distributed training. Groundbreaking AI research. Publish papers. $180-250k + RSUs.\"},\n    {\"id\": \"J11\", \"text\": \"Executive Assistant to CEO - Onsite NYC. 5+ years supporting C-level. 60+ hour weeks expected. Discretion essential. Salary commensurate with experience.\"},\n    {\"id\": \"J12\", \"text\": \"Mid-Level Backend Engineer - Remote-first. Go or Rust experience. 3-5 years. $110-140k, equity, full benefits. Async communication, no meetings Wednesdays.\"},\n]"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# Select your input data based on track\n# Modify this cell based on your chosen track\n\ninputs = []  # TODO: Set this to your track's data or your own data\n\n# Example:\n# inputs = TRACK_A_DATA\n# inputs = TRACK_B_DATA\n# inputs = TRACK_C_DATA\n# inputs = TRACK_D_DATA\n# OR define your own:\n# inputs = [{\"id\": \"X1\", \"text\": \"...\"}, ...]\n\nprint(f\"Input items: {len(inputs)}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Step 2: Define Your Schema\n",
    "\n",
    "Create a Pydantic schema that captures the key information you want to extract.\n",
    "\n",
    "**Requirements:**\n",
    "- At least one `Literal` field for classification\n",
    "- An urgency/priority field\n",
    "- A summary field\n",
    "- At least one action/recommendation field"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# TODO: Define your schema\n",
    "# This is an example for Track A - modify for your track!\n",
    "\n",
    "class ExtractedItem(BaseModel):\n",
    "    \"\"\"Extracted information from a single item.\"\"\"\n",
    "    id: str = Field(description=\"ID of the input item\")\n",
    "    \n",
    "    # TODO: Add your classification field with Literal type\n",
    "    # Example:\n",
    "    # category: Literal[\"Bug\", \"Request\", \"Policy\", \"Complaint\", \"Other\"]\n",
    "    \n",
    "    # TODO: Add urgency/priority field\n",
    "    # Example:\n",
    "    # urgency: Literal[\"low\", \"medium\", \"high\"]\n",
    "    \n",
    "    # TODO: Add summary field\n",
    "    # Example:\n",
    "    # summary: str = Field(description=\"One-sentence summary (max 20 words)\")\n",
    "    \n",
    "    # TODO: Add action/recommendation field\n",
    "    # Example:\n",
    "    # next_step: str = Field(description=\"Recommended action\")\n",
    "    \n",
    "    # Optional: Add any additional fields relevant to your track\n",
    "    # missing_info: List[str] = Field(description=\"Information needed to proceed\")\n",
    "    pass  # Remove this line when you add your fields\n",
    "\n",
    "\n",
    "class ExtractedBatch(BaseModel):\n",
    "    \"\"\"Batch of extracted items.\"\"\"\n",
    "    items: List[ExtractedItem]"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Step 3: Write Prompt v1 (Baseline)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# TODO: Write your v1 extraction prompt\n",
    "# Keep it simple - this is your baseline\n",
    "\n",
    "PROMPT_V1 = \"\"\"\n",
    "Extract information from each item into the specified schema.\n",
    "\n",
    "Rules:\n",
    "- Do NOT invent facts not present in the text\n",
    "- If information is missing, note it appropriately\n",
    "- Keep summaries concise\n",
    "\n",
    "Items:\n",
    "{items}\n",
    "\"\"\"\n",
    "\n",
    "def format_items(items):\n",
    "    \"\"\"Format items for the prompt.\"\"\"\n",
    "    return \"\\n\".join([f\"{it['id']}: {it['text']}\" for it in items])\n",
    "\n",
    "def run_extraction(prompt_template, items, label):\n",
    "    \"\"\"Run extraction with a prompt template.\"\"\"\n",
    "    prompt = prompt_template.format(items=format_items(items))\n",
    "    return generate_structured(prompt, ExtractedBatch, label=label)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Run v1 extraction\n",
    "# result_v1 = run_extraction(PROMPT_V1, inputs, \"extraction_v1\")\n",
    "# print(f\"Extracted: {len(result_v1.items)} items\")\n",
    "\n",
    "# View results\n",
    "# df_v1 = pd.DataFrame([item.model_dump() for item in result_v1.items])\n",
    "# df_v1"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Step 4: Create Your Golden Set (8+ items)\n",
    "\n",
    "Manually label at least 8 items with the correct values for your classification and urgency fields.\n",
    "\n",
    "**This is critical for evaluation!**"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# TODO: Create your golden set\n",
    "# Label at least 8 items with ground truth\n",
    "\n",
    "GOLDEN = {\n",
    "    # Format: \"id\": {\"field1\": \"value\", \"field2\": \"value\"}\n",
    "    # Example for Track A:\n",
    "    # \"T1\": {\"category\": \"Bug\", \"urgency\": \"high\"},\n",
    "    # \"T2\": {\"category\": \"Request\", \"urgency\": \"low\"},\n",
    "    # ... at least 8 items\n",
    "}\n",
    "\n",
    "print(f\"Golden set size: {len(GOLDEN)} items\")\n",
    "assert len(GOLDEN) >= 8, \"Need at least 8 labeled items!\""
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Step 5: Evaluate v1"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Create predictions dict for evaluation\n",
    "# pred_v1 = {item.id: item for item in result_v1.items}\n",
    "\n",
    "# Evaluate - update field names to match your schema!\n",
    "# cat_eval = evaluate(pred_v1, GOLDEN, \"category\")  # Change field name as needed\n",
    "# urg_eval = evaluate(pred_v1, GOLDEN, \"urgency\")   # Change field name as needed\n",
    "\n",
    "# print(f\"V1 Category Accuracy: {cat_eval['correct']}/{cat_eval['total']} = {cat_eval['accuracy']:.1%}\")\n",
    "# print(f\"V1 Urgency Accuracy:  {urg_eval['correct']}/{urg_eval['total']} = {urg_eval['accuracy']:.1%}\")\n",
    "\n",
    "# Show errors\n",
    "# if cat_eval['errors']:\n",
    "#     print(\"\\nCategory errors:\")\n",
    "#     for err in cat_eval['errors']:\n",
    "#         print(f\"  {err['id']}: expected '{err['expected']}', got '{err['got']}'\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Step 6: Improve Prompt (v2)\n",
    "\n",
    "Based on your v1 errors, improve your prompt. Consider:\n",
    "- Adding explicit decision rules for each category\n",
    "- Adding examples (few-shot)\n",
    "- Clarifying urgency definitions\n",
    "- Adding edge case handling"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# TODO: Write your improved v2 prompt\n",
    "\n",
    "PROMPT_V2 = PROMPT_V1 + \"\"\"\n",
    "\n",
    "# TODO: Add your improvements here\n",
    "# Examples:\n",
    "\n",
    "# CATEGORY RULES:\n",
    "# - Bug: ...\n",
    "# - Request: ...\n",
    "# etc.\n",
    "\n",
    "# URGENCY RULES:\n",
    "# - high: ...\n",
    "# - medium: ...\n",
    "# - low: ...\n",
    "\n",
    "# Or add few-shot examples:\n",
    "# EXAMPLES:\n",
    "# \"App crashes\" → Bug, high\n",
    "# \"Add dark mode\" → Request, low\n",
    "\"\"\""
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Run v2 extraction\n",
    "# result_v2 = run_extraction(PROMPT_V2, inputs, \"extraction_v2\")\n",
    "# pred_v2 = {item.id: item for item in result_v2.items}\n",
    "\n",
    "# Evaluate v2\n",
    "# cat_eval_v2 = evaluate(pred_v2, GOLDEN, \"category\")\n",
    "# urg_eval_v2 = evaluate(pred_v2, GOLDEN, \"urgency\")\n",
    "\n",
    "# print(f\"V2 Category Accuracy: {cat_eval_v2['correct']}/{cat_eval_v2['total']} = {cat_eval_v2['accuracy']:.1%}\")\n",
    "# print(f\"V2 Urgency Accuracy:  {urg_eval_v2['correct']}/{urg_eval_v2['total']} = {urg_eval_v2['accuracy']:.1%}\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Compare versions\n",
    "# comparison = compare_versions(pred_v1, pred_v2, GOLDEN, [\"category\", \"urgency\"])\n",
    "# print(\"\\n📊 Version Comparison:\")\n",
    "# print(comparison.to_string(index=False))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Step 7: Error Analysis\n",
    "\n",
    "Write your analysis in the markdown cell below."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Error Analysis\n",
    "\n",
    "#### Top 3 Error Patterns in v1\n",
    "\n",
    "1. **[Error type 1]**: [Description and example]\n",
    "\n",
    "2. **[Error type 2]**: [Description and example]\n",
    "\n",
    "3. **[Error type 3]**: [Description and example]\n",
    "\n",
    "#### Changes from v1 → v2\n",
    "\n",
    "1. **[Change 1]**: [What you changed and why]\n",
    "\n",
    "2. **[Change 2]**: [What you changed and why]\n",
    "\n",
    "#### What You Would Try for v3\n",
    "\n",
    "[Describe what improvements you would make with more time]\n",
    "\n",
    "#### Remaining Risks\n",
    "\n",
    "[What edge cases or risks remain? How would you mitigate them in production?]"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Step 8: Export Results"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Export prompt log\n",
    "if PROMPT_LOG:\n",
    "    df_log = pd.DataFrame(PROMPT_LOG)\n",
    "    df_log.to_csv(\"day2_independent_lab_log.csv\", index=False)\n",
    "    print(f\"✓ Saved {len(PROMPT_LOG)} prompts to day2_independent_lab_log.csv\")\n",
    "\n",
    "# Export extracted results\n",
    "# if result_v2:\n",
    "#     with open(\"day2_extracted_results.json\", \"w\") as f:\n",
    "#         json.dump([item.model_dump() for item in result_v2.items], f, indent=2)\n",
    "#     print(\"✓ Saved extracted results to day2_extracted_results.json\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Checklist Before Proceeding\n",
    "\n",
    "- [ ] Chose a track and defined input data (12+ items)\n",
    "- [ ] Created Pydantic schema with Literal types\n",
    "- [ ] Wrote baseline prompt (v1) and ran extraction\n",
    "- [ ] Created golden set (8+ labeled items)\n",
    "- [ ] Evaluated v1 accuracy\n",
    "- [ ] Improved prompt (v2) with rules/examples\n",
    "- [ ] Evaluated v2 and compared to v1\n",
    "- [ ] Wrote error analysis\n",
    "- [ ] Exported logs and results\n",
    "\n",
    "---\n",
    "\n",
    "**Proceed to: Day 2 Assignment →**"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python",
   "version": "3.10.0"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 4
}