{
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# Day 2: Guided Lab — Prompt Engineering with Structured Outputs\n",
    "\n",
    "## Learning Objectives\n",
    "\n",
    "By the end of this lab, you will be able to:\n",
    "\n",
    "1. **Structure prompts** using the 5-component framework\n",
    "2. **Apply few-shot learning** to improve output consistency\n",
    "3. **Use chain-of-thought** prompting for complex reasoning\n",
    "4. **Generate guaranteed-valid JSON** using Pydantic schemas\n",
    "5. **Compare prompt variations** systematically with evaluation metrics\n",
    "\n",
    "## Prerequisites\n",
    "\n",
    "- Completed Day 1 labs\n",
    "- Basic understanding of LLM parameters (temperature, tokens)\n",
    "- Google AI Studio API key"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 0: Setup and Infrastructure"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Install required packages\n",
    "!pip install -q -U google-genai"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# Import libraries\nimport os\nimport time\nimport json\nimport pandas as pd\nfrom datetime import datetime, timezone\nfrom typing import List, Optional, Literal\nfrom pydantic import BaseModel, Field\nfrom google import genai\nfrom google.genai import types\n\n# Configure API — reads the GEMINI_API_KEY you set up on Day 1\ntry:\n    from google.colab import userdata\n    API_KEY = userdata.get(\"GEMINI_API_KEY\")\nexcept:\n    API_KEY = None\n\nif not API_KEY:\n    import getpass\n    API_KEY = getpass.getpass(\"Enter your Gemini API key: \")\n\n# Initialize client\nclient = genai.Client(api_key=API_KEY)\nMODEL_ID = \"gemini-2.5-flash-lite\"\n\nprint(f\"✓ API key loaded\")\nprint(f\"✓ Using model: {MODEL_ID}\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Prompt logging infrastructure\n",
    "PROMPT_LOG = []\n",
    "\n",
    "def _now():\n",
    "    \"\"\"Get current timestamp (timezone-aware).\"\"\"\n",
    "    return datetime.now(timezone.utc).isoformat().replace('+00:00', 'Z')\n",
    "\n",
    "def generate(prompt, temperature=0.7, max_tokens=1000, log=True, label=None):\n",
    "    \"\"\"\n",
    "    Generate text using Gemini (free-form output).\n",
    "    \n",
    "    Args:\n",
    "        prompt: The prompt text\n",
    "        temperature: Creativity (0=deterministic, 1=creative)\n",
    "        max_tokens: Maximum response length\n",
    "        log: Whether to log this call\n",
    "        label: Optional label for this experiment\n",
    "    \n",
    "    Returns:\n",
    "        The generated text\n",
    "    \"\"\"\n",
    "    start_time = time.time()\n",
    "    \n",
    "    response = client.models.generate_content(\n",
    "        model=MODEL_ID,\n",
    "        contents=prompt,\n",
    "        config=types.GenerateContentConfig(\n",
    "            temperature=temperature,\n",
    "            max_output_tokens=max_tokens,\n",
    "        ),\n",
    "    )\n",
    "    \n",
    "    result = response.text\n",
    "    latency = time.time() - start_time\n",
    "    \n",
    "    if log:\n",
    "        PROMPT_LOG.append({\n",
    "            \"timestamp\": _now(),\n",
    "            \"label\": label,\n",
    "            \"type\": \"free_form\",\n",
    "            \"prompt\": prompt[:500] + \"...\" if len(prompt) > 500 else prompt,\n",
    "            \"prompt_length\": len(prompt),\n",
    "            \"temperature\": temperature,\n",
    "            \"response\": result[:500] + \"...\" if len(result) > 500 else result,\n",
    "            \"response_length\": len(result),\n",
    "            \"latency_s\": round(latency, 3)\n",
    "        })\n",
    "    \n",
    "    return result\n",
    "\n",
    "def generate_structured(prompt, schema_model, temperature=0.2, log=True, label=None):\n",
    "    \"\"\"\n",
    "    Generate structured output using Pydantic schema.\n",
    "    \n",
    "    The API GUARANTEES the output matches your schema - no parsing errors!\n",
    "    \n",
    "    Args:\n",
    "        prompt: The prompt text\n",
    "        schema_model: Pydantic BaseModel class defining the output structure\n",
    "        temperature: Creativity (lower is better for structured output)\n",
    "        log: Whether to log this call\n",
    "        label: Optional label for this experiment\n",
    "    \n",
    "    Returns:\n",
    "        Parsed Pydantic model instance\n",
    "    \"\"\"\n",
    "    start_time = time.time()\n",
    "    \n",
    "    response = client.models.generate_content(\n",
    "        model=MODEL_ID,\n",
    "        contents=prompt,\n",
    "        config={\n",
    "            \"temperature\": temperature,\n",
    "            \"response_mime_type\": \"application/json\",\n",
    "            \"response_json_schema\": schema_model.model_json_schema(),\n",
    "        },\n",
    "    )\n",
    "    \n",
    "    latency = time.time() - start_time\n",
    "    raw_text = response.text or \"\"\n",
    "    \n",
    "    # Parse into Pydantic model (guaranteed to work with structured outputs)\n",
    "    result = schema_model.model_validate_json(raw_text)\n",
    "    \n",
    "    if log:\n",
    "        PROMPT_LOG.append({\n",
    "            \"timestamp\": _now(),\n",
    "            \"label\": label,\n",
    "            \"type\": \"structured\",\n",
    "            \"schema\": schema_model.__name__,\n",
    "            \"prompt\": prompt[:500] + \"...\" if len(prompt) > 500 else prompt,\n",
    "            \"prompt_length\": len(prompt),\n",
    "            \"temperature\": temperature,\n",
    "            \"response\": raw_text[:500] + \"...\" if len(raw_text) > 500 else raw_text,\n",
    "            \"response_length\": len(raw_text),\n",
    "            \"latency_s\": round(latency, 3)\n",
    "        })\n",
    "    \n",
    "    return result\n",
    "\n",
    "def show_log():\n",
    "    \"\"\"Display the prompt log as a DataFrame.\"\"\"\n",
    "    if not PROMPT_LOG:\n",
    "        print(\"No prompts logged yet.\")\n",
    "        return None\n",
    "    return pd.DataFrame(PROMPT_LOG)\n",
    "\n",
    "print(\"✓ Logging infrastructure ready\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 1: The Anatomy of a Prompt\n",
    "\n",
    "A well-structured prompt can have up to **5 components**:\n",
    "\n",
    "1. **Context** — Background information, role assignment\n",
    "2. **Instructions** — What task to perform\n",
    "3. **Input** — The specific data to process\n",
    "4. **Examples** — Demonstrations of desired output\n",
    "5. **Constraints** — Boundaries, format requirements\n",
    "\n",
    "Let's see how adding components improves output quality."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Exercise 1.1: Basic vs. Structured Prompts\n",
    "\n",
    "Let's compare a basic prompt with a structured one for the same task."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# The task: Summarize a customer review\n",
    "review = \"\"\"\n",
    "I bought this wireless keyboard last month and have mixed feelings. The typing \n",
    "experience is excellent - the keys are responsive and quiet, which is perfect \n",
    "for my home office. Battery life has been impressive too, still on the original \n",
    "batteries after 4 weeks of daily use. However, the Bluetooth connection drops \n",
    "occasionally, maybe once or twice a day, which is frustrating during video calls \n",
    "when I'm trying to type in chat. The build quality feels a bit cheap for the $80 \n",
    "price point. Overall, it's a decent keyboard but not without its quirks.\n",
    "\"\"\"\n",
    "\n",
    "# Version 1: Basic prompt (Instructions + Input only)\n",
    "basic_prompt = f\"Summarize this review:\\n\\n{review}\"\n",
    "\n",
    "# Version 2: Structured prompt (all 5 components)\n",
    "structured_prompt = f\"\"\"You are a product analyst at an e-commerce company.\n",
    "\n",
    "Task: Analyze the following customer review and provide a structured summary.\n",
    "\n",
    "Review:\n",
    "---\n",
    "{review}\n",
    "---\n",
    "\n",
    "Provide your analysis in this format:\n",
    "- Sentiment: [positive/negative/mixed]\n",
    "- Pros: [bullet points]\n",
    "- Cons: [bullet points]\n",
    "- Key insight: [one sentence]\n",
    "\n",
    "Keep the summary under 100 words.\"\"\"\n",
    "\n",
    "# Compare them\n",
    "print(\"=\" * 60)\n",
    "print(\"BASIC PROMPT RESPONSE:\")\n",
    "print(\"=\" * 60)\n",
    "print(generate(basic_prompt, temperature=0.3, label=\"basic\"))\n",
    "\n",
    "print(\"\\n\" + \"=\" * 60)\n",
    "print(\"STRUCTURED PROMPT RESPONSE:\")\n",
    "print(\"=\" * 60)\n",
    "print(generate(structured_prompt, temperature=0.3, label=\"structured\"))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### 💡 Discussion\n",
    "\n",
    "Notice how the structured prompt produces:\n",
    "- Consistent format\n",
    "- Specific categories (pros/cons)\n",
    "- Actionable insights\n",
    "- Appropriate length\n",
    "\n",
    "The basic prompt works, but results vary more between runs."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Exercise 1.2: The Power of Role Assignment\n",
    "\n",
    "Let's see how different roles change the output for the same question."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "question = \"What should I consider when choosing a cloud provider for my startup?\"\n",
    "\n",
    "roles = {\n",
    "    \"no_role\": f\"{question}\",\n",
    "    \n",
    "    \"cto\": f\"\"\"You are an experienced CTO who has built multiple startups from scratch.\n",
    "{question}\"\"\",\n",
    "    \n",
    "    \"cfo\": f\"\"\"You are a CFO focused on cost optimization and financial planning.\n",
    "{question}\"\"\",\n",
    "    \n",
    "    \"security\": f\"\"\"You are a cybersecurity expert specializing in cloud security.\n",
    "{question}\"\"\"\n",
    "}\n",
    "\n",
    "for role, prompt in roles.items():\n",
    "    print(f\"\\n{'='*60}\")\n",
    "    print(f\"ROLE: {role.upper()}\")\n",
    "    print(\"=\"*60)\n",
    "    response = generate(prompt, temperature=0.5, max_tokens=300, label=f\"role_{role}\")\n",
    "    print(response[:600])"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 2: Few-Shot Prompting\n",
    "\n",
    "**Few-shot prompting** provides examples of desired input-output pairs. This is often more effective than detailed instructions.\n",
    "\n",
    "> \"Show, don't just tell\""
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Exercise 2.1: Zero-Shot vs. Few-Shot Classification"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Task: Classify customer support tickets\n",
    "ticket = \"My order #12345 arrived but one item was missing from the package.\"\n",
    "\n",
    "# Zero-shot: Just instructions\n",
    "zero_shot = f\"\"\"Classify this customer support ticket into one of these categories:\n",
    "- billing\n",
    "- shipping\n",
    "- product_issue\n",
    "- account\n",
    "- other\n",
    "\n",
    "Ticket: {ticket}\n",
    "\n",
    "Category:\"\"\"\n",
    "\n",
    "# Few-shot: Instructions + examples\n",
    "few_shot = f\"\"\"Classify customer support tickets into categories.\n",
    "\n",
    "Ticket: \"I was charged twice for my subscription this month.\"\n",
    "Category: billing\n",
    "\n",
    "Ticket: \"The laptop screen has dead pixels.\"\n",
    "Category: product_issue\n",
    "\n",
    "Ticket: \"I can't reset my password, the email never arrives.\"\n",
    "Category: account\n",
    "\n",
    "Ticket: \"When will my package arrive? It's been 2 weeks.\"\n",
    "Category: shipping\n",
    "\n",
    "Ticket: \"{ticket}\"\n",
    "Category:\"\"\"\n",
    "\n",
    "print(\"Zero-shot result:\", generate(zero_shot, temperature=0, label=\"zero_shot_classify\"))\n",
    "print(\"Few-shot result:\", generate(few_shot, temperature=0, label=\"few_shot_classify\"))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 3: Chain-of-Thought Prompting\n",
    "\n",
    "For complex reasoning tasks, asking the model to \"think step by step\" dramatically improves accuracy.\n",
    "\n",
    "**Why it works:** More output tokens = more computation = better reasoning"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Exercise 3.1: Direct vs. Chain-of-Thought"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Math problem\n",
    "problem = \"\"\"\n",
    "A store sells notebooks for $3 each. If you buy 10 or more, you get 15% off.\n",
    "Sales tax is 8%. How much does it cost to buy 12 notebooks?\n",
    "\"\"\"\n",
    "\n",
    "# Direct approach\n",
    "direct_prompt = f\"{problem}\\nAnswer:\"\n",
    "\n",
    "# Chain-of-thought approach\n",
    "cot_prompt = f\"{problem}\\nLet's solve this step by step:\"\n",
    "\n",
    "print(\"DIRECT ANSWER:\")\n",
    "print(generate(direct_prompt, temperature=0, label=\"math_direct\"))\n",
    "\n",
    "print(\"\\n\" + \"=\"*60)\n",
    "print(\"\\nCHAIN-OF-THOUGHT:\")\n",
    "print(generate(cot_prompt, temperature=0, label=\"math_cot\"))"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 4: Structured Outputs with Pydantic 🔥\n",
    "\n",
    "This is the **key technique** for production applications!\n",
    "\n",
    "Instead of hoping the model returns valid JSON, we use **Pydantic schemas** with the API's `response_json_schema` parameter. The API **guarantees** the output matches our schema."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Exercise 4.1: Define a Schema"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Define the output structure using Pydantic\n",
    "class PersonInfo(BaseModel):\n",
    "    \"\"\"Extracted information about a person.\"\"\"\n",
    "    name: str = Field(description=\"Full name of the person\")\n",
    "    age: Optional[int] = Field(description=\"Age in years, or null if not mentioned\")\n",
    "    occupation: Optional[str] = Field(description=\"Job title or profession\")\n",
    "    location: Optional[str] = Field(description=\"City or location mentioned\")\n",
    "    employer: Optional[str] = Field(description=\"Company or employer name\")\n",
    "\n",
    "# Test extraction\n",
    "text = \"John Smith is a 35-year-old software engineer living in San Francisco. He works at Google.\"\n",
    "\n",
    "prompt = f\"\"\"Extract information about the person from this text.\n",
    "If information is not mentioned, use null.\n",
    "\n",
    "Text: {text}\"\"\"\n",
    "\n",
    "result = generate_structured(prompt, PersonInfo, label=\"person_extraction\")\n",
    "\n",
    "# Result is already a Pydantic object - no parsing needed!\n",
    "print(f\"Name: {result.name}\")\n",
    "print(f\"Age: {result.age}\")\n",
    "print(f\"Occupation: {result.occupation}\")\n",
    "print(f\"Location: {result.location}\")\n",
    "print(f\"Employer: {result.employer}\")\n",
    "print(f\"\\nFull object: {result.model_dump_json(indent=2)}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Exercise 4.2: Classification with Constrained Categories\n",
    "\n",
    "Use `Literal` types to constrain outputs to specific values."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Define schema with constrained categories\n",
    "class TicketClassification(BaseModel):\n",
    "    \"\"\"Classification result for a support ticket.\"\"\"\n",
    "    ticket_id: str = Field(description=\"ID of the ticket\")\n",
    "    category: Literal[\"billing\", \"shipping\", \"product_issue\", \"account\", \"other\"] = Field(\n",
    "        description=\"Primary category\"\n",
    "    )\n",
    "    urgency: Literal[\"low\", \"medium\", \"high\"] = Field(\n",
    "        description=\"Urgency level\"\n",
    "    )\n",
    "    summary: str = Field(description=\"One-sentence summary of the issue\")\n",
    "    suggested_action: str = Field(description=\"Recommended next step\")\n",
    "\n",
    "# Test tickets\n",
    "tickets = [\n",
    "    {\"id\": \"T001\", \"text\": \"I was charged twice for my order last week!\"},\n",
    "    {\"id\": \"T002\", \"text\": \"Package says delivered but I never received it.\"},\n",
    "    {\"id\": \"T003\", \"text\": \"How do I update my email address on my account?\"},\n",
    "]\n",
    "\n",
    "for ticket in tickets:\n",
    "    prompt = f\"\"\"Classify this support ticket.\n",
    "\n",
    "Urgency guidelines:\n",
    "- high: Financial issues, security concerns, service outages\n",
    "- medium: Delayed orders, missing items, access issues\n",
    "- low: General questions, feedback, minor issues\n",
    "\n",
    "Ticket ID: {ticket['id']}\n",
    "Ticket: {ticket['text']}\"\"\"\n",
    "    \n",
    "    result = generate_structured(prompt, TicketClassification, label=f\"classify_{ticket['id']}\")\n",
    "    print(f\"\\n{ticket['id']}: {result.category} ({result.urgency})\")\n",
    "    print(f\"  → {result.suggested_action}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Exercise 4.3: Batch Extraction\n",
    "\n",
    "Process multiple items in a single API call using a batch schema."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Schema for a single review analysis\n",
    "class ReviewAnalysis(BaseModel):\n",
    "    \"\"\"Analysis of a single product review.\"\"\"\n",
    "    review_id: str\n",
    "    sentiment: Literal[\"positive\", \"negative\", \"neutral\", \"mixed\"]\n",
    "    confidence: float = Field(description=\"Confidence score 0-1\")\n",
    "    key_points: List[str] = Field(description=\"Key points from the review (max 3)\")\n",
    "\n",
    "# Batch schema wraps multiple analyses\n",
    "class ReviewBatch(BaseModel):\n",
    "    \"\"\"Batch of review analyses.\"\"\"\n",
    "    reviews: List[ReviewAnalysis]\n",
    "\n",
    "# Sample reviews\n",
    "reviews = [\n",
    "    {\"id\": \"R1\", \"text\": \"Amazing product! Works perfectly and arrived fast.\"},\n",
    "    {\"id\": \"R2\", \"text\": \"Terrible quality. Broke after one week. Waste of money.\"},\n",
    "    {\"id\": \"R3\", \"text\": \"It's okay. Does what it says, nothing special.\"},\n",
    "    {\"id\": \"R4\", \"text\": \"Love the features but the battery life is disappointing.\"},\n",
    "]\n",
    "\n",
    "# Format reviews for prompt\n",
    "reviews_text = \"\\n\".join([f\"{r['id']}: {r['text']}\" for r in reviews])\n",
    "\n",
    "prompt = f\"\"\"Analyze these product reviews.\n",
    "\n",
    "Reviews:\n",
    "{reviews_text}\n",
    "\n",
    "For each review, determine sentiment, confidence, and extract up to 3 key points.\"\"\"\n",
    "\n",
    "batch_result = generate_structured(prompt, ReviewBatch, label=\"batch_reviews\")\n",
    "\n",
    "# Convert to DataFrame for nice display\n",
    "df = pd.DataFrame([r.model_dump() for r in batch_result.reviews])\n",
    "print(df)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Exercise 4.4: Sanity Checks\n",
    "\n",
    "Even with structured outputs, validate business rules!"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "def validate_reviews(reviews: List[ReviewAnalysis]) -> List[tuple]:\n",
    "    \"\"\"Run sanity checks on review analyses.\"\"\"\n",
    "    problems = []\n",
    "    \n",
    "    for r in reviews:\n",
    "        # Check confidence is in valid range\n",
    "        if not 0 <= r.confidence <= 1:\n",
    "            problems.append((r.review_id, \"invalid_confidence\", r.confidence))\n",
    "        \n",
    "        # Check key_points isn't too long\n",
    "        if len(r.key_points) > 3:\n",
    "            problems.append((r.review_id, \"too_many_key_points\", len(r.key_points)))\n",
    "        \n",
    "        # Check for empty key points on non-neutral reviews\n",
    "        if r.sentiment != \"neutral\" and len(r.key_points) == 0:\n",
    "            problems.append((r.review_id, \"missing_key_points\", r.sentiment))\n",
    "    \n",
    "    return problems\n",
    "\n",
    "# Validate our batch results\n",
    "problems = validate_reviews(batch_result.reviews)\n",
    "print(f\"Validation problems found: {len(problems)}\")\n",
    "for p in problems:\n",
    "    print(f\"  {p}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Part 5: Prompt Iteration with Evaluation\n",
    "\n",
    "Let's build a complete workflow: extract data, evaluate against golden labels, and iterate."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Define schema for support ticket triage\n",
    "class TriageResult(BaseModel):\n",
    "    \"\"\"Triage result for a support ticket.\"\"\"\n",
    "    id: str = Field(description=\"Ticket ID\")\n",
    "    category: Literal[\"Bug\", \"Policy\", \"Request\", \"Complaint\", \"Other\"] = Field(\n",
    "        description=\"Primary category\"\n",
    "    )\n",
    "    urgency: Literal[\"low\", \"medium\", \"high\"] = Field(\n",
    "        description=\"Urgency level\"\n",
    "    )\n",
    "    summary: str = Field(description=\"One-sentence summary (max 20 words)\")\n",
    "    next_step: str = Field(description=\"Recommended action\")\n",
    "\n",
    "class TriageBatch(BaseModel):\n",
    "    \"\"\"Batch of triage results.\"\"\"\n",
    "    tickets: List[TriageResult]"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Test data\n",
    "support_tickets = [\n",
    "    {\"id\": \"T1\", \"text\": \"App crashes when I try to upload photos. Tried reinstalling, still broken.\"},\n",
    "    {\"id\": \"T2\", \"text\": \"Can you add dark mode? Would really help with eye strain.\"},\n",
    "    {\"id\": \"T3\", \"text\": \"What's your refund policy for annual subscriptions?\"},\n",
    "    {\"id\": \"T4\", \"text\": \"Your service has been down for 3 hours! I'm losing business!\"},\n",
    "    {\"id\": \"T5\", \"text\": \"How do I export my data to CSV?\"},\n",
    "    {\"id\": \"T6\", \"text\": \"I was promised a discount but was charged full price.\"},\n",
    "]\n",
    "\n",
    "# Golden set (human-labeled ground truth)\n",
    "GOLDEN = {\n",
    "    \"T1\": {\"category\": \"Bug\", \"urgency\": \"high\"},\n",
    "    \"T2\": {\"category\": \"Request\", \"urgency\": \"low\"},\n",
    "    \"T3\": {\"category\": \"Policy\", \"urgency\": \"low\"},\n",
    "    \"T4\": {\"category\": \"Bug\", \"urgency\": \"high\"},\n",
    "    \"T5\": {\"category\": \"Policy\", \"urgency\": \"low\"},\n",
    "    \"T6\": {\"category\": \"Complaint\", \"urgency\": \"medium\"},\n",
    "}"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Version 1: Basic prompt\n",
    "PROMPT_V1 = \"\"\"You are a customer support assistant.\n",
    "\n",
    "Triage these support tickets into categories and urgency levels.\n",
    "\n",
    "Tickets:\n",
    "{tickets}\n",
    "\"\"\"\n",
    "\n",
    "def format_tickets(tickets):\n",
    "    return \"\\n\".join([f\"{t['id']}: {t['text']}\" for t in tickets])\n",
    "\n",
    "def run_triage(prompt_template, tickets, label):\n",
    "    prompt = prompt_template.format(tickets=format_tickets(tickets))\n",
    "    return generate_structured(prompt, TriageBatch, label=label)\n",
    "\n",
    "# Run v1\n",
    "result_v1 = run_triage(PROMPT_V1, support_tickets, \"triage_v1\")\n",
    "pred_v1 = {t.id: t for t in result_v1.tickets}\n",
    "\n",
    "# Show results\n",
    "df_v1 = pd.DataFrame([t.model_dump() for t in result_v1.tickets])\n",
    "print(\"V1 Results:\")\n",
    "print(df_v1[[\"id\", \"category\", \"urgency\", \"summary\"]])"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Evaluation function\n",
    "def evaluate(predictions, golden, field):\n",
    "    \"\"\"Calculate accuracy for a field.\"\"\"\n",
    "    correct = 0\n",
    "    total = 0\n",
    "    errors = []\n",
    "    \n",
    "    for id, expected in golden.items():\n",
    "        if id in predictions:\n",
    "            total += 1\n",
    "            pred_value = getattr(predictions[id], field)\n",
    "            if pred_value == expected[field]:\n",
    "                correct += 1\n",
    "            else:\n",
    "                errors.append((id, expected[field], pred_value))\n",
    "    \n",
    "    accuracy = correct / total if total > 0 else 0\n",
    "    return {\"correct\": correct, \"total\": total, \"accuracy\": accuracy, \"errors\": errors}\n",
    "\n",
    "# Evaluate v1\n",
    "cat_eval = evaluate(pred_v1, GOLDEN, \"category\")\n",
    "urg_eval = evaluate(pred_v1, GOLDEN, \"urgency\")\n",
    "\n",
    "print(f\"V1 Category Accuracy: {cat_eval['correct']}/{cat_eval['total']} = {cat_eval['accuracy']:.1%}\")\n",
    "print(f\"V1 Urgency Accuracy:  {urg_eval['correct']}/{urg_eval['total']} = {urg_eval['accuracy']:.1%}\")\n",
    "\n",
    "if cat_eval['errors']:\n",
    "    print(f\"\\nCategory errors:\")\n",
    "    for id, expected, got in cat_eval['errors']:\n",
    "        print(f\"  {id}: expected {expected}, got {got}\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Version 2: Improved prompt with explicit rules\n",
    "PROMPT_V2 = \"\"\"You are a customer support assistant.\n",
    "\n",
    "Triage these support tickets into categories and urgency levels.\n",
    "\n",
    "CATEGORY RULES:\n",
    "- Bug: System errors, crashes, broken features\n",
    "- Request: Feature requests, enhancement suggestions\n",
    "- Policy: Questions about pricing, refunds, how things work\n",
    "- Complaint: Billing disputes, service issues, frustration\n",
    "- Other: Anything else\n",
    "\n",
    "URGENCY RULES:\n",
    "- high: Service outages, data loss, financial impact, angry customers\n",
    "- medium: Broken features affecting workflow, billing issues\n",
    "- low: Questions, suggestions, minor issues\n",
    "\n",
    "Tickets:\n",
    "{tickets}\n",
    "\"\"\"\n",
    "\n",
    "# Run v2\n",
    "result_v2 = run_triage(PROMPT_V2, support_tickets, \"triage_v2\")\n",
    "pred_v2 = {t.id: t for t in result_v2.tickets}\n",
    "\n",
    "# Evaluate v2\n",
    "cat_eval_v2 = evaluate(pred_v2, GOLDEN, \"category\")\n",
    "urg_eval_v2 = evaluate(pred_v2, GOLDEN, \"urgency\")\n",
    "\n",
    "print(f\"V2 Category Accuracy: {cat_eval_v2['correct']}/{cat_eval_v2['total']} = {cat_eval_v2['accuracy']:.1%}\")\n",
    "print(f\"V2 Urgency Accuracy:  {urg_eval_v2['correct']}/{urg_eval_v2['total']} = {urg_eval_v2['accuracy']:.1%}\")\n",
    "\n",
    "# Compare\n",
    "print(f\"\\n📈 Category improvement: {cat_eval['accuracy']:.1%} → {cat_eval_v2['accuracy']:.1%}\")\n",
    "print(f\"📈 Urgency improvement:  {urg_eval['accuracy']:.1%} → {urg_eval_v2['accuracy']:.1%}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Summary and Key Takeaways\n",
    "\n",
    "### What We Learned\n",
    "\n",
    "| Technique | When to Use | Key Benefit |\n",
    "|-----------|-------------|-------------|\n",
    "| **Structured prompts** | Always | Consistent, predictable outputs |\n",
    "| **Role assignment** | Domain expertise needed | Appropriate tone and focus |\n",
    "| **Few-shot examples** | Format control, classification | Shows rather than tells |\n",
    "| **Chain-of-thought** | Complex reasoning | More accurate answers |\n",
    "| **Pydantic schemas** | Production applications | Guaranteed valid JSON |\n",
    "| **Golden set evaluation** | Iteration | Measurable improvement |\n",
    "\n",
    "### The Prompt Engineering Checklist\n",
    "\n",
    "- [ ] Is there a clear **role/context**?\n",
    "- [ ] Are **instructions** specific and unambiguous?\n",
    "- [ ] Would **examples** help clarify the expected output?\n",
    "- [ ] For complex reasoning, did I ask for **step-by-step** thinking?\n",
    "- [ ] Is the **output schema** defined with Pydantic?\n",
    "- [ ] Are **constraints** (categories, lengths) enforced in the schema?\n",
    "- [ ] Do I have a **golden set** for evaluation?"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# Export your experiment log\n",
    "if PROMPT_LOG:\n",
    "    df = pd.DataFrame(PROMPT_LOG)\n",
    "    df.to_csv(\"day2_guided_lab_log.csv\", index=False)\n",
    "    print(f\"✓ Saved {len(PROMPT_LOG)} experiments to day2_guided_lab_log.csv\")\n",
    "    print(df[[\"label\", \"type\", \"latency_s\", \"response_length\"]].to_string())"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "\n",
    "## Next Steps\n",
    "\n",
    "In the **Independent Lab**, you will:\n",
    "- Build your own extraction schema for a business domain\n",
    "- Create a golden set with 8+ labeled examples\n",
    "- Iterate on prompts and measure improvement\n",
    "- Write error analysis\n",
    "\n",
    "**Proceed to: Day 2 Independent Lab →**"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python",
   "version": "3.10.0"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 4
}