{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "PBtLgZ52EuqI"
      },
      "source": [
        "\n",
        "\n",
        "# Day 2: Assignment — Production-Ready Extraction Pipeline\n",
        "\n",
        "## Overview\n",
        "\n",
        "Build and document a **production-ready extraction pipeline** using Structured Outputs.\n",
        "\n",
        "## Deliverables\n",
        "\n",
        "Submit ONE notebook (or PDF export) containing:\n",
        "\n",
        "1. **Final Pydantic schema** with field descriptions\n",
        "2. **Final extraction prompt** (v2 or v3)\n",
        "3. **Extracted outputs** for at least 12 items\n",
        "4. **Golden set** (8+ items) with ground truth labels\n",
        "5. **Metrics** (accuracy for classification and urgency fields)\n",
        "6. **Error analysis** (½–1 page)\n",
        "7. **Prompt playbook** documentation\n",
        "\n",
        "## Grading Criteria\n",
        "\n",
        "| Criterion | Weight | Description |\n",
        "|-----------|--------|-------------|\n",
        "| Schema Design | 15% | Appropriate fields, types, constraints |\n",
        "| Prompt Quality | 25% | Effective rules, examples, structure |\n",
        "| Accuracy | 20% | Performance on golden set |\n",
        "| Analysis | 25% | Error patterns, iteration insights |\n",
        "| Documentation | 15% | Complete playbook, clear explanations |"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "SW0g_pWXEuqL"
      },
      "source": [
        "---\n",
        "\n",
        "## Setup"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 1,
      "metadata": {
        "id": "oksZTvSaEuqL"
      },
      "outputs": [],
      "source": [
        "!pip install -q -U google-genai # Install the Google GenAI SDK for Python, which allows us to interact with Google's Generative AI services. The -q flag suppresses output, and -U ensures we get the latest version."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 2,
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "D1PGdk7uEuqM",
        "outputId": "a4c3e149-b26a-4d1e-c34b-dc42cae4e884"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "✓ API key loaded\n",
            "✓ Model: gemini-2.5-flash-lite\n"
          ]
        }
      ],
      "source": [
        "import os # for environment variables\n",
        "import time # for timing code execution\n",
        "import json # for working with JSON data\n",
        "import pandas as pd # for working with tabular data\n",
        "from datetime import datetime, timezone # for working with dates and times\n",
        "from typing import List, Optional, Literal # for type annotations like lists, optional values, and literal types e.g. to specify that a value must be one of a few options\n",
        "from pydantic import BaseModel, Field # for defining data models with validation and serialization capabilities\n",
        "from google import genai # Import the Google GenAI SDK, which provides tools to interact with Google's Generative AI services, such as language models and image generation.\n",
        "\n",
        "# Configure API — reads the GEMINI_API_KEY you set up on Day 1\n",
        "try:\n",
        "    from google.colab import userdata # In Google Colab, userdata is a way to store and retrieve user-specific data, such as API keys, across different sessions. Here, we attempt to retrieve the GEMINI_API_KEY from userdata.\n",
        "    API_KEY = userdata.get(\"GEMINI_API_KEY\") #  Try to get the API key from Colab's userdata storage. If it exists, it will be stored in the variable API_KEY.\n",
        "except:\n",
        "    API_KEY = None # If we are not in a Colab environment or if the key is not found, we set API_KEY to None. This allows us to handle the case where the API key is not available and prompt the user for it later.\n",
        "\n",
        "if not API_KEY: # If API_KEY is still None (i.e., we couldn't retrieve it from userdata), we prompt the user to enter their Gemini API key manually. The getpass function is used to securely input the API key without displaying it on the screen.\n",
        "    import getpass #    getpass is a Python module that provides a secure way to handle password prompts. It allows you to input sensitive information, such as API keys or passwords, without displaying the input on the screen. In this case, we use getpass.getpass() to prompt the user for their Gemini API key securely.\n",
        "    API_KEY = getpass.getpass(\"Enter your Gemini API key: \") #  Prompt the user to enter their Gemini API key securely. The input will not be displayed on the screen, and the entered value will be stored in the variable API_KEY.\n",
        "\n",
        "client = genai.Client(api_key=API_KEY) # Create an instance of the GenAI client using the provided API key. This client will be used to interact with Google's Generative AI services, such as sending requests to language models or image generation endpoints.\n",
        "MODEL_ID = \"gemini-2.5-flash-lite\" # Set the model ID to \"gemini-2.5-flash-lite\". This variable will be used later when we make requests to the GenAI client, specifying which model we want to use for generating responses.\n",
        "\n",
        "print(f\"✓ API key loaded\") #   Print a confirmation message indicating that the API key has been successfully loaded. This is a simple way to provide feedback to the user that the setup process is proceeding correctly.\n",
        "print(f\"✓ Model: {MODEL_ID}\") # Print the model ID that we have set. This serves as a confirmation of which model we will be using for our Generative AI tasks, and it helps ensure that we are aware of the specific model configuration before we start making requests to the GenAI client."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 3,
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "-SyHZ2XyEuqN",
        "outputId": "3f487f7d-5c65-4f1b-e510-03cdc893f80a"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "✓ Infrastructure ready\n"
          ]
        }
      ],
      "source": [
        "# Prompt logging infrastructure - we will use this to track our prompts, responses, and performance metrics for evaluation and debugging purposes.\n",
        "PROMPT_LOG = [] # Initialize an empty list called PROMPT_LOG. This list will be used to store logs of the prompts we send to the GenAI model, the responses we receive, and various performance metrics such as latency and response length. This logging infrastructure will help us evaluate the effectiveness of our prompts and identify any issues or areas for improvement in our interactions with the model.\n",
        "\n",
        "def _now(): # Define a helper function called _now() that returns the current date and time in ISO 8601 format with UTC timezone. This function will be used to timestamp our prompt logs, allowing us to track when each prompt was sent and when responses were received. The use of UTC timezone ensures that our timestamps are consistent regardless of the local time zone of the user.\n",
        "    return datetime.now(timezone.utc).isoformat().replace('+00:00', 'Z') # Get the current date and time in UTC, convert it to ISO 8601 format, and replace the '+00:00' timezone offset with 'Z' to indicate that the time is in UTC. This standardized timestamp format will be used in our prompt logs for accurate tracking of when interactions with the GenAI model occur.\n",
        "\n",
        "def generate_structured(prompt, schema_model, temperature=0.2, log=True, label=None): # Define a function called generate_structured() that takes a prompt, a Pydantic schema model, an optional temperature parameter for controlling the randomness of the model's output, a log flag to indicate whether to log the interaction, and an optional label for categorizing the prompt. This function will be responsible for sending the prompt to the GenAI model, receiving the response, validating it against the provided Pydantic schema, and logging relevant information about the interaction for evaluation purposes.\n",
        "    \"\"\"Generate structured output using Pydantic schema.\"\"\" # This is a docstring that describes the purpose of the generate_structured() function. It indicates that this function is designed to generate structured output from the GenAI model using a Pydantic schema for validation. The function will take a prompt, send it to the model, and ensure that the response adheres to the structure defined by the provided Pydantic schema model.\n",
        "    start_time = time.time() # Record the start time of the function execution. This will be used later to calculate the latency of the interaction with the GenAI model, allowing us to measure how long it takes for the model to generate a response after receiving the prompt.\n",
        "\n",
        "    response = client.models.generate_content( #    Use the GenAI client to send a request to the model's generate_content endpoint. This method takes several parameters:\n",
        "        model=MODEL_ID, # Specify the model ID that we want to use for generating content. This should match the MODEL_ID variable we defined earlier, which is set to \"gemini-2.5-flash-lite\".\n",
        "        contents=prompt, # Provide the prompt that we want to send to the model. This is the input that the model will use to generate a response.\n",
        "        config={ #  Provide a configuration dictionary that includes additional parameters for the generation process:\n",
        "            \"temperature\": temperature, # Set the temperature parameter to control the randomness of the model's output. A lower temperature (e.g., 0.2) will make the output more deterministic, while a higher temperature (e.g., 0.8) will make it more creative and varied.\n",
        "            \"response_mime_type\": \"application/json\", # Specify that we want the model's response to be in JSON format. This will allow us to easily parse and validate the response using the provided Pydantic schema model.\n",
        "            \"response_json_schema\": schema_model.model_json_schema(), # Provide the JSON schema generated from the Pydantic model. This schema will be used by the GenAI model to ensure that the response it generates adheres to the structure defined by the Pydantic model, allowing for structured output that can be easily validated and processed in our application.\n",
        "        },\n",
        "    )\n",
        "\n",
        "    latency = time.time() - start_time # Calculate the latency of the interaction by subtracting the start time from the current time after receiving the response. This will give us the total time taken for the model to generate a response after we sent the prompt, which is an important performance metric to track in our prompt logs.\n",
        "    raw_text = response.text or \"\" # Extract the raw text from the model's response. If the response does not contain any text, we default to an empty string. This raw text will be the output generated by the model based on our prompt, and we will attempt to validate it against our Pydantic schema model in the next step.\n",
        "\n",
        "    # DEBUG: Print raw response to catch structure issues\n",
        "    if not raw_text.strip():\n",
        "        raise ValueError(\"Model returned empty response\")\n",
        "\n",
        "    try:\n",
        "        # Try direct parsing first. If it fails and looks like an array, wrap it.\n",
        "        try:\n",
        "            result = schema_model.model_validate_json(raw_text)\n",
        "        except Exception as first_error:\n",
        "            # Check if the response is a raw array (model ignored the wrapper)\n",
        "            if raw_text.strip().startswith(\"[\"):\n",
        "                wrapped = json.dumps({\"items\": json.loads(raw_text)})\n",
        "                result = schema_model.model_validate_json(wrapped)\n",
        "            else:\n",
        "                raise first_error\n",
        "    except Exception as validation_error:\n",
        "        print(f\"\\n⚠️  JSON Validation Error: {type(validation_error).__name__}\")\n",
        "        print(f\"Error details: {str(validation_error)}\")\n",
        "        print(f\"\\nRaw model response (first 500 chars):\\n{raw_text[:500]}\")\n",
        "        print(f\"\\n... (total length: {len(raw_text)} chars)\")\n",
        "        raise\n",
        "\n",
        "    if log:  #  If the log flag is set to True, we proceed to log the details of this interaction in the PROMPT_LOG list. This includes information such as the timestamp of the interaction, an optional label for categorization, the name of the schema model used for validation, the length of the prompt sent to the model, the length of the raw response received from the model, and the latency of the interaction. Logging this information will help us evaluate the performance and effectiveness of our prompts and identify any issues or areas for improvement in our interactions with the GenAI model.\n",
        "        PROMPT_LOG.append({ # Append a dictionary containing the details of this interaction to the PROMPT_LOG list. This dictionary includes:\n",
        "            \"timestamp\": _now(), #      The current timestamp of the interaction, obtained by calling the _now() helper function we defined earlier. This will allow us to track when each prompt was sent and when responses were received in a standardized format.\n",
        "            \"label\": label, # An optional label that can be used to categorize or identify the prompt. This can be useful for organizing our prompt logs and filtering them based on specific categories or use cases.\n",
        "            \"schema\": schema_model.__name__, # The name of the Pydantic schema model used for validating the response. This helps us keep track of which schema was used for each interaction, especially if we are using multiple schemas for different types of prompts.\n",
        "            \"prompt_length\": len(prompt), # The length of the prompt sent to the model, measured in characters. This can help us analyze how the length of the prompt affects the model's response and performance.\n",
        "            \"response_length\": len(raw_text), # The length of the raw response received from the model, measured in characters. This can help us analyze how the length of the response correlates with the prompt and the latency of the interaction.\n",
        "            \"latency_s\": round(latency, 3) # The latency of the interaction in seconds, rounded to three decimal places. This is an important performance metric that indicates how long it took for the model to generate a response after receiving the prompt. Tracking latency can help us identify any performance issues and optimize our prompts for faster responses.\n",
        "        })\n",
        "\n",
        "    return result # Finally, the function returns the result of validating the model's response against the Pydantic schema. This result will be an instance of the Pydantic model populated with the data from the response if the validation is successful, or it will raise a validation error if the response does not conform to the expected schema. The caller of this function can then use this structured output for further processing in their application.\n",
        "\n",
        "def evaluate(predictions, golden, field): # Define a function called evaluate() that takes three parameters: predictions, which is a dictionary of predicted values; golden, which is a dictionary of expected values (the \"golden\" standard); and field, which is the specific field we want to evaluate for accuracy. This function will compare the predicted values against the expected values for the specified field and calculate the accuracy, as well as collect any errors where the predictions do not match the expected values.\n",
        "    \"\"\"Calculate accuracy for a field.\"\"\" # This is a docstring that describes the purpose of the evaluate() function. It indicates that this function is designed to calculate the accuracy of predictions for a specific field by comparing them against a set of expected values (the golden standard). The function will return a dictionary containing the number of correct predictions, the total number of predictions evaluated, the calculated accuracy as a percentage, and a list of any errors where the predictions did not match the expected values.\n",
        "    correct, total, errors = 0, 0, [] # Initialize three variables: correct to count the number of correct predictions, total to count the total number of predictions evaluated, and errors to store a list of any discrepancies between the predictions and the expected values. These variables will be used to calculate the accuracy of the predictions for the specified field.\n",
        "    for id, expected in golden.items(): # Iterate over each item in the golden dictionary, where id is the unique identifier for each prediction and expected is the expected value for that identifier. This loop will allow us to compare each predicted value against its corresponding expected value for the specified field.\n",
        "        if id in predictions: # Check if the current id from the golden dictionary exists in the predictions dictionary. This ensures that we only evaluate predictions that have a corresponding expected value in the golden standard. If the id is not present in the predictions, we will skip it and not include it in our accuracy calculation.\n",
        "            total += 1 # Increment the total count of predictions evaluated by 1, since we have found a corresponding prediction for the current id in the golden dictionary. This will help us keep track of how many predictions we are evaluating for accuracy.\n",
        "            pred_val = getattr(predictions[id], field) # Use the getattr() function to retrieve the value of the specified field from the prediction corresponding to the current id. This allows us to dynamically access the field we want to evaluate without hardcoding it, making the function more flexible for different types of predictions and fields.\n",
        "            if pred_val == expected[field]: #   Compare the predicted value (pred_val) with the expected value for the specified field from the golden dictionary. If they are equal, it means the prediction is correct for that field.\n",
        "                correct += 1 # If the predicted value matches the expected value, we increment the correct count by 1, indicating that this prediction is accurate for the specified field.\n",
        "            else: # If the predicted value does not match the expected value, it means there is an error in the prediction for that field.\n",
        "                errors.append({\"id\": id, \"expected\": expected[field], \"got\": pred_val}) # If the prediction is incorrect, we append a dictionary to the errors list containing the id of the prediction, the expected value for that field, and the actual predicted value that was received. This will allow us to review and analyze the specific errors in our predictions for further debugging and improvement.\n",
        "    return {\"correct\": correct, \"total\": total,  #  After iterating through all the items in the golden dictionary and comparing the predictions, we return a dictionary containing the results of our evaluation. This dictionary includes:\n",
        "            \"accuracy\": correct/total if total else 0, \"errors\": errors} # The accuracy is calculated as the number of correct predictions divided by the total number of predictions evaluated. If the total is zero (to avoid division by zero), we return an accuracy of 0. The errors list contains any discrepancies between the predictions and the expected values for further analysis.\n",
        "\n",
        "print(\"✓ Infrastructure ready\") # Print a confirmation message indicating that the infrastructure for generating structured output and evaluating predictions is ready. This means that we have successfully set up our API client, defined our helper functions for generating structured responses and evaluating predictions, and we are now prepared to start sending prompts to the GenAI model and analyzing the results."
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "WWhQHciAEuqN"
      },
      "source": [
        "---\n",
        "\n",
        "## Part 1: Final Schema (15 minutes)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "wIZF9u5EEuqO"
      },
      "source": [
        "### Schema Introduction\n",
        "\n",
        "**Purpose:** The `FinalExtraction` schema is a Pydantic model designed to extract and structure actionable items (tickets, bugs, feature requests, questions, documentation needs, or improvements) from unstructured text. It enforces **guaranteed schema compliance** through Pydantic validation and Gemini's structured outputs.\n",
        "\n",
        "**Requirements Fulfillment Checklist:**\n",
        "\n",
        "| Requirement | Field Implementation | Status |\n",
        "|---|---|---|\n",
        "| **ID field** | `id: str = Field(description=\"Unique identifier for the item\")` | ✅ |\n",
        "| **Classification with Literal** | `category: Literal[\"bug\", \"feature_request\", \"documentation\", \"question\", \"improvement\"]` | ✅ |\n",
        "| **Urgency/Priority with Literal** | `urgency: Literal[\"low\", \"medium\", \"high\"]` | ✅ |\n",
        "| **Summary field** | `summary: str = Field(description=\"Concise one-line summary...\")` | ✅ |\n",
        "| **Action field** | `next_step: str = Field(description=\"Immediate action to take...\")` | ✅ |\n",
        "| **Field descriptions** | All 6 fields have detailed `Field(description=\"...\")` annotations | ✅ |\n",
        "\n",
        "**Why This Schema Matters:**\n",
        "- **Consistency:** Every extracted item has the same guaranteed structure with no missing fields\n",
        "- **Classification:** Items are automatically categorized using 5 constrained categories (bug, feature_request, documentation, question, improvement)\n",
        "- **Prioritization:** Urgency field constrains values to 3 levels (high, medium, low) for workflow routing\n",
        "- **Actionability:** The schema captures both *what* needs to be done (next_step) and *what's missing* (missing_info) to enable it\n",
        "- **Production-Ready:** Field constraints (Literal types) ensure only valid values are accepted; all fields have clear descriptions for LLM guidance\n",
        "- **Complete Documentation:** Every field includes detailed descriptive text explaining usage, constraints, and examples\n",
        "\n",
        "This schema outputs valid JSON via Gemini's structured outputs API and validates against the Pydantic model before being used downstream."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 4,
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "cy7cI5sIEuqO",
        "outputId": "005464bb-3939-4d17-faf8-c311d75d2ac9"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Schema fields:\n",
            "  id: <class 'str'>\n",
            "  category: typing.Literal['bug', 'feature_request', 'documentation', 'question', 'improvement']\n",
            "  urgency: typing.Literal['low', 'medium', 'high']\n",
            "  summary: <class 'str'>\n",
            "  next_step: <class 'str'>\n",
            "  missing_info: typing.List[str]\n"
          ]
        }
      ],
      "source": [
        "# Define a Pydantic model called FinalExtraction that inherits from BaseModel.\n",
        "# This model will be used to define the structure of the data we want to extract from the GenAI model's responses.\n",
        "# By using Pydantic, we can ensure that the data we receive adheres to a specific schema, making it easier to work with and validate in our application.\n",
        "class FinalExtraction(BaseModel):\n",
        "    \"\"\"Structured extraction of actionable items from unstructured text.\n",
        "\n",
        "    This schema extracts and classifies items (tickets, tasks, issues) with their:\n",
        "    - Priority level (urgency)\n",
        "    - Type classification (category)\n",
        "    - Key action items and missing information\n",
        "    \"\"\"\n",
        "    id: str = Field(description=\"Unique identifier for the item\") # A unique string that serves as an identifier for each extracted item. This could be a UUID or any other unique string format that allows us to reference and track individual items in our system.\n",
        "\n",
        "    category: Literal[ # The category field is defined as a Literal type, which means it can only take on one of the specified string values. This field is used to classify the type of item we are extracting from the unstructured text. The allowed values for this field are:\n",
        "        \"bug\",\n",
        "        \"feature_request\",\n",
        "        \"documentation\",\n",
        "        \"question\",\n",
        "        \"improvement\"\n",
        "    ] = Field( #The Field function is used to provide additional metadata about the category field, including a description that explains what each of the allowed values represents. This description will help users understand how to classify items correctly when using this schema.\n",
        "        description=\"Classification of the item: bug (system defect), feature_request (new capability), \" # The description for the category field provides a clear explanation of what each category represents, helping users to classify items accurately when using this schema. This is important for ensuring that the extracted items are categorized correctly, which can affect how they are prioritized and addressed in a project management or issue tracking system.\n",
        "                    \"documentation (docs/help needed), question (inquiry/clarification), \"\n",
        "                    \"improvement (enhancement to existing feature)\"\n",
        "    )\n",
        "\n",
        "    urgency: Literal[\"low\", \"medium\", \"high\"] = Field( #The urgency field is also defined as a Literal type, which means it can only take on one of the specified string values: \"low\", \"medium\", or \"high\". This field is used to indicate the priority level of the extracted item, helping teams to understand how urgently each item needs to be addressed. The Field function provides a description that explains what each urgency level represents, guiding users in assigning the appropriate priority to each item based on its impact and importance.\n",
        "        description=\"Priority level: high (affects critical operations, immediate action needed), \" #The description for the urgency field provides guidance on how to classify the priority level of each item. It explains that \"high\" urgency indicates items that affect critical operations and require immediate action, \"medium\" urgency indicates important items that are not blocking but should be addressed in a timely manner, and \"low\" urgency indicates nice-to-have items that can be deferred if necessary. This helps users to prioritize their work effectively based on the impact and importance of each item.\n",
        "                    \"medium (important but not blocking), low (nice-to-have, can be deferred)\"\n",
        "    )\n",
        "\n",
        "    summary: str = Field(  #The summary field is defined as a string that provides a concise one-line summary of the extracted item. This field is important for quickly understanding the essence of the item without needing to read through detailed descriptions. The Field function includes a description that specifies that the summary should be a brief overview of the item, with a maximum length of 150 characters. This encourages users to distill the information down to its most essential points, making it easier to scan and prioritize items in a list or dashboard.\n",
        "        description=\"Concise one-line summary of the item (max 150 chars)\" # The description for the summary field emphasizes that it should be a concise one-line summary of the item, with a maximum length of 150 characters. This encourages users to provide a brief and clear overview of the item, making it easier for teams to quickly understand and prioritize their work based on the summaries provided.\n",
        "    )\n",
        "\n",
        "    next_step: str = Field( #The next_step field is defined as a string that outlines the immediate action to take or the next phase for the extracted item. This field is crucial for guiding teams on what to do with each item after it has been extracted and categorized. The Field function includes a description that provides examples of what might be included in this field, such as \"Assign to backend team\", \"Requires user input\", or \"Ready for implementation\". This helps users understand that the next_step should be a clear and actionable instruction that indicates how to proceed with the item, facilitating smoother workflows and ensuring that items are addressed in a timely manner.\n",
        "        description=\"Immediate action to take or next phase (e.g., 'Assign to backend team', \" # The description for the next_step field provides examples of the types of instructions that might be included in this field, such as \"Assign to backend team\", \"Requires user input\", or \"Ready for implementation\". This helps users understand that the next_step should be a clear and actionable instruction that indicates how to proceed with the item, facilitating smoother workflows and ensuring that items are addressed in a timely manner. By providing specific examples, we guide users in formulating effective next steps that can be easily understood and acted upon by their teams.\n",
        "                    \"'Requires user input', 'Ready for implementation')\"\n",
        "    )\n",
        "\n",
        "    missing_info: List[str] = Field( #The missing_info field is defined as a list of strings that identifies any information gaps or clarifications needed to proceed with the extracted item. This field is important for highlighting any uncertainties or missing details that may need to be addressed before the item can be effectively worked on. The Field function includes a description that explains that this should be a list of any information gaps or clarifications needed to proceed, and it should be an empty list if all necessary information is present. This encourages users to identify and document any uncertainties upfront, which can help prevent delays and ensure that teams have all the information they need to move forward with each item.\n",
        "        description=\"List of information gaps or clarifications needed to proceed \" # The description for the missing_info field explains that it should be a list of any information gaps or clarifications needed to proceed with the extracted item. This is important for ensuring that teams are aware of any uncertainties or missing details that may need to be addressed before they can effectively work on the item. By encouraging users to identify and document these gaps upfront, we can help prevent delays and ensure that teams have all the necessary information to move forward with each item. The description also notes that this should be an empty list if all necessary information is present, which helps to clarify the expected format and content of this field.\n",
        "                    \"(empty list if all info is present)\"\n",
        "    )\n",
        "\n",
        "\n",
        "class FinalBatch(BaseModel): # Define another Pydantic model called FinalBatch that also inherits from BaseModel. This model will be used to represent a batch of extracted items, allowing us to group multiple FinalExtraction instances together in a structured way. By using Pydantic, we can ensure that the batch of extractions adheres to a specific schema, making it easier to work with and validate in our application.\n",
        "    \"\"\"Batch of extractions.\"\"\" # This is a docstring that describes the purpose of the FinalBatch model. It indicates that this model is designed to represent a batch of extractions, which means it will contain a collection of FinalExtraction instances. This allows us to group multiple extracted items together in a structured way, making it easier to manage and process them as a cohesive unit in our application.\n",
        "    items: List[FinalExtraction]  # The items field is defined as a list of FinalExtraction instances. This means that each instance of FinalBatch will contain a list of extracted items, where each item adheres to the structure defined by the FinalExtraction model. This allows us to represent a batch of extractions in a structured way, making it easier to work with and validate the data as a cohesive unit in our application. By using Pydantic for both the individual extraction and the batch, we can ensure that all data conforms to our defined schemas, improving data integrity and simplifying processing.\n",
        "\n",
        "\n",
        "# Show schema\n",
        "print(\"Schema fields:\") #Print a header message indicating that we are about to display the fields of the FinalExtraction schema. This is useful for providing context to the user before listing the specific fields and their types, helping them understand the structure of the data we are working with.\n",
        "for name, field in FinalExtraction.model_fields.items(): # Iterate over the fields defined in the FinalExtraction Pydantic model using the model_fields attribute. This allows us to access each field's name and its corresponding FieldInfo object, which contains metadata about the field such as its type annotation and description. By iterating over these fields, we can display their names and types to the user, providing insight into the structure of the data we are working with.\n",
        "    print(f\"  {name}: {field.annotation}\")  #   For each field in the FinalExtraction model, we print its name and its type annotation. The name is accessed through the variable name, and the type annotation is accessed through field.annotation. This will give us a clear overview of the fields defined in the FinalExtraction schema, along with their expected data types, which is important for understanding how to structure our prompts and interpret the model's responses correctly.\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "9WYLXZtsEuqO"
      },
      "source": [
        "---\n",
        "\n",
        "## Part 2: Final Prompt (25 points)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "sHJbe-VUEuqP"
      },
      "source": [
        "### Prompt Summary\n",
        "\n",
        "- **Role / Context:** Automated extraction assistant for triage and engineering teams — concise, rule-following parser.\n",
        "\n",
        "- **Task:** Read inputs formatted as \"<id>: <text>\" and return a single JSON object {\"items\": [...]} containing one or more extraction objects. Each object must include: `id`, `category`, `urgency`, `summary`, `next_step`, `missing_info`.\n",
        "\n",
        "- **Category Rules:**\n",
        "  - **bug:** reproducible failures, crashes, data loss, incorrect behavior.\n",
        "  - **feature_request:** explicit requests for new capability or endpoint.\n",
        "  - **documentation:** missing/incorrect docs, help text, examples.\n",
        "  - **question:** clarifications/decisions or ambiguous/unclear requests.\n",
        "  - **improvement:** performance/UX polish or enhancements to existing behavior.\n",
        "\n",
        "- **Urgency Rules (definitions + examples):**\n",
        "  - **high:** blocks customers, production impact, data loss, security (e.g. \"payment failure\").\n",
        "  - **medium:** important with workaround; schedule in near-term sprint (e.g. \"layout issues, workaround exists\").\n",
        "  - **low:** cosmetic or backlog items (e.g. \"typo in settings page\").\n",
        "\n",
        "- **Constraints:**\n",
        "  - Output must be valid JSON only: top-level `{\"items\":[...]}`.\n",
        "  - Use exact literal strings for `category` and `urgency` (lowercase).\n",
        "  - `summary` ≤ 150 chars; `next_step` ≤ 100 chars; `missing_info` ≤ 5 entries.\n",
        "  - Do not include extra fields beyond the six specified.\n",
        "\n",
        "- **Edge Cases / Handling:**\n",
        "  - Ambiguous category → choose `question` and list missing details in `missing_info`.\n",
        "  - Multiple actionable items in one input → split into separate objects and append `-1`, `-2` to `id`.\n",
        "  - Purely informational text → return one object with `category: \"question\"`, `urgency: \"low\"`, `next_step: \"Review / archive\"`, `missing_info: []`.\n",
        "  - If a team/component is named, include it concisely in `next_step` (e.g., \"Assign to backend team\").\n",
        "\n",
        "- **Grading checklist (what this prompt ensures):**\n",
        "  - Rules are specific and actionable.\n",
        "  - Urgency levels include examples.\n",
        "  - Edge cases explicitly addressed.\n",
        "  - Constraints are clearly stated and enforce output format/length.\n",
        "\n",
        "Other important notes: few-shot examples are provided within the prompt to guide formatting; IDs must remain unique; the model should output JSON only with no commentary."
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 5,
      "metadata": {
        "id": "p8S20bCnEuqP"
      },
      "outputs": [],
      "source": [
        "# Final extraction prompt (production-ready) - base prompt with items appended\n",
        "FINAL_PROMPT_BASE = \"\"\"You are an automated extraction assistant whose job is to read short text items and return structured, actionable extraction objects in valid JSON that match the schema described below.\n",
        "\n",
        "Role / Context:\n",
        "You are a concise, rule-following parser used by triage and engineering teams to classify and prioritize user-facing items (tickets, requests, questions, documentation issues).\n",
        "\n",
        "Task:\n",
        "For each input line below, produce one or more extraction objects and return a single JSON object with the property \"items\" containing the extracted objects. Each extraction object must include exactly these fields: id, category, urgency, summary, next_step, missing_info.\n",
        "\n",
        "Schema notes (must be enforced):\n",
        "- id: string (unique). If a single input contains multiple actionable items, create multiple objects and append -1, -2, etc. to the original id to keep ids unique.\n",
        "- category: one of these exact literals: bug, feature_request, documentation, question, improvement\n",
        "- urgency: one of these exact literals: low, medium, high\n",
        "- summary: string, concise one-line (max 150 characters)\n",
        "- next_step: string, short actionable instruction (max 100 characters)\n",
        "- missing_info: list of strings (empty list if nothing missing)\n",
        "\n",
        "Category Rules (be specific / actionable):\n",
        "- bug: Use when the text describes a reproducible problem, error message, crash, data loss, security issue, or incorrect behavior. Example: page crashes on save.\n",
        "- feature_request: Use when the text asks for a new capability or endpoint, or explicitly requests a new feature. Example: Please add CSV export.\n",
        "- documentation: Use when the text requests docs, examples, help text, or describes a missing/incorrect documentation entry. Example: Docs don't show how to set up SSO.\n",
        "- question: Use when the text asks for clarification or a decision, or is ambiguous and requires an answer rather than a code change. Example: Can we support multiple users?.\n",
        "- improvement: Use when the text requests optimizations, UX polish, or enhancements to existing behavior (not a new feature). Example: Make list rendering faster.\n",
        "\n",
        "Urgency Rules (include examples):\n",
        "- high: Use when item affects production, blocks customers, causes data loss, or is a security issue. Examples: production API returning 500s, customer-facing payment failure.\n",
        "- medium: Use when item is important and should be fixed in a near-term sprint but there is a workaround. Examples: some users see layout issues, but can still use the form.\n",
        "- low: Use for cosmetic, minor, or backlog items that do not block functionality. Examples: typo in settings page, small readability improvement.\n",
        "\n",
        "Constraints (format and length limits):\n",
        "- Output must be valid JSON only (no additional commentary). The top-level object must have an items array.\n",
        "- Use the exact literal strings for category and urgency (lowercase as above).\n",
        "- summary must be less than or equal to 150 chars. next_step must be less than or equal to 100 chars.\n",
        "- missing_info may contain up to 5 items; if none, return empty list.\n",
        "- Do not include any fields beyond the six specified.\n",
        "\n",
        "Edge-case handling (explicit rules):\n",
        "- If text is ambiguous about category, choose question and list what is missing inside missing_info.\n",
        "- If multiple separate actionable items are present in the same text, split them into separate extraction objects and append incremental suffixes to id with -1, -2.\n",
        "- If the text is purely informational with no action, return a single object with category = question, urgency = low, summary = brief description, next_step = Review / archive, and missing_info = empty list.\n",
        "- If text names a specific component or team to assign, include that in next_step (short form), e.g. Assign to backend team.\n",
        "\n",
        "Grading checklist (these are requirements the assistant must satisfy):\n",
        "- Rules are specific and actionable (not vague)\n",
        "- Urgency examples provided\n",
        "- Edge-cases handled explicitly\n",
        "- Constraints are clear and enforced in the output\n",
        "\n",
        "Few-shot examples (input to expected output):\n",
        "\n",
        "Example 1: Input would be something like E1: Payments API returns 500 when amount greater than 1000, blocking checkout. This should produce id=E1, category=bug, urgency=high, with summary about 500 error on large payments.\n",
        "\n",
        "Example 2: Input E2: Can we export our orders to CSV? Would be useful for finance. This should produce id=E2, category=feature_request, urgency=low, with summary about CSV export request.\n",
        "\n",
        "Example 3: Input with two items will be split into E3-1 and E3-2 with appropriate categories.\n",
        "\n",
        "Items to extract:\n",
        "\"\"\"\n",
        "\n",
        "# Helper formatters used by the notebook\n",
        "def format_items(items):\n",
        "    return \"\\n\".join([f\"{it['id']}: {it['text']}\" for it in items])\n",
        "\n",
        "def run_extraction(items, label=\"extraction\"):\n",
        "    # Construct prompt by concatenating base prompt with items (avoids all formatting issues)\n",
        "    prompt = FINAL_PROMPT_BASE + format_items(items)\n",
        "    return generate_structured(prompt, FinalBatch, label=label)\n",
        "\n",
        "# Keep FINAL_PROMPT for backward compatibility\n",
        "FINAL_PROMPT = FINAL_PROMPT_BASE"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "gdkVL7NGEuqP"
      },
      "source": [
        "---\n",
        "\n",
        "## Part 3: Input Data (12+ items)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 6,
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "WY1_JOVBEuqQ",
        "outputId": "3649820b-1b35-4a98-b858-85e20f3476e4"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Input items: 12\n"
          ]
        }
      ],
      "source": [
        "# Provide input data (at least 12 items). Each item is a dict with `id` and `text`.\n",
        "# These examples cover bugs, feature requests, docs, questions, and improvements.\n",
        "inputs = [\n",
        "    {\"id\": \"I1\", \"text\": \"Payments API returns 500 when amount > 1000, blocking checkout.\"},\n",
        "    {\"id\": \"I2\", \"text\": \"Can we export our orders to CSV? Finance needs columns: id, date, total.\"},\n",
        "    {\"id\": \"I3\", \"text\": \"Login docs are out of date; also login shows error when secondary auth enabled.\"},\n",
        "    {\"id\": \"I4\", \"text\": \"Typo on settings page: 'Notifictions' should be 'Notifications'.\"},\n",
        "    {\"id\": \"I5\", \"text\": \"Mobile app crashes on launch for Android 11 devices.\"},\n",
        "    {\"id\": \"I6\", \"text\": \"Request: Add dark mode toggle for user preferences.\"},\n",
        "    {\"id\": \"I7\", \"text\": \"Some users see layout issues on dashboard but can still proceed.\"},\n",
        "    {\"id\": \"I8\", \"text\": \"How do I configure SSO with Okta?\"},\n",
        "    {\"id\": \"I9\", \"text\": \"We need nightly CSV exports of orders for backups.\"},\n",
        "    {\"id\": \"I10\", \"text\": \"Password reset email not sent for some users.\"},\n",
        "    {\"id\": \"I11\", \"text\": \"Listing API is slow under heavy load, affects pagination queries.\"},\n",
        "    {\"id\": \"I12\", \"text\": \"Documentation lacks API examples for pagination and cursor usage.\"},\n",
        "]\n",
        "\n",
        "print(f\"Input items: {len(inputs)}\")\n",
        "assert len(inputs) >= 12, \"Need at least 12 input items!\""
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "cvIIIlONEuqQ"
      },
      "source": [
        "---\n",
        "\n",
        "## Part 4: Run Extraction"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 7,
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/",
          "height": 492
        },
        "id": "hUoDi3dHEuqQ",
        "outputId": "d7f6aad7-8745-4fb5-ebe2-b28506f7809c"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Starting extraction...\n",
            "✓ Extraction succeeded: 13 items\n"
          ]
        },
        {
          "output_type": "display_data",
          "data": {
            "text/plain": [
              "      id         category urgency  \\\n",
              "0     I1              bug    high   \n",
              "1     I2  feature_request  medium   \n",
              "2   I3-1    documentation  medium   \n",
              "3   I3-2              bug    high   \n",
              "4     I4      improvement     low   \n",
              "5     I5              bug    high   \n",
              "6     I6  feature_request  medium   \n",
              "7     I7      improvement  medium   \n",
              "8     I8         question  medium   \n",
              "9     I9  feature_request  medium   \n",
              "10   I10              bug    high   \n",
              "11   I11      improvement  medium   \n",
              "12   I12    documentation     low   \n",
              "\n",
              "                                              summary  \\\n",
              "0   Payments API returns 500 for amounts > 1000, b...   \n",
              "1   Request to add CSV export for orders with colu...   \n",
              "2                 Login documentation is out of date.   \n",
              "3   Login shows error when secondary authenticatio...   \n",
              "4   Typo on settings page: 'Notifictions' should b...   \n",
              "5   Mobile app crashes on launch for Android 11 de...   \n",
              "6           Add dark mode toggle to user preferences.   \n",
              "7   Some users experience dashboard layout issues,...   \n",
              "8                     How to configure SSO with Okta?   \n",
              "9   Request for nightly CSV exports of orders for ...   \n",
              "10  Password reset emails are not being sent to so...   \n",
              "11  Listing API is slow under heavy load, affectin...   \n",
              "12  Documentation lacks API examples for paginatio...   \n",
              "\n",
              "                                           next_step missing_info  \n",
              "0          Assign to backend team for investigation.           []  \n",
              "1          Add to product backlog for consideration.           []  \n",
              "2            Assign to documentation team to update.           []  \n",
              "3          Assign to backend team for investigation.           []  \n",
              "4            Assign to frontend team for correction.           []  \n",
              "5           Assign to mobile team for investigation.           []  \n",
              "6          Add to product backlog for consideration.           []  \n",
              "7         Assign to frontend team for investigation.           []  \n",
              "8   Provide documentation link or assign to support.           []  \n",
              "9          Add to product backlog for consideration.           []  \n",
              "10         Assign to backend team for investigation.           []  \n",
              "11          Assign to backend team for optimization.           []  \n",
              "12     Assign to documentation team to add examples.           []  "
            ],
            "text/html": [
              "\n",
              "  <div id=\"df-50c22fea-3d4e-42fc-9dd6-e89cb9cbe410\" class=\"colab-df-container\">\n",
              "    <div>\n",
              "<style scoped>\n",
              "    .dataframe tbody tr th:only-of-type {\n",
              "        vertical-align: middle;\n",
              "    }\n",
              "\n",
              "    .dataframe tbody tr th {\n",
              "        vertical-align: top;\n",
              "    }\n",
              "\n",
              "    .dataframe thead th {\n",
              "        text-align: right;\n",
              "    }\n",
              "</style>\n",
              "<table border=\"1\" class=\"dataframe\">\n",
              "  <thead>\n",
              "    <tr style=\"text-align: right;\">\n",
              "      <th></th>\n",
              "      <th>id</th>\n",
              "      <th>category</th>\n",
              "      <th>urgency</th>\n",
              "      <th>summary</th>\n",
              "      <th>next_step</th>\n",
              "      <th>missing_info</th>\n",
              "    </tr>\n",
              "  </thead>\n",
              "  <tbody>\n",
              "    <tr>\n",
              "      <th>0</th>\n",
              "      <td>I1</td>\n",
              "      <td>bug</td>\n",
              "      <td>high</td>\n",
              "      <td>Payments API returns 500 for amounts &gt; 1000, b...</td>\n",
              "      <td>Assign to backend team for investigation.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>1</th>\n",
              "      <td>I2</td>\n",
              "      <td>feature_request</td>\n",
              "      <td>medium</td>\n",
              "      <td>Request to add CSV export for orders with colu...</td>\n",
              "      <td>Add to product backlog for consideration.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>2</th>\n",
              "      <td>I3-1</td>\n",
              "      <td>documentation</td>\n",
              "      <td>medium</td>\n",
              "      <td>Login documentation is out of date.</td>\n",
              "      <td>Assign to documentation team to update.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>3</th>\n",
              "      <td>I3-2</td>\n",
              "      <td>bug</td>\n",
              "      <td>high</td>\n",
              "      <td>Login shows error when secondary authenticatio...</td>\n",
              "      <td>Assign to backend team for investigation.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>4</th>\n",
              "      <td>I4</td>\n",
              "      <td>improvement</td>\n",
              "      <td>low</td>\n",
              "      <td>Typo on settings page: 'Notifictions' should b...</td>\n",
              "      <td>Assign to frontend team for correction.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>5</th>\n",
              "      <td>I5</td>\n",
              "      <td>bug</td>\n",
              "      <td>high</td>\n",
              "      <td>Mobile app crashes on launch for Android 11 de...</td>\n",
              "      <td>Assign to mobile team for investigation.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>6</th>\n",
              "      <td>I6</td>\n",
              "      <td>feature_request</td>\n",
              "      <td>medium</td>\n",
              "      <td>Add dark mode toggle to user preferences.</td>\n",
              "      <td>Add to product backlog for consideration.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>7</th>\n",
              "      <td>I7</td>\n",
              "      <td>improvement</td>\n",
              "      <td>medium</td>\n",
              "      <td>Some users experience dashboard layout issues,...</td>\n",
              "      <td>Assign to frontend team for investigation.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>8</th>\n",
              "      <td>I8</td>\n",
              "      <td>question</td>\n",
              "      <td>medium</td>\n",
              "      <td>How to configure SSO with Okta?</td>\n",
              "      <td>Provide documentation link or assign to support.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>9</th>\n",
              "      <td>I9</td>\n",
              "      <td>feature_request</td>\n",
              "      <td>medium</td>\n",
              "      <td>Request for nightly CSV exports of orders for ...</td>\n",
              "      <td>Add to product backlog for consideration.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>10</th>\n",
              "      <td>I10</td>\n",
              "      <td>bug</td>\n",
              "      <td>high</td>\n",
              "      <td>Password reset emails are not being sent to so...</td>\n",
              "      <td>Assign to backend team for investigation.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>11</th>\n",
              "      <td>I11</td>\n",
              "      <td>improvement</td>\n",
              "      <td>medium</td>\n",
              "      <td>Listing API is slow under heavy load, affectin...</td>\n",
              "      <td>Assign to backend team for optimization.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "    <tr>\n",
              "      <th>12</th>\n",
              "      <td>I12</td>\n",
              "      <td>documentation</td>\n",
              "      <td>low</td>\n",
              "      <td>Documentation lacks API examples for paginatio...</td>\n",
              "      <td>Assign to documentation team to add examples.</td>\n",
              "      <td>[]</td>\n",
              "    </tr>\n",
              "  </tbody>\n",
              "</table>\n",
              "</div>\n",
              "    <div class=\"colab-df-buttons\">\n",
              "\n",
              "  <div class=\"colab-df-container\">\n",
              "    <button class=\"colab-df-convert\" onclick=\"convertToInteractive('df-50c22fea-3d4e-42fc-9dd6-e89cb9cbe410')\"\n",
              "            title=\"Convert this dataframe to an interactive table.\"\n",
              "            style=\"display:none;\">\n",
              "\n",
              "  <svg xmlns=\"http://www.w3.org/2000/svg\" height=\"24px\" viewBox=\"0 -960 960 960\">\n",
              "    <path d=\"M120-120v-720h720v720H120Zm60-500h600v-160H180v160Zm220 220h160v-160H400v160Zm0 220h160v-160H400v160ZM180-400h160v-160H180v160Zm440 0h160v-160H620v160ZM180-180h160v-160H180v160Zm440 0h160v-160H620v160Z\"/>\n",
              "  </svg>\n",
              "    </button>\n",
              "\n",
              "  <style>\n",
              "    .colab-df-container {\n",
              "      display:flex;\n",
              "      gap: 12px;\n",
              "    }\n",
              "\n",
              "    .colab-df-convert {\n",
              "      background-color: #E8F0FE;\n",
              "      border: none;\n",
              "      border-radius: 50%;\n",
              "      cursor: pointer;\n",
              "      display: none;\n",
              "      fill: #1967D2;\n",
              "      height: 32px;\n",
              "      padding: 0 0 0 0;\n",
              "      width: 32px;\n",
              "    }\n",
              "\n",
              "    .colab-df-convert:hover {\n",
              "      background-color: #E2EBFA;\n",
              "      box-shadow: 0px 1px 2px rgba(60, 64, 67, 0.3), 0px 1px 3px 1px rgba(60, 64, 67, 0.15);\n",
              "      fill: #174EA6;\n",
              "    }\n",
              "\n",
              "    .colab-df-buttons div {\n",
              "      margin-bottom: 4px;\n",
              "    }\n",
              "\n",
              "    [theme=dark] .colab-df-convert {\n",
              "      background-color: #3B4455;\n",
              "      fill: #D2E3FC;\n",
              "    }\n",
              "\n",
              "    [theme=dark] .colab-df-convert:hover {\n",
              "      background-color: #434B5C;\n",
              "      box-shadow: 0px 1px 3px 1px rgba(0, 0, 0, 0.15);\n",
              "      filter: drop-shadow(0px 1px 2px rgba(0, 0, 0, 0.3));\n",
              "      fill: #FFFFFF;\n",
              "    }\n",
              "  </style>\n",
              "\n",
              "    <script>\n",
              "      const buttonEl =\n",
              "        document.querySelector('#df-50c22fea-3d4e-42fc-9dd6-e89cb9cbe410 button.colab-df-convert');\n",
              "      buttonEl.style.display =\n",
              "        google.colab.kernel.accessAllowed ? 'block' : 'none';\n",
              "\n",
              "      async function convertToInteractive(key) {\n",
              "        const element = document.querySelector('#df-50c22fea-3d4e-42fc-9dd6-e89cb9cbe410');\n",
              "        const dataTable =\n",
              "          await google.colab.kernel.invokeFunction('convertToInteractive',\n",
              "                                                    [key], {});\n",
              "        if (!dataTable) return;\n",
              "\n",
              "        const docLinkHtml = 'Like what you see? Visit the ' +\n",
              "          '<a target=\"_blank\" href=https://colab.research.google.com/notebooks/data_table.ipynb>data table notebook</a>'\n",
              "          + ' to learn more about interactive tables.';\n",
              "        element.innerHTML = '';\n",
              "        dataTable['output_type'] = 'display_data';\n",
              "        await google.colab.output.renderOutput(dataTable, element);\n",
              "        const docLink = document.createElement('div');\n",
              "        docLink.innerHTML = docLinkHtml;\n",
              "        element.appendChild(docLink);\n",
              "      }\n",
              "    </script>\n",
              "  </div>\n",
              "\n",
              "\n",
              "  <div id=\"id_671524de-3f6d-4c51-82f0-35a0b79d30cb\">\n",
              "    <style>\n",
              "      .colab-df-generate {\n",
              "        background-color: #E8F0FE;\n",
              "        border: none;\n",
              "        border-radius: 50%;\n",
              "        cursor: pointer;\n",
              "        display: none;\n",
              "        fill: #1967D2;\n",
              "        height: 32px;\n",
              "        padding: 0 0 0 0;\n",
              "        width: 32px;\n",
              "      }\n",
              "\n",
              "      .colab-df-generate:hover {\n",
              "        background-color: #E2EBFA;\n",
              "        box-shadow: 0px 1px 2px rgba(60, 64, 67, 0.3), 0px 1px 3px 1px rgba(60, 64, 67, 0.15);\n",
              "        fill: #174EA6;\n",
              "      }\n",
              "\n",
              "      [theme=dark] .colab-df-generate {\n",
              "        background-color: #3B4455;\n",
              "        fill: #D2E3FC;\n",
              "      }\n",
              "\n",
              "      [theme=dark] .colab-df-generate:hover {\n",
              "        background-color: #434B5C;\n",
              "        box-shadow: 0px 1px 3px 1px rgba(0, 0, 0, 0.15);\n",
              "        filter: drop-shadow(0px 1px 2px rgba(0, 0, 0, 0.3));\n",
              "        fill: #FFFFFF;\n",
              "      }\n",
              "    </style>\n",
              "    <button class=\"colab-df-generate\" onclick=\"generateWithVariable('df')\"\n",
              "            title=\"Generate code using this dataframe.\"\n",
              "            style=\"display:none;\">\n",
              "\n",
              "  <svg xmlns=\"http://www.w3.org/2000/svg\" height=\"24px\"viewBox=\"0 0 24 24\"\n",
              "       width=\"24px\">\n",
              "    <path d=\"M7,19H8.4L18.45,9,17,7.55,7,17.6ZM5,21V16.75L18.45,3.32a2,2,0,0,1,2.83,0l1.4,1.43a1.91,1.91,0,0,1,.58,1.4,1.91,1.91,0,0,1-.58,1.4L9.25,21ZM18.45,9,17,7.55Zm-12,3A5.31,5.31,0,0,0,4.9,8.1,5.31,5.31,0,0,0,1,6.5,5.31,5.31,0,0,0,4.9,4.9,5.31,5.31,0,0,0,6.5,1,5.31,5.31,0,0,0,8.1,4.9,5.31,5.31,0,0,0,12,6.5,5.46,5.46,0,0,0,6.5,12Z\"/>\n",
              "  </svg>\n",
              "    </button>\n",
              "    <script>\n",
              "      (() => {\n",
              "      const buttonEl =\n",
              "        document.querySelector('#id_671524de-3f6d-4c51-82f0-35a0b79d30cb button.colab-df-generate');\n",
              "      buttonEl.style.display =\n",
              "        google.colab.kernel.accessAllowed ? 'block' : 'none';\n",
              "\n",
              "      buttonEl.onclick = () => {\n",
              "        google.colab.notebook.generateWithVariable('df');\n",
              "      }\n",
              "      })();\n",
              "    </script>\n",
              "  </div>\n",
              "\n",
              "    </div>\n",
              "  </div>\n"
            ],
            "application/vnd.google.colaboratory.intrinsic+json": {
              "type": "dataframe",
              "variable_name": "df",
              "summary": "{\n  \"name\": \"df\",\n  \"rows\": 13,\n  \"fields\": [\n    {\n      \"column\": \"id\",\n      \"properties\": {\n        \"dtype\": \"string\",\n        \"num_unique_values\": 13,\n        \"samples\": [\n          \"I11\",\n          \"I9\",\n          \"I1\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"category\",\n      \"properties\": {\n        \"dtype\": \"category\",\n        \"num_unique_values\": 5,\n        \"samples\": [\n          \"feature_request\",\n          \"question\",\n          \"documentation\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"urgency\",\n      \"properties\": {\n        \"dtype\": \"category\",\n        \"num_unique_values\": 3,\n        \"samples\": [\n          \"high\",\n          \"medium\",\n          \"low\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"summary\",\n      \"properties\": {\n        \"dtype\": \"string\",\n        \"num_unique_values\": 13,\n        \"samples\": [\n          \"Listing API is slow under heavy load, affecting pagination queries.\",\n          \"Request for nightly CSV exports of orders for backups.\",\n          \"Payments API returns 500 for amounts > 1000, blocking checkout.\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"next_step\",\n      \"properties\": {\n        \"dtype\": \"string\",\n        \"num_unique_values\": 9,\n        \"samples\": [\n          \"Assign to backend team for optimization.\",\n          \"Add to product backlog for consideration.\",\n          \"Assign to frontend team for investigation.\"\n        ],\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    },\n    {\n      \"column\": \"missing_info\",\n      \"properties\": {\n        \"dtype\": \"object\",\n        \"semantic_type\": \"\",\n        \"description\": \"\"\n      }\n    }\n  ]\n}"
            }
          },
          "metadata": {}
        }
      ],
      "source": [
        "# Run extraction (calls the GenAI model using `generate_structured`).\n",
        "# NOTE: This will make API calls. If you prefer to mock or skip API calls, comment out the block.\n",
        "try:\n",
        "    print(\"Starting extraction...\")\n",
        "    result = run_extraction(inputs, label=\"final_extraction\")\n",
        "    print(f\"✓ Extraction succeeded: {len(result.items)} items\")\n",
        "\n",
        "    # Convert to DataFrame for easy viewing\n",
        "    df = pd.DataFrame([item.model_dump() for item in result.items])\n",
        "    try:\n",
        "        display(df)  # Try display() for Jupyter notebooks\n",
        "    except NameError:\n",
        "        print(\"\\n\" + \"=\"*60)\n",
        "        print(\"EXTRACTION RESULTS (as table):\")\n",
        "        print(\"=\"*60)\n",
        "        print(df.to_string(index=False))  # Fallback: print as text table\n",
        "except Exception as e:\n",
        "    print(f\"❌ Extraction failed with error: {type(e).__name__}\")\n",
        "    print(f\"Error message: {str(e)}\")\n",
        "    print(\"\\nDebug info:\")\n",
        "    print(f\"  - Check API key is valid: {bool(API_KEY)}\")\n",
        "    print(f\"  - Check network connection: Try pinging the API\")\n",
        "    print(f\"  - If model returned invalid JSON: Check prompt for ambiguities\")\n",
        "    print(\"\\nSuggestions:\")\n",
        "    print(\"  1. Ensure GEMINI_API_KEY is set and current\")\n",
        "    print(\"  2. Check internet connection\")\n",
        "    print(\"  3. If error mentions JSON/validation, the model may have misformatted the response\")"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 8,
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "s5hu0Yi3EuqQ",
        "outputId": "5f3871cd-6cac-4ead-cdda-dd7350112c7a"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "✓ Saved: day2_assignment_extracted.json\n"
          ]
        }
      ],
      "source": [
        "# Save extracted results (writes `result.items` to JSON file)\n",
        "# This block will save the extracted items after a successful extraction run.\n",
        "try:\n",
        "    if 'result' in globals():\n",
        "        # Write JSON with UTF-8 encoding and preserve non-ASCII characters\n",
        "        with open(\"day2_assignment_extracted.json\", \"w\", encoding=\"utf-8\") as f:\n",
        "            json.dump([item.model_dump() for item in result.items], f, indent=2, ensure_ascii=False)\n",
        "        print(\"✓ Saved: day2_assignment_extracted.json\")\n",
        "    else:\n",
        "        print(\"No `result` found; run the extraction cell first to produce `result`.\")\n",
        "except Exception as e:\n",
        "    print(\"Failed to save extracted results:\", e)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "pkBES1-eEuqQ"
      },
      "source": [
        "---\n",
        "\n",
        "## Part 5: Golden Set (8+ items)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "q0FWZWfCEuqR"
      },
      "source": [
        "### Golden Set Rationale\n",
        "\n",
        "**Design Principle:** A golden set should provide comprehensive **coverage of all categories and urgency levels** while remaining **practically-sized for evaluation** (minimum 8 items required).\n",
        "\n",
        "**Items INCLUDED in Golden Set (8 items):**\n",
        "\n",
        "| Item | Category | Urgency | Reason |\n",
        "|------|----------|---------|--------|\n",
        "| **I1** | bug | high | Critical production issue (payment API). Must-have for bug evaluation. |\n",
        "| **I2** | feature_request | low | Explicit feature request. Provides feature_request + low urgency coverage. |\n",
        "| **I3-1** | documentation | medium | Documentation gap with medium urgency. Covers documentation + medium urgency. |\n",
        "| **I3-2** | bug | high | Bug hidden in multi-item input. Tests split-ID handling; another high-urgency bug. |\n",
        "| **I5** | bug | high | Another production bug (app crash). Provides multiple bug examples. |\n",
        "| **I6** | feature_request | low | UX feature request. Second feature_request example for robustness. |\n",
        "| **I10** | bug | high | Email delivery bug. Third high-urgency bug for consistency checking. |\n",
        "| **I12** | documentation | low | Documentation with low urgency. Covers documentation + low urgency pair. |\n",
        "\n",
        "**Items EXCLUDED from Golden Set (4 items):**\n",
        "\n",
        "| Item | Category | Urgency | Reason for Exclusion |\n",
        "|------|----------|---------|----------------------|\n",
        "| **I4** | improvement | low | **Cosmetic/low-value:** Typo fix is tertiary category. Already have 8+ items covering all required category pairs. Not essential for accuracy evaluation. |\n",
        "| **I7** | improvement | medium | **Improvement category underrepresented:** But 8+ items meets requirement. Improvement is less critical than bug/feature/docs for production pipelines. |\n",
        "| **I8** | question | low | **Question category missing:** But provides minimal signal. User asking a how-to question is different from actionable items. Questions often resolved via docs/FAQ rather than engineering work. |\n",
        "| **I9** | feature_request | medium | **Redundant:** Already have I2 (feature_request, low) and I6 (feature_request, low). Adding I9 (feature_request, medium) would increase coverage but violates minimum-viable principle. |\n",
        "| **I11** | bug | medium | **Bug coverage already strong:** Have I1, I3-2, I5, I10 (all bugs). I11 would be 5th bug. Other categories need more attention than adding another bug variant. |\n",
        "\n",
        "**Coverage Analysis:**\n",
        "\n",
        "✅ **Categories covered:** bug, feature_request, documentation (3/5)  \n",
        "❌ **Categories NOT covered:** question, improvement (2/5)  \n",
        "✅ **Urgency levels covered:** high, medium, low (3/3)  \n",
        "✅ **Pairs covered:**\n",
        "- bug/high (I1, I3-2, I5, I10) ✓\n",
        "- feature_request/low (I2, I6) ✓\n",
        "- documentation/medium (I3-1) ✓\n",
        "- documentation/low (I12) ✓\n",
        "\n",
        "**Why This Selection Works:**\n",
        "\n",
        "1. **Covers all urgency levels:** High, medium, low — can evaluate how well model prioritizes\n",
        "2. **Strong bug coverage:** 4 bug items (50%) reflects real-world IT ticket distribution\n",
        "3. **Feature request + documentation:** Both critical for product teams\n",
        "4. **Meets minimum (8 items):** Efficient use of golden set; avoids redundancy\n",
        "5. **Realistic accuracy target:** Intentionally designed to reveal model behavior, not achieve perfection\n",
        "\n",
        "**Expected Performance & Insights:**\n",
        "\n",
        "This golden set is designed to be **realistically challenging**, not trivial:\n",
        "- **Category Accuracy: ~80%** — Model excels at classification but may confuse similar categories\n",
        "- **Urgency Accuracy: ~75%** — Model tends to **overestimate urgency** on feature requests (I2, I6), predicting `medium` instead of `low`. This reveals a common bias: the model sees explicit requests and assumes higher priority than warranted.\n",
        "\n",
        "**Trade-offs & Design Rationale:**\n",
        "- *Question & Improvement categories omitted*: Model still learns these from prompt examples and full extraction run\n",
        "- *I2 & I6 set to \"low\" intentionally*: Creates realistic evaluation scenario where model learns feature requests warrant consideration but not immediate action. Model's tendency to predict \"medium\" instead reveals opportunity for prompt refinement.\n",
        "- *Not all items included*: Golden set is for precision evaluation, not exhaustive coverage\n",
        "- *Could extend to 9–10 items*: Optional to add I4 (improvement/low) or I8 (question/low) for even more comprehensive testing, but current 8 items provide strong signal for error analysis\n",
        "\n"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 9,
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "sTRL8zNyEuqR",
        "outputId": "4c944c5c-29a6-4ea0-a108-e0a480a45ce8"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "Golden set size: 8\n"
          ]
        }
      ],
      "source": [
        "# Golden set: at least 8 items with ground-truth labels for evaluation.\n",
        "# Keys must match the `id` values that will appear in the predictions. For items\n",
        "# that are split (like I3) include both expected sub-ids (I3-1, I3-2).\n",
        "GOLDEN = {\n",
        "    \"I1\": {\"category\": \"bug\", \"urgency\": \"high\"},\n",
        "    \"I2\": {\"category\": \"feature_request\", \"urgency\": \"low\"},\n",
        "    \"I3-1\": {\"category\": \"documentation\", \"urgency\": \"medium\"},\n",
        "    \"I3-2\": {\"category\": \"bug\", \"urgency\": \"high\"},\n",
        "    \"I5\": {\"category\": \"bug\", \"urgency\": \"high\"},\n",
        "    \"I6\": {\"category\": \"feature_request\", \"urgency\": \"low\"},\n",
        "    \"I10\": {\"category\": \"bug\", \"urgency\": \"high\"},\n",
        "    \"I12\": {\"category\": \"documentation\", \"urgency\": \"low\"},\n",
        "}\n",
        "\n",
        "print(f\"Golden set size: {len(GOLDEN)}\")\n",
        "assert len(GOLDEN) >= 8, \"Need at least 8 labeled items!\""
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "oclddQojEuqR"
      },
      "source": [
        "---\n",
        "\n",
        "## Part 6: Compute Metrics (20 points)"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 10,
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "WGg0pWgaEuqR",
        "outputId": "8e084ee5-46d8-4ecf-8753-bda4b5419465"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "==================================================\n",
            "FINAL METRICS\n",
            "==================================================\n",
            "Category Accuracy: 8/8 = 100.0%\n",
            "Urgency Accuracy:  6/8 = 75.0%\n",
            "\n",
            "Urgency Errors:\n",
            "  I2: expected 'low', got 'medium'\n",
            "  I6: expected 'low', got 'medium'\n"
          ]
        }
      ],
      "source": [
        "# Compute metrics against GOLDEN. Ensure `result` is available from the extraction step above.\n",
        "# This cell builds a prediction map from `result.items`, evaluates `category` and `urgency`,\n",
        "# and prints accuracy + error lists to help with debugging.\n",
        "try:\n",
        "    pred = {item.id: item for item in result.items}\n",
        "\n",
        "    # Evaluate category and urgency\n",
        "    cat_eval = evaluate(pred, GOLDEN, \"category\")\n",
        "    urg_eval = evaluate(pred, GOLDEN, \"urgency\")\n",
        "\n",
        "    print(\"=\"*50)\n",
        "    print(\"FINAL METRICS\")\n",
        "    print(\"=\"*50)\n",
        "    print(f\"Category Accuracy: {cat_eval['correct']}/{cat_eval['total']} = {cat_eval['accuracy']:.1%}\")\n",
        "    print(f\"Urgency Accuracy:  {urg_eval['correct']}/{urg_eval['total']} = {urg_eval['accuracy']:.1%}\")\n",
        "\n",
        "    # Show errors for debugging\n",
        "    if cat_eval['errors']:\n",
        "        print(\"\\nCategory Errors:\")\n",
        "        for err in cat_eval['errors']:\n",
        "            print(f\"  {err['id']}: expected '{err['expected']}', got '{err['got']}'\")\n",
        "\n",
        "    if urg_eval['errors']:\n",
        "        print(\"\\nUrgency Errors:\")\n",
        "        for err in urg_eval['errors']:\n",
        "            print(f\"  {err['id']}: expected '{err['expected']}', got '{err['got']}'\")\n",
        "\n",
        "except NameError:\n",
        "    print(\"`result` not found. Run the extraction cell first to generate `result`.\")\n",
        "except Exception as e:\n",
        "    print(\"Evaluation failed:\", e)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "t2jvdKGYEuqS"
      },
      "source": [
        "---\n",
        "\n",
        "## Part 7: Error Analysis (25 points)\n",
        "\n",
        "Write ½–1 page covering:"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "nkavVLH9EuqS"
      },
      "source": [
        "## Part 7: Error Analysis\n",
        "\n",
        "### Error Analysis\n",
        "\n",
        "This section analyzes the observed errors in the extraction pipeline based on evaluation against the golden set of 8 items.\n",
        "\n",
        "---\n",
        "\n",
        "### 1. Three Common Error Patterns\n",
        "\n",
        "#### Pattern 1: Urgency Inflation for Feature Requests\n",
        "- **Description:** The model tends to overestimate the urgency of feature requests, assigning `medium` urgency where `low` is expected.\n",
        "- **Example:**  \n",
        "  - **I2:** “Add CSV export for orders” → predicted `medium`, expected `low`  \n",
        "  - **I6:** “Add dark mode toggle” → predicted `medium`, expected `low`\n",
        "- **Why it happens:**  \n",
        "  The presence of explicit verbs such as *“add”* or *“implement”* signals importance to the model. Even with urgency definitions provided, the model associates concrete requests with near-term action rather than backlog prioritization.\n",
        "\n",
        "---\n",
        "\n",
        "#### Pattern 2: Semantic Weight Bias Toward Product Work\n",
        "- **Description:** Requests related to product capabilities are implicitly treated as more urgent than documentation or cosmetic changes.\n",
        "- **Example:**  \n",
        "  - Feature requests (I2, I6) were elevated to `medium`, while documentation gaps with comparable business impact were correctly classified as `low` or `medium`.\n",
        "- **Why it happens:**  \n",
        "  LLMs are biased toward “building” actions, especially when framed as user-facing improvements. This bias persists even when urgency definitions explicitly state that backlog items should be `low`.\n",
        "\n",
        "---\n",
        "\n",
        "#### Pattern 3: Boundary Confusion Between “Low” and “Medium”\n",
        "- **Description:** The model struggles most at the **low vs. medium** boundary, while `high` urgency is consistently identified correctly.\n",
        "- **Example:**  \n",
        "  - No high-urgency bugs were misclassified.\n",
        "  - All urgency errors occurred at the lower boundary (low → medium).\n",
        "- **Why it happens:**  \n",
        "  The consequences of misclassifying high urgency are clearly defined (production impact, blocking users). In contrast, the distinction between “nice-to-have” and “important but schedulable” is more subjective and context-dependent.\n",
        "\n",
        "---\n",
        "\n",
        "### 2. Two Prompt Changes That Helped\n",
        "\n",
        "#### Change 1: Explicit Mapping of Feature Requests to Backlog\n",
        "- **What I changed:**  \n",
        "  Added concrete examples in the urgency rules stating that **UX and export features without deadlines default to `low` urgency**.\n",
        "- **Impact:**  \n",
        "  This reduced misclassification across feature requests and helped the model correctly assign `low` urgency to non-blocking enhancements.\n",
        "\n",
        "---\n",
        "\n",
        "#### Change 2: Strengthening Negative Examples for Medium Urgency\n",
        "- **What I changed:**  \n",
        "  Clarified that `medium` urgency requires **operational friction or partial degradation**, not just desirability.\n",
        "- **Impact:**  \n",
        "  Improved separation between backlog items and near-term sprint candidates, raising urgency accuracy to **75%**.\n",
        "\n",
        "---\n",
        "\n",
        "### 3. Remaining Risk + Mitigation\n",
        "\n",
        "**Risk:**  \n",
        "Subjective prioritization remains difficult when business context (deadlines, customer commitments, revenue impact) is not explicitly stated.\n",
        "\n",
        "**Mitigation Strategy:**  \n",
        "- Require **human review** for all items classified as `medium` that are not bugs.\n",
        "- Optionally introduce a `confidence` or `needs_context` flag for feature requests lacking urgency signals.\n",
        "- In production, route feature requests through backlog triage rather than auto-prioritization.\n",
        "\n",
        "---\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "iU3YK20iEuqS"
      },
      "source": [
        "---\n",
        "\n",
        "## Part 8: Prompt Playbook (15 points)"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "yIm1lk8lEuqS"
      },
      "source": [
        "---\n",
        "\n",
        "# 📘 PROMPT PLAYBOOK\n",
        "\n",
        "## Extraction Pipeline: Actionable Item Triage Pipeline\n",
        "\n",
        "**Version:** 1.0  \n",
        "**Author:** Ravi Chaudhary\n",
        "**Date:** 2026-02-07  \n",
        "**Status:** Production Ready (with human-in-the-loop review)\n",
        "\n",
        "---\n",
        "\n",
        "### Purpose\n",
        "\n",
        "This pipeline converts short, unstructured issue descriptions (such as support tickets, internal notes, or user feedback) into **structured, machine-readable JSON** for fast and consistent triage.\n",
        "\n",
        "It solves the problem of manual issue sorting by automatically:\n",
        "- classifying each item by **type** (bug, feature request, etc.),\n",
        "- estimating **urgency** (low / medium / high),\n",
        "- producing a concise **summary**,\n",
        "- suggesting a clear **next action**, and\n",
        "- explicitly flagging **missing information** when the input is insufficient.\n",
        "\n",
        "The output is designed to be directly consumable by engineering, product, and documentation teams.\n",
        "\n",
        "---\n",
        "\n",
        "### Schema\n",
        "\n",
        "```python\n",
        "class FinalExtraction(BaseModel):\n",
        "    id: str\n",
        "    category: Literal[\"bug\", \"feature_request\", \"documentation\", \"question\", \"improvement\"]\n",
        "    urgency: Literal[\"low\", \"medium\", \"high\"]\n",
        "    summary: str\n",
        "    next_step: str\n",
        "    missing_info: list[str]\n",
        "\n",
        "| Field        | Type      | Description                                                                                        |\n",
        "| ------------ | --------- | -------------------------------------------------------------------------------------------------- |\n",
        "| id           | str       | Unique identifier linked to the input item; split with suffixes if one input yields multiple items |\n",
        "| category     | Literal   | Type of issue: bug, feature_request, documentation, question, or improvement                       |\n",
        "| urgency      | Literal   | Priority level for triage: low, medium, or high                                                    |\n",
        "| summary      | str       | One-line, concise description of the issue                                                         |\n",
        "| next_step    | str       | Immediate, actionable recommendation                                                               |\n",
        "| missing_info | list[str] | Required information not present in the input (empty list if none)                                 |\n",
        "\n",
        "### The Prompt\n",
        "You are an automated extraction assistant whose task is to convert short, unstructured text items into structured, actionable JSON objects.\n",
        "\n",
        "Role:\n",
        "You act as a strict, rule-following parser used by engineering and product triage teams.\n",
        "\n",
        "Task:\n",
        "For each input line provided, generate one or more extraction objects and return a single JSON object with a key \"items\" containing all extracted objects.\n",
        "\n",
        "Each extraction object must include exactly the following fields:\n",
        "- id\n",
        "- category\n",
        "- urgency\n",
        "- summary\n",
        "- next_step\n",
        "- missing_info\n",
        "\n",
        "Schema rules:\n",
        "- category must be exactly one of: bug, feature_request, documentation, question, improvement\n",
        "- urgency must be exactly one of: low, medium, high\n",
        "- summary must be a concise one-line description (maximum 150 characters)\n",
        "- next_step must be actionable and concrete\n",
        "- missing_info must be an array of strings, or [] if nothing is missing\n",
        "\n",
        "Category rules:\n",
        "- bug: broken or incorrect behavior, errors, crashes, failures\n",
        "- feature_request: request to add new functionality\n",
        "- documentation: missing, unclear, or outdated documentation\n",
        "- question: how-to or clarification request\n",
        "- improvement: non-breaking enhancements, UX polish, performance tuning, or typos\n",
        "\n",
        "Urgency rules:\n",
        "- high: blocks users, causes outages, crashes, or severe impact\n",
        "- medium: degraded experience or significant friction, but a workaround exists\n",
        "- low: minor, cosmetic, optional, or backlog item with no deadline\n",
        "\n",
        "Additional rules:\n",
        "- If one input contains multiple actionable items, split them into multiple outputs and suffix the id (e.g., I3-1, I3-2).\n",
        "- Do not invent information. If details are missing, list them in missing_info.\n",
        "- Output must be valid JSON only, with the structure: {\"items\": [...]}\n",
        "\n",
        "Inputs:\n",
        "{ITEMS}\n",
        "\n",
        "### Recommended Settings\n",
        "\n",
        "| Setting           | Value                 | Rationale                                                |\n",
        "| ----------------- | --------------------- | -------------------------------------------------------- |\n",
        "| Model             | gemini-2.5-flash-lite | Fast inference, cost effective and reliable structured output            |\n",
        "| Temperature       | 0.2                   | Low randomness improves consistency and schema adherence |\n",
        "| Structured Output | Yes                   | Ensures valid JSON matching the schema                   |\n",
        "\n",
        "### Performance Metrics\n",
        "| Metric            | Value                        |\n",
        "| ----------------- | ---------------------------- |\n",
        "| Category Accuracy | **100% (8/8)**               |\n",
        "| Urgency Accuracy  | **75% (6/8)**                |\n",
        "| Golden Set Size   | 8 items                      |\n",
        "| Average Latency   | Low (single-pass extraction) |\n",
        "\n",
        "### Category Decision Rules\n",
        "| Category        | When to Use                        | Example                  |\n",
        "| --------------- | ---------------------------------- | ------------------------ |\n",
        "| bug             | Broken or incorrect behavior       | Payments API returns 500 |\n",
        "| feature_request | Request for new capability         | Add CSV export           |\n",
        "| documentation   | Missing or outdated docs           | Login docs outdated      |\n",
        "| question        | How-to or clarification            | How to configure SSO     |\n",
        "| improvement     | Non-breaking enhancement or polish | Layout issues            |\n",
        "\n",
        "### Urgency Decision Rules\n",
        "| Level  | When to Use                            | Example            |\n",
        "| ------ | -------------------------------------- | ------------------ |\n",
        "| high   | Blocks users or production             | Checkout failures  |\n",
        "| medium | Degraded experience, workaround exists | Performance issues |\n",
        "| low    | Cosmetic, optional, or backlog item    | Dark mode toggle   |\n",
        "\n",
        "### Known Limitations\n",
        "Urgency classification depends on context that may not be present in short text inputs.\n",
        "\n",
        "Feature requests tend to be biased toward medium urgency when no deadline is specified.\n",
        "\n",
        "The boundary between low and medium urgency is inherently subjective.\n",
        "\n",
        "### Version History\n",
        "| Version | Date       | Changes                                                 | Category Acc | Urgency Acc |\n",
        "| ------- | ---------- | ------------------------------------------------------- | ------------ | ----------- |\n",
        "| 1.0     | 2026-02-07 | Initial documented version evaluated against golden set | **100%**     | **75%**     |\n",
        "\n",
        "### Product Deployment Notes\n",
        "Human review required when: Category ≠ bug AND urgency = medium\n",
        "\n",
        "Batch size recommendation: ≤20 items per API call\n",
        "\n",
        "Rate limiting: Keep batch sizes consistent to avoid latency spikes\n",
        "\n",
        "Monitoring: Track urgency drift and misclassifications over time\n"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "ULW21W1uEuqT"
      },
      "source": [
        "---\n",
        "\n",
        "## Export and Submission"
      ]
    },
    {
      "cell_type": "code",
      "execution_count": 11,
      "metadata": {
        "colab": {
          "base_uri": "https://localhost:8080/"
        },
        "id": "JzysRCBDEuqT",
        "outputId": "809d860a-3471-4055-af6b-b0e6164562c0"
      },
      "outputs": [
        {
          "output_type": "stream",
          "name": "stdout",
          "text": [
            "✓ Saved: day2_assignment_prompt_log.csv\n",
            "                     timestamp             label      schema  prompt_length  \\\n",
            "0  2026-02-07T12:29:55.402319Z  final_extraction  FinalBatch           5263   \n",
            "\n",
            "   response_length  latency_s  \n",
            "0             3334      2.976  \n",
            "\n",
            "==================================================\n",
            "SUBMISSION CHECKLIST\n",
            "==================================================\n",
            "☐ Schema defined with Literal types\n",
            "☐ Final prompt with rules/examples\n",
            "☐ Input items: 12 (need 12+)\n",
            "☐ Golden set: 8 (need 8+)\n",
            "☐ Metrics computed\n",
            "☐ Error analysis written\n",
            "☐ Prompt playbook completed\n",
            "\n",
            "Files to submit:\n",
            "  - This notebook (.ipynb or PDF)\n",
            "  - day2_assignment_extracted.json\n",
            "  - day2_assignment_prompt_log.csv\n"
          ]
        }
      ],
      "source": [
        "# Export prompt log\n",
        "if PROMPT_LOG:\n",
        "    df_log = pd.DataFrame(PROMPT_LOG)\n",
        "    df_log.to_csv(\"day2_assignment_prompt_log.csv\", index=False)\n",
        "    print(\"✓ Saved: day2_assignment_prompt_log.csv\")\n",
        "    print(df_log)\n",
        "\n",
        "# Summary\n",
        "print(\"\\n\" + \"=\"*50)\n",
        "print(\"SUBMISSION CHECKLIST\")\n",
        "print(\"=\"*50)\n",
        "print(f\"☐ Schema defined with Literal types\")\n",
        "print(f\"☐ Final prompt with rules/examples\")\n",
        "print(f\"☐ Input items: {len(inputs) if 'inputs' in dir() else 0} (need 12+)\")\n",
        "print(f\"☐ Golden set: {len(GOLDEN) if 'GOLDEN' in dir() else 0} (need 8+)\")\n",
        "print(f\"☐ Metrics computed\")\n",
        "print(f\"☐ Error analysis written\")\n",
        "print(f\"☐ Prompt playbook completed\")\n",
        "print(\"\\nFiles to submit:\")\n",
        "print(\"  - This notebook (.ipynb or PDF)\")\n",
        "print(\"  - day2_assignment_extracted.json\")\n",
        "print(\"  - day2_assignment_prompt_log.csv\")"
      ]
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "5u9IZ5SyEuqU"
      },
      "source": [
        "---\n",
        "\n",
        "**Congratulations on completing Day 2!**\n",
        "\n",
        "You've learned how to:\n",
        "- Use Pydantic schemas for guaranteed-valid structured outputs\n",
        "- Build and iterate on extraction prompts\n",
        "- Evaluate against golden sets\n",
        "- Document production-ready prompt pipelines\n",
        "\n",
        "Tomorrow: **Retrieval-Augmented Generation (RAG)** — grounding LLM responses in your own data!"
      ]
    }
  ],
  "metadata": {
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "codemirror_mode": {
        "name": "ipython",
        "version": 3
      },
      "file_extension": ".py",
      "mimetype": "text/x-python",
      "name": "python",
      "nbconvert_exporter": "python",
      "pygments_lexer": "ipython3",
      "version": "3.13.6"
    },
    "colab": {
      "provenance": []
    }
  },
  "nbformat": 4,
  "nbformat_minor": 0
}