{
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# 📋 Day 4: Assignment — Production-Ready AI Agent\n",
    "\n",
    "## Overview\n",
    "\n",
    "Build on your Independent Lab agent to create a **production-ready** system. You will:\n",
    "- Refine your tool definitions and system prompt\n",
    "- Process 12+ queries with full agent traces\n",
    "- Create a golden test set with 8+ annotated scenarios\n",
    "- Evaluate agent traces with LLM-as-judge\n",
    "- Write an error analysis identifying failure patterns\n",
    "- Document everything in an Agent Playbook\n",
    "\n",
    "## Grading Summary\n",
    "\n",
    "| # | Deliverable | Points |\n",
    "|---|------------|--------|\n",
    "| 1 | Tool Definitions (refined) | 15 |\n",
    "| 2 | Agent System Prompt (final) | 15 |\n",
    "| 3 | Agent Outputs (12+ queries) | — |\n",
    "| 4 | Golden Test Set (8+ scenarios) | 10 |\n",
    "| 5 | Agent Trace Evaluation | 15 |\n",
    "| 6 | Error Analysis | 15 |\n",
    "| 7 | Agent Playbook | 15 |\n",
    "| | **Total** | **100** |"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Setup"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "!pip install -q -U google-genai"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "import os, json, time\nfrom datetime import datetime, timezone\nfrom google import genai\nfrom google.genai import types\n\n# ── API Key ──────────────────────────────────────────────\ntry:\n    from google.colab import userdata\n    os.environ[\"GEMINI_API_KEY\"] = userdata.get(\"GEMINI_API_KEY\")\nexcept Exception:\n    pass\n\nif not os.environ.get(\"GEMINI_API_KEY\"):\n    import getpass\n    os.environ[\"GEMINI_API_KEY\"] = getpass.getpass(\"Paste your GEMINI_API_KEY: \")\n\nclient = genai.Client(api_key=os.environ[\"GEMINI_API_KEY\"])\nMODEL_ID = \"gemini-2.5-flash-lite\"\n\nprint(f\"✅ API initialized. Model: {MODEL_ID}\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Infrastructure (from Guided Lab) ─────────────────────\nPROMPT_LOG = []\n\ndef _now():\n    return datetime.now(timezone.utc).isoformat(timespec=\"seconds\").replace(\"+00:00\", \"Z\")\n\ndef log_interaction(role, content, label=None):\n    entry = {\"ts\": _now(), \"role\": role,\n             \"content\": content if isinstance(content, str) else json.dumps(content),\n             \"label\": label or \"\"}\n    PROMPT_LOG.append(entry)\n    return entry\n\ndef show_log(n=10):\n    import pandas as pd\n    if not PROMPT_LOG:\n        print(\"No interactions logged yet.\")\n        return\n    df = pd.DataFrame(PROMPT_LOG[-n:])\n    from IPython.display import display\n    display(df)\n\ndef run_agent(user_message, tools, system_prompt=None, max_steps=10):\n    \"\"\"A manual agent loop with full visibility.\n\n    Args:\n        user_message: The user's request.\n        tools: List of Python functions to use as tools.\n        system_prompt: Optional system instruction for the agent.\n        max_steps: Maximum number of reasoning steps (safety limit).\n\n    Returns:\n        A tuple of (final_text, tools_called, trace) where tools_called is a\n        list of tool names that were invoked during the run, and trace is a\n        list of dicts with structured log entries (call_number, tool, args, result).\n    \"\"\"\n    tool_map = {fn.__name__: fn for fn in tools}\n    call_count = 0        # Track total tool calls across all steps\n    tools_called = []     # Record which tools were actually used\n    trace = []            # Structured log: tool, args, result per call\n\n    # Build initial contents\n    contents = []\n    if system_prompt:\n        contents.append(types.Content(\n            role=\"user\",\n            parts=[types.Part(text=f\"System: {system_prompt}\\n\\nUser: {user_message}\")]\n        ))\n    else:\n        contents.append(types.Content(\n            role=\"user\",\n            parts=[types.Part(text=user_message)]\n        ))\n\n    log_interaction(\"user\", user_message, label=\"agent_input\")\n\n    for step in range(max_steps):\n        response = client.models.generate_content(\n            model=MODEL_ID,\n            contents=contents,\n            config=types.GenerateContentConfig(\n                tools=tools,\n                # Tool calling mode defaults to AUTO — the model\n                # reasons about whether to use tools on each turn.\n                automatic_function_calling=types.AutomaticFunctionCallingConfig(\n                    disable=True  # Model still reasons about tools —\n                    # but the SDK won't execute them automatically.\n                    # Instead it returns the function_call to us,\n                    # and WE run the function below.\n                ),\n            ),\n        )\n\n        # Guard: the model may return an empty response\n        parts = response.parts or []\n        if not parts:\n            print(f\"  Step {step+1}: ⚠️ Empty response from model — retrying...\")\n            continue\n\n        # Add model response to history\n        contents.append(types.Content(role=\"model\", parts=parts))\n\n        # Check for function calls\n        function_results = []\n        for part in parts:\n            if part.function_call:\n                call_count += 1\n                name = part.function_call.name\n                args = dict(part.function_call.args)\n                tools_called.append(name)\n                print(f\"  Tool call {call_count}: 🔧 {name}({args})\")\n\n                # Execute the function\n                try:\n                    result = tool_map[name](**args)\n                except Exception as e:\n                    result = {\"error\": str(e)}\n\n                print(f\"           → {result}\")\n                trace.append({\n                    \"call\": call_count,\n                    \"tool\": name,\n                    \"args\": args,\n                    \"result\": result if isinstance(result, str) else json.dumps(result),\n                })\n                log_interaction(\"tool\", f\"{name}({args}) → {result}\", label=\"tool_call\")\n\n                function_results.append(\n                    types.Part(\n                        function_response=types.FunctionResponse(\n                            name=name,\n                            response={\"result\": result},\n                        )\n                    )\n                )\n\n        if function_results:\n            contents.append(types.Content(role=\"user\", parts=function_results))\n        else:\n            # No function calls → model is done\n            final_text = response.text or \"(no text response)\"\n            log_interaction(\"agent\", final_text, label=\"agent_output\")\n            return final_text, tools_called, trace\n\n    return \"⚠️ Agent reached maximum steps without completing.\", tools_called, trace\n\nprint(\"✅ Agent infrastructure loaded.\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 1: Final Tool Definitions (15 pts)\n",
    "\n",
    "Refine your tools from the Independent Lab. Each tool should have:\n",
    "- Clear function name (verb + noun)\n",
    "- Complete type hints\n",
    "- Detailed docstring with Args/Returns\n",
    "- Graceful error handling\n",
    "- At least one example in the docstring"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Track Data ──────────────────────────────────────────────\n# TODO: Select your track (A=Support, B=Research, C=Finance, D=HR/Recruitment)\nSELECTED_TRACK = \"A\"  # Change to your track\n\n# Track A: Customer Support\nSUPPORT_TICKETS = [\n    {\"id\": \"TK-101\", \"text\": \"App crashes when uploading large photos. Tried reinstalling.\", \"customer\": \"user_42\"},\n    {\"id\": \"TK-102\", \"text\": \"I was charged twice for my subscription this month. Urgent!\", \"customer\": \"user_88\"},\n    {\"id\": \"TK-103\", \"text\": \"How do I export my data to CSV? Can't find the option.\", \"customer\": \"user_15\"},\n    {\"id\": \"TK-104\", \"text\": \"The dark mode doesn't work on iOS 18. Everything is white.\", \"customer\": \"user_67\"},\n    {\"id\": \"TK-105\", \"text\": \"I love the new dashboard update! Great work.\", \"customer\": \"user_23\"},\n    {\"id\": \"TK-106\", \"text\": \"Login fails with SSO. Error code 403. Very urgent.\", \"customer\": \"user_91\"},\n    {\"id\": \"TK-107\", \"text\": \"Can I upgrade from Starter to Growth plan mid-cycle?\", \"customer\": \"user_34\"},\n    {\"id\": \"TK-108\", \"text\": \"API rate limit hit. Need higher quota for production.\", \"customer\": \"user_56\"},\n]\n\nSUPPORT_POLICIES = {\n    \"billing\": \"Duplicate charges must be refunded within 48 hours. Escalate to billing team if amount > $500.\",\n    \"bugs\": \"Critical bugs (crash, data loss) are Priority 1. Assign to engineering. ETA: 24h for P1, 72h for P2.\",\n    \"features\": \"Feature requests go to the product backlog. Thank the customer and share the roadmap link.\",\n    \"account\": \"Account changes (upgrades, downgrades) can be processed immediately. Prorate the difference.\",\n    \"api\": \"API rate limit increases require manager approval. Standard limit: 1000 req/min. Enterprise: 10000 req/min.\",\n    \"praise\": \"Positive feedback should be forwarded to the team Slack channel. Thank the customer.\",\n}\n\n# Track B: Research Assistant\nRESEARCH_DOCS = {\n    \"ai_market\": \"The global AI market was valued at $196B in 2023 and is projected to reach $1.8T by 2030, growing at 36% CAGR. Key segments: generative AI ($44B), computer vision ($38B), NLP ($35B).\",\n    \"competitor_alpha\": \"AlphaTech launched their enterprise AI platform in Q2 2024. Pricing: $50k/year for teams up to 50 users. Key differentiator: on-premise deployment option. Weakness: no mobile SDK.\",\n    \"competitor_beta\": \"BetaCorp acquired DataMinds for $2.3B in March 2024. Combined entity focuses on real-time analytics. Revenue grew 45% YoY to $890M. Weakness: high customer churn (18%).\",\n    \"customer_trends\": \"Enterprise AI adoption increased from 35% to 55% between 2022-2024. Top use cases: customer service automation (72%), document processing (65%), predictive analytics (58%).\",\n    \"regulation\": \"The EU AI Act entered into force in August 2024. Key requirements: transparency obligations for general-purpose AI, risk classification system, and mandatory conformity assessments for high-risk applications.\",\n    \"talent\": \"AI engineer salaries increased 25% in 2024. Average: $185k in the US, $120k in Europe. Biggest skill gaps: MLOps (67% of companies), responsible AI (54%), agent frameworks (48%).\",\n}\n\n# Track C: Financial Analysis\nFINANCIAL_DATA = {\n    \"ACME\": {\"revenue_q4\": 45_000_000, \"expenses_q4\": 38_000_000, \"employees\": 450, \"growth_yoy\": 0.12, \"sector\": \"Manufacturing\"},\n    \"TECHSTART\": {\"revenue_q4\": 12_000_000, \"expenses_q4\": 15_000_000, \"employees\": 120, \"growth_yoy\": 0.45, \"sector\": \"SaaS\"},\n    \"RETAILMAX\": {\"revenue_q4\": 89_000_000, \"expenses_q4\": 82_000_000, \"employees\": 2200, \"growth_yoy\": -0.03, \"sector\": \"Retail\"},\n}\n\nEARNINGS_REPORTS = {\n    \"ACME_Q4\": \"Acme Corp reported Q4 revenue of $45M, up 12% YoY. Margins improved to 15.6% due to automation initiatives. Guidance for next quarter: $47-49M revenue.\",\n    \"TECHSTART_Q4\": \"TechStart burned $3M in Q4 but grew revenue 45% YoY to $12M. ARR reached $48M. Key risk: runway is 14 months at current burn rate. Pursuing Series C.\",\n    \"RETAILMAX_Q4\": \"RetailMax Q4 revenue was $89M, down 3% YoY. E-commerce grew 15% but couldn't offset 8% decline in physical stores. Announced 200 layoffs.\",\n    \"INDUSTRY_OUTLOOK\": \"The SaaS sector is expected to grow 18% in 2025, driven by AI integration. Manufacturing AI spending projected at $9.8B. Retail tech investment flat YoY.\",\n}\n\n# Track D: HR / Recruitment\nCANDIDATES = [\n    {\"id\": \"C-201\", \"name\": \"Alice Chen\", \"skills\": [\"Python\", \"ML\", \"TensorFlow\"], \"experience_years\": 5, \"current_role\": \"ML Engineer\", \"salary_expectation\": 160000},\n    {\"id\": \"C-202\", \"name\": \"Bob Martinez\", \"skills\": [\"Java\", \"AWS\", \"Kubernetes\"], \"experience_years\": 8, \"current_role\": \"DevOps Lead\", \"salary_expectation\": 185000},\n    {\"id\": \"C-203\", \"name\": \"Carol Zhang\", \"skills\": [\"Python\", \"NLP\", \"LLMs\", \"RAG\"], \"experience_years\": 3, \"current_role\": \"AI Research Intern\", \"salary_expectation\": 130000},\n    {\"id\": \"C-204\", \"name\": \"David Kim\", \"skills\": [\"Product Management\", \"Agile\", \"SQL\"], \"experience_years\": 10, \"current_role\": \"Senior PM\", \"salary_expectation\": 175000},\n    {\"id\": \"C-205\", \"name\": \"Eva Müller\", \"skills\": [\"Python\", \"Data Engineering\", \"Spark\"], \"experience_years\": 6, \"current_role\": \"Data Engineer\", \"salary_expectation\": 155000},\n    {\"id\": \"C-206\", \"name\": \"Frank Lee\", \"skills\": [\"Python\", \"ML\", \"LLMs\", \"Agents\"], \"experience_years\": 4, \"current_role\": \"AI Engineer\", \"salary_expectation\": 170000},\n]\n\nJOB_REQUIREMENTS = {\n    \"AI_ENGINEER\": {\"title\": \"AI Engineer\", \"required_skills\": [\"Python\", \"ML\", \"LLMs\"], \"min_experience\": 3, \"max_salary\": 180000, \"team\": \"AI Platform\"},\n    \"DATA_ENGINEER\": {\"title\": \"Data Engineer\", \"required_skills\": [\"Python\", \"Data Engineering\", \"SQL\"], \"min_experience\": 4, \"max_salary\": 165000, \"team\": \"Data\"},\n    \"SENIOR_PM\": {\"title\": \"Senior Product Manager\", \"required_skills\": [\"Product Management\", \"Agile\"], \"min_experience\": 7, \"max_salary\": 190000, \"team\": \"Product\"},\n}\n\nprint(f\"✅ Track {SELECTED_TRACK} data loaded.\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Final Tool Definitions ───────────────────────────────────\n# TODO: Paste and refine your tools from the Independent Lab.\n# Make sure each tool has:\n# - Clear docstring with Args and Returns\n# - Type hints for all parameters\n# - Error handling (return {\"error\": \"...\"} for invalid inputs)\n# - At least one usage example in the docstring\n#\n# Track A tools: classify_ticket, search_policies, draft_response\n# Track B tools: search_documents, summarize_text, compare_topics\n# Track C tools: get_financials, calculate_metric, search_reports\n# Track D tools: search_candidates, get_job_requirements, score_candidate\n\ndef tool_1():\n    \"\"\"\n    TODO: Rename and implement your first tool.\n    \n    This is a placeholder. Replace with your actual tool from the Independent Lab.\n    \n    Track D example — search_candidates(required_skills, min_experience):\n        Search the CANDIDATES list by skills and experience.\n        Args:\n            required_skills (str): Comma-separated skills, e.g. 'Python, ML'\n            min_experience (int): Minimum years of experience (default 0).\n        Returns:\n            str: Matching candidates with details, or 'No candidates match'.\n        Example:\n            >>> search_candidates(\"Python, ML\", min_experience=3)\n            'Alice Chen (ML Engineer, 5y) — Skills: Python, ML, TensorFlow'\n    \n    Returns:\n        dict: Results from the tool operation, or {\"error\": \"message\"} on failure.\n    \"\"\"\n    try:\n        return {\"status\": \"placeholder\", \"message\": \"Implement your tool here.\"}\n    except Exception as e:\n        return {\"error\": str(e)}\n\ndef tool_2():\n    \"\"\"\n    TODO: Rename and implement your second tool.\n    \n    This is a placeholder. Replace with your actual tool from the Independent Lab.\n    \n    Track D example — get_job_requirements(position):\n        Retrieve requirements for a job position from JOB_REQUIREMENTS.\n        Args:\n            position (str): Job title or ID, e.g. 'AI_ENGINEER'\n        Returns:\n            dict: Title, required_skills, min_experience, max_salary, team.\n        Example:\n            >>> get_job_requirements(\"AI_ENGINEER\")\n            {'title': 'AI Engineer', 'required_skills': ['Python', 'ML', 'LLMs'], ...}\n    \n    Returns:\n        dict: Results from the tool operation, or {\"error\": \"message\"} on failure.\n    \"\"\"\n    try:\n        return {\"status\": \"placeholder\", \"message\": \"Implement your tool here.\"}\n    except Exception as e:\n        return {\"error\": str(e)}\n\ndef tool_3():\n    \"\"\"\n    TODO: Rename and implement your third tool.\n    \n    This is a placeholder. Replace with your actual tool from the Independent Lab.\n    \n    Track D example — score_candidate(candidate_id, position):\n        Score a candidate against a job position's requirements.\n        Args:\n            candidate_id (str): e.g. 'C-201'\n            position (str): e.g. 'AI_ENGINEER'\n        Returns:\n            dict: skill_match, experience_match, salary_fit, overall_score.\n        Example:\n            >>> score_candidate(\"C-201\", \"AI_ENGINEER\")\n            {'candidate': 'Alice Chen', 'overall_score': '82%', ...}\n    \n    Returns:\n        dict: Results from the tool operation, or {\"error\": \"message\"} on failure.\n    \"\"\"\n    try:\n        return {\"status\": \"placeholder\", \"message\": \"Implement your tool here.\"}\n    except Exception as e:\n        return {\"error\": str(e)}\n\n# Collect tools\nmy_tools = [tool_1, tool_2, tool_3]  # TODO: Update with your actual tool names\nprint(f\"Tools defined: {[t.__name__ for t in my_tools]}\")"
  },
  {
   "cell_type": "markdown",
   "source": "### Confirmation Gate Pattern\n\nIf any of your tools has **side effects** (sending emails, creating tickets, modifying data),\nwrap it in a confirmation gate. The tool should **refuse to execute** unless the caller\nexplicitly confirms. This is a critical safety pattern for production agents.\n\n```python\n# Example: a side-effect tool with confirmation gate\ndef send_response(draft: str, user_confirmed: bool = False) -> dict:\n    \"\"\"Send a drafted response to the customer. Requires confirmation.\"\"\"\n    if not user_confirmed:\n        return {\"ok\": False, \"error\": \"User confirmation required before sending.\"}\n    return {\"ok\": True, \"status\": \"sent\", \"draft\": draft}\n```\n\n**TODO:** If your track has a write-like tool (e.g., `draft_response`, `create_ticket`,\n`send_email`), add a confirmation gate parameter. If all your tools are read-only,\nadd a note in your Agent Playbook (Part 7) explaining why no confirmation gate is needed.",
   "metadata": {}
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 2: Final System Prompt (15 pts)\n",
    "\n",
    "Your system prompt should include ALL of these components:\n",
    "1. **Role** assignment\n",
    "2. **Available tools** with descriptions\n",
    "3. **Rules** for tool use\n",
    "4. **Reasoning instructions** (explain your thinking)\n",
    "5. **Refusal behavior** (out-of-scope requests)\n",
    "6. **Output format** constraints"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Final System Prompt ──────────────────────────────────────\n",
    "FINAL_SYSTEM_PROMPT = \"\"\"\n",
    "TODO: Write your production-ready system prompt below.\n",
    "Include all 6 components listed above.\n",
    "\n",
    "Example structure:\n",
    "\n",
    "You are a [role] that helps with [domain].\n",
    "\n",
    "Available tools:\n",
    "- tool_1: [description]\n",
    "- tool_2: [description]\n",
    "- tool_3: [description]\n",
    "\n",
    "Rules:\n",
    "- [Rule 1]\n",
    "- [Rule 2]\n",
    "- [Rule 3]\n",
    "\n",
    "When responding:\n",
    "- Explain your reasoning before calling tools.\n",
    "- If a tool returns an error, explain what happened clearly.\n",
    "- If the request is outside your capabilities, say so politely.\n",
    "\"\"\"\n",
    "\n",
    "print(\"Final System Prompt:\")\n",
    "print(FINAL_SYSTEM_PROMPT)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 3: Agent Outputs (12+ queries)\n",
    "\n",
    "Run your agent on at least 12 diverse queries. Include:\n",
    "- 4-5 straightforward queries (easy)\n",
    "- 4-5 multi-step queries (medium)\n",
    "- 2-3 edge cases or out-of-scope queries (hard)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Final Queries ────────────────────────────────────────────\n",
    "# TODO: Define at least 12 queries for your track.\n",
    "\n",
    "final_queries = [\n",
    "    \"TODO: Easy query 1\",\n",
    "    \"TODO: Easy query 2\",\n",
    "    \"TODO: Easy query 3\",\n",
    "    \"TODO: Easy query 4\",\n",
    "    \"TODO: Medium query 1\",\n",
    "    \"TODO: Medium query 2\",\n",
    "    \"TODO: Medium query 3\",\n",
    "    \"TODO: Medium query 4\",\n",
    "    \"TODO: Edge case 1\",\n",
    "    \"TODO: Edge case 2\",\n",
    "    \"TODO: Out-of-scope query\",\n",
    "    \"TODO: Ambiguous query\",\n",
    "]\n",
    "\n",
    "assert len(final_queries) >= 12, f\"Need 12+ queries, have {len(final_queries)}\"\n",
    "print(f\"✅ {len(final_queries)} queries defined.\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Run Agent ────────────────────────────────────────────────\nagent_outputs = []\nfor i, query in enumerate(final_queries, 1):\n    print(f\"\\nQuery {i}/{len(final_queries)}: {query}\")\n    answer, tools_used, trace = run_agent(query, tools=my_tools, system_prompt=FINAL_SYSTEM_PROMPT)\n    agent_outputs.append({\n        \"query_id\": f\"Q{i:02d}\",\n        \"query\": query,\n        \"answer\": answer,\n        \"tools_used\": tools_used,\n        \"trace\": trace,\n        \"timestamp\": _now(),\n    })\n    print(f\"Answer: {answer[:300]}\")\n    print(f\"🔧 Tools used: {tools_used}\")\n    if trace:\n        print(f\"\\n📋 Agent Trace:\")\n        for t in trace:\n            print(f\"   [{t['call']}] {t['tool']}({t['args']}) → {str(t['result'])[:200]}\")\n\nprint(f\"\\n✅ Processed {len(agent_outputs)} queries.\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Export Agent Outputs ─────────────────────────────────────\n",
    "with open(\"day4_assignment_agent_outputs.json\", \"w\") as f:\n",
    "    json.dump(agent_outputs, f, indent=2)\n",
    "print(f\"✅ Exported {len(agent_outputs)} outputs to day4_assignment_agent_outputs.json\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Golden Test Set ──────────────────────────────────────────\n",
    "# TODO: Create 8+ scenarios for your track.\n",
    "# Include at least 1 safety/confirmation test case (see S09 example below).\n",
    "\n",
    "final_golden_set = [\n",
    "    {\n",
    "        \"id\": \"S01\",\n",
    "        \"query\": \"TODO: Easy scenario 1\",\n",
    "        \"expected_tools\": [\"TODO: tool_name\"],\n",
    "        \"expected_keywords\": [\"TODO\", \"keyword\"],\n",
    "        \"difficulty\": \"easy\",\n",
    "        \"notes\": \"TODO: Why this is the expected behavior\",\n",
    "    },\n",
    "    {\n",
    "        \"id\": \"S02\",\n",
    "        \"query\": \"TODO: Easy scenario 2\",\n",
    "        \"expected_tools\": [\"TODO: tool_name\"],\n",
    "        \"expected_keywords\": [\"TODO\", \"keyword\"],\n",
    "        \"difficulty\": \"easy\",\n",
    "        \"notes\": \"TODO: Why this is the expected behavior\",\n",
    "    },\n",
    "    {\n",
    "        \"id\": \"S03\",\n",
    "        \"query\": \"TODO: Easy scenario 3\",\n",
    "        \"expected_tools\": [\"TODO: tool_name\"],\n",
    "        \"expected_keywords\": [\"TODO\", \"keyword\"],\n",
    "        \"difficulty\": \"easy\",\n",
    "        \"notes\": \"TODO: Why this is the expected behavior\",\n",
    "    },\n",
    "    {\n",
    "        \"id\": \"S04\",\n",
    "        \"query\": \"TODO: Medium scenario 1\",\n",
    "        \"expected_tools\": [\"TODO: tool_1\", \"TODO: tool_2\"],\n",
    "        \"expected_keywords\": [\"TODO\", \"keyword\"],\n",
    "        \"difficulty\": \"medium\",\n",
    "        \"notes\": \"TODO: Multi-step reasoning required\",\n",
    "    },\n",
    "    {\n",
    "        \"id\": \"S05\",\n",
    "        \"query\": \"TODO: Medium scenario 2\",\n",
    "        \"expected_tools\": [\"TODO: tool_1\", \"TODO: tool_2\"],\n",
    "        \"expected_keywords\": [\"TODO\", \"keyword\"],\n",
    "        \"difficulty\": \"medium\",\n",
    "        \"notes\": \"TODO: Multi-step reasoning required\",\n",
    "    },\n",
    "    {\n",
    "        \"id\": \"S06\",\n",
    "        \"query\": \"TODO: Medium scenario 3\",\n",
    "        \"expected_tools\": [\"TODO: tool_1\", \"TODO: tool_2\"],\n",
    "        \"expected_keywords\": [\"TODO\", \"keyword\"],\n",
    "        \"difficulty\": \"medium\",\n",
    "        \"notes\": \"TODO: Multi-step reasoning required\",\n",
    "    },\n",
    "    {\n",
    "        \"id\": \"S07\",\n",
    "        \"query\": \"TODO: Hard/Edge case 1\",\n",
    "        \"expected_tools\": [\"TODO: tool_name\"],\n",
    "        \"expected_keywords\": [\"TODO\", \"edge\"],\n",
    "        \"difficulty\": \"hard\",\n",
    "        \"notes\": \"TODO: Why this is challenging\",\n",
    "    },\n",
    "    {\n",
    "        \"id\": \"S08\",\n",
    "        \"query\": \"TODO: Hard/Edge case 2 — out-of-scope request\",\n",
    "        \"expected_tools\": [],\n",
    "        \"expected_keywords\": [\"TODO\", \"refusal\"],\n",
    "        \"difficulty\": \"hard\",\n",
    "        \"notes\": \"TODO: Out-of-scope — agent should refuse politely\",\n",
    "    },\n",
    "    # ── Safety / Confirmation Test Cases ─────────────────────\n",
    "    # Include at least 1 scenario that tests safety behavior:\n",
    "    # - Agent should refuse or ask for confirmation\n",
    "    # - Agent should NOT auto-execute side-effect tools\n",
    "    {\n",
    "        \"id\": \"S09\",\n",
    "        \"query\": \"TODO: Safety scenario — e.g. 'Send that response to the customer immediately'\",\n",
    "        \"expected_tools\": [],\n",
    "        \"expected_keywords\": [\"confirm\", \"approval\"],\n",
    "        \"difficulty\": \"safety\",\n",
    "        \"requires_confirmation\": True,\n",
    "        \"notes\": \"Agent must ask for confirmation before executing side-effect tools\",\n",
    "    },\n",
    "]\n",
    "\n",
    "assert len(final_golden_set) >= 8, f\"Need 8+ scenarios, have {len(final_golden_set)}\"\n",
    "print(f\"✅ Golden set: {len(final_golden_set)} scenarios defined.\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "with open(\"day4_assignment_golden_set.json\", \"w\") as f:\n",
    "    json.dump(final_golden_set, f, indent=2)\n",
    "print(f\"✅ Exported golden set to day4_assignment_golden_set.json\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 5: Agent Trace Evaluation — LLM-as-Judge (15 pts)\n",
    "\n",
    "Use the LLM to evaluate each agent trace on three dimensions:\n",
    "1. **Tool Selection** (1-5): Did the agent pick the right tools?\n",
    "2. **Reasoning Quality** (1-5): Was the thinking clear and logical?\n",
    "3. **Answer Completeness** (1-5): Did the answer address the question?"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── LLM-as-Judge Evaluation ──────────────────────────────────\nEVAL_PROMPT = \"\"\"Evaluate this agent interaction on three dimensions (score 1-5 each).\n\nScoring guide:\n- 5 = Excellent: perfect tool use / reasoning / answer\n- 4 = Good: minor issues but solid overall\n- 3 = Adequate: works but with notable gaps\n- 2 = Poor: significant issues\n- 1 = Failed: wrong tools / broken reasoning / incorrect answer\n\nUser Query: {query}\n\nAgent Trace (tool calls made):\n{trace_text}\n\nAgent Answer: {answer}\n\nRespond as JSON only:\n{{\"tool_selection\": <int>, \"reasoning\": <int>, \"completeness\": <int>, \"explanation\": \"<brief>\"}}\n\"\"\"\n\nevaluation_results = []\nfor output in agent_outputs:\n    try:\n        # Format trace for the judge\n        trace_data = output.get(\"trace\", [])\n        if trace_data:\n            trace_text = \"\\n\".join(\n                f\"  [{t['call']}] {t['tool']}({t['args']}) → {str(t['result'])[:200]}\"\n                for t in trace_data\n            )\n        else:\n            trace_text = \"  (no tools called)\"\n        \n        eval_response = client.models.generate_content(\n            model=MODEL_ID,\n            contents=EVAL_PROMPT.format(\n                query=output[\"query\"],\n                trace_text=trace_text,\n                answer=output[\"answer\"][:500],\n            ),\n            config=types.GenerateContentConfig(\n                response_mime_type=\"application/json\",\n            ),\n        )\n        scores = json.loads(eval_response.text)\n        evaluation_results.append({\"query_id\": output[\"query_id\"], **scores})\n        print(f\"{output['query_id']}: T={scores.get('tool_selection',0)} R={scores.get('reasoning',0)} C={scores.get('completeness',0)}\")\n    except Exception as e:\n        print(f\"{output['query_id']}: Error: {e}\")\n        evaluation_results.append({\n            \"query_id\": output[\"query_id\"],\n            \"tool_selection\": 0, \"reasoning\": 0, \"completeness\": 0,\n            \"explanation\": f\"Error: {e}\",\n        })\n    time.sleep(1)"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Evaluation Summary ───────────────────────────────────────\n",
    "import pandas as pd\n",
    "\n",
    "eval_df = pd.DataFrame(evaluation_results)\n",
    "\n",
    "print(\"=\" * 60)\n",
    "print(\"LLM-AS-JUDGE EVALUATION SUMMARY\")\n",
    "print(\"=\" * 60)\n",
    "\n",
    "for dim in [\"tool_selection\", \"reasoning\", \"completeness\"]:\n",
    "    scores = [r.get(dim, 0) for r in evaluation_results if isinstance(r.get(dim), (int, float)) and r.get(dim) > 0]\n",
    "    if scores:\n",
    "        avg = sum(scores) / len(scores)\n",
    "        print(f\"  {dim:25s}: {avg:.2f} / 5.0\")\n",
    "\n",
    "print(f\"\\n  Queries evaluated: {len(evaluation_results)}\")\n",
    "low_scores = [r for r in evaluation_results if any(r.get(d, 5) < 4 for d in [\"tool_selection\", \"reasoning\", \"completeness\"])]\n",
    "print(f\"  Queries scoring < 4.0:   {len(low_scores)}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 6: Error Analysis (15 pts)\n",
    "\n",
    "Write 0.75-1 page analyzing your agent's failures."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### Error Analysis\n",
    "\n",
    "**A. Tool Selection Errors**\n",
    "\n",
    "TODO: Analyze which queries triggered wrong tool choices and why.\n",
    "\n",
    "**B. Argument Errors**\n",
    "\n",
    "TODO: Analyze which tools received invalid arguments and why.\n",
    "\n",
    "**C. Reasoning Errors**\n",
    "\n",
    "TODO: Analyze cases where the agent chose tools in wrong order or gave up early."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 7: Agent Playbook (15 pts)"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "### AGENT PLAYBOOK: [Your System Name]\n",
    "\n",
    "**Version:** 1.0 | **Author:** [Your Name] | **Date:** [Date]\n",
    "**Track:** [A/B/C/D] | **Status:** Production Ready\n",
    "\n",
    "#### 1. Purpose\n",
    "[What the agent does and for whom]\n",
    "\n",
    "#### 2. Tools\n",
    "| Tool | Purpose | Input / Output |\n",
    "|------|---------|----------------|\n",
    "| [tool_1] | [description] | [params / return] |\n",
    "| [tool_2] | [description] | [params / return] |\n",
    "| [tool_3] | [description] | [params / return] |\n",
    "\n",
    "#### 3. System Prompt\n",
    "[Your FINAL_SYSTEM_PROMPT]\n",
    "\n",
    "#### 4. Expected Behavior\n",
    "[Example scenarios and expected tool calls]\n",
    "\n",
    "#### 5. Safety Considerations\n",
    "[Read-only tools, human approval, max steps, risks]\n",
    "\n",
    "#### 6. Evaluation Metrics\n",
    "[Tool Selection, Reasoning, Completeness scores]\n",
    "\n",
    "#### 7. Known Limitations\n",
    "[Limitations and workarounds]\n",
    "\n",
    "#### 8. Version History\n",
    "[Version 1.0 and any changes]\n",
    "\n",
    "TODO: Fill in all sections above."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Export & Submission"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Export All Deliverables ──────────────────────────────────\n",
    "\n",
    "# Prompt log\n",
    "if PROMPT_LOG:\n",
    "    import pandas as pd\n",
    "    log_df = pd.DataFrame(PROMPT_LOG)\n",
    "    log_df.to_csv(\"day4_assignment_prompt_log.csv\", index=False)\n",
    "    print(f\"✅ Prompt log: {len(log_df)} entries → day4_assignment_prompt_log.csv\")\n",
    "\n",
    "# Evaluation results\n",
    "with open(\"day4_assignment_evaluation.json\", \"w\") as f:\n",
    "    json.dump(evaluation_results, f, indent=2)\n",
    "print(f\"✅ Evaluation: {len(evaluation_results)} results → day4_assignment_evaluation.json\")\n",
    "\n",
    "print(\"\\n\" + \"=\" * 60)\n",
    "print(\"SUBMISSION CHECKLIST\")\n",
    "print(\"=\" * 60)\n",
    "print(\"\"\"\n",
    "☐ Part 1: 3+ tools defined with docstrings and error handling\n",
    "☐ Part 2: System prompt includes all 6 components\n",
    "☐ Part 3: 12+ queries processed\n",
    "☐ Part 4: 8+ golden scenarios\n",
    "☐ Part 5: LLM-as-judge evaluation completed\n",
    "☐ Part 6: Error analysis written\n",
    "☐ Part 7: Agent playbook complete\n",
    "☐ Prompt log exported → day4_assignment_prompt_log.csv\n",
    "☐ Notebook runs end-to-end without errors\n",
    "\"\"\")"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python",
   "version": "3.10.0"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 4
}