{
 "cells": [
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "# 🚀 Day 4: Guided Lab — Your First AI Agent with Function Calling\n",
    "\n",
    "## Learning Objectives\n",
    "\n",
    "By the end of this lab, you will be able to:\n",
    "- Define Python functions as tools for the Gemini API\n",
    "- Use function calling to let the model invoke your tools\n",
    "- Control tool behavior with AUTO, ANY, and NONE modes\n",
    "- Build a manual agent loop with conversation history\n",
    "- Create a multi-tool agent combining CRM, calculator, and search\n",
    "- Evaluate agent performance with test cases"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 0: Setup & Connection Test (10 min)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "!pip install -q -U google-genai"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "import os, json, time\n",
    "from datetime import datetime, timezone\n",
    "from google import genai\n",
    "from google.genai import types\n",
    "\n",
    "# ── API Key ──────────────────────────────────────────────\n",
    "try:\n",
    "    from google.colab import userdata\n",
    "    os.environ[\"GEMINI_API_KEY\"] = userdata.get(\"GEMINI_API_KEY\")\n",
    "except Exception:\n",
    "    pass\n",
    "\n",
    "if not os.environ.get(\"GEMINI_API_KEY\"):\n",
    "    import getpass\n",
    "    os.environ[\"GEMINI_API_KEY\"] = getpass.getpass(\"Paste your GEMINI_API_KEY: \")\n",
    "\n",
    "client = genai.Client(api_key=os.environ[\"GEMINI_API_KEY\"])\n",
    "MODEL_ID = \"gemini-2.5-flash-lite\""
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Logging Infrastructure ───────────────────────────────\n",
    "PROMPT_LOG = []\n",
    "\n",
    "def _now():\n",
    "    return datetime.now(timezone.utc).isoformat(timespec=\"seconds\").replace(\"+00:00\", \"Z\")\n",
    "\n",
    "def log_interaction(role, content, label=None):\n",
    "    \"\"\"Log an agent interaction.\"\"\"\n",
    "    entry = {\n",
    "        \"ts\": _now(),\n",
    "        \"role\": role,\n",
    "        \"content\": content if isinstance(content, str) else json.dumps(content),\n",
    "        \"label\": label or \"\",\n",
    "    }\n",
    "    PROMPT_LOG.append(entry)\n",
    "    return entry\n",
    "\n",
    "def show_log(n=10):\n",
    "    \"\"\"Display the last n log entries.\"\"\"\n",
    "    import pandas as pd\n",
    "    if not PROMPT_LOG:\n",
    "        print(\"No interactions logged yet.\")\n",
    "        return\n",
    "    df = pd.DataFrame(PROMPT_LOG[-n:])\n",
    "    from IPython.display import display\n",
    "    display(df)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Test Connection ──────────────────────────────────────\n",
    "response = client.models.generate_content(\n",
    "    model=MODEL_ID,\n",
    "    contents=\"Say 'Hello from Gemini!' in one sentence.\",\n",
    ")\n",
    "print(f\"✅ Connected to {MODEL_ID}\")\n",
    "print(f\"   Response: {response.text}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 1: Your First Tool — Single Function Calling (10 min)\n",
    "\n",
    "In function calling, you:\n",
    "1. **Define a Python function** with type hints and a docstring\n",
    "2. **Pass it as a tool** to `generate_content()`\n",
    "3. **The model decides** whether and how to call it\n",
    "\n",
    "> The model never executes your code — it only *requests* a call. Your code runs the function and returns the result."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Define a Tool ────────────────────────────────────────\n",
    "def get_weather(location: str) -> dict:\n",
    "    \"\"\"Get the current weather for a location.\n",
    "\n",
    "    Args:\n",
    "        location: City name, e.g. 'New York', 'London', 'Tokyo'\n",
    "\n",
    "    Returns:\n",
    "        A dict with temperature, condition, and humidity.\n",
    "    \"\"\"\n",
    "    weather_db = {\n",
    "        \"New York\": {\"temp_celsius\": 22, \"condition\": \"Sunny\", \"humidity\": 65},\n",
    "        \"London\":   {\"temp_celsius\": 13, \"condition\": \"Rainy\", \"humidity\": 85},\n",
    "        \"Tokyo\":    {\"temp_celsius\": 20, \"condition\": \"Cloudy\", \"humidity\": 70},\n",
    "    }\n",
    "    result = weather_db.get(location)\n",
    "    if result:\n",
    "        return {**result, \"location\": location, \"status\": \"ok\"}\n",
    "    return {\"error\": f\"No weather data for '{location}'\", \"status\": \"not_found\"}\n",
    "\n",
    "# Test the function directly (no LLM yet)\n",
    "print(get_weather(\"London\"))\n",
    "print(get_weather(\"Paris\"))   # Not in our database"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Call the Model with a Tool ───────────────────────────\n",
    "response = client.models.generate_content(\n",
    "    model=MODEL_ID,\n",
    "    contents=\"What's the weather like in London today?\",\n",
    "    config=types.GenerateContentConfig(\n",
    "        tools=[get_weather],\n",
    "    ),\n",
    ")\n",
    "\n",
    "print(\"Response text:\", response.text)"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Inspect the Response Parts ───────────────────────────\n# With automatic function calling (the default), the SDK runs\n# the entire loop behind the scenes: it calls get_weather,\n# feeds the result back, and returns ONLY the final text.\n# That's why we see just 1 part — the intermediate function_call\n# and function_response are consumed internally.\n\nprint(\"With automatic function calling (default):\")\nprint(f\"  Number of parts: {len(response.candidates[0].content.parts)}\")\nfor i, part in enumerate(response.candidates[0].content.parts):\n    print(f\"  Part {i}: Text = {part.text[:200]}\")\n\nprint(\"\\n--- Now let's DISABLE automatic calling to see what really happens ---\\n\")\n\n# Same request, but this time the SDK stops and hands us the function_call\nraw_response = client.models.generate_content(\n    model=MODEL_ID,\n    contents=\"What's the weather like in London today?\",\n    config=types.GenerateContentConfig(\n        tools=[get_weather],\n        automatic_function_calling=types.AutomaticFunctionCallingConfig(\n            disable=True  # SDK returns the function_call instead of executing it\n        ),\n    ),\n)\n\nprint(\"With automatic function calling DISABLED:\")\nprint(f\"  Number of parts: {len(raw_response.candidates[0].content.parts)}\")\nfor i, part in enumerate(raw_response.candidates[0].content.parts):\n    print(f\"\\n  Part {i}:\")\n    if part.text:\n        print(f\"    Text: {part.text[:200]}\")\n    if part.function_call:\n        print(f\"    Function call: {part.function_call.name}\")\n        print(f\"    Arguments: {dict(part.function_call.args)}\")\n        print(\"    → The model REQUESTED this call. It did NOT execute it.\")\n        print(\"      In Part 3, we'll build a loop that handles this ourselves.\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 2: Controlling Tool Behavior (10 min)\n",
    "\n",
    "Tool calling **modes** control how the model interacts with tools:\n",
    "\n",
    "| Mode | Behavior | Use When |\n",
    "|------|----------|----------|\n",
    "| `AUTO` (default) | Model decides whether to call a tool or respond directly | General use |\n",
    "| `ANY` | Model **must** call a tool | Force structured output or routing |\n",
    "| `NONE` | Model cannot call any tools | Pure text response |"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── AUTO Mode: Model Decides ─────────────────────────────\n",
    "print(\"=== AUTO Mode ===\")\n",
    "print(\"Query about a city IN our database:\")\n",
    "r1 = client.models.generate_content(\n",
    "    model=MODEL_ID,\n",
    "    contents=\"What's the weather in Tokyo?\",\n",
    "    config=types.GenerateContentConfig(\n",
    "        tools=[get_weather],\n",
    "        tool_config=types.ToolConfig(\n",
    "            function_calling_config=types.FunctionCallingConfig(mode=\"AUTO\")\n",
    "        ),\n",
    "    ),\n",
    ")\n",
    "print(f\"  → {r1.text}\\n\")\n",
    "\n",
    "print(\"Query about something unrelated to weather:\")\n",
    "r2 = client.models.generate_content(\n",
    "    model=MODEL_ID,\n",
    "    contents=\"What is the capital of France?\",\n",
    "    config=types.GenerateContentConfig(\n",
    "        tools=[get_weather],\n",
    "        tool_config=types.ToolConfig(\n",
    "            function_calling_config=types.FunctionCallingConfig(mode=\"AUTO\")\n",
    "        ),\n",
    "    ),\n",
    ")\n",
    "print(f\"  → {r2.text}\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── ANY Mode: Must Call a Tool ────────────────────────────\n",
    "print(\"=== ANY Mode ===\")\n",
    "print(\"The model is FORCED to call get_weather, even for an unrelated question:\\n\")\n",
    "\n",
    "r3 = client.models.generate_content(\n",
    "    model=MODEL_ID,\n",
    "    contents=\"What is the capital of France?\",\n",
    "    config=types.GenerateContentConfig(\n",
    "        tools=[get_weather],\n",
    "        tool_config=types.ToolConfig(\n",
    "            function_calling_config=types.FunctionCallingConfig(mode=\"ANY\")\n",
    "        ),\n",
    "        # Disable auto-execution so we can see the raw function call\n",
    "        automatic_function_calling=types.AutomaticFunctionCallingConfig(disable=True),\n",
    "    ),\n",
    ")\n",
    "\n",
    "for part in r3.candidates[0].content.parts:\n",
    "    if part.function_call:\n",
    "        print(f\"  Tool called: {part.function_call.name}\")\n",
    "        print(f\"  Arguments:   {dict(part.function_call.args)}\")"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── NONE Mode: No Tools Allowed ──────────────────────────\n",
    "print(\"=== NONE Mode ===\")\n",
    "print(\"Even for a weather question, the model must answer from its own knowledge:\\n\")\n",
    "\n",
    "r4 = client.models.generate_content(\n",
    "    model=MODEL_ID,\n",
    "    contents=\"What's the weather in London?\",\n",
    "    config=types.GenerateContentConfig(\n",
    "        tools=[get_weather],\n",
    "        tool_config=types.ToolConfig(\n",
    "            function_calling_config=types.FunctionCallingConfig(mode=\"NONE\")\n",
    "        ),\n",
    "    ),\n",
    ")\n",
    "print(f\"  → {r4.text}\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 3: Building the Agent Loop (15 min)\n",
    "\n",
    "Automatic function calling is convenient, but for real agents you want **full control**. The manual agent loop:\n",
    "\n",
    "1. Send user message to the model\n",
    "2. If the model requests a tool call → execute it, send the result back\n",
    "3. Repeat until the model gives a final text answer\n",
    "4. Stop after `max_steps` as a safety limit\n",
    "\n",
    "This is the **core pattern** for all agent architectures (ReAct, Plan-and-Execute, etc.)."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── The Agent Loop ───────────────────────────────────────\ndef run_agent(user_message, tools, system_prompt=None, max_steps=10):\n    \"\"\"A manual agent loop with full visibility.\n\n    Args:\n        user_message: The user's request.\n        tools: List of Python functions to use as tools.\n        system_prompt: Optional system instruction for the agent.\n        max_steps: Maximum number of reasoning steps (safety limit).\n\n    Returns:\n        A tuple of (final_text, tools_called, trace) where:\n        - final_text: The agent's final response string\n        - tools_called: List of tool names invoked during the run\n        - trace: List of dicts with 'call', 'tool', 'args', 'result' per tool call\n    \"\"\"\n    tool_map = {fn.__name__: fn for fn in tools}\n    call_count = 0        # Track total tool calls across all steps\n    tools_called = []     # Record which tools were actually used\n    trace = []            # Structured log: tool, args, result per call\n\n    # Build initial contents\n    contents = []\n    if system_prompt:\n        contents.append(types.Content(\n            role=\"user\",\n            parts=[types.Part(text=f\"System: {system_prompt}\\n\\nUser: {user_message}\")]\n        ))\n    else:\n        contents.append(types.Content(\n            role=\"user\",\n            parts=[types.Part(text=user_message)]\n        ))\n\n    log_interaction(\"user\", user_message, label=\"agent_input\")\n\n    for step in range(max_steps):\n        response = client.models.generate_content(\n            model=MODEL_ID,\n            contents=contents,\n            config=types.GenerateContentConfig(\n                tools=tools,\n                # Tool calling mode defaults to AUTO — the model\n                # reasons about whether to use tools on each turn.\n                automatic_function_calling=types.AutomaticFunctionCallingConfig(\n                    disable=True  # Model still reasons about tools —\n                    # but the SDK won't execute them automatically.\n                    # Instead it returns the function_call to us,\n                    # and WE run the function below.\n                ),\n            ),\n        )\n\n        # Guard: the model may return an empty response\n        # (e.g. rate limit, safety filter, transient error)\n        parts = response.parts or []\n        if not parts:\n            print(f\"  Step {step+1}: ⚠️ Empty response from model — retrying...\")\n            continue\n\n        # Add model response to history\n        contents.append(types.Content(role=\"model\", parts=parts))\n\n        # Check for function calls\n        function_results = []\n        for part in parts:\n            if part.function_call:\n                call_count += 1\n                name = part.function_call.name\n                args = dict(part.function_call.args)\n                tools_called.append(name)\n                print(f\"  Tool call {call_count}: 🔧 {name}({args})\")\n\n                # Execute the function\n                try:\n                    result = tool_map[name](**args)\n                except Exception as e:\n                    result = {\"error\": str(e)}\n\n                print(f\"           → {result}\")\n                trace.append({\n                    \"call\": call_count,\n                    \"tool\": name,\n                    \"args\": args,\n                    \"result\": result if isinstance(result, str) else json.dumps(result),\n                })\n                log_interaction(\"tool\", f\"{name}({args}) → {result}\", label=\"tool_call\")\n\n                function_results.append(\n                    types.Part(\n                        function_response=types.FunctionResponse(\n                            name=name,\n                            response={\"result\": result},\n                        )\n                    )\n                )\n\n        if function_results:\n            contents.append(types.Content(role=\"user\", parts=function_results))\n        else:\n            # No function calls → model is done\n            final_text = response.text or \"(no text response)\"\n            log_interaction(\"agent\", final_text, label=\"agent_output\")\n            return final_text, tools_called, trace\n\n    return \"⚠️ Agent reached maximum steps without completing.\", tools_called, trace\n\nprint(\"✅ run_agent() defined.\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Test: Simple Query ───────────────────────────────────\nprint(\"Query: 'What's the weather in London?'\\n\")\nanswer, tools_used, trace = run_agent(\n    \"What's the weather in London?\",\n    tools=[get_weather],\n)\nprint(f\"\\n📝 Final Answer:\\n{answer}\")\nprint(f\"🔧 Tools used: {tools_used}\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Test: Query Requiring Reasoning ──────────────────────\nprint(\"Query: 'Compare weather in New York and Tokyo. Where's better for a picnic?'\\n\")\nanswer, tools_used, trace = run_agent(\n    \"Compare the weather in New York and Tokyo. \"\n    \"Which city would you recommend for an outdoor picnic today?\",\n    tools=[get_weather],\n)\nprint(f\"\\n📝 Final Answer:\\n{answer}\")\nprint(f\"🔧 Tools used: {tools_used}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 4: Multi-Tool Agent (15 min)\n",
    "\n",
    "Real agents combine multiple tools. The model decides **which** tool to call (or whether to call one at all) based on the tool descriptions.\n",
    "\n",
    "We'll add two more tools:\n",
    "- **`get_customer_info`**: Look up customer account details (simulating a CRM)\n",
    "- **`calculate`**: Evaluate mathematical expressions"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Tool: Customer Lookup ────────────────────────────────\n",
    "def get_customer_info(customer_id: str) -> dict:\n",
    "    \"\"\"Look up a customer's account details from the CRM.\n",
    "\n",
    "    Args:\n",
    "        customer_id: The unique customer ID, e.g. 'CUST-1234'\n",
    "\n",
    "    Returns:\n",
    "        A dict with name, plan, and monthly recurring revenue (MRR).\n",
    "    \"\"\"\n",
    "    crm_database = {\n",
    "        \"CUST-1234\": {\"name\": \"Acme Corp\",  \"plan\": \"Enterprise\", \"mrr\": 12000, \"industry\": \"Manufacturing\"},\n",
    "        \"CUST-5678\": {\"name\": \"TechStart\",  \"plan\": \"Growth\",     \"mrr\": 3500,  \"industry\": \"SaaS\"},\n",
    "        \"CUST-9012\": {\"name\": \"RetailMax\",  \"plan\": \"Starter\",    \"mrr\": 800,   \"industry\": \"Retail\"},\n",
    "    }\n",
    "    result = crm_database.get(customer_id)\n",
    "    if result:\n",
    "        return {**result, \"customer_id\": customer_id, \"status\": \"found\"}\n",
    "    return {\"error\": f\"Customer '{customer_id}' not found\", \"status\": \"not_found\"}\n",
    "\n",
    "print(get_customer_info(\"CUST-1234\"))"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Tool: Calculator ─────────────────────────────────────\nimport re\n\ndef calculate(expression: str) -> str:\n    \"\"\"Evaluate a mathematical expression using actual numbers and operators.\n\n    Args:\n        expression: A math expression with numeric values and operators only.\n                    Do NOT use variable names — substitute the actual numbers.\n                    Supports +, -, *, /, **, (). Example: '12000 * 12 * 0.85'\n\n    Returns:\n        The numeric result as a string, or an error message.\n    \"\"\"\n    # Check for variable names (letters that aren't part of numbers)\n    # Allow digits, operators, whitespace, parentheses, and decimal points\n    if re.search(r'[a-zA-Z_]', expression):\n        return (\n            f\"Error: expression contains variable names: '{expression}'. \"\n            f\"Please substitute the actual numeric values and try again.\"\n        )\n    try:\n        result = eval(expression)\n        return str(result)\n    except Exception as e:\n        return f\"Error: {e}. Use only numbers and operators (+, -, *, /, **, parentheses).\"\n\nprint(calculate(\"12000 * 12 * 0.85\"))\nprint(calculate(\"MRR * 12\"))  # Shows the error the model would see"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Multi-Tool Agent ─────────────────────────────────────\nprint(\"Query: 'What plan is CUST-1234 on? What would their annual cost be with 15% discount?'\\n\")\n\nanswer, tools_used, trace = run_agent(\n    \"What plan is customer CUST-1234 on? \"\n    \"What would their annual cost be if we gave them a 15% discount?\",\n    tools=[get_customer_info, calculate],\n)\nprint(f\"\\n📝 Final Answer:\\n{answer}\")\nprint(f\"🔧 Tools used: {tools_used}\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Complex Multi-Tool Query ─────────────────────────────\nprint(\"Query: 'Compare MRR of CUST-1234 and CUST-5678. What's combined annual revenue?'\\n\")\n\nanswer, tools_used, trace = run_agent(\n    \"Compare the MRR of customers CUST-1234 and CUST-5678. \"\n    \"Which one pays more? What's their combined annual revenue?\",\n    tools=[get_customer_info, calculate],\n)\nprint(f\"\\n📝 Final Answer:\\n{answer}\")\nprint(f\"🔧 Tools used: {tools_used}\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Agent with System Prompt ─────────────────────────────\nprint(\"Query: 'Tell me about CUST-9012 and check the weather in New York.'\\n\")\nprint(\"This time we add a system prompt to guide the agent's behavior.\\n\")\n\nanswer, tools_used, trace = run_agent(\n    \"Tell me about customer CUST-9012 and check the weather in their area (New York).\",\n    tools=[get_weather, get_customer_info, calculate],\n    system_prompt=(\n        \"You are a helpful business assistant. \"\n        \"Use tools to look up information. \"\n        \"Always cite the source of your data (tool name). \"\n        \"If a tool returns an error, tell the user clearly.\"\n    ),\n)\nprint(f\"\\n📝 Final Answer:\\n{answer}\")\nprint(f\"🔧 Tools used: {tools_used}\")"
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Part 5: Agent Evaluation Basics (15 min)\n",
    "\n",
    "How do you know your agent is working correctly? Build **test cases** with:\n",
    "- An input query\n",
    "- The tools you expect the agent to call\n",
    "- Keywords the answer should contain\n",
    "- A maximum step limit"
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Define Test Cases ────────────────────────────────────\n# Note: Keywords are checked against the model's FINAL ANSWER (case-insensitive).\n# The model rephrases tool results, so use words it's likely to say,\n# not the exact strings from the tool output.\n\ntest_cases = [\n    {\n        \"id\": \"T1\",\n        \"input\": \"What's the weather in New York?\",\n        \"expected_tools\": [\"get_weather\"],\n        \"expected_keywords\": [\"sunny\", \"22\"],\n        \"max_steps\": 3,\n        \"description\": \"Simple single-tool query\",\n    },\n    {\n        \"id\": \"T2\",\n        \"input\": \"What's the weather in Paris?\",\n        \"expected_tools\": [\"get_weather\"],\n        \"expected_keywords\": [\"paris\"],  # Model should mention Paris; it will rephrase the error\n        \"max_steps\": 3,\n        \"description\": \"Query for missing data — should handle gracefully\",\n    },\n    {\n        \"id\": \"T3\",\n        \"input\": \"What plan is customer CUST-5678 on?\",\n        \"expected_tools\": [\"get_customer_info\"],\n        \"expected_keywords\": [\"growth\"],\n        \"max_steps\": 3,\n        \"description\": \"CRM lookup\",\n    },\n    {\n        \"id\": \"T4\",\n        \"input\": \"How much would CUST-1234 pay annually with a 20% discount?\",\n        \"expected_tools\": [\"get_customer_info\", \"calculate\"],\n        \"expected_keywords\": [\"115200\", \"115,200\"],  # 12000 * 12 * 0.8 — model may format with comma\n        \"max_steps\": 5,\n        \"description\": \"Multi-tool: lookup + calculation (may fail — discuss why!)\",\n    },\n    {\n        \"id\": \"T5\",\n        \"input\": \"What is the meaning of life?\",\n        \"expected_tools\": [],\n        \"expected_keywords\": [],\n        \"max_steps\": 2,\n        \"description\": \"No tools needed — pure text response\",\n    },\n]\n\nprint(f\"Defined {len(test_cases)} test cases.\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": "# ── Run Test Cases ───────────────────────────────────────\nall_tools = [get_weather, get_customer_info, calculate]\nresults = []\n\nfor test in test_cases:\n    print(f\"\\n{'='*60}\")\n    print(f\"Test {test['id']}: {test['description']}\")\n    print(f\"Input: {test['input']}\")\n\n    answer, tools_called, trace = run_agent(\n        test[\"input\"],\n        tools=all_tools,\n        max_steps=test[\"max_steps\"],\n    )\n\n    # Show the final answer\n    print(f\"\\n  Answer: {(answer or '(none)')[:200]}\")\n\n    # Check tools used\n    expected_tools = set(test[\"expected_tools\"])\n    actual_tools = set(tools_called)\n    tools_match = expected_tools == actual_tools\n    if not tools_match:\n        print(f\"  Tools expected: {sorted(expected_tools)}\")\n        print(f\"  Tools actual:   {sorted(actual_tools)}\")\n\n    # Check keywords (at least one match is enough)\n    answer_lower = (answer or \"\").lower()\n    keywords_found = [kw for kw in test[\"expected_keywords\"] if kw.lower() in answer_lower]\n    keywords_missing = [kw for kw in test[\"expected_keywords\"] if kw.lower() not in answer_lower]\n\n    if test[\"expected_keywords\"]:\n        passed = len(keywords_found) >= 1  # At least one keyword match\n    else:\n        passed = True  # No keywords to check\n\n    status = \"✅ PASS\" if passed else \"❌ FAIL\"\n\n    results.append({\n        \"test_id\": test[\"id\"],\n        \"description\": test[\"description\"],\n        \"passed\": passed,\n        \"tools_match\": tools_match,\n        \"tools_called\": sorted(actual_tools),\n        \"keywords_found\": len(keywords_found),\n        \"keywords_total\": len(test[\"expected_keywords\"]),\n        \"answer_preview\": (answer or \"\")[:100],\n    })\n\n    print(f\"\\n  {status} | Keywords: {len(keywords_found)}/{len(test['expected_keywords'])} | Tools match: {'✅' if tools_match else '❌'}\")\n    if keywords_missing:\n        print(f\"  Missing keywords: {keywords_missing}\")"
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Evaluation Summary ───────────────────────────────────\n",
    "import pandas as pd\n",
    "\n",
    "eval_df = pd.DataFrame(results)\n",
    "print(\"\\n\" + \"=\"*60)\n",
    "print(\"EVALUATION SUMMARY\")\n",
    "print(\"=\"*60)\n",
    "print(eval_df[[\"test_id\", \"description\", \"passed\", \"keywords_found\", \"keywords_total\"]].to_string(index=False))\n",
    "\n",
    "accuracy = eval_df[\"passed\"].mean() * 100\n",
    "print(f\"\\nOverall Pass Rate: {accuracy:.0f}% ({eval_df['passed'].sum()}/{len(eval_df)})\")"
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## 💬 Discussion Questions\n",
    "\n",
    "1. **Why did the agent call `get_weather` even for \"Paris\" (not in our database)?** What does this tell you about tool selection?\n",
    "\n",
    "2. **In the multi-tool test (T4), how many steps did the agent take?** Could it have been more efficient?\n",
    "\n",
    "3. **How would you improve the test cases?** What scenarios are missing?\n",
    "\n",
    "4. **What happens if you give vague tool descriptions?** Try changing a docstring and re-running."
   ]
  },
  {
   "cell_type": "markdown",
   "metadata": {},
   "source": [
    "---\n",
    "## Wrap-up & Key Takeaways\n",
    "\n",
    "**What you learned:**\n",
    "- Function calling = model requests a call, your code executes it\n",
    "- AUTO/ANY/NONE modes control tool usage aggressiveness\n",
    "- The agent loop: history → generate → execute tools → repeat\n",
    "- Multi-tool agents: the model picks the right tool(s) for each query\n",
    "- Test cases with expected keywords and tools catch regressions\n",
    "\n",
    "**What's next:** In the **Independent Lab**, you'll build your own domain-specific agent with 3-4 custom tools, test it, and improve it."
   ]
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Export Prompt Log ────────────────────────────────────\n",
    "import pandas as pd\n",
    "\n",
    "if PROMPT_LOG:\n",
    "    log_df = pd.DataFrame(PROMPT_LOG)\n",
    "    log_df.to_csv(\"day4_lab1_prompt_log.csv\", index=False)\n",
    "    print(f\"✅ Exported {len(log_df)} log entries to day4_lab1_prompt_log.csv\")\n",
    "else:\n",
    "    print(\"No interactions logged.\")"
   ]
  },
  {
   "cell_type": "code",
   "source": "# ── Export Agent Trace as JSONL ───────────────────────────\n# JSONL (one JSON object per line) is ideal for debugging agent behavior:\n# - Easy to grep for specific tool calls or errors\n# - Can be loaded with pd.read_json(path, lines=True)\n# - Each line is self-contained (no array structure to break)\n\ndef save_trace(trace_data, path):\n    \"\"\"Save a list of dicts as JSONL (one JSON object per line).\"\"\"\n    with open(path, \"w\", encoding=\"utf-8\") as f:\n        for row in trace_data:\n            f.write(json.dumps(row, ensure_ascii=False) + \"\\n\")\n    print(f\"✅ Saved trace → {path}\")\n\n# Save all logged interactions as a JSONL trace\nif PROMPT_LOG:\n    save_trace(PROMPT_LOG, \"day4_lab1_trace.jsonl\")\nelse:\n    print(\"No interactions to save.\")",
   "metadata": {},
   "execution_count": null,
   "outputs": []
  },
  {
   "cell_type": "code",
   "execution_count": null,
   "metadata": {},
   "outputs": [],
   "source": [
    "# ── Session Statistics ───────────────────────────────────\n",
    "tool_calls = sum(1 for log in PROMPT_LOG if log.get(\"label\") == \"tool_call\")\n",
    "agent_outputs = sum(1 for log in PROMPT_LOG if log.get(\"label\") == \"agent_output\")\n",
    "\n",
    "print(f\"Session Statistics:\")\n",
    "print(f\"  Total log entries:  {len(PROMPT_LOG)}\")\n",
    "print(f\"  Tool calls:         {tool_calls}\")\n",
    "print(f\"  Agent completions:  {agent_outputs}\")\n",
    "print(f\"  Test cases run:     {len(test_cases)}\")\n",
    "print(f\"  Pass rate:          {accuracy:.0f}%\")"
   ]
  }
 ],
 "metadata": {
  "kernelspec": {
   "display_name": "Python 3",
   "language": "python",
   "name": "python3"
  },
  "language_info": {
   "name": "python"
  }
 },
 "nbformat": 4,
 "nbformat_minor": 4
}