Day 4 - GenAI Agents
Generative AI for Managers - TUHH Institute of Entrepreneurship · part of my Technology Management MBA · study notes for revision.
For three days the model could only ever talk back. I asked, it answered - one turn, done. Useful, but oddly passive: it could tell me how to find our churned customers and draft re-engagement emails, but it couldn’t actually go and do it.
Day 4 closes that gap. An agent is what you get when you wrap an LLM in a loop and hand it a few tools - functions it can ask to run. Now the same model can look up the customers, read their support tickets, draft an email for each, and keep going until the job is finished. Same brain, but it can now reach out and act on the world.
1 · Why agents - the automation gap
Section titled “1 · Why agents - the automation gap”Look at what we built so far, and the wall each version hits:
| Day | What it could do | Where it stopped |
|---|---|---|
| Day 1 | Understand how an LLM generates text | Only produces words - no actions |
| Day 2 | Prompt well, pull out structured data | One shot: every call stands alone |
| Day 3 | Ground answers in real documents (RAG) | Fixed pipeline: always retrieve → generate |
Every one of these is a single-turn interaction: one request in, one response out. But plenty of real work isn’t one turn. It needs several steps, decisions taken along the way depending on what the last step found, and reaching into outside systems (a CRM, a database, the web). That’s the gap an agent fills.
Three things made agents practical only recently:
- Reliable function calling - models can now reliably output a clean, structured “please run this function with these arguments” request, consistent enough to trust in a real product.
- Bigger context windows - the agent can pile up tool results and its own intermediate reasoning without losing the thread.
- Stronger reasoning - models got much better at planning multi-step work and at recovering when a step goes wrong.
2 · What an agent actually is
Section titled “2 · What an agent actually is”An AI agent is a system where an LLM runs inside a loop and is allowed to:
- Reason about what to do next,
- Call tools (functions, APIs, databases),
- Observe the results those tools give back,
- Decide whether to keep going or stop with a final answer.
The word “tool” just means a piece of your code the model is permitted to request - a function that looks up a customer, searches your docs, does a calculation. The “loop” is what makes it an agent rather than a one-off call: it can go round again, using what it just learned.
The agent spectrum - not everything needs an agent
Section titled “The agent spectrum - not everything needs an agent”“Agent” isn’t all-or-nothing. It’s a ladder of autonomy, and most of the business value sits in the middle:
| Level | What it is | Who decides the steps |
|---|---|---|
| 0 · Direct call | ”Translate this sentence” - one plain generate call | You, in one shot |
| 1 · Fixed pipeline | RAG, extraction chains | Your code, fixed order |
| 2 · Tool-augmented LLM | Model picks from tools you offer, one call per turn | Model, one step |
| 3 · Autonomous agent | Model plans, calls several tools, loops until done | Model, many steps |
| 4 · Multi-agent system | Several specialised agents collaborate | An orchestrator |
The labs live at Levels 2-3 - enough autonomy to be genuinely useful, little enough to stay understandable and safe.
Chatbot vs pipeline vs agent
Section titled “Chatbot vs pipeline vs agent”| Chatbot | Pipeline (RAG) | Agent | |
|---|---|---|---|
| Interaction | Single turn | Fixed sequence | Dynamic loop |
| Tools | None | Retrieval only | Several, model’s choice |
| Who’s in control | Human drives | Code drives | LLM drives |
| Good for | Simple Q&A | Knowledge lookup | Multi-step tasks needing judgement |
3 · Function calling - teaching the model to use tools
Section titled “3 · Function calling - teaching the model to use tools”Function calling is the mechanism that makes all of this possible. Here’s the single most important thing to hold onto, because it’s the source of most confusion:
The three layers
Section titled “The three layers”With the Gemini SDK (google.genai), a tool is built from three layers - and the model only ever sees the middle one:
| Layer | What it is | Its job |
|---|---|---|
| Python function | Your real, executable code | Does the actual work |
FunctionDeclaration | A schema: the function’s name, its parameters, what it’s for | The only thing the model sees |
Tool | A container grouping one or more declarations | What you pass to the API via tools=[...] |
The model reads the description and decides whether to call the function and with what arguments. It never sees your source code - so a clear name, good docstring and honest parameter descriptions are the instructions the model uses to choose well.
The easy way: hand the SDK a plain Python function
Section titled “The easy way: hand the SDK a plain Python function”The simplest approach is to pass a normal function. The SDK reads its name, docstring and type hints and builds the schema for you:
from google import genaifrom google.genai import types
client = genai.Client() # reads the API key from the environment - never hard-code it
def get_customer_info(customer_id: str) -> dict: """Look up a customer's account details.
Args: customer_id: The unique customer identifier, e.g. 'CUST-1234' """ database = { "CUST-1234": {"name": "Acme Corp", "plan": "Enterprise", "mrr": 12000}, "CUST-5678": {"name": "TechStart", "plan": "Growth", "mrr": 3500}, } return database.get(customer_id, {"error": "Customer not found"})
response = client.models.generate_content( model="gemini-2.5-flash-lite", contents="What plan is customer CUST-1234 on?", config=types.GenerateContentConfig( tools=[get_customer_info], # SDK wraps this into a Tool for you ),)print(response.text)There’s also a manual route where you write the FunctionDeclaration yourself - useful when the function lives on a remote server, or when you want fine control over the descriptions. Most projects start with the easy way and only reach for the manual one when they actually need it. The full code for both is in the notebooks.
Tool-calling modes - how much freedom the model gets
Section titled “Tool-calling modes - how much freedom the model gets”When you offer tools, the model still reasons about whether it even needs one. The mode setting controls that freedom:
| Mode | Behaviour | Use when |
|---|---|---|
AUTO (default) | Model decides: call a tool, or just answer in text | General use - let it judge |
ANY | Model must call at least one tool (it still picks which) | You want to force tool use |
NONE | Model may not call any tool (descriptions still visible) | You want a pure text reply |
AUTO is the one to understand - it’s what turns the model into an active reasoner about tools rather than a blind executor. It weighs your request against the tool descriptions and decides on its own whether a call would actually help.
4 · The agent loop
Section titled “4 · The agent loop”This is the heart of the day. Strip away frameworks and providers, and every agent runs the same cycle. The model does exactly one thing in it - the Reason step - and your code does everything else.
Walking the five steps once:
- History (your code) - assemble the whole conversation so far: the user’s message, the model’s earlier replies, and any tool results. This is the only thing the model gets to see.
- Reason (the model) - it reads the history and decides: request a tool call, or produce a final text answer. This is the model’s only move.
- Act (your code) - if it asked for a tool, you look up the matching function and run it for real.
- Observe (your code) - you wrap the result in a
FunctionResponseand append it to the history, so the model can see it next time round. - Repeat or stop - still missing information? Go back to step 2. Enough to answer? It returns text and the loop ends. Always cap the rounds with a
max_stepslimit so it can’t spin forever.
Automatic vs manual execution
Section titled “Automatic vs manual execution”There are two ways to actually run this loop, and they’re a separate choice from the AUTO/ANY/NONE modes above (that was whether the model may ask; this is who executes once it asks):
- Automatic function calling - the SDK runs the whole loop inside a single
generate_content()call: it executes your function, feeds the result back, repeats. Great for a quick prototype, but you can’t see what happened. You’d still set a safety ceiling likemaximum_remote_calls=5. - Manual loop - you write the
forloop yourself, executing each tool call and printing every step. More code, but total visibility. The labs use this so you can watch every decision.
Here’s the manual loop, trimmed to its bones:
def run_agent(user_message, tools, max_steps=10): tool_map = {fn.__name__: fn for fn in tools} history = [types.Content(role="user", parts=[types.Part(text=user_message)])]
for step in range(max_steps): # REASON - send full history; model may ask for a tool response = client.models.generate_content( model="gemini-2.5-flash-lite", contents=history, config=types.GenerateContentConfig( tools=tools, automatic_function_calling=types.AutomaticFunctionCallingConfig( disable=True # we execute tools ourselves, in Step ACT below ), ), ) history.append(types.Content(role="model", parts=response.parts))
# ACT + OBSERVE - run any requested tool, append its result function_results = [] for part in response.parts: if part.function_call: name = part.function_call.name args = dict(part.function_call.args) result = tool_map[name](**args) # run the real function function_results.append(types.Part( function_response=types.FunctionResponse( name=name, response={"result": result})))
# REPEAT or STOP if function_results: history.append(types.Content(role="user", parts=function_results)) else: return response.text # no tool call → this is the final answer
return "Agent reached maximum steps without completing."5 · How agents think - ReAct and plan-first
Section titled “5 · How agents think - ReAct and plan-first”ReAct - Reason + Act
Section titled “ReAct - Reason + Act”The most common pattern is ReAct: the model openly alternates between a Thought (what should I do next?) and an Action (calling a tool), then reads the Observation and thinks again.
A real trace reads almost like someone reasoning aloud:
Thought: I need the customer's current plan first.Action: get_customer_info(customer_id="CUST-1234")Observation: {"name": "Acme Corp", "plan": "Enterprise", "mrr": 12000}
Thought: Now check our discount policy.Action: search_knowledge_base(query="enterprise discount policy")Observation: Enterprise customers may receive up to 20% discount...
Thought: They qualify. Compute the discounted annual price.Action: calculate(expression="12000 * 12 * 0.85")Observation: 122400.0
Answer: Acme Corp is on Enterprise at $12k/month. A 15% annual discount gives $122,400/year - within our 20% policy limit.You encourage this simply by asking for it in the system prompt (“explain your reasoning in a Thought step before each action; reflect on each result”).
Plan-and-Execute - plan first, then do
Section titled “Plan-and-Execute - plan first, then do”For well-defined tasks, it’s often better to make the model write a plan up front, then execute it step by step. The win is transparency: you can show a human the plan before anything runs, and revise it if a step fails.
| Pattern | Best for | Trade-off |
|---|---|---|
| ReAct | Exploratory tasks, unknown number of steps | Flexible, but harder to predict |
| Plan-and-Execute | Well-defined multi-step tasks | Predictable, but less adaptive |
| Single tool call | One-step lookups | No planning overhead, limited |
Many production agents do both: plan first, then reason ReAct-style within each step. The labs use ReAct because once you get it, plan-first is a small extension.
6 · Memory - the context window as a scratchpad
Section titled “6 · Memory - the context window as a scratchpad”An agent’s working memory is just its conversation history. Every tool call and every result gets appended, so by turn four the model can reason over everything it gathered in turns one to three. Powerful - but the context window is finite, so long-running agents eventually have to summarise older steps to stay within the limit.
| Type | Scope | How it’s done |
|---|---|---|
| Short-term (working) | Current task | The conversation history itself |
| Session | Current session | A running summary of previous steps |
| Long-term | Across sessions | A vector database of past interactions |
That last row reuses Day 3 directly: the same embeddings-and-vector-search idea, but storing summaries of past agent work so it can recall and build on what it did before.
7 · Designing good tools
Section titled “7 · Designing good tools”Because tool descriptions are prompts, tool design is where a lot of agent quality is won or lost. The principles are common-sense once stated:
| Principle | Why it matters |
|---|---|
| Single responsibility | One tool does one thing well - separate search_docs from search_tickets |
| Clear boundaries | The model can tell which tool fits (“customer data” vs “product info”) |
| Descriptive parameters | It sends the right arguments - spell out the format, e.g. customer_id: 'CUST-1234' |
| Graceful errors | Return {"error": "Not found"} instead of crashing, so the agent can recover |
| Minimal scope | Give narrow, specific tools - never a “run any SQL” tool |
And for what a tool hands back: return structured data (a dict) with stable keys like status, data, error; keep it short so it doesn’t clog the context; and always include error info so the agent can react to a failure instead of guessing past it.
8 · Safety and guardrails
Section titled “8 · Safety and guardrails”This is the part that matters most for a manager, because an agent with tools has real power - and therefore real ways to cause harm. A chatbot can only say a wrong thing; an agent can do a wrong thing.
| Risk | What it looks like | Guardrail |
|---|---|---|
| Unintended actions | Sends an email it shouldn’t | Human approval for irreversible actions |
| Runaway loops | Keeps calling tools, never converges | A hard max_steps ceiling |
| Data leakage | Exposes confidential info via a tool | Scope tool access to authorised data only |
| Prompt injection | Malicious input hijacks the agent | Validate inputs; keep data separate from instructions |
| Hallucinated calls | Calls a tool with made-up arguments | Validate arguments before executing |
Match permission to risk
Section titled “Match permission to risk”A clean, practical pattern is to sort tools by how much damage they can do, and gate them accordingly:
| Tool type | Examples | Permission |
|---|---|---|
| Read-only | Search docs, look up a customer, get status | Allow freely |
| Write | Create a ticket, update the CRM, save a draft | Require human approval |
| Irreversible | Send an email, delete records, place an order | Require explicit confirmation |
The human-in-the-loop pattern implements this: before a sensitive tool runs, the agent pauses and asks a person to approve (user_confirmed=True). And the aim isn’t to prevent all failure - agents will fail - it’s to make failure cheap and recoverable: timeouts, retries with backoff, fallbacks (“if search fails, ask the user for the document”), and always a step ceiling.
9 · When to use an agent (and when not)
Section titled “9 · When to use an agent (and when not)”The honest answer is: often you shouldn’t. Reach for an agent only when a task genuinely needs the loop.
| Use an agent when… | Use something simpler when… |
|---|---|
| The task needs multiple steps and decisions | A single model call is enough |
| Different tools are needed depending on the query | The workflow is always identical → use a pipeline |
| User intent is varied and unpredictable | Queries are predictable → use templates |
| Actions must be taken (create, update, send) | You only need to look things up → use RAG |
| Transparency and an audit trail matter | Speed is the only priority |
Where it pays off in business
Section titled “Where it pays off in business”Customer-support triage (knowledge base + CRM + ticketing), sales research (CRM + market data), financial analysis (database + calculator + report templates), IT helpdesk, content pipelines - all share the same shape: naturally iterative work where the agent may search, refine, and combine several times before answering. That’s exactly where a fixed pipeline breaks, because you’d have to anticipate every step in advance.
For bigger workflows, agents even compose into multi-agent systems - an orchestrator routing to specialists, a research→analysis→writing pipeline, or a draft-then-critique pair where one agent writes and another reviews. Useful, but a warning: if every agent shares the same wrong assumption, more agents just amplifies the error. Always ground the key facts in trusted tools or human input.
Hands-on: labs & assignment
Section titled “Hands-on: labs & assignment”The theory above is the map; the notebooks are where I actually built a working agent from scratch and watched every step of its loop.
- Guided Lab - Your first AI agent. Define a Python function as a tool, see the model request a call, compare
AUTO/ANY/NONEmodes, then build the manual agent loop by hand. Add a second and third tool (customer lookup, calculator) so the model has to choose, and finish by writing test cases with expected keywords to measure how well it did. - Independent Lab - Build your own agentic system. Pick a business track (support, research, finance, or HR), design three domain tools with docstrings and graceful error handling, write a system prompt (role, tools, rules, refusal behaviour), then evaluate a v1, do error analysis, and improve it to a measurably better v2. The track you pick here carries into the assignment.
- Assignment - A production-ready agent. The full package: refined tools, a final system prompt with all six components, 12+ queries run through the loop, a golden test set, LLM-as-judge scoring on tool selection / reasoning / completeness, a proper error analysis, and a deployment “playbook” for handing the agent to a team.
Download the notebooks (open in Google Colab):
My submitted solution: Assignment 4 - my solved notebook (.ipynb) - my own work from the course.
Revision summary
Section titled “Revision summary”| Must-know | One-line recall |
|---|---|
| What an agent is | Agent = LLM + Tools + Loop - it reasons, calls tools, observes, repeats until done. |
| Agent vs plain LLM | An LLM only thinks; an agent thinks and acts over multiple steps. |
| Function calling | The model requests a named function with arguments - a tool. |
| The model never runs code | It only asks; your code executes the function and returns the result. |
| The three layers | Python function → FunctionDeclaration (all the model sees) → Tool. |
| Tool-calling modes | AUTO = decide, ANY = must call one, NONE = no calls. |
| The agent loop | History → Reason → Act → Observe → repeat or stop. |
| Only Reason is the model | Every other step is your code, including running the tool. |
| ReAct | Alternate Thought → Action → Observation; the visible Thought keeps it deliberate. |
| Plan-and-Execute | Plan first, then run steps - predictable and reviewable. |
| Tool descriptions are prompts | Clear, specific docstrings → better tool selection. |
| Memory | History is working memory; summarise or use a vector DB when it grows. |
| Safety scales with power | Read-only free, write needs approval, irreversible needs confirmation. |
| The golden rule | Never give an agent a tool you wouldn’t trust an intern with unsupervised. |
| Always cap steps | max_steps prevents runaway loops. |
| When not to | If one call with no tools works, do that - add agents only on real pain. |
| RAG becomes a tool | Everything from Days 1-3 is now a capability the agent can invoke. |