Skip to content

Day 4 - GenAI Agents

Generative AI for Managers - TUHH Institute of Entrepreneurship · part of my Technology Management MBA · study notes for revision.


For three days the model could only ever talk back. I asked, it answered - one turn, done. Useful, but oddly passive: it could tell me how to find our churned customers and draft re-engagement emails, but it couldn’t actually go and do it.

Day 4 closes that gap. An agent is what you get when you wrap an LLM in a loop and hand it a few tools - functions it can ask to run. Now the same model can look up the customers, read their support tickets, draft an email for each, and keep going until the job is finished. Same brain, but it can now reach out and act on the world.

Look at what we built so far, and the wall each version hits:

DayWhat it could doWhere it stopped
Day 1Understand how an LLM generates textOnly produces words - no actions
Day 2Prompt well, pull out structured dataOne shot: every call stands alone
Day 3Ground answers in real documents (RAG)Fixed pipeline: always retrieve → generate

Every one of these is a single-turn interaction: one request in, one response out. But plenty of real work isn’t one turn. It needs several steps, decisions taken along the way depending on what the last step found, and reaching into outside systems (a CRM, a database, the web). That’s the gap an agent fills.

Three things made agents practical only recently:

  • Reliable function calling - models can now reliably output a clean, structured “please run this function with these arguments” request, consistent enough to trust in a real product.
  • Bigger context windows - the agent can pile up tool results and its own intermediate reasoning without losing the thread.
  • Stronger reasoning - models got much better at planning multi-step work and at recovering when a step goes wrong.

An AI agent is a system where an LLM runs inside a loop and is allowed to:

  • Reason about what to do next,
  • Call tools (functions, APIs, databases),
  • Observe the results those tools give back,
  • Decide whether to keep going or stop with a final answer.

The word “tool” just means a piece of your code the model is permitted to request - a function that looks up a customer, searches your docs, does a calculation. The “loop” is what makes it an agent rather than a one-off call: it can go round again, using what it just learned.

The agent spectrum - not everything needs an agent

Section titled “The agent spectrum - not everything needs an agent”

“Agent” isn’t all-or-nothing. It’s a ladder of autonomy, and most of the business value sits in the middle:

LevelWhat it isWho decides the steps
0 · Direct call”Translate this sentence” - one plain generate callYou, in one shot
1 · Fixed pipelineRAG, extraction chainsYour code, fixed order
2 · Tool-augmented LLMModel picks from tools you offer, one call per turnModel, one step
3 · Autonomous agentModel plans, calls several tools, loops until doneModel, many steps
4 · Multi-agent systemSeveral specialised agents collaborateAn orchestrator

The labs live at Levels 2-3 - enough autonomy to be genuinely useful, little enough to stay understandable and safe.

ChatbotPipeline (RAG)Agent
InteractionSingle turnFixed sequenceDynamic loop
ToolsNoneRetrieval onlySeveral, model’s choice
Who’s in controlHuman drivesCode drivesLLM drives
Good forSimple Q&AKnowledge lookupMulti-step tasks needing judgement

3 · Function calling - teaching the model to use tools

Section titled “3 · Function calling - teaching the model to use tools”

Function calling is the mechanism that makes all of this possible. Here’s the single most important thing to hold onto, because it’s the source of most confusion:

With the Gemini SDK (google.genai), a tool is built from three layers - and the model only ever sees the middle one:

LayerWhat it isIts job
Python functionYour real, executable codeDoes the actual work
FunctionDeclarationA schema: the function’s name, its parameters, what it’s forThe only thing the model sees
ToolA container grouping one or more declarationsWhat you pass to the API via tools=[...]

The model reads the description and decides whether to call the function and with what arguments. It never sees your source code - so a clear name, good docstring and honest parameter descriptions are the instructions the model uses to choose well.

The easy way: hand the SDK a plain Python function

Section titled “The easy way: hand the SDK a plain Python function”

The simplest approach is to pass a normal function. The SDK reads its name, docstring and type hints and builds the schema for you:

from google import genai
from google.genai import types
client = genai.Client() # reads the API key from the environment - never hard-code it
def get_customer_info(customer_id: str) -> dict:
"""Look up a customer's account details.
Args:
customer_id: The unique customer identifier, e.g. 'CUST-1234'
"""
database = {
"CUST-1234": {"name": "Acme Corp", "plan": "Enterprise", "mrr": 12000},
"CUST-5678": {"name": "TechStart", "plan": "Growth", "mrr": 3500},
}
return database.get(customer_id, {"error": "Customer not found"})
response = client.models.generate_content(
model="gemini-2.5-flash-lite",
contents="What plan is customer CUST-1234 on?",
config=types.GenerateContentConfig(
tools=[get_customer_info], # SDK wraps this into a Tool for you
),
)
print(response.text)

There’s also a manual route where you write the FunctionDeclaration yourself - useful when the function lives on a remote server, or when you want fine control over the descriptions. Most projects start with the easy way and only reach for the manual one when they actually need it. The full code for both is in the notebooks.

Tool-calling modes - how much freedom the model gets

Section titled “Tool-calling modes - how much freedom the model gets”

When you offer tools, the model still reasons about whether it even needs one. The mode setting controls that freedom:

ModeBehaviourUse when
AUTO (default)Model decides: call a tool, or just answer in textGeneral use - let it judge
ANYModel must call at least one tool (it still picks which)You want to force tool use
NONEModel may not call any tool (descriptions still visible)You want a pure text reply

AUTO is the one to understand - it’s what turns the model into an active reasoner about tools rather than a blind executor. It weighs your request against the tool descriptions and decides on its own whether a call would actually help.

This is the heart of the day. Strip away frameworks and providers, and every agent runs the same cycle. The model does exactly one thing in it - the Reason step - and your code does everything else.

Historybuild the context so far
→
Reasonmodel: tool call or final answer?
→
Actyour code runs the function
→
Observeresult added back to history
The loop: assemble the conversation, let the model reason, run any tool it asks for, feed the result back - then round again. If the model returns text instead of a tool request, the loop stops and that text is the answer.

Walking the five steps once:

  1. History (your code) - assemble the whole conversation so far: the user’s message, the model’s earlier replies, and any tool results. This is the only thing the model gets to see.
  2. Reason (the model) - it reads the history and decides: request a tool call, or produce a final text answer. This is the model’s only move.
  3. Act (your code) - if it asked for a tool, you look up the matching function and run it for real.
  4. Observe (your code) - you wrap the result in a FunctionResponse and append it to the history, so the model can see it next time round.
  5. Repeat or stop - still missing information? Go back to step 2. Enough to answer? It returns text and the loop ends. Always cap the rounds with a max_steps limit so it can’t spin forever.

There are two ways to actually run this loop, and they’re a separate choice from the AUTO/ANY/NONE modes above (that was whether the model may ask; this is who executes once it asks):

  • Automatic function calling - the SDK runs the whole loop inside a single generate_content() call: it executes your function, feeds the result back, repeats. Great for a quick prototype, but you can’t see what happened. You’d still set a safety ceiling like maximum_remote_calls=5.
  • Manual loop - you write the for loop yourself, executing each tool call and printing every step. More code, but total visibility. The labs use this so you can watch every decision.

Here’s the manual loop, trimmed to its bones:

def run_agent(user_message, tools, max_steps=10):
tool_map = {fn.__name__: fn for fn in tools}
history = [types.Content(role="user", parts=[types.Part(text=user_message)])]
for step in range(max_steps):
# REASON - send full history; model may ask for a tool
response = client.models.generate_content(
model="gemini-2.5-flash-lite",
contents=history,
config=types.GenerateContentConfig(
tools=tools,
automatic_function_calling=types.AutomaticFunctionCallingConfig(
disable=True # we execute tools ourselves, in Step ACT below
),
),
)
history.append(types.Content(role="model", parts=response.parts))
# ACT + OBSERVE - run any requested tool, append its result
function_results = []
for part in response.parts:
if part.function_call:
name = part.function_call.name
args = dict(part.function_call.args)
result = tool_map[name](**args) # run the real function
function_results.append(types.Part(
function_response=types.FunctionResponse(
name=name, response={"result": result})))
# REPEAT or STOP
if function_results:
history.append(types.Content(role="user", parts=function_results))
else:
return response.text # no tool call → this is the final answer
return "Agent reached maximum steps without completing."

5 · How agents think - ReAct and plan-first

Section titled “5 · How agents think - ReAct and plan-first”

The most common pattern is ReAct: the model openly alternates between a Thought (what should I do next?) and an Action (calling a tool), then reads the Observation and thinks again.

Thoughtwhat do I need?
→
Actioncall a tool
→
Observationread the result
→
…until enoughthen answer
A ReAct trace, in plain terms: think about the next move, take it, look at what came back, repeat - the visible “Thought” step is what keeps the model deliberate instead of just reacting.

A real trace reads almost like someone reasoning aloud:

Thought: I need the customer's current plan first.
Action: get_customer_info(customer_id="CUST-1234")
Observation: {"name": "Acme Corp", "plan": "Enterprise", "mrr": 12000}
Thought: Now check our discount policy.
Action: search_knowledge_base(query="enterprise discount policy")
Observation: Enterprise customers may receive up to 20% discount...
Thought: They qualify. Compute the discounted annual price.
Action: calculate(expression="12000 * 12 * 0.85")
Observation: 122400.0
Answer: Acme Corp is on Enterprise at $12k/month. A 15% annual discount
gives $122,400/year - within our 20% policy limit.

You encourage this simply by asking for it in the system prompt (“explain your reasoning in a Thought step before each action; reflect on each result”).

For well-defined tasks, it’s often better to make the model write a plan up front, then execute it step by step. The win is transparency: you can show a human the plan before anything runs, and revise it if a step fails.

PatternBest forTrade-off
ReActExploratory tasks, unknown number of stepsFlexible, but harder to predict
Plan-and-ExecuteWell-defined multi-step tasksPredictable, but less adaptive
Single tool callOne-step lookupsNo planning overhead, limited

Many production agents do both: plan first, then reason ReAct-style within each step. The labs use ReAct because once you get it, plan-first is a small extension.

6 · Memory - the context window as a scratchpad

Section titled “6 · Memory - the context window as a scratchpad”

An agent’s working memory is just its conversation history. Every tool call and every result gets appended, so by turn four the model can reason over everything it gathered in turns one to three. Powerful - but the context window is finite, so long-running agents eventually have to summarise older steps to stay within the limit.

TypeScopeHow it’s done
Short-term (working)Current taskThe conversation history itself
SessionCurrent sessionA running summary of previous steps
Long-termAcross sessionsA vector database of past interactions

That last row reuses Day 3 directly: the same embeddings-and-vector-search idea, but storing summaries of past agent work so it can recall and build on what it did before.

Because tool descriptions are prompts, tool design is where a lot of agent quality is won or lost. The principles are common-sense once stated:

PrincipleWhy it matters
Single responsibilityOne tool does one thing well - separate search_docs from search_tickets
Clear boundariesThe model can tell which tool fits (“customer data” vs “product info”)
Descriptive parametersIt sends the right arguments - spell out the format, e.g. customer_id: 'CUST-1234'
Graceful errorsReturn {"error": "Not found"} instead of crashing, so the agent can recover
Minimal scopeGive narrow, specific tools - never a “run any SQL” tool

And for what a tool hands back: return structured data (a dict) with stable keys like status, data, error; keep it short so it doesn’t clog the context; and always include error info so the agent can react to a failure instead of guessing past it.

This is the part that matters most for a manager, because an agent with tools has real power - and therefore real ways to cause harm. A chatbot can only say a wrong thing; an agent can do a wrong thing.

RiskWhat it looks likeGuardrail
Unintended actionsSends an email it shouldn’tHuman approval for irreversible actions
Runaway loopsKeeps calling tools, never convergesA hard max_steps ceiling
Data leakageExposes confidential info via a toolScope tool access to authorised data only
Prompt injectionMalicious input hijacks the agentValidate inputs; keep data separate from instructions
Hallucinated callsCalls a tool with made-up argumentsValidate arguments before executing

A clean, practical pattern is to sort tools by how much damage they can do, and gate them accordingly:

Tool typeExamplesPermission
Read-onlySearch docs, look up a customer, get statusAllow freely
WriteCreate a ticket, update the CRM, save a draftRequire human approval
IrreversibleSend an email, delete records, place an orderRequire explicit confirmation

The human-in-the-loop pattern implements this: before a sensitive tool runs, the agent pauses and asks a person to approve (user_confirmed=True). And the aim isn’t to prevent all failure - agents will fail - it’s to make failure cheap and recoverable: timeouts, retries with backoff, fallbacks (“if search fails, ask the user for the document”), and always a step ceiling.

The honest answer is: often you shouldn’t. Reach for an agent only when a task genuinely needs the loop.

Use an agent when…Use something simpler when…
The task needs multiple steps and decisionsA single model call is enough
Different tools are needed depending on the queryThe workflow is always identical → use a pipeline
User intent is varied and unpredictableQueries are predictable → use templates
Actions must be taken (create, update, send)You only need to look things up → use RAG
Transparency and an audit trail matterSpeed is the only priority

Customer-support triage (knowledge base + CRM + ticketing), sales research (CRM + market data), financial analysis (database + calculator + report templates), IT helpdesk, content pipelines - all share the same shape: naturally iterative work where the agent may search, refine, and combine several times before answering. That’s exactly where a fixed pipeline breaks, because you’d have to anticipate every step in advance.

For bigger workflows, agents even compose into multi-agent systems - an orchestrator routing to specialists, a research→analysis→writing pipeline, or a draft-then-critique pair where one agent writes and another reviews. Useful, but a warning: if every agent shares the same wrong assumption, more agents just amplifies the error. Always ground the key facts in trusted tools or human input.

The theory above is the map; the notebooks are where I actually built a working agent from scratch and watched every step of its loop.

  • Guided Lab - Your first AI agent. Define a Python function as a tool, see the model request a call, compare AUTO/ANY/NONE modes, then build the manual agent loop by hand. Add a second and third tool (customer lookup, calculator) so the model has to choose, and finish by writing test cases with expected keywords to measure how well it did.
  • Independent Lab - Build your own agentic system. Pick a business track (support, research, finance, or HR), design three domain tools with docstrings and graceful error handling, write a system prompt (role, tools, rules, refusal behaviour), then evaluate a v1, do error analysis, and improve it to a measurably better v2. The track you pick here carries into the assignment.
  • Assignment - A production-ready agent. The full package: refined tools, a final system prompt with all six components, 12+ queries run through the loop, a golden test set, LLM-as-judge scoring on tool selection / reasoning / completeness, a proper error analysis, and a deployment “playbook” for handing the agent to a team.

Download the notebooks (open in Google Colab):

My submitted solution: Assignment 4 - my solved notebook (.ipynb) - my own work from the course.

Must-knowOne-line recall
What an agent isAgent = LLM + Tools + Loop - it reasons, calls tools, observes, repeats until done.
Agent vs plain LLMAn LLM only thinks; an agent thinks and acts over multiple steps.
Function callingThe model requests a named function with arguments - a tool.
The model never runs codeIt only asks; your code executes the function and returns the result.
The three layersPython function → FunctionDeclaration (all the model sees) → Tool.
Tool-calling modesAUTO = decide, ANY = must call one, NONE = no calls.
The agent loopHistory → Reason → Act → Observe → repeat or stop.
Only Reason is the modelEvery other step is your code, including running the tool.
ReActAlternate Thought → Action → Observation; the visible Thought keeps it deliberate.
Plan-and-ExecutePlan first, then run steps - predictable and reviewable.
Tool descriptions are promptsClear, specific docstrings → better tool selection.
MemoryHistory is working memory; summarise or use a vector DB when it grows.
Safety scales with powerRead-only free, write needs approval, irreversible needs confirmation.
The golden ruleNever give an agent a tool you wouldn’t trust an intern with unsupervised.
Always cap stepsmax_steps prevents runaway loops.
When not toIf one call with no tools works, do that - add agents only on real pain.
RAG becomes a toolEverything from Days 1-3 is now a capability the agent can invoke.