Skip to content

Day 3 - Retrieval-Augmented Generation (RAG)

Generative AI for Managers - TUHH Institute of Entrepreneurship · part of my Technology Management MBA · study notes for revision.


Day 1 taught me what a large language model (LLM) actually is, and Day 2 taught me how to talk to it well through prompting. But there was always an elephant in the room: the model only knows what it saw during training. It has never read my company’s handbook, never seen last quarter’s sales, and has no idea what changed in the news yesterday. Day 3 is about fixing exactly that - feeding the model the right information at the moment I ask, so its answers are grounded in real data instead of guesswork. That technique is called Retrieval-Augmented Generation, or RAG.

An LLM is trained once, on a giant snapshot of text, and then frozen. That creates three problems that no amount of clever prompting can solve:

ProblemWhat it meansExample question that breaks
Training cutoffThe model’s knowledge stops at a fixed date”Who is the current CEO of that company?” (may have changed)
Private knowledgeIt never saw your internal data”What were our Q4 2025 sales?”
HallucinationWhen it doesn’t know, it invents a confident, wrong answer”What does our handbook say about remote work?”

That last one - hallucination (the model producing plausible-sounding but false statements) - is the dangerous one for business. The model doesn’t say “I don’t know”; it fills the gap with something that reads perfectly and is simply made up.

A simple analogy: a closed-book exam versus an open-book exam. A plain LLM sits a closed-book exam - it answers from memory, and when memory fails it bluffs. RAG turns it into an open-book exam: before answering, it flips to the exact right pages of your textbook and reads from them. Same student, far more reliable answers.

Section titled “2. From keyword search to meaning-based search”

RAG’s first job is finding the right pages. To appreciate how it does that, it helps to see the old way first.

Traditional search (think of a classic search box, or algorithms with names like TF-IDF and BM25) treats each document as a bag of words - literally a tally of which words appear and how often, ignoring word order. Common words like “the” count for little; rare, distinctive words count for a lot. To answer a query it finds documents that share the exact same words.

The fatal weakness is the vocabulary-mismatch problem: it only matches literal words. Ask about “remote work” and it will miss a document that says “working from home” or “telecommuting”, because the letters don’t line up - even though the meaning is identical.

Semantic search matches by meaning instead of by spelling. Here is the difference in practice:

I search for…Keyword search findsSemantic search also finds
”remote work productivity”only pages with those exact wordspages about telecommuting, WFH, distributed teams
”how to reduce costs”only “reduce costs” verbatimefficiency, budget optimisation, lean operations
”customer complaints”only “customer complaints”negative reviews, support tickets, angry feedback

3. Embeddings - turning meaning into numbers

Section titled “3. Embeddings - turning meaning into numbers”

How can software possibly compare meaning? Through embeddings - the idea we first met on Day 1.

An embedding is a list of numbers (a vector) that represents a piece of text as a point in a huge multi-dimensional space. The trick that makes it useful: texts with similar meaning land close together in that space, and unrelated texts land far apart. “Car” and “automobile” end up as near-neighbours; “car” and “banana” end up far apart. The model learned this arrangement from reading enormous amounts of text.

On Day 1 we saw this for single words. For RAG we need it for whole passages - sentences, paragraphs, chunks - collapsed into one vector each:

LevelWhat it capturesTypical use
Word embeddinga single wordanalogies, word similarity
Sentence embeddingone sentencesemantic search, clustering
Passage embeddinga paragraph / chunkRAG retrieval
Document embeddinga whole documentdocument classification

A modern embedding model (the course uses Google’s gemini-embedding-001) reads a passage and hands back a single long vector - in this case 3,072 numbers - that stands for the meaning of that entire passage.

Measuring “closeness”: cosine similarity

Section titled “Measuring “closeness”: cosine similarity”

Once two texts are vectors, we need one number for how similar they are. The standard measure is cosine similarity - it looks at the angle between two vectors rather than their length. Same direction means same meaning.

Cosine scoreWhat it means
1.0identical meaning
0.7 - 0.9very similar
0.3 - 0.6loosely related
below 0.3unrelated
import numpy as np
def cosine_similarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
sim = cosine_similarity(embedding_1, embedding_2)
print(f"Similarity: {sim:.3f}")

The current google-genai SDK makes this a couple of lines. (Never paste a real key into code - store it in Colab Secrets, as the Setup notes show.)

from google import genai
from google.genai import types
client = genai.Client() # reads the key from the environment / Colab Secrets
response = client.models.embed_content(
model="gemini-embedding-001",
contents="How does remote work affect team productivity?",
)
embedding = response.embeddings[0].values
print(len(embedding)) # 3072
response = client.models.embed_content(
model="gemini-embedding-001",
contents=["Remote work increases flexibility but needs clear communication.",
"Office environments foster spontaneous collaboration."],
config=types.EmbedContentConfig(task_type="RETRIEVAL_DOCUMENT"),
)
for i, emb in enumerate(response.embeddings):
print(f"Document {i}: {len(emb.values)} dimensions")

4. Chunking - cutting documents into bite-size pieces

Section titled “4. Chunking - cutting documents into bite-size pieces”

Real documents are long - a 100-page handbook, a financial report, a product manual. We can’t just embed the whole thing as one vector, for four reasons:

ProblemWhy it bites
Input limitsembedding models accept only ~8,000 tokens at a time
Diluted meaningone vector for 100 pages captures nothing sharply
Finite contextwe can only fit so much text into the LLM’s prompt later
Precisionwe want the one paragraph that answers the question, not the whole book

So we chunk: split each document into smaller pieces, embed each piece separately, and later retrieve only the handful of chunks that matter. (A token is roughly a word-fragment; ~750 words ≈ 1,000 tokens.)

StrategyHow it splitsTrade-off
Fixed-sizeevery N characters/tokenssimple, but may cut mid-sentence
Sentence-basedat sentence boundariespreserves meaning; sentences vary in length
Paragraph-basedat paragraph breaksnatural units; some paragraphs are huge
Recursivetry big chunks, split smaller if too longwell-balanced; a bit more complex
Semanticwherever the topic shiftsbest relevance; most expensive to compute

Chunk size is one of the most consequential design choices in a RAG system. Small chunks (~100 tokens) are laser-precise but can miss surrounding context; large chunks (~500 tokens) carry more context but dilute the meaning and cost more to search.

Why overlap? Without it, a fact sitting on a chunk boundary gets sliced in half and may never be retrieved. Imagine a sentence - “Q3 revenue grew 12% driven by the new product line, launched in July” - cut so that “the new product line” lands in chunk 1 and “launched in July” in chunk 2. Neither chunk now holds the whole fact. Overlapping the chunks by a few sentences means each one still carries the complete thought across the seam.

Smart systems attach a little record of metadata to every chunk: which document it came from, the page, the section, the date. This is what lets you later cite sources, filter (“only search HR docs”), and debug (“which chunk produced this wrong answer?“).

chunk = {
"text": "Remote work policy allows up to 3 days per week...",
"source": "employee_handbook_v4.pdf",
"page": 23,
"section": "HR Policies",
"last_updated": "2025-09-01",
}

Now the whole thing fits together. A RAG system runs in two phases: an indexing phase you do once, offline to prepare the knowledge, and a querying phase that runs every time someone asks a question.

DocumentsPDFs, wikis, DBs
→
Chunksplit + overlap
→
Embedchunk → vector
→
Vector storesave vectors
Phase 1 - Indexing (offline, done once): turn your documents into a searchable library of vectors.
Questionfrom user
→
Embed query
→
Similarity searchtop-k chunks
→
Augment promptpaste chunks in
→
Grounded answer
Phase 2 - Querying (online, per question): find the most relevant chunks and let the LLM answer from them.

Reading the two phases as one component table:

StepPhaseWhat happens
1. IngestIndexingload documents (PDFs, web pages, databases, emails)
2. ChunkIndexingsplit into right-sized pieces with overlap
3. EmbedIndexingconvert each chunk to a vector
4. StoreIndexingsave vectors + original text + metadata
5. Embed queryQueryingturn the user’s question into a vector
6. RetrieveQueryingfind the k most similar chunks
7. AugmentQueryinginsert those chunks into the prompt as context
8. GenerateQueryingthe LLM writes an answer grounded in that context

The step that retrieves the k closest chunks is called nearest-neighbour search (k is just “how many chunks to grab” - often 3 to 5).

The augmented prompt - where RAG meets Day 2

Section titled “The augmented prompt - where RAG meets Day 2”

The heart of RAG is how the retrieved chunks get wired into the prompt. And this is pure Day 2 prompting: system/user separation, delimiters to fence off the context, and an explicit grounding instruction telling the model to answer only from what it was given.

System: You answer questions using ONLY the provided context.
If the answer isn't in the context, say
"I don't have enough information to answer this."
User: Answer the question based ONLY on the context below.
<context>
[Retrieved chunk 1]
[Retrieved chunk 2]
[Retrieved chunk 3]
</context>
Question: What is our remote work policy?

Citations - turning a black box into a research assistant

Section titled “Citations - turning a black box into a research assistant”

For business use, a correct answer isn’t enough - people need to know where it came from so they can verify it. You instruct the model to tag each claim with its source:

Rules:
- Cite sources as [Source N] after each claim.
- If the answer isn't in the sources, say so.
- If sources conflict, surface the conflict and prefer the most recent/authoritative.

The model then answers like: “The current policy allows up to 3 days per week [Source 1], but a more recent update raises this to 4 days from Q4 2025 [Source 2].” In regulated fields - legal, healthcare, finance - this traceability is the difference between a toy and a tool. When documents genuinely disagree, a good grounding prompt defines an authority hierarchy (official policy > standard procedures > FAQ > informal wiki) and, within a tier, prefers the newest document - and always surfaces the conflict rather than silently guessing.

The lab builds a complete, working RAG system with nothing more than Python lists and the Gemini API - no database needed. The shape of it:

  1. Prepare documents - a small “knowledge base” of short text snippets standing in for real company facts (founding story, remote-work policy, Q3 revenue, leave policy, product details).

  2. Embed everything - one batched call embeds all documents at once with task_type="RETRIEVAL_DOCUMENT", and we keep the vectors as NumPy arrays.

  3. Search by meaning - embed the incoming question with RETRIEVAL_QUERY, compare it to every stored vector via cosine similarity, and return the top-k.

  4. Generate a grounded answer - paste those top-k chunks into the augmented prompt and call the chat model.

def ask(question, top_k=3):
# A. retrieve the most relevant chunks
results = search(question, top_k=top_k)
context = "\n\n".join(documents[i] for i, _ in results)
# B. build the augmented prompt
prompt = f"""Answer based ONLY on the context. If it's not there,
say "I don't have enough information."
<context>
{context}
</context>
Question: {question}"""
# C. generate
response = client.models.generate_content(
model="gemini-2.5-flash-lite",
contents=prompt,
)
return response.text

7. Vector databases - when lists stop scaling

Section titled “7. Vector databases - when lists stop scaling”

The Python-list approach is perfect for learning, but it scans every vector on every query and forgets everything on restart. Real deployments use a vector database - storage purpose-built for embeddings.

ChallengeIn-memory listsVector database
100,000+ documentsslow (checks every one)fast (indexed)
Persistencelost on restartsaved to disk
Concurrent usersnot supportedbuilt-in
Metadata filteringmanualnative queries
Updatesrebuild everythingadd/remove one at a time

The speed comes from indexed search: instead of comparing the query against every vector, algorithms (with names like HNSW or IVF) pre-organise the vectors into clusters or graphs so a query only checks a small, promising subset. It’s the vector-space cousin of the old keyword “inverted index”. Popular options range from ChromaDB (easiest to start), through FAISS and pgvector (open-source, self-hosted), to Pinecone, Weaviate, and Qdrant (managed / full-featured). A sensible path: prototype with ChromaDB, move to a managed service when you need scale.

8. RAG vs fine-tuning - why RAG wins for enterprise Q&A

Section titled “8. RAG vs fine-tuning - why RAG wins for enterprise Q&A”

A natural question: if the model doesn’t know my data, why not just fine-tune it - retrain it on company documents until it “knows” them? For enterprise question-answering, RAG almost always wins. Here’s the honest comparison:

DimensionRAG (retrieve at query time)Fine-tuning (retrain on your data)
Updating knowledgeadd/replace a document - instantretrain the model - slow, costly
Source citationsyes - points to the exact chunkno - knowledge is baked in, untraceable
Hallucination controlstrong - answers are grounded in retrieved textweaker - still generating from memory
Cost & skilllow - no training pipeline neededhigh - GPUs, ML expertise, data prep
Access controlfilter chunks by user permissionshard - everyone gets the same baked-in model
Best atinjecting facts/knowledgeteaching style, format, or tone

The key insight: fine-tuning changes how the model behaves; RAG changes what the model knows right now. Enterprise Q&A is a knowledge problem - the facts change weekly, must be citable, and must respect who’s allowed to see what. That’s RAG’s home turf. (The two aren’t rivals, really: fine-tune to fix tone, use RAG to supply facts.)

9. Making RAG better - advanced techniques

Section titled “9. Making RAG better - advanced techniques”

Basic RAG (embed → retrieve → generate) works impressively well. Production systems then layer on refinements - worth recognising even if the labs stick to the basics:

TechniqueProblem it fixes
Rerankingfast retrieval grabs roughly relevant chunks; a slower, sharper second-pass model re-scores the top ~20 down to the best 3
Hybrid searchpure semantic search can miss exact terms (an acronym, “ISO 27001”); combine it with keyword/BM25 to catch both
Query transformationrewrite a vague question before searching - expand terms, or ask a broader “step-back” question first
Multi-hop RAGchain several retrievals for questions that span documents (“which of our top-3 customers has the longest contract?”)
Query routinglet the LLM decide which source to search - knowledge base vs CRM vs calendar
Agentic RAGlet the LLM decide whether, what, where, and when to retrieve - the doorway to Day 4

10. Evaluation - is my RAG actually any good?

Section titled “10. Evaluation - is my RAG actually any good?”

A RAG system can fail in two different places: retrieval (fetched the wrong chunks) or generation (had the right chunks but answered badly). You have to measure both. The course’s framework is the RAG triad - three complementary scores:

Context RelevanceGroundednessAnswer Relevance
MetricThe question it answersHow to check
Context relevanceDid we retrieve the right chunks?score each retrieved chunk against the query
GroundednessIs the answer supported by those chunks?trace every claim back to a chunk
Answer relevanceDoes the answer actually address the question?score the answer against the original question

You need all three because each catches a different failure: perfect chunks can still produce a wrong answer (low groundedness), and a beautifully written answer can miss the question entirely (low answer relevance). To score at scale, the labs use LLM-as-judge - a second LLM rates each dimension 1-5 and returns JSON. For retrieval specifically, classic metrics help too: Precision@k (what fraction of retrieved chunks were relevant) and Recall@k (what fraction of the relevant chunks we managed to find) - both need a golden test set of real questions with known correct answers.

PitfallSymptomFix
Chunks too smallretrieved text lacks contextbigger chunks or more overlap
Chunks too largeretrieved text is padded with noisesmaller chunks, recursive splitting
No grounding rulemodel ignores context, hallucinatesadd explicit “answer ONLY from context”
Wrong kanswer incomplete, or prompt floodedtune how many chunks you retrieve
Stale dataknowledge base out of dateschedule regular re-indexing
Access-control leakageuser sees restricted answersfilter by permissions before retrieval

RAG shines wherever valuable knowledge is trapped in documents that are hard to search:

Use caseKnowledge sourceBusiness impact
Internal knowledge basewikis, handbooks, SOPsanswers in seconds, not hours
Customer supportproduct docs, FAQs, ticketsfaster, more consistent replies
Legal & compliancecontracts, regulationsinstant clause lookup, less risk
Research & analysisreports, papers, market datasynthesise across hundreds of docs
Sales enablementcase studies, pricingreps find the right material instantly

When to reach for RAG (and when not to): use it when knowledge changes often, spans many documents, gets queried unpredictably, and demands accuracy with citations. Skip it when the data is tiny and static (just use the context window), the queries are highly predictable (use templates), or general knowledge already suffices.

Day 3’s practice takes the theory all the way to a documented, evaluated RAG system.

Building a RAG pipeline (instructor-led, ~60-75 min). Step by step: create embeddings with the Gemini API, implement cosine similarity and nearest-neighbour search, chunk a document (fixed-size vs sentence-based, with overlap and metadata), assemble the full retrieve → augment → generate pipeline, then evaluate it - the RAG triad, LLM-as-judge scoring, and Precision@k / Recall@k. Includes the telling “with vs without context” hallucination demo and a refusal test on out-of-scope questions.

Download the notebooks (open in Google Colab):

My submitted solution: Assignment 3 - my solved notebook (.ipynb) - my own work from the course.

Must-knowOne-line recall
Why RAG existsLLMs are frozen at training - they miss private, current, and niche facts, and hallucinate to fill the gap
What RAG doesretrieve the relevant chunks from your data, then let the LLM answer using them (open-book exam)
Embeddingstext turned into vectors where similar meaning sits close together
Cosine similarityone number for how alike two vectors are; ~1.0 identical, below 0.3 unrelated
Task typesRETRIEVAL_DOCUMENT for stored docs, RETRIEVAL_QUERY for questions - matching them lifts accuracy
Chunkingsplit long docs into 200-500-token pieces with 50-100-token overlap so no fact is cut in half
Metadatatag each chunk (source, page, date) to enable citations, filtering, and debugging
The pipelineindexing offline (chunk → embed → store) + querying online (embed → retrieve → augment → generate)
Augmented promptfence retrieved chunks in delimiters and instruct “answer ONLY from context” - pure Day 2 prompting
Citationsmake the model cite [Source N] so answers are verifiable; define an authority hierarchy for conflicts
Vector databaseneeded at scale for speed, persistence, filtering, and easy updates (ChromaDB → Pinecone)
RAG vs fine-tuningRAG injects facts (cheap, instant, citable, permission-aware); fine-tuning teaches style
RAG triadcontext relevance + groundedness + answer relevance - measure all three to catch every failure
Biggest mistakenot testing with real, messy, out-of-scope user questions