Skip to content

Day 1 - LLMs & Transformers: The Engine Behind GenAI

Generative AI for Managers - TUHH Institute of Entrepreneurship · part of my Technology Management MBA · study notes for revision.


Day 1 was the “lift the hood and look at the engine” session. No maths, no coding - just building an honest mental model of what a large language model (LLM) really is, so that everything we do later (prompting, retrieval, agents, business cases) rests on solid ground. These are my revision notes, written the way I’d explain it to a colleague over coffee.

The single most useful idea to hold on to before we start: an LLM is not a database of facts and it is not a thinking person. It is a very, very good next-word guesser. Almost every strength and every weakness we’ll see comes straight out of that one sentence.

AI didn’t appear overnight. It moved through a few clearly different eras, and each one changed who did the hard work of figuring out the rules.

EraApproachWhat it meant in practice
1950s-1990sHand-coded rulesA human wrote every “if this, then that” by hand - e.g. if the message says “refund”, send it to the returns team.
1990s-2010sMachine LearningHumans picked useful features; the computer learned the weights from labelled examples (classic spam filters).
2010sDeep LearningNeural networks learned the patterns themselves from labelled data (image recognition, voice).
2017-nowLarge Language ModelsNetworks learn language directly from enormous amounts of raw text - no hand-written rules, no manual features.

The key shift is about effort. Older systems needed a human to define the rules or hand-pick the important signals. LLMs skip that: they read raw text at a scale no human ever could and work out the patterns on their own. That’s the leap that made today’s chatbots possible.

2 · The core idea: predicting the next token

Section titled “2 · The core idea: predicting the next token”

Strip away the hype and an LLM does exactly one thing: it looks at some text and predicts what comes next.

Give it “The capital of France is” and it predicts “Paris” with high confidence. That’s it. But here’s the clever part - to get good at that one game, the model is forced to absorb grammar, facts, reasoning patterns, writing styles, even the structure of computer code. You can’t reliably guess the next word across billions of sentences without picking all of that up along the way.

So how does “guess one word” turn into a full paragraph? Through a loop called autoregressive generation - a fancy term for a simple habit: predict a token, stick it onto the end, then predict again using the longer text. The model keeps eating its own output.

Prompt”The capital of France is”
→
Predict 1 token”Paris”
→
Append & repeat”…Paris. It is…”
→
Stopend token or length limit
Text is generated one token at a time; each new token is fed back in before the next is chosen. The loop stops at a natural end point or a maximum length.
  1. Start with your input: “The capital of France is”.
  2. Predict the next token and add it → “The capital of France is Paris”.
  3. Feed the whole thing back and predict again → ”…Paris.”
  4. Keep looping → ”…Paris. It is” → ”…Paris. It is the largest…”
  5. Stop when the model emits an “end” signal or hits the length limit you set.

This is worth internalising because it explains a lot of everyday behaviour: the model has no plan for the whole answer up front, it’s committing to one token at a time. That’s why a small wrong turn early can snowball, and why the same prompt can produce different answers on different runs.

3 · Tokens: the chunks the model actually reads

Section titled “3 · Tokens: the chunks the model actually reads”

We keep saying “token” instead of “word” on purpose, because LLMs don’t read whole words. They read tokens - pieces of text that are often, but not always, a full word.

TextHow it splits into tokens
Hello["Hello"] - one common word, one token
ChatGPT["Chat", "G", "PT"] - an unusual string, split into pieces
unhappiness["un", "happiness"] - a prefix plus a known chunk

Why chop words up like this? Because common pieces show up far more often than rare whole words. Working in reusable chunks keeps the model’s vocabulary manageable and lets it handle words it has never seen before by assembling them from familiar parts.

3.1 How tokenising works - BPE in one line

Section titled “3.1 How tokenising works - BPE in one line”

Most modern models use Byte Pair Encoding (BPE). The idea is refreshingly simple: start with individual characters, then repeatedly glue together the pair that appears most often, over and over. Frequent combinations (like th, ing, or whole common words) end up as single tokens; rare stuff stays broken into smaller pieces.

Tokens are not an academic detail - they’re the unit that gets billed, capped, and timed, which is exactly why they come back later as a “practitioner knob.”

Here’s a problem that sounds obvious once stated: computers can’t do anything with the letters “c-a-t”. They only work with numbers. So every token has to become a number - actually a list of numbers - before the model can process it. The trick is doing that meaningfully.

The simplest idea is to give each word its own slot - this is called one-hot encoding:

  • cat = [1, 0, 0, 0, …]
  • dog = [0, 1, 0, 0, …]
  • bank = [0, 0, 1, 0, …]

It works, but it’s dumb in one specific way: it treats every word as equally unrelated to every other. By this scheme, “cat” is no closer to “dog” than it is to “bank” - which is nonsense. There’s no notion of meaning anywhere in these numbers.

Instead of arbitrary slots, we learn a rich list of numbers (a vector) for each token, arranged so that words with similar meanings sit close together and related pairs point in similar directions. That learned representation is called an embedding.

The famous demonstration of how much meaning is captured:

king
−
man
+
woman
≈
queen
The step from “man” to “woman” captures the idea of gender; adding that same step to “king” lands you next to “queen”. Meaning becomes arithmetic.

That only works because the direction from man to woman encodes “gender”, and adding it to king moves you to the female equivalent. The same pattern holds for verb tense (walk → walked) and country → capital (Germany → Berlin). Meaning has literally become geometry.

Embeddings aren’t a party trick - they quietly power a lot of practical tools, because “close together in number-space” is a machine-friendly stand-in for “similar in meaning”.

ApplicationHow embeddings help
SearchFind documents that mean the same thing, not just ones sharing keywords
Recommendations”Customers who liked X also liked Y” - X and Y sit near each other
ClusteringAutomatically group similar customer feedback or tickets
RAG (Day 3)Fetch the most relevant context to feed the model before it answers

We now have tokens turned into meaningful numbers. The transformer is the architecture that processes them - and it’s the “T” in GPT. It arrived with a 2017 paper titled “Attention Is All You Need”, and it’s genuinely the reason modern GenAI works.

Before 2017, models read text strictly left to right, one word at a time. That had three nagging problems:

  • Slow - each word had to wait for the one before it; you couldn’t process them together.
  • Forgetful - details from early in a long passage faded by the time the model reached the end.
  • Short-sighted - long documents were a real struggle.

The transformer fixed all three with two moves: it looks at all the words at once (so it’s parallel and fast), and it uses self-attention to let every word decide which other words are relevant to it.

Attention answers one question for each word: “which other words should I pay attention to in order to understand myself?”

The classic example: “The cat sat on the mat because it was tired.” What does “it” refer to? You knew instantly - the cat, not the mat. Attention is the mechanism that lets the model learn and apply exactly that kind of connection.

Each word produces three things, and the easiest way to remember them is the library analogy:

RoleLibrary analogyIn attention
Query (Q)Your question: “I need books on climate change”What this word is looking for
Key (K)The label on each book’s spineWhat each word advertises about itself
Value (V)The actual contents of the bookThe information each word can hand over

You walk in with a Query, scan every book’s Key (label) to see what’s relevant, then take Value (content) from the books in proportion to how well they matched. Word-processing works the same way: compare, weight, then blend.

5.3 Worked example: resolving “it” → “cat”

Section titled “5.3 Worked example: resolving “it” → “cat””

Let me trace it for the word “it” in “The cat sat because it was tired”:

  1. “it” forms a Query - essentially asking “what do I refer to?”
  2. Every word offers a Key - “cat” advertises “I’m an animal, a noun”, “sat” advertises “I’m an action”, and so on.
  3. Compare the Query against every Key to score relevance:
    • it vs “The” → 0.05 (low)
    • it vs “cat” → 0.72 (high - best match!)
    • it vs “sat” → 0.15 (medium)
    • it vs “because” → 0.08 (low)
  4. Normalise the scores (a step called softmax) so they add up to 1.0 - now they’re clean weights.
  5. Blend the Values using those weights. Since “cat” carries 0.72 of the weight, the result is mostly information from “cat”.

Result: the model’s internal representation of “it” now effectively contains “cat”. That’s a contextual embedding being built in real time - the same word would blend differently in a different sentence.

5.4 Why three separate versions of each word?

Section titled “5.4 Why three separate versions of each word?”

A fair question: why not just use the word’s embedding directly for all three roles? Because a word plays different roles depending on context. Take “bank”:

  • As a Query it needs to ask: “am I money-bank or river-bank?”
  • As a Key it advertises: “I could be financial or geographical.”
  • As a Value it provides: “here’s my actual meaning.”

Asking, advertising, and delivering are genuinely different jobs, so the model learns three separate transformations - one for each.

5.5 Multi-head attention and stacking blocks

Section titled “5.5 Multi-head attention and stacking blocks”

Real transformers don’t run attention just once. They run several attention heads in parallel, each a different “perspective” on the same sentence - one head might track grammar, another meaning, another simple word-position. Running many heads lets the model spot many kinds of relationship at once (a large model might use dozens of heads).

And attention is only part of one transformer block. A block bundles together:

Self-attentionwho relates to whom
→
Feed-forwardprocess the result
→
Residual + normkeep info stable
One transformer block: attention finds the relationships, a feed-forward network digests them, and residual connections plus normalisation keep earlier information intact and training stable.

Then these blocks are stacked - modern models pile up dozens of them. Each layer adds a level of abstraction: early layers handle grammar and word relationships, middle layers build up meaning and facts, and later layers do the heavier reasoning. Depth is where sophistication comes from.

A finished, helpful assistant is built in three distinct phases. Skipping any one of them gives you a very different (and worse) product.

1 · Pre-traininglearn language
→
2 · Instruction tuninglearn to be helpful
→
3 · RLHFlearn human preferences
Raw language ability comes first, then the habit of following instructions, then fine polish on tone, safety and helpfulness.

The model reads an enormous pile of text - books, Wikipedia, websites, forums, academic papers, code, news - and plays the next-token game over and over. Through nothing but that game it picks up grammar, facts, reasoning patterns, style, and code structure.

Scale to appreciate: trillions of words, thousands of specialised chips running for months, and an estimated cost of roughly $10M-$100M+ for a frontier model. This is why only a handful of organisations train models from scratch.

Analogy: it’s like learning to write by reading every book in every library, millions of times over - you’d absorb grammar and facts without anyone teaching you a single rule.

These are the handful of terms you’ll actually touch when working with a model. Understand these four and you can hold your own in any GenAI conversation.

Because everything is measured in tokens, tokens are the currency of working with LLMs:

  • Pricing - APIs charge per token, counting both your input and the model’s output.
  • Limits - every model has a maximum it can handle at once (see context window below).
  • Speed - more tokens means more processing time.

The money angle is real: a verbose 1,000-token prompt costs 5× a tight 200-token one for the same job. Concise prompts, caching repeated requests, and using a smaller model for simple tasks are the easy wins.

7.2 Context window = the size of the model’s desk

Section titled “7.2 Context window = the size of the model’s desk”

The context window is the maximum amount of text the model can “see” at one moment - and crucially it covers your input and the generated output together.

ModelContext windowRoughly
GPT-3.516K tokens~12,000 words
GPT-4128K tokens~96,000 words
Claude 3200K tokens~150,000 words
Gemini 1.51M+ tokens~750,000 words

Temperature controls how random the output is. Low temperature makes the model play it safe and pick the most likely next token every time; high temperature lets it take chances.

Use caseRecommended temperature
Data extraction0.0 - 0.2 most precise
Summarisation0.2 - 0.4
General Q&A0.5 - 0.7
Creative writing0.7 - 0.9
Brainstorming0.9 - 1.0 most varied

At 0.0 the model is deterministic - same input, same output. Around 0.7-0.9 you get useful creative variety. Push past 1.0 and it often turns incoherent.

Parameters are the internal learned values - think of them as millions or billions of tiny knobs that training adjusted. Rough sizing:

SizeParametersTypical use
Small7-13 billionFast and cheap; good for simple tasks
Medium30-70 billionA balance of capability and cost
Large100B+Frontier capability, highest cost and latency

More parameters generally means more knowledge, more nuance, and better reasoning - but also more money, slower responses, and heavier hardware.

8 · What LLMs do well - and where they trip

Section titled “8 · What LLMs do well - and where they trip”

Being clear-eyed here is the whole point of the day. LLMs are astonishing at some things and quietly unreliable at others, and a manager needs to know which is which.

Good atStruggles with
Generating fluent text (emails, reports, copy)Precise factual accuracy
Following simple instructionsComplex multi-constraint tasks
Understanding contextReal-time / recent information
Summarising and extractingExact counting and arithmetic
Creative variationsVery long, perfectly consistent output

The pattern: LLMs excel at tasks humans do with language. If a job can be phrased as “read this, then write that,” an LLM can probably help. Now the four failure modes worth naming - each with its fix:

What: the model states confident but false information - e.g. inventing plausible-sounding research paper titles that don’t exist.

Why: it predicts likely-sounding text; it is not looking anything up in a database of facts. Fluent and correct are not the same thing.

Mitigation: Retrieval-Augmented Generation (RAG) - give the model the real source documents to answer from. That’s the whole of Day 3.

The field moves fast, but the main players and the three architecture families are worth knowing.

CompanyModel familyNotable for
OpenAIGPT-4, GPT-4oPioneered modern LLMs; strong reasoning
GoogleGeminiMultimodal (text + images); tied into Google
AnthropicClaudeEmphasis on safety and helpfulness
MetaLlamaOpen weights - you can run it yourself
MistralMistral, MixtralEuropean; efficient open models

Under the hood, transformer-based models come in three shapes depending on the job:

ArchitectureExamplesBest for
Decoder-onlyGPT, Claude, LlamaGenerating text - chat, writing
Encoder-onlyBERT, RoBERTaUnderstanding text - classification, search
Encoder-decoderT5, BARTTranslation, summarisation

Almost every chatbot you’ve used (GPT-4, Claude, Gemini) is decoder-only - built to generate, one token at a time, exactly as we traced back in section 2.

Day 1 pairs the theory above with time in Google Colab, calling the Gemini API for real. Full working code lives in the notebooks - download them and open in Colab to run along.

The guided lab - Your First LLM Interactions - is instructor-led and walks through: setting up the environment and API key safely, making your first calls and reading the response, experimenting with temperature (comparing 0.0 vs 1.0 outputs) and max tokens, building a multi-turn conversation that remembers context, and assembling two small business tools - an email-response generator and a meeting-summary generator. It ends by exporting a prompt log so every experiment is tracked.

The independent lab and assignment then let you practise on your own: pick a business use case, apply the same parameters and multi-turn patterns yourself, and log and review what you build - reinforcing the four ground rules from the lab (if unsure, say so; specify the output format; don’t invent facts; verify important outputs).

Download the notebooks (open in Google Colab):

My submitted solution: Assignment 1 - my solved notebook (.ipynb) - my own work from the course.

Must-knowOne-line recall
What an LLM isA very good next-token predictor - not a fact database, not a person.
Autoregressive generationPredict one token, append it, repeat - no full-answer plan up front.
TokensThe chunks models read; 1 token ≈ 4 chars ≈ ¾ word; they drive cost, limits, speed.
BPEBuild a vocabulary by repeatedly merging the most frequent character pairs.
EmbeddingsWords as meaningful vectors - king − man + woman ≈ queen; power search, recs, clustering, RAG.
Attention (Q/K/V)Query asks, Keys answer, Values deliver - how “it” learns it means “cat”.
Multi-head + stackingMany perspectives per layer; stacked blocks build grammar → meaning → reasoning.
Three training stagesPre-training → instruction tuning → RLHF (language → helpfulness → human taste).
The four knobsTokens (cost), context window (capacity), temperature (creativity), parameters (size).
TemperatureLow (0-0.2) for facts; high (0.7-1.0) for creativity; 0.0 = deterministic.
LimitationsHallucination, maths/counting, real-time info, complex constraints - fix with RAG, code, search, step-by-step.
ArchitecturesDecoder-only (generate) · encoder-only (understand) · encoder-decoder (translate).
The golden ruleLLMs augment, not replace - always verify high-stakes outputs.