Skip to content

Agile, Scrum & Working with AI

SAP Implementation Consulting & ERP - Startupistan Germany · Block I · study notes for revision.


Until now I’ve mostly coded alone: my branch, my pace, my mess. Real software is built by teams, over years, for customers who change their minds. That creates two separate problems, and it helps to keep them apart:

  • The code problem - two people editing the same files. Solved by version control (branches, pull requests). I already know this.
  • The coordination problem - deciding what to build, in what order, and how we know it’s done. Solved by a shared working method.

Git tells me how changes merge; it says nothing about what should exist or when it’s finished. That second half is what this whole chapter is about.

Before agile, software was built like a bridge: plan everything, then execute. That’s the waterfall model - fixed phases in strict sequence, each finished completely before the next starts (Requirements → Design → Implementation → Testing → Delivery). Water only flows one way; going back up is expensive and rare.

It has one structural flaw for software: real feedback arrives at the end, when change costs the most. Requirements drift over a long project, and software is invisible until it runs - so a customer can’t react to a half-built system the way they can walk through a half-built house. The misunderstanding survives until delivery.

Agile fixes this with the obvious move: make the loop short. Work in small iterations (typically 1-4 weeks), each ending in a piece of working software shown to real users, whose reactions shape the next iteration. Each small delivery is really an experiment - it tests whether the team understood the need, months before waterfall would find out.

AxisWaterfallAgile
Working software appearsOnce, at the endEnd of every iteration
Changing the planException (re-open a signed-off spec)Normal, expected between iterations
Risk discoveredLate (test/delivery phase)A little each iteration, while cheap to fix
Progress meansDocuments & phases completedWorking software delivered
The core betWe’ll get it right by planningWe’ll get it right by feedback

Waterfall isn’t dead: where requirements are genuinely fixed and a defect costs lives - medical device firmware, aerospace control - waterfall-like process still fits. It’s wrong for most business software, which is exactly where requirements never stop moving.

Agile has an official birthplace: seventeen practitioners met in 2001 and wrote down what their competing methods had in common - four values and twelve principles. Unchanged since, it’s still the definition of agile.

The core is four “X over Y” comparisons. The most misread word is over: it does not mean instead of. The manifesto’s own closing line settles it - there is value in the right-hand items; the left-hand ones are valued more when the two conflict.

Valued moreoverStill valuable
Individuals and interactionsoverprocesses and tools
Working softwareovercomprehensive documentation
Customer collaborationovercontract negotiation
Responding to changeoverfollowing a plan

The twelve principles turn those values into observable daily behaviour. Twelve is a lot to hold, so I group them into five themes:

ThemeWhat it asks of a team
Deliver early & oftenShip valuable working software frequently (weeks, not months); working software is the primary measure of progress
Welcome changeEmbrace changing requirements even late - a change is information, not an enemy
People & trustBuild around motivated people, prefer face-to-face talk, business + devs work together daily
Sustainable pace & qualityA pace you can hold forever, plus technical excellence and simplicity (maximise work not done)
Reflect & adjustAt regular intervals the team tunes its own behaviour - the process improves by iteration too

Values are like a constitution (short, stable, abstract); principles are the laws that show what it means in daily situations. Neither is a checklist you pass or fail - they’re directions of travel.

Waterfall captured requirements in long specs. Agile needed something one iteration can swallow, written from the point of view of the person the software serves. That’s the user story - the most common work item on real teams. It’s one sentence with three slots:

As a [role] → WHO wants this (a real kind of user, not "the system")
I want [capability] → WHAT they need to be able to do
so that [benefit] → WHY it matters (the justification)

Example: As a learner, I want to save my quiz score, so that I can see my progress over time. One sentence, and the team already knows who benefits, what to build, and how to judge design choices.

The so that slot looks like decoration but is the most powerful of the three: if no honest benefit can be written, maybe the story isn’t worth building; and knowing the goal lets the team propose a better solution (maybe a progress chart beats a raw score list). A story states a need - it deliberately doesn’t lock in a design. It’s “a placeholder for a conversation,” not a mini spec.

The sentence says why and what, but not when it’s done. That’s the job of acceptance criteria - a short list of concrete, binary (true/false) statements that must all hold for the story to be accepted:

  • The score is stored after every completed quiz.
  • The latest score appears on the learner’s profile page.
  • A failed save shows the learner an error message.

Each is something a stranger could verify with a plain yes or no - that’s what turns “I think it’s finished” into a checklist. Acceptance criteria describe observable behaviour, never internal design.

An iteration is a fixed box of time, so a team must judge how much fits. Humans are famously bad at predicting hours for creative work - but very good at comparing sizes. So agile teams stop estimating time and estimate relative size instead. (I can’t tell you a shelter dog’s weight, but I can reliably line ten dogs up lightest to heaviest.)

  • Story points = a number for the overall size of a story - effort, complexity and uncertainty combined - on a shared team scale. Two rules: points are relative (an 8 is roughly 4× a 2; nothing maps to hours) and team-local (one team’s 3 ≠ another’s; comparing between teams is meaningless).
  • Over a few iterations a team learns how many points it actually completes per iteration. That observed number - a forecast, not a promise - is what makes the next plan realistic.
  • Planning poker = everyone privately picks an estimate card, then all reveal at once. The simultaneous reveal stops the first loud voice from anchoring everyone, and the disagreements it exposes are usually hidden assumptions - exactly the conversation worth having.

Scrum is the most widely used framework built on agile values - a deliberately incomplete one. It doesn’t tell you how to design, code or test; it defines just enough structure to make agile operational: 3 accountabilities, 3 artifacts, 5 events, all wrapped around one container, the Sprint. Everything else the team adds itself.

One small team, usually ten people or fewer, no hierarchy among the three. (The 2020 guide renamed “roles” to “accountabilities” - same content; “accountability” names what you answer for, not a job title.)

AccountabilityOwnsIn one line
Product OwnerThe what & whyMaximises product value by owning and ordering the one Product Backlog - the hard part is saying “not yet” to most things
Scrum MasterThe processMakes the team effective, coaches good Scrum, and removes impediments - holds no authority over people or priorities
DevelopersThe howEveryone building the Increment (coding, testing, design…); they organise their own work - nobody assigns them tasks

The three artifacts (+ Definition of Done)

Section titled “The three artifacts (+ Definition of Done)”
ArtifactWhat it isOwned by
Product BacklogThe single, ordered, never-finished list of everything the product might need (mostly user stories); top items small & clear, lower ones allowed to stay vagueProduct Owner
Sprint BacklogThe Developers’ plan for this Sprint - the items they picked plus how they’ll deliver them; a living picture updated dailyDevelopers
IncrementThe accumulated, usable result - product that works, whether or not it’s releasedDevelopers

The Definition of Done is the team’s shared written quality checklist (e.g. code reviewed, tests passing, docs updated). Until an item meets every entry, it isn’t done and doesn’t join the Increment. Key distinction: acceptance criteria are per-story (what this feature must do); the Definition of Done applies to all work (what quality every feature must meet). A story must satisfy both.

The Sprint is itself the first event - the fixed time-box (one month or less, commonly two weeks) that contains the other four. Sprints run back-to-back, giving the team a steady heartbeat.

EventMax length (2-wk Sprint)Purpose
SprintThe containerTurn selected backlog items into a done Increment
Sprint Planning~half a dayCraft the Sprint Goal, pull items into the Sprint Backlog, sketch the plan (why / what / how)
Daily Scrum15 minDevelopers inspect progress toward the goal and re-plan the next 24h - by them, for them
Sprint Review~2 hoursDemo the working Increment to stakeholders; the discussion re-orders the backlog
Sprint Retrospective~1.5 hoursThe team inspects itself - what to improve next Sprint

The Review inspects the product; the Retrospective inspects the process and people. None of these are status reports to a manager - they’re the crew steering the ship.

Product Backlogordered wish list
→
Sprint Planninggoal + selection
→
Sprintdaily scrums
→
Incrementdone & usable
→
Review + Retroinspect & adapt
The Scrum loop - Review and Retrospective feed straight back into the Product Backlog, and the next Sprint starts the moment this one ends.

Scrum isn’t the only agile method. Kanban thinks in continuous flow instead of fixed time-boxes. A board makes work visible with columns (stages) and cards (work items) - minimally To Do → In Progress → Done, one card in exactly one place at a time.

Its sharp edge is the work-in-progress (WIP) limit: a hard cap on how many cards a column may hold. If “In Progress” is limited to two and both slots are full, nobody starts a third - the job becomes finish something, not start something. Why so strict? Five half-done tasks deliver nothing, and switching between them burns time; the limit converts that invisible waste into a visible constraint, and forces the question “why is this card stuck?” the moment it sticks.

ScrumKanban
RhythmFixed Sprints with a goalContinuous flow, one item at a time
Roles / eventsRequired (3 + 5)None required
Change mid-cycleBetween SprintsAny time capacity frees up
Best fitGoal-shaped product developmentInterrupt-driven work: support, ops, maintenance

Useful first question on any team: does our work arrive as a stream (lean Kanban) or can it be shaped into goals (lean Scrum)? Many teams blend both (“Scrumban” - Sprints + WIP limits). Without WIP limits, Kanban is just a board with sticky notes; the limit is what makes it a method.

Zooming all the way in: the single unit of work a developer holds is a ticket - one unit of work in a tracking tool, carrying a title (one line), a description (the why and context), and acceptance criteria (the checkable done). Before coding, I check: do I understand the why, and could I verify every criterion? If not, the right move is a question to the author, not a guess - guessed requirements are how wrong features get built with perfect code.

The path from ticket to done is the same industry-standard flow I already used per-exercise in the JS course:

  1. Ticket picked up → a branch is created for it, named after the ticket (one branch per task, however small).
  2. Work happens as small, described commits on that branch - safely away from main.
  3. A pull request is opened: the formal request to merge, presenting all changes for inspection.
  4. A colleague reviews it; feedback flows back as new commits on the same branch.
  5. Approved work is merged, and the ticket moves to Done - meaning acceptance criteria pass and the Definition of Done is met.

Everyone gets stuck; what separates people is what happens next. A question someone can actually answer carries three things:

  1. The goal - what I’m trying to achieve, one level above the bug (the real problem, not the fix I imagined).
  2. What I tried - the steps and code, so nobody suggests what I already ruled out.
  3. The exact error text - copied and pasted, never paraphrased. “It says null pointer somewhere” destroys the single most diagnostic thing I own.

Naming the goal first avoids the classic trap of asking about my attempted solution instead of my actual problem: ask “how do I parse this date string with regex?” and people debug my regex; say “I need the year from this date value” and someone hands me a one-line date method that erases the whole struggle.

Most errors carry four readable parts - type (e.g. TypeError), message (read word by word - it means what it says), location (file + line), and stack trace (top line = where it broke, lines below = how it got there). The gold-standard attachment is a minimal reproducible example: the smallest complete code that still fails. Building one often solves the problem itself (the rubber-duck effect). Ask order: docs → search the error → a teammate → a public forum - and asking a teammate after ~15 minutes of honest effort is the correct professional move, not weakness.

The trick is matching the type of document to the question:

TypeForRead it how
TutorialI’m new to the whole areaStart to finish, learning by doing
Guide (how-to)I have one concrete task, know the basicsJump to the task
ReferenceI need exact behaviour of one function/optionLook up the piece - never read cover to cover

Most documentation frustration is a mismatch - a tutorial is maddening when I need one detail; a reference is impenetrable when I needed a path. When jargon appears, resolve it immediately (a personal glossary helps), and remember the senior habit: skim changelogs / release notes when something that worked yesterday breaks after an update.

Review is the most frequent collaboration ritual in software. A reviewer reads a pull request with three questions, in this order:

  1. Correctness - does it do what the ticket claims, meet the criteria, handle edge cases (empty input, failed request)?
  2. Clarity - will the next person understand it in a year without the author? Clear names, small functions.
  3. Consistency - does it follow team conventions and the Definition of Done? (Formatting is automated away so reviews spend attention on what machines can’t judge.)

The whole craft of giving feedback: comment on the code, never the coder. Ask, don’t command (“what happens here if the list is empty?” beats “this is wrong”); explain the why; say what’s good; mark must vs could. Receiving it: read everything before replying, treat questions as questions, thank the catch (a bug found now is cheapest), and disagree openly and professionally when you do. A review with zero comments usually means a rushed reviewer, not flawless code.

Everything above is for teams - but the one project I’m guaranteed to manage is my own learning. Same tool, team of one: To Do → In Progress → Done, honest cards (a card is well-formed if I can say yes/no whether it’s done). To Do ordered by value is my personal backlog; the WIP limit should be brutal - two at most. Both slots full and tempted by something new? Finish something first. The Done column is the motivation instrument on a hard day. GitHub Projects is the natural home since it lives where my work lives.

A large language model (LLM) - ChatGPT, Claude, Gemini and relatives - is a program trained on enormous amounts of text to do one thing: given some text, predict what text plausibly comes next. That’s the whole trick, at colossal scale. Two consequences drive everything else:

  • It has no database of facts it looks up. It produces text that’s statistically plausible given training. Plausible and true overlap heavily (that’s why it’s useful) but aren’t the same.
  • It’s genuinely excellent at language-shaped work: explaining, summarising, rephrasing, translating between levels, and generating structured text including code.

A hallucination is generated content presented as fact but false or invented - and it’s the most important habit in this whole topic. In code, the classic specimens: a method that doesn’t exist but sounds right; config options no version accepts; a confident wrong explanation of why code fails; invented sources and version numbers. They aren’t random noise - they’re maximally plausible falsehoods, engineered by the very mechanism to sail past my defences.

The professional response fits in one sentence: no claim from an assistant is trusted until verified against an authoritative source (official docs, MDN, the tool’s own site).

  1. Ask, and read the answer.
  2. Check the load-bearing claims against official documentation.
  3. If confirmed → use it with understanding, because I just read the real source too.
  4. If wrong → tell the assistant what the docs actually say, and re-ask. The corrected conversation usually lands, and now I know this corner better than either of us did.

The reframe that makes this click: verification is the learning. The assistant narrows an open search to a single yes/no lookup; the docs visit is where durable knowledge is built. Learners who verify build real knowledge from every exchange; learners who paste unverified answers build a pile of code they can’t explain.

The assistant can only respond to the text I give it - the prompt - and it’s a learnable structure I already know from user stories and good questions. A strong prompt carries four parts:

  1. Context - who I am, what I already know, the situation. The assistant can’t see me; this slot shows it my level, and an answer pitched at my level is the whole point.
  2. Goal - what I actually want produced, stated directly.
  3. Constraints - the shape of a good answer: length, format, what to include/avoid, which tech is in or out of bounds.
  4. Example - where format matters, one small sample of good output. Nothing communicates a format better than one instance.

Then iterate, don’t restart: when an answer misses, steer (“too advanced - explain as if I’d never seen a loop”; “compress to five sentences”). Each correction narrows the next answer; restarting throws away everything the conversation established. The single highest-value constraint costs one sentence: state what I know. “I know HTML/CSS but no JavaScript yet” changes vocabulary, pace, and examples entirely. Prompts that land are worth saving into a personal prompt library.

AI as a learning partner - and the copy trap

Section titled “AI as a learning partner - and the copy trap”

Four workflows that actually build skill:

WorkflowThe move
Explainer”Explain this line by line at my level; I know X but not Y” - then push: why this way? what breaks if line 3 moves?
Exercise generator”Generate five exercises slightly harder than this” - solve them myself, then return for review
Debugging partnerGoal + attempt + exact error - then “explain the error and where to look; don’t fix it yet”
Quizmaster”Quiz me, one question at a time, then tell me what my wrong answers reveal I misunderstood” (retrieval is elite study)

The moment I work for an employer, a fourth party enters every prompt: the company whose information I’m holding. What I type is sent to the provider’s servers and may be stored or reviewed. Fine for a question about closures - a serious professional failure for three categories:

Never paste without explicit permissionExample
Secrets & credentialsPasswords, access keys/tokens, connection strings, config files with keys inside
Personal dataNames, emails, addresses, customer/employee records - anything identifying a real person
Company code & documentsSource code, internal docs, strategy, unreleased plans

One principle closes it: whatever I submit is mine. If an AI drafted the code I merge or the analysis I present, I adopted it - “the AI wrote it” is no defence in a review, an audit, or a legal dispute. The professional version of the copy trap: understand and verify before I adopt, with career-grade stakes. Employers increasingly expect fluent AI use - the skill being tested is using it within the lines.

Must-knowOne-line recall
Two problemsVersion control merges code; a working method coordinates who builds what, when
SDLCPlan → Build → Test → Release → Maintain, then loop - software is never finished
Waterfall flawFeedback arrives at the end, when change costs most
Agile core betGet it right by short iterations of feedback, not one big up-front plan
Manifesto “over”Valued more when they conflict - never instead of
4 valuesIndividuals · working software · customer collaboration · responding to change
User storyAs a [role], I want [capability], so that [benefit] - the “so that” justifies it
Acceptance criteriaBinary, checkable statements that define done for one story
Story pointsRelative, team-local size - a forecast, never a productivity score
Planning pokerSimultaneous reveal → no anchoring, surfaces hidden assumptions
EmpiricismTransparency + Inspection + Adaptation - artifacts show, events adapt
Scrum rolesPO = what/why · Scrum Master = process/impediments · Developers = how
Scrum artifactsProduct Backlog · Sprint Backlog · Increment (+ Definition of Done)
Acceptance vs DoDCriteria = per-story behaviour; DoD = quality bar on all work
Scrum eventsSprint · Planning · Daily Scrum · Review (product) · Retro (team)
KanbanCards flow through columns; the WIP limit forces finishing over starting
Scrum vs KanbanGoal-shaped product work vs continuous interrupt-driven flow
Ticket → doneTicket → branch → commits → PR → review → merge; done is public, defined
Good questionGoal + what I tried + exact pasted error; name the real goal
Read docsTutorial (learn) · Guide (task) · Reference (lookup) - match type to question
Code reviewCorrectness → Clarity → Consistency; on the code, never the coder
Personal boardTo Do / In Progress / Done, honest cards, brutal WIP limit of two
LLM in one linePredicts plausible next text - it does not look facts up
Fluency ≠ truthConfident tone is a style property, carries no info about correctness
HallucinationPlausible invented falsehood - verify every claim against official docs
VerificationVerification is the learning - narrows an open search to one yes/no
Prompt partsContext + Goal + Constraints + Example; state my level; iterate, don’t restart
Copy trapAI explains/reviews/hints - the attempt is mine; struggle is the mechanism
AI bright lineNever paste secrets, personal data, or company code - GDPR makes it law
Own the outputWhatever I submit is mine; “the AI wrote it” is no defence