Skip to content

Day 6 - GenAI Business Cases

Generative AI for Managers - TUHH Institute of Entrepreneurship · part of my Technology Management MBA · study notes for revision.


Days 1 to 5 were about how the technology works - prompting, retrieval, extraction, agents, and a little machine learning. Day 6 flips the whole question around. For the first five days I kept asking “can we build this?” Today the only question that matters is the manager’s question: “should we build this - and if so, why, and how would I prove it?”

So this is the day where the technical stuff turns into a business case: a short, honest document that convinces a decision-maker to fund a small pilot. Below are my revision notes on how to build one, in the order the course walks through it.

Jargon check. GenAI = generative AI, tools like large language models (LLMs) that produce text, code, or images. MVP = Minimum Viable Product, the smallest version that is still genuinely useful. Pilot = a small, time-boxed trial with a handful of real users before any big rollout.

The deliverable: one Capstone Proposal Pack

Section titled “The deliverable: one Capstone Proposal Pack”

The whole day builds toward a single package a real team would take to a steering committee (the group of senior people who approve budgets). It comes together in four blocks, each feeding the next:

A · Pickthe right idea
→
B · Provethe value
→
C · Designthe product
→
D · Deploy& measure
The four blocks of Day 6 - each one hands its output to the next.

By the end you have six connected pieces:

#PieceWhat it answers
1Use-case one-pagerWhich problem, framed in business terms
2Process map + value sizingHow GenAI changes the work, and what it saves
3MVP requirements (PRD)What to build, for whom, in what order
4Build/buy + architecture sketchHow it works, and make-vs-buy
5Rollout plan + KPIsHow you deploy it and prove success
6Governance-lite checklistHow you keep it safe and responsible

The rest of this chapter is really just the thinking behind each of these.

Not every task is a good GenAI candidate. The fastest way to waste money is to point the technology at the wrong job. Two quick lenses help.

Good fit vs poor fit. GenAI shines at repetitive language tasks and at turning messy, unstructured text into something structured or useful.

Good fit for GenAIPoor fit - rethink it
Repetitive drafting (emails, replies, summaries)One-off decisions done a handful of times a year
Turning messy text into structure (invoices → data)Tasks where nobody can judge if the output is good
Answering questions over your own documents”Replace an entire department”
Classifying or routing high-volume requestsAnything needing data you cannot legally access

The MVP-friendly filter. A strong candidate passes all five of these checks:

FilterThe question to askPasses if…
High volumeDoes this happen hundreds/thousands of times a year?e.g. 500 tickets/week
Clear input/outputCan I define what goes in and what “good” out looks like?Fixed reply categories; a data schema
Evaluable qualityCan I measure whether the output is any good?Human-check 50 samples; ≥85% accuracy
Low MVP complexityCan a useful v1 exist without complex integrations?A standalone tool before wiring into legacy systems
Feasible data accessCan I get the data given policy and law?Internal docs - not customer personal data without controls

And the anti-patterns to avoid on sight - vague, unmeasurable, or impossibly broad ideas:

“A chatbot for everything""Replace a whole team”No path to the dataSuccess can’t be measured

Once you have a list of candidate ideas, you rank them by impact first, with a realistic path to a pilot. Two tools do the job.

The value × time-to-value 2×2. Plot each idea by how much value it creates against how fast that value arrives (“time-to-value” = how long until you see results).

Fast to valueSlow to value
High valueQuick wins - launch first, build momentumStrategic bets - plan carefully, need exec sponsorship
Low valueFill-ins - only if resources are idleAvoid - don’t waste effort

The 1-5 scoring rubric. Score each shortlisted idea on five criteria, for a total out of 25. Note that two criteria are reversed - for time-to-value and risk, a high score means “fast” and “low risk,” so higher is always better.

CriterionScore 1 (low)Score 3 (medium)Score 5 (high)
Value potentialMinor convenience (under 5% saved)Meaningful (15-25% less manual work)Major (>€500k/yr or big risk cut)
Time-to-value reverse6+ months to ROI2-3 months to pilot resultsunder 4 weeks to first results
FeasibilityMajor integration, shaky data1-2 systems, structured dataPlug-and-play, standard formats
Risk/constraints reverseRegulated, sensitive personal dataSome sensitive data, policies existInternal data only, low risk
EvaluabilityNo metric, subjectiveProxy metrics existClear ground truth, easy to audit

Read the total like this: 18+ is usually MVP-ready; 12-17 needs more validation; below 12, reconsider. One human note from the course that I like - after the maths, also ask: does the team actually care about this domain? The best pick is often the process someone has been personally frustrated by for years.

Jargon check. ROI = return on investment (the value you get back versus what you spend). Ground truth = a known-correct answer you can compare the AI’s output against.

The output of this step is the use-case one-pager - your business case boiled onto a single page:

  1. Use-case name - clear, not cute (e.g. “Automated invoice line-item extraction”).
  2. Primary users - who, in which department, and how often they do the task.
  3. Problem statement - 2-3 sentences: what manual work happens, the pain, why it matters.
  4. Why GenAI (not simple rules or classic machine learning) - three bullets.
  5. Constraints & dependencies - data sensitivity, systems, timeline, approvals needed.
  6. Expected value - quantified where possible, given as a range.
  7. Score & next steps - the rubric total, and what to validate before designing the MVP.

A one-pager gets attention; a process map plus value sizing turns it into a business case. This is Block B.

Map one workflow, before and after. Pick a single workflow (6-10 steps - not the whole company). First draw the as-is map: every step, who does it, how long it takes, and where errors or rework creep in. Then draw the to-be map with GenAI inserted only where it genuinely helps, keeping a human in the loop for any real decision.

Read ticket
→
Search KB
→
Draft reply
→
Manager review
→
Send · ~37 min
AI drafts reply
→
Agent reviews
→
Edit & send
→
Auto-logged · ~6 min
A support-ticket example: the AI drafts, but the human still reviews and sends. The judgement stays human; the grunt work shrinks.

Notice what stayed human - the review and send steps. GenAI assists; it does not replace judgement on anything customer-facing.

Size the value with transparent maths. You are not building a fancy financial model (no DCF - discounted cash flow, a detailed multi-year valuation). You just need “good enough,” auditable numbers to justify a pilot:

Annual hours saved = (volume per year × minutes saved per case) ÷ 60 Annual value = annual hours saved × loaded hourly cost

“Loaded hourly cost” means the fully-costed hourly rate of the person (salary plus overhead). Always give a low / likely / high range and write down your assumptions - the honesty matters more than false precision. Source your volumes from real data (“we log ~500 tickets a week”), never from “a lot.”

AssumptionLowLikelyHigh
Tickets per year26,00026,00026,000
Minutes saved per ticket152531
Annual hours saved6,50010,83313,433
Loaded hourly cost€45€45€45
Annual time savings€292,500€487,500€604,500
+ quality savings (less rework)€39,000€78,000€117,000
Total annual value€331,500€565,500€721,500

Then a 6-line business-case narrative ties it together: where the value comes from, why it is measurable, your biggest assumption plus how you’ll test it, the main risks, what success looks like at 30 days, and what would make you stop.

Block C turns the idea into something buildable: a PRD, an architecture sketch, and a build-vs-buy call.

The MVP PRD (Product Requirements Document - the spec of what you’re building). For GenAI, the PRD’s real job is scope discipline: the biggest risk in any AI project is scope creep (endlessly adding features). It must contain:

  1. Target users + job-to-be-done (JTBD) - the outcome the user actually wants.
  2. 3-5 user stories - “As a [role], I want [thing] so that [benefit].”
  3. Output format + constraints - what “good” looks like (tone, latency, accuracy).
  4. Non-features - an explicit list of what you are not building. This is the most valuable part; saying no is what separates a 12-week MVP from an 18-month slog.
  5. Acceptance criteria - testable pass/fail bars (e.g. “80%+ of suggestions accepted over a 2-week pilot”).

Pick one architecture pattern and sketch it. Every sketch must show four things: data sources (what you read), outputs (what you write), human approval points (where risk is managed), and logging (what you measure). The four manager-level patterns map neatly onto Days 1-5:

Suggestions inside an existing tool. The AI drafts, the human approves.

User works in toolAI suggestsHuman edits/rejectsAction taken
Best for: efficiency inside current workflows. Builds on: prompting (Day 1).

Build vs buy. Do you build custom, buy a platform, or go hybrid? Trim the comparison to what matters for a pilot:

CriterionBuild customBuy platform
Speed to MVPSlow (8-16 weeks)Fast (2-4 weeks)
Data & securityFull controlVendor-dependent
Year-1 costHigh (engineers)Medium (license)
DifferentiationHigh (custom logic)Low (same as rivals)
Vendor lock-inNoneHigh

Responsible AI in practice: governance-lite

Section titled “Responsible AI in practice: governance-lite”

The course is refreshingly blunt here: this is not an ethics lecture, it is operational pragmatics. Skip it and you invite data breaches, user resistance, and lost credibility. “Governance-lite” is the minimum you must settle before any pilot goes live.

AreaKey questionPilot minimum
Data classificationWhat data does the model see?Label every input public / internal / confidential
Access controlWho can use the tool?Define roles and permissions; log who accesses it
Logging & auditWhat must be recorded?Log inputs, outputs, user actions, timestamps
Evaluation planHow do we know it works?50-100 labelled test cases; a human-review cadence
Human-in-the-loop (HITL)Who approves high-impact actions?Define the review scope and the escalation path

Alongside this, name your failure modes honestly - hallucination (the model inventing false facts), bias, data leaks, user rejection - and decide, in advance, when a human steps in and who can pause the pilot.

Block D is deployment. You don’t flip a switch for everyone at once; you roll out in three phases, each with a “success gate” that must be cleared before expanding.

  1. Days 1-30 - Pilot. Pick a small pilot group (say 5-10 users) and say why them. Train them, give a prompt library and usage guide, run a weekly feedback loop. Gate: is this “good enough to continue”?
  2. Days 31-60 - Expand. Broaden the group using pilot learnings, fix the top three issues, refine prompts and approval gates. Gate: is this “ready to scale”?
  3. Days 61-90 - Scale. Full team/department rollout, integrate with existing systems, set an ongoing monitoring cadence, and do the first formal ROI readout.

KPIs: leading and lagging. A KPI (Key Performance Indicator) is a number you track to judge success. You need two kinds - leading indicators tell you early whether people are adopting it; lagging indicators tell you later whether real value appeared.

TypeWhat it measuresWhen you see it
LeadingAdoption & engagementImmediately (e.g. weekly active users)
LaggingBusiness outcomesAfter weeks/months (e.g. cost per case)

Spread your KPIs across four categories, and give each one a definition, a data source, and a target:

CategoryWhat it tracksExample KPIs
AdoptionAre people using it?Weekly active users; tasks per user
EfficiencyIs the work faster?Time-to-complete (before vs after); throughput
QualityIs the output good?Rework rate; human override rate
Business outcomesDid it move the needle?Cost per case; CSAT or cycle time

“Human override rate” - how often users reject the AI’s suggestion - is my favourite quiet metric: a high one is an early warning that quality isn’t there yet.

The neat thing about this day is that nothing from the earlier days is wasted - the technical fluency becomes the credibility behind the business case.

From an earlier dayWhat it gave meHow Day 6 uses it
Prompting (Day 1)Getting reliable output from an LLMThe copilot pattern; writing prompt libraries for rollout
RAG (Day 2)Grounding answers in real documentsThe RAG-assistant architecture pattern
Extraction (Day 3)Messy text → structured dataThe extractor→workflow pattern; invoice-style use cases
Agents (Day 4)Multi-step, tool-using AIThe agentic pattern, with human approval gates
Machine learning (Day 5)When classic ML beats GenAI (and vice versa)Justifying “why GenAI, not rules or classic ML”

The technical skills give you the right to be believed; the business thinking gives you the impact.

Day 6 has no notebooks to download - the deliverable is a team capstone, not a coding exercise. In teams, you produce the full Capstone Proposal Pack and pitch it, as if to a CEO deciding whether to fund the pilot.

There are two things to submit:

  1. The Capstone Proposal Pack - one PDF (~9-10 pages) containing all six components: the use-case one-pager, process map + value sizing, MVP PRD, build/buy + architecture sketch, rollout plan + KPIs, and the governance-lite checklist.
  2. The pitch - a slide deck (8-12 slides) plus a recorded 10-minute video (link on the first or last slide). The pitch runs: Problem → Why now → MVP → Value → Plan → Ask, and should lead with the business problem, not the technology.

Marking rewards specificity above ambition - a detailed plan for a simple use case beats a vague plan for a grand one. The golden test throughout: if a reader can’t tell who does what, when, and how success is measured, it isn’t finished. (The pack is also designed to extend naturally into a master thesis, if you want to carry it further.)

My capstone submission: Business-case proposal (PDF) · Pitch deck (PDF) - my team’s final deliverables for the course.

Must-knowOne-line recall
The core shiftMove from “can we build it?” to “should we - and can I prove it?”
Good use caseBounded + measurable + feasible as a ~4-week MVP.
GenAI sweet spotRepetitive language tasks and messy-text-to-structure, done at high volume.
MVP filterHigh volume · clear I/O · evaluable · low complexity · data you can access.
PrioritisationValue × time-to-value 2×2, then score 1-5 across five criteria (18+ = ready).
Value sizing(volume × minutes saved ÷ 60) × loaded cost, always as low/likely/high.
Human stays inGenAI drafts; the human reviews and approves anything that matters.
PRD’s real jobScope discipline - the non-features list is the most valuable part.
Architecture patternsCopilot · RAG · extractor · agentic - each maps to a Day 1-5 skill.
Build vs buyCore advantage → build; commodity efficiency → buy.
Governance-liteClassify data, control access, log everything, plan evals, keep HITL.
Rollout & KPIs30/60/90 phases with gates; track leading and lagging KPIs.
The capstoneA 6-part proposal pack + a 10-minute pitch that leads with the problem.