Skip to content

Day 5 - GenAI-Assisted Machine Learning

Generative AI for Managers - TUHH Institute of Entrepreneurship · part of my Technology Management MBA · study notes for revision.


Days 1 to 4 were all about text: understanding LLMs, prompting them, grounding them in documents (RAG), and letting them take actions (agents). Day 5 is a different beast. It’s about prediction from tables of numbers - the kind of question a spreadsheet full of customers raises: Which of these people will cancel next quarter? How much is each one worth? Where should we spend our limited attention?

That job belongs to machine learning (ML), not to a chatbot. But here’s the twist that makes this a GenAI course: GenAI doesn’t do the predicting - it makes the whole ML process accessible to a non-coder. It helps you frame the problem, clean the data, write the code, read the results, and explain them to your boss. So the mental model for the whole day is:

Human manager + ML model + GenAI copilot. The ML model produces the number. GenAI helps you decide what number to compute, how to compute it, and how to explain it.


The single most useful idea to start with: GenAI and predictive ML solve different problems. Confusing them is the classic beginner mistake.

If your problem sounds like…The right toolWhat comes out
”Write / summarise / rewrite this”LLM (GenAI)Natural-language text
”Find the answer in our documents”RAG (Day 3)A grounded, sourced answer
”Decide something from a table of data”Supervised MLA number: a probability or a prediction
”Rank thousands of customers, target the best”ML + decision rulesA ranked list + who to act on

GenAI is brilliant at unstructured work (language). Predictive ML is brilliant at structured decisions (rows and columns - numbers, categories, dates): ranking, targeting, forecasting, risk scoring. They’re teammates, not rivals.

Why this matters now: ML has traditionally been locked behind data scientists - expensive to hire, slow to iterate with, and their outputs need translating before a manager can act. GenAI + AutoML (more on that later) lowers that wall so people like me can meaningfully take part in an ML project instead of just receiving a slide at the end.


Normal programming: a human writes the rules, feeds in data, and gets answers. Machine learning flips the middle and the end around:

TraditionalRules + Data → Answers
↔
Machine learningData + Answers → Rules (patterns)
You don’t teach a child to recognise dogs by listing every breed. You show them hundreds of dogs and they learn the pattern. ML learns from examples, not from hand-written rules.

For business, two families matter most:

FamilyWhat it doesBusiness exampleOutput
SupervisedLearn from labelled examples (we already know the answer)Churn prediction, fraud detection, forecasting”This customer has a 78% chance of leaving”
UnsupervisedFind patterns in unlabelled dataCustomer segmentation, anomaly detection”These 5 groups behave differently”

Today is almost entirely supervised learning, which itself splits in two - and we’ll use both:

Question: which category? The answer is a class, usually delivered as a probability. Our example: Will this customer churn - yes or no? → p(churn), e.g. 0.78.


The business case: churn + value = revenue at risk

Section titled “The business case: churn + value = revenue at risk”

Our running example all day is customer churn (“churn” = a customer leaving/cancelling) for a telecom company. But here’s the crucial managerial point:

A churn model is only useful if it connects to an action - offer a discount, make a call, improve service, or do nothing. Prediction without a decision is trivia.

And to decide well, one number isn’t enough. A high-risk customer might be cheap to lose; a medium-risk customer might be your most valuable account. So we combine two predictions:

Riskp(churn) - classification
×
Valuemonthly charges - regression
=
Revenue at Riskwho to prioritise
Revenue at Risk = p(churn) × predicted monthly value. This one formula is the bridge between the classification model and the regression model - and it’s what managers actually care about: where the money is, not just who is nervous.

A tiny decision-rule exercise that shows why probability alone drives action. Say a retention offer costs €10, and losing a churner costs you €150. It’s worth making the offer when the expected saved loss beats the cost:

Offer if p(churn) × 150 > 10, i.e. when p(churn) > 10/150 ≈ 0.067.

So even a 7% churn probability justifies a €10 offer here. The “right” threshold is a business number (costs and capacity), not a technical default - a theme we return to constantly.


The data - and the traps a manager must know

Section titled “The data - and the traps a manager must know”

The Telco dataset has ~7,000 customers and columns like: demographics (senior citizen, partner), services (internet, streaming, phone), contract & billing (contract type, payment method), an age proxy (tenure = months as a customer), money (MonthlyCharges, TotalCharges), and the target (Churn: Yes/No).

Real data is messy, and you should expect these issues:

  • Missing values (e.g. TotalCharges blank for brand-new customers).
  • Mixed types (text and numbers together).
  • Class imbalance - usually far fewer churners than stayers (here ~27% churn). This quietly breaks naïve accuracy (explained below).
  • Leakage - the single most dangerous trap.

Every ML project - churn, demand, fraud - follows roughly the same six steps:

Define problembusiness Q → ML task
→
Collect & cleanfix, merge, impute
→
Engineer featurescreate predictors
→
Train & tunefit the model
→
Evaluateon held-out data
→
Deploy & monitorwatch for drift
The ML workflow. GenAI can assist at every single one of these steps - and AutoML automates the training-and-tuning part in the middle.

Train/test split, and over- vs under-fitting

Section titled “Train/test split, and over- vs under-fitting”

We must know whether a model generalises to new customers, not just the ones it memorised. So we split the data:

  • Training set - used to fit the model.
  • Test set - locked away until the very end, then used once as a realistic report card.

Every model lives on a tension between two failure modes:

FailureNicknameWhat it looks like
Underfitting (high bias)Too simplePoor on both training and test data - e.g. predicting the same churn risk for everyone
Overfitting (high variance)Too complexGreat on training, poor on test - memorised the noise

The goal is the sweet spot: complex enough to catch real patterns, simple enough to travel. This is exactly why we build models in a ladder from simple to complex (below) and only add complexity when the test set proves it helps.

A reliable ML system isn’t “a model” - it’s a pipeline that always does the same steps in the same order: fill missing values → encode categories into numbers → (scale if needed) → predict. Bundling this into one object is what prevents leakage from the test set and makes results reproducible.

# A minimal scikit-learn pipeline: preprocessing + model in one object
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
preprocess = ColumnTransformer([
("num", SimpleImputer(strategy="median"), numeric_cols),
("cat", Pipeline([
("impute", SimpleImputer(strategy="most_frequent")),
("onehot", OneHotEncoder(handle_unknown="ignore")),
]), categorical_cols),
])
model = Pipeline([("prep", preprocess),
("clf", LogisticRegression(max_iter=1000))])

GenAI meets ML: two engines working together

Section titled “GenAI meets ML: two engines working together”

This is the heart of Day 5. GenAI does not replace the model - it removes the friction at every step. Think of two engines:

  • AutoML handles the mechanical complexity - trying algorithms, tuning settings, comparing models.
  • GenAI handles the cognitive complexity - understanding the data, suggesting features, explaining results in plain English.
Workflow stepThe old wayWith a GenAI copilot
Define problemLong meetings between business & data teamsLLM helps translate a business question into an ML task
Understand dataWrite code, eyeball chartsDescribe the dataset; LLM flags quality issues & leakage
Engineer featuresExpert brainstorm + manual codingLLM suggests features from domain knowledge + writes the code
Train & tuneData scientist hand-picks & tunesAutoML searches; LLM drafts the pipeline code
EvaluateData scientist reads metricsLLM explains metrics in business terms, tied to cost
CommunicateAnalyst writes the exec reportLLM drafts the summary, caveats, and recommendations

GenAI for understanding data and inventing features

Section titled “GenAI for understanding data and inventing features”

Before building anything, you describe your dataset to the LLM - column names, types, a few sample rows - and ask what it notices. It’ll often spot that TotalCharges may be blank for new customers, that a wide MonthlyCharges range hints at service tiers, or that too much class imbalance means “predicting no-churn for everyone” will look deceptively accurate.

Feature engineering - inventing better predictor columns from the raw ones - is traditionally the hardest, most creative part of ML. It’s also where an LLM shines, because it has “read” millions of business documents. You ask for ideas as structured JSON (a direct callback to Day 2) so you can use them programmatically:

# Ask the model for derived-feature ideas, returned as clean JSON
resp = client.models.generate_content(
model="gemini-2.5-flash-lite",
contents="""I'm predicting telecom churn. Columns: tenure, monthly_charges,
total_charges, contract, payment_method, internet_service, num_support_tickets.
Suggest 8 derived features. Return JSON objects with:
"feature_name", "formula", "rationale".""",
config=types.GenerateContentConfig(response_mime_type="application/json"),
)

Typical suggestions: is_new_customer (tenure under 6 months - new customers churn more), support_ticket_rate (tickets ÷ tenure), or is_month_to_month (no lock-in = easy to leave).


The model ladder: simple → strong → best

Section titled “The model ladder: simple → strong → best”

A key managerial principle: start transparent, add complexity only if it earns its place. For churn (classification) we climb a three-rung ladder. First, always beat a dumb baseline (e.g. “always predict no churn”, or “month-to-month = churn”) - if a fancy model can’t beat a naïve rule, it shouldn’t ship.

ModelHow it works (plain English)InterpretabilityTypical performanceWatch for
Logistic regressionWeighted sum of features → squeezed into a 0-1 probabilityHigh - read each feature’s weightGood baselineOnly finds simple patterns
Random forestHundreds of decision trees on random data slices, averagedMediumStrong out-of-the-boxLittle tuning needed
Gradient boostingTrees built one after another, each fixing the last one’s mistakesMedium (use SHAP)Often the best on tablesNeeds tuning; can overfit

A few intuitions worth carrying:

  • Logistic regression gives every feature one readable coefficient: “month-to-month contract adds +1.2 to the churn score; a two-year contract subtracts 0.8.” That transparency is why it’s the natural starting point.
  • A single decision tree is just a flowchart of yes/no questions (“Is the contract month-to-month?” → “Is tenure short?”). Intuitive, but one tree is unstable - nudge the data and you get a totally different tree.
  • A random forest fixes that instability by building many diverse trees and averaging them, so their individual mistakes cancel out (this averaging trick is called bagging).
  • Gradient boosting is cleverer still: each new tree is trained on the leftover errors of all previous trees, so the model keeps patching its own weak spots. It usually wins on tabular data - at the cost of careful tuning.

The exact same ladder works for the regression side (predicting MonthlyCharges): linear regression → random-forest regressor → gradient-boosting regressor.


Evaluating models the way a manager should

Section titled “Evaluating models the way a manager should”

This section is where managers add the most value - because the obvious metric is usually the wrong one.

Every yes/no prediction lands in one of four boxes:

Predicted: will churnPredicted: will stay
Actually churnedTrue Positive (caught it)False Negative (missed it)
Actually stayedFalse Positive (false alarm)True Negative (correctly ignored)

In manager language: a False Positive means you spent effort on someone who was never leaving (wasted cost); a False Negative means a real churner slipped through (lost revenue). Which mistake hurts more is a business question.

So we use metrics that actually reflect the goal:

MetricPlain-English meaningWhen it matters most
Precision”When we flag someone, how often are we right?”When false alarms are expensive (wasted offers)
Recall”Of all the real churners, how many did we catch?”When missing a churner is expensive
F1 scoreA single score that’s only high when both precision and recall are highWhen both error types matter, and churn is rare
ROC-AUC”Chance the model ranks a random churner above a random stayer” (1.0 = perfect, 0.5 = coin flip)Comparing models broadly; balanced classes
PR-AUCOverall ranking quality focused on the rare positive classThe honest choice for churn/fraud (imbalanced)

Why F1 uses a “harmonic mean” and not a plain average: imagine 95% precision but only 10% recall (super picky, catches almost nobody). A plain average says 52.5% - sounds fine. F1 correctly reports ~18%, exposing that the model is nearly useless for actually finding churners. F1 refuses to be fooled by one good number.

Why ROC can look too rosy, and PR is more honest for churn: ROC’s false-positive rate is diluted by the huge crowd of stayers, so a moderate number of false alarms barely moves it - the same model might show ROC-AUC 0.87 (looks great) but PR-AUC 0.68 (more sober). PR only looks at the customers you actually flagged, so every false alarm hurts. Report both; decide on PR-AUC for imbalanced problems.

The model outputs a probability; you choose the cut-off that turns it into a yes/no. The default 0.5 is arbitrary. Lower the threshold and you flag more people - catching more churners (higher recall) but with more false alarms (lower precision). Same model, different threshold → a completely different call list. Pick it from costs and capacity, not from a textbook default.

Lift and deciles - the most manager-friendly view

Section titled “Lift and deciles - the most manager-friendly view”

AUC scores are for comparing models. But a manager’s real question is concrete: “I can only call 500 customers this week - how many churners will I catch?” That’s what lift answers.

Rank everyone by predicted risk, split into ten equal groups (deciles), and compare against random dialling:

  • Contact the top 10% → the model captures, say, 52% of all churners (random would get 10%). That’s a lift of ~5.2×.
  • Contact the top 30% → capture ~88% of churners.
  • By decile 4, lift drops below 1.0 - those customers are less likely than average to churn, so calling them wastes effort.

Lift maps directly onto operational capacity, which is why executives love it: it tells you whether your limited 500 calls will find 5× more churners than random, or barely beat a coin.

MetricManager translationBest for
MAE (mean absolute error)“On average we’re off by €8.50”Communicating to stakeholders - most intuitive
RMSE (root mean squared error)“Typical error, but punishes big misses”Flagging occasional wild predictions
R²”The model explains 75% of the variation”Comparing models (less intuitive to non-experts)

Instead of handing an exec a table of metrics, you can ask the LLM to explain them - but only after you’ve computed them for real:

# Feed REAL metrics; ask for a plain-language, cost-aware explanation
resp = client.models.generate_content(
model="gemini-2.5-flash-lite",
contents="""Explain these churn results to a non-technical VP. Cover what the
model catches, what it misses, and the precision/recall trade-off, in concrete
numbers. Accuracy 82%, Precision 71%, Recall 65%, 1,400 test customers.""",
)

A model that predicts churn is useful. A model that explains why is far more valuable - because it tells you what to do, and it lets you defend the decision (vital in regulated settings).

SHAP (SHapley Additive exPlanations) borrows a fair-credit idea from game theory: if a team produces a result, how do you divide the credit fairly among players? SHAP treats each feature as a player and the prediction as the team’s result, and works out how much each feature pushed this specific prediction away from the average.

For one customer, the prediction is built additively from the baseline:

prediction = average prediction + contribution₁ + contribution₂ + … + contributionₙ

Each contribution is a signed number - positive pushes toward churn, negative pulls toward staying.

Aggregate SHAP across all customers and you get a ranked picture of what the model relies on, e.g.:

month-to-month contract · 0.18tenure · 0.15support tickets · 0.12monthly charges · 0.09

It answers: “What drives churn risk in our data - and are those drivers sensible?”

Then GenAI turns SHAP into words - global drivers into a recommended strategy, or one customer’s contributions into a two-line note for the customer-success rep: “Flagged because of a month-to-month contract, very short tenure, and four recent support tickets; suggested action: offer a discounted annual plan.”


AutoML automates the mechanical parts - trying several algorithms, tuning their settings, and combining them - to get a strong model fast. In the labs you can run it with a tiny time budget and compare it to your hand-built pipeline:

from flaml import AutoML
automl = AutoML()
automl.fit(X_train, y_train, task="classification",
metric="ap", # average precision ≈ PR-AUC
time_budget=60) # give it 60 seconds

Course rule of thumb: use AutoML as a benchmark after you’ve built a transparent baseline (logistic/linear) and a strong manual model (forest/boosting) - not as a replacement for thinking.


From prediction to action (and keeping it alive)

Section titled “From prediction to action (and keeping it alive)”

A model earns its keep only when it changes what the business does. A simple retention playbook:

  1. Score everyone - compute p(churn) for all customers.
  2. Value everyone - predict monthly value (regression).
  3. Rank by Revenue at Risk = p(churn) × MonthlyCharges.
  4. Apply real-world constraints - contact capacity (e.g. 500 calls/week), eligibility rules (don’t offer discounts to customers already on a promo).
  5. Run an experiment - A/B test the retention offer and measure the incremental save, because the model shows who’s at risk, not what fixes it.

A model trained on last year’s world silently degrades as the world moves on. This is drift, and a manager should demand a monitoring plan before deployment.

Drift typeWhat changesHow to spot it
Data driftFeature distributions shift (a campaign floods in new customers)Compare feature histograms: training vs. recent
Performance driftAccuracy/AUC slowly fallsTrack metrics on a rolling window of fresh data
Calibration drift”30% risk” no longer means 30% churn in realityCompare predicted vs. observed rates
Concept driftThe feature→outcome relationship itself changesWatch whether the top SHAP drivers move over time

The usual fix when drift appears: retrain on more recent data.


You built the model, the SHAP makes sense - but now you face the VP, the CFO, or the board. They don’t care about F1 scores; they care about revenue and what to do. This is where GenAI closes the loop, turning outputs into an executive-ready story. What executives actually ask:

They ask…You give them…
”What did you find?""We can spot 65% of churners before they leave."
"How much is at stake?""Those predicted churners represent €1.2M of annual revenue."
"What should we do?""Three prioritised retention plays, by segment."
"How confident are you?""High for the top 50 customers, moderate for the next 100."
"What could go wrong?""Trained on 2024 data; a big price change could reduce accuracy.”

You can draft this memo with GenAI - feeding it your real numbers and asking for a structured one-pager (key finding → business impact → recommended actions → confidence & limitations → next steps) aimed at a non-technical reader.


The manager’s real job (and GenAI’s limits)

Section titled “The manager’s real job (and GenAI’s limits)”

Pulling the whole day together: the human isn’t automated away - the human sets the questions and guards the quality.


The theory lands only once you run it on real data. All three exercises use the Telco churn dataset in Google Colab, with GenAI as a copilot for coding, debugging, and interpretation - never for inventing results.

Guided Lab Churn + value modelling with explainable ML. Load and clean the Telco data, build a leak-proof preprocessing pipeline, train and compare all three churn models with manager-friendly metrics (lift + PR-AUC), add a MonthlyCharges regression model, generate global and local SHAP explanations, and produce a ranked Revenue-at-Risk call list.

Independent Lab From predictions to decisions. Pick one extension track and take it further: (A) cost & capacity targeting (choose a threshold by expected value), (B) calibration & probability quality (calibration curve + Brier score), (C) an AutoML benchmark vs. your manual models, or (D) a segment stress-test (does the model hold up across contract types and tenure bands?). Deliver improved reports, a final call list, and a short manager recommendation - plus a Copilot Log.

Assignment 5 Model card + manager recommendation. The graded deliverable, continuing your chosen track: a defensible workflow with an explicit leakage scan, 5-fold cross-validation across all models, business-facing evaluation (lift, threshold chosen by a cost curve, capacity plan), SHAP global + three local stories, a Revenue-at-Risk call list with reason codes, and a monitoring/risk plan. Submit a notebook, a call-list CSV, a metrics JSON, and a one-page manager memo in plain business language.

Download the notebooks (open in Google Colab):


My submitted solution: Assignment 5 - my solved notebook (.ipynb) - my own work from the course.

Must-knowOne-line recall
Two AI jobsGenAI handles unstructured text; predictive ML handles structured decisions - teammates, not rivals
Supervised learningLearn from labelled examples: classification (which category) or regression (what number)
Features vs. labelFeatures are the input columns; the label is the answer we predict
Revenue at Riskp(churn) × predicted value - the formula that bridges the two models and drives targeting
LeakageA feature the model wouldn’t have at decision time; looks great, then fails in production
The 80/20 rule~80% of ML effort is data work (steps 1-3), only ~20% is modelling
Train/test splitTest set = a locked exam; a big train-vs-test gap means overfitting
Over/underfittingOverfit = memorised noise (great train, poor test); underfit = too simple (poor on both)
Model ladderLogistic (transparent) → random forest (strong default) → gradient boosting (often best)
Accuracy trapOn imbalanced data, “predict no churn” scores high and is useless - use precision/recall/PR-AUC
Precision vs. recallPrecision = right when we flag; recall = share of churners caught; F1 balances both
ThresholdA business choice from costs/capacity - same model, different threshold, different call list
Lift”How many times better than random” for your top decile - maps to call capacity
SHAPFairly splits a prediction into per-feature contributions; global (drivers) + local (this customer)
AutoMLAutomates model search & tuning - a fast benchmark, never a replacement for judgement
DriftModels decay as the world shifts; demand a monitoring & retraining plan before deployment
Correlation ≠ causationA model says who’s at risk, not what fixes it - prove impact with an experiment
GenAI is copilot, not pilotIt frames, codes, interprets, and communicates - but every claim must be validated with real outputs