Conjoint Analysis
Innovation & New Business Proposal - TUHH Institute of Entrepreneurship & Institute of Innovation Marketing, Hamburg · part of my Technology Management MBA · study notes for revision.
Kano told me which kind of satisfaction each attribute produces. Conjoint analysis answers the next question, and it is a quantitative one: which product attributes generate which value for customers? Not “is this feature nice”, but “how many units of preference is this level of this attribute worth, compared with that level of that other attribute”. It is the tool that finally puts numbers on the trade-offs a customer makes in their head.
The class notes describe conjoint analyses as tools that measure the importance of product attributes holistically, and they attach four goals to that. Which attributes are how important? Which attribute levels provide which utility to the customer? Are there anomalies - jumps in the utility curve, knock-out criteria, minimum levels below which nothing is acceptable? And what are the trade-offs between different attributes? The last one is really the point of the whole method: every product decision is a trade, and this is the only market research tool in the module that measures the exchange rate.
What makes it feel different from a survey is that respondents are never asked about attributes at all. They are shown whole products and asked which one they prefer. The importance numbers are then worked backwards out of those judgements. That reversal is the entire trick, and it is why the technique survives the problem I describe in section 2.
1 · The underlying question and the four goals
Section titled “1 · The underlying question and the four goals”The class makes the fourth goal concrete with four questions any product manager would recognise. How much more would customers prefer a product if it offered a UHD resolution rather than HD on a television, an e-bike weighing 20 kg rather than 25 kg, an electric car with 500 km of range on one charge rather than 300 km, or an online meeting tool with a breakout session function rather than without one? Each is a “how much more”, not a “yes or no”. Only a method that measures preference on a scale can answer them.
2 · Why asking directly does not work
Section titled “2 · Why asking directly does not work”The build-up in the session is done with a holiday hotel, and it is worth reproducing because the failure happens in two stages.
Stage one: rate the importance. How important to you personally are the following attributes of a holiday hotel - the walking time from the hotel to the beach, the size of your room, the frequency of room cleaning, the speed of the Wi-Fi connection, and the availability of leisure and sports activities? Everyone answers happily, and almost everyone says most of them are quite important. Nothing has been given up, so nothing has been revealed. The answers are flattering rather than informative.
Stage two: distribute 100 points. So the question is tightened: please spread exactly 100 points across those same five attributes, so that the points represent how important each one is for you. This is better, because the budget forces some discipline. But it still fails, for a reason worth stating precisely: the respondent is being asked to introspect about abstract categories, not to make a decision. Nobody has ever stood at a hotel booking page choosing between “room size” and “cleaning frequency” as such. They choose between hotels, each of which bundles a particular level of each attribute at a particular price.
- Asks about attributes in the abstract, detached from any product
- No real sacrifice is involved, so answers are polite and inflated
- People are poor at introspecting about their own weightings
- Cannot capture levels - only that an attribute matters, not how much 80 km beats 50 km
- Asks the respondent to rank, rate or choose complete products
- Every product is good at some things and weak at others, so a real trade-off is forced
- The interview situation is natural, close to how buying actually feels
- Importance is derived afterwards from the pattern of judgements, never asked
3 · The second trap: interactions
Section titled “3 · The second trap: interactions”There is a further problem that no amount of careful wording fixes. The session illustrates it with a single question asked three times: how important is the size of the window in a hotel room? Then: how important is it in this room? And in this one? The answer changes completely depending on what the rest of the room looks like. A large window matters enormously in a small dark room and hardly at all in a bright, generously sized one.
The principle stated in the notes is that the importance or preference of one attribute may depend on another attribute. That is what an interaction is: an interaction occurs when the combined effect of two attributes differs from the sum of their two separate main-effect utilities. Asked in isolation, the window question has no single correct answer, because importance is not a property of the attribute alone.
4 · The idea hidden in the name
Section titled “4 · The idea hidden in the name”The name is a contraction, and unpacking it is the fastest way to remember what the method does. Products and services are CONsidered JOINTly - the respondent sees the whole bundle, never the pieces - and the analysis then treats that bundle as an aggregation of single components called part-worth utilities.
2. part-worth utilities of attribute levels
This is why conjoint is called a decomposition method, in contrast to the compositional logic of a points-allocation survey. A compositional method asks for the weights and builds the product score up from them. A decompositional method takes the product score as given and pulls the weights out of it.
5 · The additive utility model
Section titled “5 · The additive utility model”The model underlying a conjoint analysis is an additive utility function. The total utility of a product is the sum of the part-worth utilities of the levels it happens to have, plus an error term that absorbs everything the model does not capture.
y(k) = SUM over j of SUM over m of [ b(jm) · x(jm) ] + Error| Symbol | Meaning in the notation used in class |
|---|---|
| y(k) | the preference for product or service k, that is its total utility |
| j | the attributes, for example weight, range, price |
| m | the levels within each attribute, for example 20 kg and 25 kg |
| x(jm) | an indicator that equals 1 if the product has level m for attribute j, and 0 otherwise |
| b(jm) | the part-worth utility of level m of attribute j - the unknown the analysis solves for |
| Error | everything about the stated preference that the additive model does not explain |
The indicator variable is the mechanical heart of it. Each product profile switches on exactly one level per attribute, so the sum collapses to “one part-worth from each attribute, added together”. In the supporting paper these indicators are handled as dummy variables in a regression, with a useful economy: an attribute with three levels only needs two dummies, because the third level is identified automatically when both dummies are zero. That third level becomes the baseline and its part-worth is absorbed into the intercept.
6 · The smallest possible example: shape and colour
Section titled “6 · The smallest possible example: shape and colour”The session demonstrates the mechanics on a deliberately trivial case: two attributes (shape and colour), two levels each. Multiplying out gives four product profiles, and the respondent is asked one simple thing - please rank order the four products according to your preferences, which would you choose first, second, third and fourth?
Four holistic judgements, and from them come four part-worths. Writing the additive model out for each profile shows what the estimation is up against:
| Profile | Shape level | Colour level | Preference equation |
|---|---|---|---|
| 1 | Shape level 1 | Colour level 1 | y1 = b(shape 1) + b(colour 1) + Error |
| 2 | Shape level 2 | Colour level 2 | y2 = b(shape 2) + b(colour 2) + Error |
| 3 | Shape level 2 | Colour level 1 | y3 = b(shape 2) + b(colour 1) + Error |
| 4 | Shape level 1 | Colour level 2 | y4 = b(shape 1) + b(colour 2) + Error |
Four equations, four unknown part-worths, one shared set of stated ranks. The respondent never said a word about shape or colour, and yet the system now contains everything needed to price both.
7 · How the part-worths are estimated
Section titled “7 · How the part-worths are estimated”The estimation rule is stated in the notes in one sentence, and it is worth learning in that form. Statistical estimation determines the values of the part-worth utilities so that the calculated preference values, the estimated y values, are as similar as possible to the preferences actually stated by the customer, the real y values.
That is all it is: choose the numbers that make the model’s predicted ranking agree with the observed ranking as closely as it can. In the supporting paper this is done with ordinary least squares regression of the stated ratings on the dummy variables - the fitted equation for a public transport example with fare and waiting time came out as an intercept of 49.8 plus 26.7 and 14.0 for the two fare dummies and 25.7 and 20.0 for the two waiting-time dummies, reproducing the nine stated ratings almost exactly.
A second convenience from the paper is normalisation: rescale all the part-worths so that the smallest across the whole study becomes 0 and the largest becomes 1, using v = (u - a) / (b - a), where a and b are the smallest and largest raw values. This changes nothing about the internal relationships but makes numbers comparable within and across attributes.
8 · Relative importance of an attribute
Section titled “8 · Relative importance of an attribute”Once every level has a part-worth, the importance of the attribute that owns them follows immediately. An attribute matters to the extent that moving from its worst level to its best level moves total utility a lot. So importance is a range.
Range(j) = highest part-worth of j MINUS lowest part-worth of j (absolute value)Importance(j) = Range(j) / [ Range(1) + Range(2) + … + Range(J) ]Property the importances of all attributes add up to 1, or to 100 percent
In words, exactly as the slides put it: the relative importance of an attribute is the absolute difference between its highest and its lowest partial utility value, divided by the sum of all such differences. The supporting paper calls this the relative importance measure and computes it for the transport case as 26.7 divided by (26.7 plus 25.7), giving 51 percent for fare and therefore 49 percent for waiting time.
Two consequences that are easy to get wrong in an exam. First, importance is a property of the levels you tested, not of the attribute in the world. Widen the price range in the study and price will look more important, purely by construction. Second, an attribute whose levels are all roughly equally liked has a near-zero range and near-zero importance, no matter how loudly respondents would have insisted it mattered if you had asked them directly.
9 · Designing and running a study
Section titled “9 · Designing and running a study”The class lays out seven steps, and works them through on an e-scooter sharing system.
| # | Step | What actually happens |
|---|---|---|
| 1 | Collect product or service attributes | List what could plausibly drive the choice |
| 2 | Decide attribute levels | Fix realistic values per attribute, and count the full permutation |
| 3 | Decide model and software | Choice-based, rank-order or adaptive; Sawtooth is the standard tool |
| 4 | Design stimuli | Reduce the full permutation of combinations to a workable set of profiles |
| 5 | Collect data | Scenario construction plus background questions |
| 6 | Analyse | Average-level and individual-level analysis are both possible |
| 7 | Interpret and simulate | Relative importances, level utilities, and the utility of new combinations |
Step 4 is where most study designs live or die.
The e-scooter attributes and levels. Maximum speed of the scooters at 15, 20, 25, 30 or 35 km per hour. Average walking time to the next available scooter at 1, 2.5 or 4 minutes. Pricing as either a 1 euro unlock fee plus 30 cents per minute, or no unlock fee plus 50 cents per minute. Reservation possible or not possible. Smartphone holder available or not available.
Why step 4 exists. The full permutation is 5 times 3 times 2 times 2 times 2, which is 120 combinations. No respondent will judge 120 profiles. So the full factorial is reduced - in the class example to 30 profiles presented as 10 comparisons of 3 profiles each, relying on the assumption that the attributes are independent. The supporting paper does the same thing at larger scale for a smartphone study with five attributes at four levels each, where the full set runs past a thousand profiles and a fractional factorial design cuts it to 32 profiles rated on a zero to 100 scale.
Choosing how to ask. The alternative ways of measuring preference, all of which work as conjoint input:
Software and data collection. The class names Sawtooth Software as the world’s leading provider, free of charge for master theses through its grants programme, and mentions a web-based choice-based implementation. Data collection is not only the profiles: it also involves scenario construction, so the respondent is judging in a defined situation, and additional questions such as age, gender, residence and general mobility preferences.
Analysis and segmentation. Both average and individual-level analysis are possible, which is unusual and valuable - conjoint gives every single respondent their own set of part-worths. That opens a segmentation route the class demonstrates: run a cluster analysis on the vectors of relative importance, using the Ward algorithm with squared Euclidean distances, and the e-scooter respondents fall into three clusters of people who trade off speed, walking time and price in visibly different ways. The supporting paper makes the same point: clusters of individuals with similar importances are effectively market segments.
10 · Prediction, simulation and validity
Section titled “10 · Prediction, simulation and validity”The final step is the one that earns the study its budget. Because the model is additive, it can score any combination of levels, including combinations nobody was ever shown. That is the meaning of step 7, simulating the utility of new product combinations.
The supporting paper spells out three things about doing this responsibly:
- Interpolation between levels is acceptable. A value between two tested levels can be scored by assuming the part-worth function is linear over that gap - a fare halfway between two tested fares takes the average of their two part-worths.
- Extrapolation beyond the tested range is not. Scoring a level outside the span you tested can be badly misleading, and the remedy is to choose the attribute values sensibly when designing the profiles in the first place, so the interesting region is inside the range.
- Validate the estimated utility functions. The two named routes are predictive validation on a holdout sample of profiles kept out of the estimation, and validation against actual or intended market behaviour such as first choices, sales or market share.
The paper also summarises the technique in five roles, a compact revision list: it is a measurement technique for quantifying buyer trade-offs, an analytical technique for predicting reactions to new products, a segmentation technique for grouping buyers with similar trade-offs, a simulation technique for assessing new product ideas against competitors, and an optimisation technique for finding the profile that maximises share or return.
Choice-based conjoint specifically. The two mainstream variants are ratings-based and choice-based, and the implementation path forks early: decide the purpose, decide the approach, identify attributes and levels, then either design profiles and analyse with regression, or design choice sets and analyse with a logit model. Both converge on the same outputs. In the paper’s choice-based version of the transport study, 42 choice sets of size three were generated and each respondent simply said which of the three they would take; the answers were analysed with a conditional logit model estimated by maximum likelihood, and the part-worth pattern matched the ratings-based result. One design detail worth stealing: the sets deliberately mixed dominated options - one alternative better on every attribute, which acts as a sanity check on the respondent - with sets where a genuine trade-off is unavoidable.
11 · Strengths and limits
Section titled “11 · Strengths and limits”| Strengths | Limits |
|---|---|
| The interview situation is natural and realistic - people compare products, which is what they do anyway | Respondents must process a lot of information in each choice task |
| Yields the utility of each attribute level, not just of the attribute, for example price at 5, 10 and 15 euros | The task is tiring, and fatigue degrades the later answers |
| Yields the relative importance of attributes, for example that price outweighs the energy label in a television purchase | A practical ceiling of roughly 10 attributes can be shown |
| Can detect interaction effects, such as red being specifically popular on a Ferrari | Reduced designs usually assume attributes are independent, so most interactions go unmeasured |
| Flexible across industrial products, consumer products and services | Importance depends on the level ranges you chose, so a badly designed study gives confidently wrong weights |
Worked example
Section titled “Worked example”A small conjoint on a compact commuter e-bike. Three attributes, two levels each, so the full permutation is 2 times 2 times 2 = 8 profiles, small enough to show all of them.
| Attribute | Level A | Level B |
|---|---|---|
| Weight | 20 kg | 25 kg |
| Range on one charge | 80 km | 50 km |
| Price | 1800 euros | 2400 euros |
One respondent ranks all eight from 1 (most preferred) to 8 (least preferred). To turn ranks into a preference score I reverse them, so score = 9 minus rank, and the best profile scores 8.
| Profile | Weight | Range | Price | Rank | Score y |
|---|---|---|---|---|---|
| P1 | 20 kg | 80 km | 1800 | 1 | 8 |
| P2 | 20 kg | 80 km | 2400 | 3 | 6 |
| P3 | 20 kg | 50 km | 1800 | 4 | 5 |
| P4 | 20 kg | 50 km | 2400 | 7 | 2 |
| P5 | 25 kg | 80 km | 1800 | 2 | 7 |
| P6 | 25 kg | 80 km | 2400 | 5 | 4 |
| P7 | 25 kg | 50 km | 1800 | 6 | 3 |
| P8 | 25 kg | 50 km | 2400 | 8 | 1 |
Step 1 - the grand mean. The eight scores are 8, 6, 5, 2, 7, 4, 3, 1, summing to 36, so the grand mean is 36 / 8 = 4.5.
Step 2 - the average score of every level. Because the design is balanced, each level appears in exactly four profiles, so the level average is a fair estimate of what that level contributes.
| Attribute | Level | Profiles containing it | Sum of scores | Level average | Part-worth (average minus 4.5) |
|---|---|---|---|---|---|
| Weight | 20 kg | P1, P2, P3, P4 | 8+6+5+2 = 21 | 5.25 | +0.75 |
| Weight | 25 kg | P5, P6, P7, P8 | 7+4+3+1 = 15 | 3.75 | -0.75 |
| Range | 80 km | P1, P2, P5, P6 | 8+6+7+4 = 25 | 6.25 | +1.75 |
| Range | 50 km | P3, P4, P7, P8 | 5+2+3+1 = 11 | 2.75 | -1.75 |
| Price | 1800 | P1, P3, P5, P7 | 8+5+7+3 = 23 | 5.75 | +1.25 |
| Price | 2400 | P2, P4, P6, P8 | 6+2+4+1 = 13 | 3.25 | -1.25 |
Step 3 - check the model reproduces the ranking. Predicted score = 4.5 plus the three part-worths of that profile.
| Profile | Calculation | Predicted | Stated | Error |
|---|---|---|---|---|
| P1 | 4.5 + 0.75 + 1.75 + 1.25 | 8.25 | 8 | +0.25 |
| P5 | 4.5 - 0.75 + 1.75 + 1.25 | 6.75 | 7 | -0.25 |
| P2 | 4.5 + 0.75 + 1.75 - 1.25 | 5.75 | 6 | -0.25 |
| P3 | 4.5 + 0.75 - 1.75 + 1.25 | 4.75 | 5 | -0.25 |
| P6 | 4.5 - 0.75 + 1.75 - 1.25 | 4.25 | 4 | +0.25 |
| P7 | 4.5 - 0.75 - 1.75 + 1.25 | 3.25 | 3 | +0.25 |
| P4 | 4.5 + 0.75 - 1.75 - 1.25 | 2.25 | 2 | +0.25 |
| P8 | 4.5 - 0.75 - 1.75 - 1.25 | 0.75 | 1 | -0.25 |
Every error is a quarter of a point and the predicted order is P1, P5, P2, P3, P6, P7, P4, P8 - exactly the stated ranking. The additive model fits this respondent well, which is the licence to use it for prediction.
Step 4 - relative importance. Each attribute’s range is the gap between its best and worst part-worth.
Weight = 0.75 - (-0.75) = 1.50 · Range = 1.75 - (-1.75) = 3.50 · Price = 1.25 - (-1.25) = 2.50 · Sum = 7.50Weight = 1.50 / 7.50 = 20.0% · Range = 3.50 / 7.50 = 46.7% · Price = 2.50 / 7.50 = 33.3%Range 46.7%Price 33.3%Weight 20.0%
Battery range is roughly twice as important to this respondent as weight, and price sits between them. Note that nobody said so - it was extracted from eight rankings.
Step 5 - the trade-off in money. The price part-worth spans 2.50 utility points across a 600 euro gap, so one utility point is worth about 600 / 2.50 = 240 euros to this respondent. Therefore going from 50 km to 80 km of range (worth 3.50 points) is worth about 3.50 times 240 = 840 euros, and shedding 5 kg (worth 1.50 points) is worth about 360 euros. That is a directly usable pricing input.
Step 6 - predict between two new configurations. Manufacturing says a 2100 euro price point is achievable, a level that was never tested. Interpolating linearly, its part-worth is the average of the two tested price part-worths: (1.25 + (-1.25)) / 2 = 0.00. Two candidate builds at that price:
| Candidate | Weight | Range | Price | Total utility | Result |
|---|---|---|---|---|---|
| Config A - the light one | 20 kg | 50 km | 2100 | 4.5 + 0.75 - 1.75 + 0.00 = 3.50 | loses |
| Config B - the long-range one | 25 kg | 80 km | 2100 | 4.5 - 0.75 + 1.75 + 0.00 = 5.50 | preferred |
Config B wins by 2.00 utility points. The reasoning is transparent: the extra 30 km buys +3.50, the extra 5 kg costs -1.50, and the net gain of +2.00 is why range should be protected and weight conceded. Worth 2.00 times 240 = about 480 euros of willingness to pay. And the honest caveat from section 10 applies: 2100 euros is safely between two tested levels, so interpolating is fine. Predicting for a 3500 euro model would be extrapolation, and should not be trusted.
Apply it to your project
Section titled “Apply it to your project”-
Write down the decision the study must inform. Which two or three specification choices are you genuinely torn about, and what would you do differently depending on the answer? A conjoint that cannot change a decision is not worth the respondents’ fatigue.
-
Pick at most five attributes and justify the list. Choose the ones you believe actually drive the choice, and be able to say why these and not others. Ten attributes is the practical ceiling and five is a far more comfortable number for a student project.
-
Define levels that are realistic and bracket your real options. Two to four levels each, spanning what you could actually build and what competitors already offer. Remember that importance is relative to the ranges you choose, and that you can only safely interpolate inside them, so put your candidate specifications inside the span.
-
Count the full permutation. Multiply the level counts together. If the number is more than about ten, you need a reduced design, presented either as profiles to rank or rate, or as choice sets of two or three options.
-
Choose how you will ask, and set the scene. Ranking is easy to analyse by hand for a small design; a rating scale gives more information per profile; a choice task is the most realistic. Put the respondent in a specific scenario before the first profile, add an occasional dominated option as an attention check, and collect the background variables you would later want to segment on.
-
Collect the data, and keep a few profiles as a holdout. Do not use every profile for estimation. Hold two or three back so you can test whether the fitted part-worths predict judgements the model has not already seen.
-
Estimate the part-worths and the importances. For a small balanced design, level averages minus the grand mean is enough for a spreadsheet; for anything larger use dummy variable regression, or a logit model if the data are choices. Then compute each importance as its range over the sum of ranges, and plot the part-worths across levels to spot anomalies - jumps, plateaus, or a level so bad it acts as a knock-out.
-
Simulate your candidate configurations and check for segments. Score each realistic specification, convert the differences into money using the price part-worth, and cluster the individual importance vectors to see whether you are looking at one market or three.
Key terms
Section titled “Key terms”| Term | What it means in plain words |
|---|---|
| Conjoint analysis | A method that shows people whole products, records which they prefer, and works backwards to the value of each attribute level |
| Attribute | A dimension of the product that can be varied, such as weight, range or price |
| Attribute level | One specific value the attribute can take, such as 20 kg or 1800 euros |
| Profile (stimulus) | One complete hypothetical product, made of exactly one level from every attribute |
| Part-worth utility | The preference value contributed by one specific attribute level - the number the analysis exists to find |
| Additive utility model | The assumption that a product’s total utility is the sum of the part-worths of its levels, plus an error term |
| Relative importance | An attribute’s part-worth range divided by the sum of all attribute ranges, so all importances add to 100 percent |
| Range of an attribute | The absolute gap between its highest and lowest part-worth - the measure of how much it moves total utility |
| Interaction | When the joint effect of two attributes differs from the sum of their separate main effects, so one attribute’s importance depends on another |
| Full permutation | Every possible combination of levels - the product of the level counts, which explodes quickly |
| Reduced (fractional) design | A carefully chosen subset of profiles that keeps attributes uncorrelated so their effects can still be separated |
| Choice-based conjoint | The variant where respondents pick one option from a small choice set, analysed with a logit model |
| Holdout sample | Profiles deliberately excluded from estimation and used afterwards to test whether the model predicts well |
| Simulation | Scoring product configurations that were never shown to anyone, to compare candidate designs before building them |
Test yourself
Section titled “Test yourself”- State the underlying question of conjoint analysis and the four goals attached to it in the notes.
- Why does asking respondents to distribute 100 points across hotel attributes still fail, even though the points budget forces a choice?
- Explain an interaction in your own words using the hotel window example, and say what the standard conjoint model assumes instead.
- What does the name conjoint refer to, and what are the two outputs the decomposition produces?
- Calculation. A two-attribute study on a laptop yields part-worths of +1.2 and -1.2 for screen size, and +0.4 and -0.4 for keyboard backlight. Compute each attribute’s range and its relative importance. Then, if the base utility is 5.0, compute the total utility of a laptop with the good screen and no backlight.
- Your fitted model uses price levels of 200 and 400 euros. A colleague asks you to predict demand at 900 euros. What do you say, and why?
Revision summary
Section titled “Revision summary”Next: House of Quality → - turning what customers want into what engineers build.