Skip to content

Conjoint Analysis

Innovation & New Business Proposal - TUHH Institute of Entrepreneurship & Institute of Innovation Marketing, Hamburg · part of my Technology Management MBA · study notes for revision.


Kano told me which kind of satisfaction each attribute produces. Conjoint analysis answers the next question, and it is a quantitative one: which product attributes generate which value for customers? Not “is this feature nice”, but “how many units of preference is this level of this attribute worth, compared with that level of that other attribute”. It is the tool that finally puts numbers on the trade-offs a customer makes in their head.

The class notes describe conjoint analyses as tools that measure the importance of product attributes holistically, and they attach four goals to that. Which attributes are how important? Which attribute levels provide which utility to the customer? Are there anomalies - jumps in the utility curve, knock-out criteria, minimum levels below which nothing is acceptable? And what are the trade-offs between different attributes? The last one is really the point of the whole method: every product decision is a trade, and this is the only market research tool in the module that measures the exchange rate.

What makes it feel different from a survey is that respondents are never asked about attributes at all. They are shown whole products and asked which one they prefer. The importance numbers are then worked backwards out of those judgements. That reversal is the entire trick, and it is why the technique survives the problem I describe in section 2.

1 · The underlying question and the four goals

Section titled “1 · The underlying question and the four goals”
Which attributes are how important?the relative weight each attribute carries in the decision
Which attribute levels give which utility?not just that range matters, but how much 80 km beats 50 km
Are there anomalies?jumps in the curve, knock-out criteria, minimum levels
What are the trade-offs between attributes?how much of one attribute buys how much of another
The four goals stated in the notes. The fourth is the one that separates conjoint from every other tool in the module.

The class makes the fourth goal concrete with four questions any product manager would recognise. How much more would customers prefer a product if it offered a UHD resolution rather than HD on a television, an e-bike weighing 20 kg rather than 25 kg, an electric car with 500 km of range on one charge rather than 300 km, or an online meeting tool with a breakout session function rather than without one? Each is a “how much more”, not a “yes or no”. Only a method that measures preference on a scale can answer them.

The build-up in the session is done with a holiday hotel, and it is worth reproducing because the failure happens in two stages.

Stage one: rate the importance. How important to you personally are the following attributes of a holiday hotel - the walking time from the hotel to the beach, the size of your room, the frequency of room cleaning, the speed of the Wi-Fi connection, and the availability of leisure and sports activities? Everyone answers happily, and almost everyone says most of them are quite important. Nothing has been given up, so nothing has been revealed. The answers are flattering rather than informative.

Stage two: distribute 100 points. So the question is tightened: please spread exactly 100 points across those same five attributes, so that the points represent how important each one is for you. This is better, because the budget forces some discipline. But it still fails, for a reason worth stating precisely: the respondent is being asked to introspect about abstract categories, not to make a decision. Nobody has ever stood at a hotel booking page choosing between “room size” and “cleaning frequency” as such. They choose between hotels, each of which bundles a particular level of each attribute at a particular price.

Direct questioning rating, or 100 points
  • Asks about attributes in the abstract, detached from any product
  • No real sacrifice is involved, so answers are polite and inflated
  • People are poor at introspecting about their own weightings
  • Cannot capture levels - only that an attribute matters, not how much 80 km beats 50 km
Conjoint questioning whole products, judged holistically
  • Asks the respondent to rank, rate or choose complete products
  • Every product is good at some things and weak at others, so a real trade-off is forced
  • The interview situation is natural, close to how buying actually feels
  • Importance is derived afterwards from the pattern of judgements, never asked
The same information, obtained two ways. The left column asks the respondent to do the analysis; the right column asks them only to prefer, and does the analysis afterwards.

There is a further problem that no amount of careful wording fixes. The session illustrates it with a single question asked three times: how important is the size of the window in a hotel room? Then: how important is it in this room? And in this one? The answer changes completely depending on what the rest of the room looks like. A large window matters enormously in a small dark room and hardly at all in a bright, generously sized one.

The principle stated in the notes is that the importance or preference of one attribute may depend on another attribute. That is what an interaction is: an interaction occurs when the combined effect of two attributes differs from the sum of their two separate main-effect utilities. Asked in isolation, the window question has no single correct answer, because importance is not a property of the attribute alone.

The name is a contraction, and unpacking it is the fastest way to remember what the method does. Products and services are CONsidered JOINTly - the respondent sees the whole bundle, never the pieces - and the analysis then treats that bundle as an aggregation of single components called part-worth utilities.

Product or servicedescribed as a bundle of attribute levels
→
Measurement of preferencesthe respondent ranks, rates or chooses whole products
→
Decomposition1. relative importance of attributes
2. part-worth utilities of attribute levels
The direction of travel. Data goes in as holistic preference judgements about complete products; two separate results come out - how much each attribute matters, and what each individual level is worth.

This is why conjoint is called a decomposition method, in contrast to the compositional logic of a points-allocation survey. A compositional method asks for the weights and builds the product score up from them. A decompositional method takes the product score as given and pulls the weights out of it.

The model underlying a conjoint analysis is an additive utility function. The total utility of a product is the sum of the part-worth utilities of the levels it happens to have, plus an error term that absorbs everything the model does not capture.

Additive utilityy(k) = SUM over j of SUM over m of [ b(jm) · x(jm) ] + Error
SymbolMeaning in the notation used in class
y(k)the preference for product or service k, that is its total utility
jthe attributes, for example weight, range, price
mthe levels within each attribute, for example 20 kg and 25 kg
x(jm)an indicator that equals 1 if the product has level m for attribute j, and 0 otherwise
b(jm)the part-worth utility of level m of attribute j - the unknown the analysis solves for
Erroreverything about the stated preference that the additive model does not explain

The indicator variable is the mechanical heart of it. Each product profile switches on exactly one level per attribute, so the sum collapses to “one part-worth from each attribute, added together”. In the supporting paper these indicators are handled as dummy variables in a regression, with a useful economy: an attribute with three levels only needs two dummies, because the third level is identified automatically when both dummies are zero. That third level becomes the baseline and its part-worth is absorbed into the intercept.

6 · The smallest possible example: shape and colour

Section titled “6 · The smallest possible example: shape and colour”

The session demonstrates the mechanics on a deliberately trivial case: two attributes (shape and colour), two levels each. Multiplying out gives four product profiles, and the respondent is asked one simple thing - please rank order the four products according to your preferences, which would you choose first, second, third and fourth?

Four holistic judgements, and from them come four part-worths. Writing the additive model out for each profile shows what the estimation is up against:

ProfileShape levelColour levelPreference equation
1Shape level 1Colour level 1y1 = b(shape 1) + b(colour 1) + Error
2Shape level 2Colour level 2y2 = b(shape 2) + b(colour 2) + Error
3Shape level 2Colour level 1y3 = b(shape 2) + b(colour 1) + Error
4Shape level 1Colour level 2y4 = b(shape 1) + b(colour 2) + Error

Four equations, four unknown part-worths, one shared set of stated ranks. The respondent never said a word about shape or colour, and yet the system now contains everything needed to price both.

The estimation rule is stated in the notes in one sentence, and it is worth learning in that form. Statistical estimation determines the values of the part-worth utilities so that the calculated preference values, the estimated y values, are as similar as possible to the preferences actually stated by the customer, the real y values.

That is all it is: choose the numbers that make the model’s predicted ranking agree with the observed ranking as closely as it can. In the supporting paper this is done with ordinary least squares regression of the stated ratings on the dummy variables - the fitted equation for a public transport example with fare and waiting time came out as an intercept of 49.8 plus 26.7 and 14.0 for the two fare dummies and 25.7 and 20.0 for the two waiting-time dummies, reproducing the nine stated ratings almost exactly.

A second convenience from the paper is normalisation: rescale all the part-worths so that the smallest across the whole study becomes 0 and the largest becomes 1, using v = (u - a) / (b - a), where a and b are the smallest and largest raw values. This changes nothing about the internal relationships but makes numbers comparable within and across attributes.

Once every level has a part-worth, the importance of the attribute that owns them follows immediately. An attribute matters to the extent that moving from its worst level to its best level moves total utility a lot. So importance is a range.

Range of one attributeRange(j) = highest part-worth of j MINUS lowest part-worth of j (absolute value)
Relative importanceImportance(j) = Range(j) / [ Range(1) + Range(2) + … + Range(J) ]

Property the importances of all attributes add up to 1, or to 100 percent

In words, exactly as the slides put it: the relative importance of an attribute is the absolute difference between its highest and its lowest partial utility value, divided by the sum of all such differences. The supporting paper calls this the relative importance measure and computes it for the transport case as 26.7 divided by (26.7 plus 25.7), giving 51 percent for fare and therefore 49 percent for waiting time.

Two consequences that are easy to get wrong in an exam. First, importance is a property of the levels you tested, not of the attribute in the world. Widen the price range in the study and price will look more important, purely by construction. Second, an attribute whose levels are all roughly equally liked has a near-zero range and near-zero importance, no matter how loudly respondents would have insisted it mattered if you had asked them directly.

The class lays out seven steps, and works them through on an e-scooter sharing system.

#StepWhat actually happens
1Collect product or service attributesList what could plausibly drive the choice
2Decide attribute levelsFix realistic values per attribute, and count the full permutation
3Decide model and softwareChoice-based, rank-order or adaptive; Sawtooth is the standard tool
4Design stimuliReduce the full permutation of combinations to a workable set of profiles
5Collect dataScenario construction plus background questions
6AnalyseAverage-level and individual-level analysis are both possible
7Interpret and simulateRelative importances, level utilities, and the utility of new combinations

Step 4 is where most study designs live or die.

The e-scooter attributes and levels. Maximum speed of the scooters at 15, 20, 25, 30 or 35 km per hour. Average walking time to the next available scooter at 1, 2.5 or 4 minutes. Pricing as either a 1 euro unlock fee plus 30 cents per minute, or no unlock fee plus 50 cents per minute. Reservation possible or not possible. Smartphone holder available or not available.

Why step 4 exists. The full permutation is 5 times 3 times 2 times 2 times 2, which is 120 combinations. No respondent will judge 120 profiles. So the full factorial is reduced - in the class example to 30 profiles presented as 10 comparisons of 3 profiles each, relying on the assumption that the attributes are independent. The supporting paper does the same thing at larger scale for a smartphone study with five attributes at four levels each, where the full set runs past a thousand profiles and a fractional factorial design cuts it to 32 profiles rated on a zero to 100 scale.

Choosing how to ask. The alternative ways of measuring preference, all of which work as conjoint input:

Rating - to what extent does this product match your preferences, from very much to not at allChoice - which of these products would you chooseRanking - rank order these products, most preferred firstGraded paired comparison - how much do you prefer one of these two, from clearly prefer A through indifferent to clearly prefer B

Software and data collection. The class names Sawtooth Software as the world’s leading provider, free of charge for master theses through its grants programme, and mentions a web-based choice-based implementation. Data collection is not only the profiles: it also involves scenario construction, so the respondent is judging in a defined situation, and additional questions such as age, gender, residence and general mobility preferences.

Analysis and segmentation. Both average and individual-level analysis are possible, which is unusual and valuable - conjoint gives every single respondent their own set of part-worths. That opens a segmentation route the class demonstrates: run a cluster analysis on the vectors of relative importance, using the Ward algorithm with squared Euclidean distances, and the e-scooter respondents fall into three clusters of people who trade off speed, walking time and price in visibly different ways. The supporting paper makes the same point: clusters of individuals with similar importances are effectively market segments.

The final step is the one that earns the study its budget. Because the model is additive, it can score any combination of levels, including combinations nobody was ever shown. That is the meaning of step 7, simulating the utility of new product combinations.

The supporting paper spells out three things about doing this responsibly:

  • Interpolation between levels is acceptable. A value between two tested levels can be scored by assuming the part-worth function is linear over that gap - a fare halfway between two tested fares takes the average of their two part-worths.
  • Extrapolation beyond the tested range is not. Scoring a level outside the span you tested can be badly misleading, and the remedy is to choose the attribute values sensibly when designing the profiles in the first place, so the interesting region is inside the range.
  • Validate the estimated utility functions. The two named routes are predictive validation on a holdout sample of profiles kept out of the estimation, and validation against actual or intended market behaviour such as first choices, sales or market share.

The paper also summarises the technique in five roles, a compact revision list: it is a measurement technique for quantifying buyer trade-offs, an analytical technique for predicting reactions to new products, a segmentation technique for grouping buyers with similar trade-offs, a simulation technique for assessing new product ideas against competitors, and an optimisation technique for finding the profile that maximises share or return.

Choice-based conjoint specifically. The two mainstream variants are ratings-based and choice-based, and the implementation path forks early: decide the purpose, decide the approach, identify attributes and levels, then either design profiles and analyse with regression, or design choice sets and analyse with a logit model. Both converge on the same outputs. In the paper’s choice-based version of the transport study, 42 choice sets of size three were generated and each respondent simply said which of the three they would take; the answers were analysed with a conditional logit model estimated by maximum likelihood, and the part-worth pattern matched the ratings-based result. One design detail worth stealing: the sets deliberately mixed dominated options - one alternative better on every attribute, which acts as a sanity check on the respondent - with sets where a genuine trade-off is unavoidable.

StrengthsLimits
The interview situation is natural and realistic - people compare products, which is what they do anywayRespondents must process a lot of information in each choice task
Yields the utility of each attribute level, not just of the attribute, for example price at 5, 10 and 15 eurosThe task is tiring, and fatigue degrades the later answers
Yields the relative importance of attributes, for example that price outweighs the energy label in a television purchaseA practical ceiling of roughly 10 attributes can be shown
Can detect interaction effects, such as red being specifically popular on a FerrariReduced designs usually assume attributes are independent, so most interactions go unmeasured
Flexible across industrial products, consumer products and servicesImportance depends on the level ranges you chose, so a badly designed study gives confidently wrong weights

A small conjoint on a compact commuter e-bike. Three attributes, two levels each, so the full permutation is 2 times 2 times 2 = 8 profiles, small enough to show all of them.

AttributeLevel ALevel B
Weight20 kg25 kg
Range on one charge80 km50 km
Price1800 euros2400 euros

One respondent ranks all eight from 1 (most preferred) to 8 (least preferred). To turn ranks into a preference score I reverse them, so score = 9 minus rank, and the best profile scores 8.

ProfileWeightRangePriceRankScore y
P120 kg80 km180018
P220 kg80 km240036
P320 kg50 km180045
P420 kg50 km240072
P525 kg80 km180027
P625 kg80 km240054
P725 kg50 km180063
P825 kg50 km240081

Step 1 - the grand mean. The eight scores are 8, 6, 5, 2, 7, 4, 3, 1, summing to 36, so the grand mean is 36 / 8 = 4.5.

Step 2 - the average score of every level. Because the design is balanced, each level appears in exactly four profiles, so the level average is a fair estimate of what that level contributes.

AttributeLevelProfiles containing itSum of scoresLevel averagePart-worth (average minus 4.5)
Weight20 kgP1, P2, P3, P48+6+5+2 = 215.25+0.75
Weight25 kgP5, P6, P7, P87+4+3+1 = 153.75-0.75
Range80 kmP1, P2, P5, P68+6+7+4 = 256.25+1.75
Range50 kmP3, P4, P7, P85+2+3+1 = 112.75-1.75
Price1800P1, P3, P5, P78+5+7+3 = 235.75+1.25
Price2400P2, P4, P6, P86+2+4+1 = 133.25-1.25

Step 3 - check the model reproduces the ranking. Predicted score = 4.5 plus the three part-worths of that profile.

ProfileCalculationPredictedStatedError
P14.5 + 0.75 + 1.75 + 1.258.258+0.25
P54.5 - 0.75 + 1.75 + 1.256.757-0.25
P24.5 + 0.75 + 1.75 - 1.255.756-0.25
P34.5 + 0.75 - 1.75 + 1.254.755-0.25
P64.5 - 0.75 + 1.75 - 1.254.254+0.25
P74.5 - 0.75 - 1.75 + 1.253.253+0.25
P44.5 + 0.75 - 1.75 - 1.252.252+0.25
P84.5 - 0.75 - 1.75 - 1.250.751-0.25

Every error is a quarter of a point and the predicted order is P1, P5, P2, P3, P6, P7, P4, P8 - exactly the stated ranking. The additive model fits this respondent well, which is the licence to use it for prediction.

Step 4 - relative importance. Each attribute’s range is the gap between its best and worst part-worth.

RangesWeight = 0.75 - (-0.75) = 1.50 · Range = 1.75 - (-1.75) = 3.50 · Price = 1.25 - (-1.25) = 2.50 · Sum = 7.50
Relative importanceWeight = 1.50 / 7.50 = 20.0% · Range = 3.50 / 7.50 = 46.7% · Price = 2.50 / 7.50 = 33.3%

Range 46.7%Price 33.3%Weight 20.0%

Battery range is roughly twice as important to this respondent as weight, and price sits between them. Note that nobody said so - it was extracted from eight rankings.

Step 5 - the trade-off in money. The price part-worth spans 2.50 utility points across a 600 euro gap, so one utility point is worth about 600 / 2.50 = 240 euros to this respondent. Therefore going from 50 km to 80 km of range (worth 3.50 points) is worth about 3.50 times 240 = 840 euros, and shedding 5 kg (worth 1.50 points) is worth about 360 euros. That is a directly usable pricing input.

Step 6 - predict between two new configurations. Manufacturing says a 2100 euro price point is achievable, a level that was never tested. Interpolating linearly, its part-worth is the average of the two tested price part-worths: (1.25 + (-1.25)) / 2 = 0.00. Two candidate builds at that price:

CandidateWeightRangePriceTotal utilityResult
Config A - the light one20 kg50 km21004.5 + 0.75 - 1.75 + 0.00 = 3.50loses
Config B - the long-range one25 kg80 km21004.5 - 0.75 + 1.75 + 0.00 = 5.50preferred

Config B wins by 2.00 utility points. The reasoning is transparent: the extra 30 km buys +3.50, the extra 5 kg costs -1.50, and the net gain of +2.00 is why range should be protected and weight conceded. Worth 2.00 times 240 = about 480 euros of willingness to pay. And the honest caveat from section 10 applies: 2100 euros is safely between two tested levels, so interpolating is fine. Predicting for a 3500 euro model would be extrapolation, and should not be trusted.

  1. Write down the decision the study must inform. Which two or three specification choices are you genuinely torn about, and what would you do differently depending on the answer? A conjoint that cannot change a decision is not worth the respondents’ fatigue.

  2. Pick at most five attributes and justify the list. Choose the ones you believe actually drive the choice, and be able to say why these and not others. Ten attributes is the practical ceiling and five is a far more comfortable number for a student project.

  3. Define levels that are realistic and bracket your real options. Two to four levels each, spanning what you could actually build and what competitors already offer. Remember that importance is relative to the ranges you choose, and that you can only safely interpolate inside them, so put your candidate specifications inside the span.

  4. Count the full permutation. Multiply the level counts together. If the number is more than about ten, you need a reduced design, presented either as profiles to rank or rate, or as choice sets of two or three options.

  5. Choose how you will ask, and set the scene. Ranking is easy to analyse by hand for a small design; a rating scale gives more information per profile; a choice task is the most realistic. Put the respondent in a specific scenario before the first profile, add an occasional dominated option as an attention check, and collect the background variables you would later want to segment on.

  6. Collect the data, and keep a few profiles as a holdout. Do not use every profile for estimation. Hold two or three back so you can test whether the fitted part-worths predict judgements the model has not already seen.

  7. Estimate the part-worths and the importances. For a small balanced design, level averages minus the grand mean is enough for a spreadsheet; for anything larger use dummy variable regression, or a logit model if the data are choices. Then compute each importance as its range over the sum of ranges, and plot the part-worths across levels to spot anomalies - jumps, plateaus, or a level so bad it acts as a knock-out.

  8. Simulate your candidate configurations and check for segments. Score each realistic specification, convert the differences into money using the price part-worth, and cluster the individual importance vectors to see whether you are looking at one market or three.

TermWhat it means in plain words
Conjoint analysisA method that shows people whole products, records which they prefer, and works backwards to the value of each attribute level
AttributeA dimension of the product that can be varied, such as weight, range or price
Attribute levelOne specific value the attribute can take, such as 20 kg or 1800 euros
Profile (stimulus)One complete hypothetical product, made of exactly one level from every attribute
Part-worth utilityThe preference value contributed by one specific attribute level - the number the analysis exists to find
Additive utility modelThe assumption that a product’s total utility is the sum of the part-worths of its levels, plus an error term
Relative importanceAn attribute’s part-worth range divided by the sum of all attribute ranges, so all importances add to 100 percent
Range of an attributeThe absolute gap between its highest and lowest part-worth - the measure of how much it moves total utility
InteractionWhen the joint effect of two attributes differs from the sum of their separate main effects, so one attribute’s importance depends on another
Full permutationEvery possible combination of levels - the product of the level counts, which explodes quickly
Reduced (fractional) designA carefully chosen subset of profiles that keeps attributes uncorrelated so their effects can still be separated
Choice-based conjointThe variant where respondents pick one option from a small choice set, analysed with a logit model
Holdout sampleProfiles deliberately excluded from estimation and used afterwards to test whether the model predicts well
SimulationScoring product configurations that were never shown to anyone, to compare candidate designs before building them
  1. State the underlying question of conjoint analysis and the four goals attached to it in the notes.
  2. Why does asking respondents to distribute 100 points across hotel attributes still fail, even though the points budget forces a choice?
  3. Explain an interaction in your own words using the hotel window example, and say what the standard conjoint model assumes instead.
  4. What does the name conjoint refer to, and what are the two outputs the decomposition produces?
  5. Calculation. A two-attribute study on a laptop yields part-worths of +1.2 and -1.2 for screen size, and +0.4 and -0.4 for keyboard backlight. Compute each attribute’s range and its relative importance. Then, if the base utility is 5.0, compute the total utility of a laptop with the good screen and no backlight.
  6. Your fitted model uses price levels of 200 and 400 euros. A colleague asks you to predict demand at 900 euros. What do you say, and why?

Next: House of Quality → - turning what customers want into what engineers build.