Rory Sutherland’s 2019 book is filed under marketing. The chapter on why the need to look rational makes organizations act dumb belongs on every engineering manager’s desk.

Bottom line. Read Alchemy. Strip away the theory and three claims survive, all narrow and all testable. Averages destroy information about variation. Some decisions depend on exactly the information they destroy. And organizational incentives push people toward quantification even where the numbers cannot support the decision being made. The first two were proved with calipers in 1950. The third explains most of what goes wrong in a planning meeting.
Rory Sutherland is Vice Chairman of Ogilvy UK, the book came out in 2019, and it is filed under marketing, which is why most engineering leaders have not picked it up. That is a mistake. One chapter in particular is titled “Be Careful with Maths: Or Why the Need to Look Rational Can Make You Act Dumb.” It is the clearest treatment I have found of a failure I watch play out constantly. The book is uneven and its theory is loose, and I will get to that. The observations are worth the price on their own.
What follows is the argument as Sutherland makes it, the evidence he brings, where I think it applies to running an engineering organization, and where he overreaches. If the argument convinces you, buy the book rather than settling for this summary.
His opening move on the subject is a story about cockpits.
The cockpit that fit nobody
Flying for the US Air Force in the late 1940s was dangerous work. Planes were faster and more complicated, and pilots were losing control of them at an alarming rate. The problems spanned many different aircraft. On the worst day, seventeen pilots crashed.
Blame first went to pilot error and poor training. Attention then turned to the cockpit, which had been built to a single specification: the average pilot of 1926. Seat dimensions, distance to the pedals and stick, windshield height, even the helmets all conformed to measurements taken a quarter century earlier.
The reasonable next step was to update the average. In 1950 researchers at Wright Air Force Base ran the biggest study of pilots yet attempted. They measured more than 4,000 of them across 140 dimensions, including thumb length, crotch height, and the distance from eye to ear. Recalculate the average, rebuild the cockpit around it, and the crashes should stop.
A 23-year-old lieutenant named Gilbert Daniels had a different question. He wondered whether the average pilot existed at all.
Sutherland lays this out on page 99. Todd Rose covers the same study at greater length in The End of Average, and the primary source is Daniels’ own 1952 Air Force technical note. Alchemy is where most business readers will meet it.
What Daniels found
Daniels took ten dimensions relevant to cockpit design, including height, chest circumference, and sleeve length. For each he computed the average, then defined an “average” range generously as the middle 30 percent of the distribution. For height that meant anyone between 5 feet 7 and 5 feet 11 counted.
Then he asked how many of the 4,063 pilots fell in the average range on all ten dimensions.
Zero.
Not one pilot out of four thousand. Daniels then cut the requirement to any three dimensions, picking something like neck, thigh, and wrist circumference. Fewer than 3.5 percent qualified.
What he found instead was that every pilot had an uneven profile. A man with long arms often had short legs. A big chest came with narrow hips. Two pilots of identical height had different head sizes and different reach. The average described a person who did not exist, and the cockpit built around that person fit no one who had to fly it.
Daniels wrote it up in a 1952 report. The tendency to think in terms of the average man, he said, is a pitfall, and it is virtually impossible to find an average man in the Air Force population. That came from inside the institution rather than from a critic outside it.
The fix was adjustability
Air Force engineers did not respond by computing a better average. They abandoned the idea that a single specification could work and required cockpits to adjust to the pilot.
Adjustable seats came out of this. So did adjustable straps, pedals, and helmet fittings. The design target moved from the middle of the distribution to the range of it, and the equipment flexed to meet whoever sat down. Every car you have driven inherits that decision.
Air Force accident rates fell substantially over the following two decades. Aircraft destroyed per 100,000 flying hours went from around 23.6 in the late 1940s to roughly 4.3 by the end of the 1960s. Many things changed across those twenty years, and crediting adjustability alone would overstate the case. The design principle survived, spread far beyond aviation, and has not been seriously challenged since.
The averages a company runs on
Sutherland works this territory in marketing terms. Here is where it goes if you run an engineering organization, which is an extension the book invites rather than makes.
Count the averages your organization treats as real.
Handle time. Velocity per sprint. The customer in a persona document. Engagement score, time to hire, deal size, tenure. Each of those is reported as an average, computed across a population, and each gets used as though a representative individual sits behind it.
Take average velocity. A team completes 34 points a sprint on average, so the plan assumes 34 next sprint. But the sprints behind that mean might be 12, 51, 29, and 44. No sprint was average. The mean describes a sprint that never happened, and planning against it produces commitments that miss in both directions.
Personas repeat the failure in a friendlier format. A composite customer assembled from survey averages gets the mean age, the mean budget, the mean job title, and the mean complaint. Daniels would recognize the arithmetic. That customer does not exist, and the product built for them fits the actual customers approximately or not at all.
Performance ratings do it too. Rate a hundred engineers and average the scores. The middle of that distribution is someone with average code quality, average communication, average mentoring, and average judgment. Nobody on the team looks like that. Strong engineers are strong unevenly, which is the profile Daniels measured with calipers.
Sutherland’s larger complaint
Sutherland draws two consequences from the cockpit study that go well past cockpit design.
The first concerns where innovation comes from. Averages pull attention to the middle of a market, he argues, while innovation happens at the extremes. One outlier customer teaches you more than ten average ones. The people who use a product strangely drive more development than the people who use it as intended. A team studying its median user is studying the person least likely to show it something new.
The second concerns hiring, and it is the sharpest thing in the chapter for anyone who manages. Apply identical criteria to every candidate in the name of fairness, Sutherland writes, and you end up recruiting identical people. He argues that complementary talent is worth more than conformist talent, and that recruitment involves an unavoidable trade-off between fairness and variety.
That sits badly with how most hiring loops are built. A standardized rubric scored consistently across candidates is defensible, auditable, and legally safer. It also selects for the candidate who scores acceptably on every dimension, which describes someone average across the board. The uneven candidate, strong in two areas and weak in a third, gets filtered out by the same consistency that makes the process fair.
His broader claim is that organizations have become trapped in a preference for decisions that can be justified numerically. In his words: if we allow the world to be run by logical people, we will only discover logical things, when most real things are psycho-logical rather than logical. Logic should be a tool rather than a rule, he argues. He also warns that many self-described data-driven people lack insight into the model they are using or the context of their data, which boxes them in.
His sharpest observation for anyone inside a company is about incentives. Sutherland points out that it is much easier to be fired for being illogical than for being unimaginative. That asymmetry explains a great deal about why meetings fill with numbers. A quantified decision that fails is defensible, since the analysis was sound and the world misbehaved. An intuitive decision that fails is a firing. Given those payoffs, people will manufacture a metric for a decision that does not have one, and the metric will usually be an average.
190,000 taste tests
Sutherland does not use this case, and it is the one I would have added. Coca-Cola shows what that incentive produces at scale.
By the early 1980s Pepsi had spent years running the Pepsi Challenge, blind taste tests that consistently favored their sweeter formula. Coca-Cola ran its own tests and got the same result. Executives concluded that taste, rather than Pepsi’s advertising, explained their declining share.
What followed was not a hunch. Under the internal codename Project Kansas, the company spent two years and roughly $4 million on market research, including about 190,000 blind taste tests across the US and Canada. The reformulated product beat both the original and Pepsi. On the numbers the decision was sound.
New Coke launched in April 1985 and was withdrawn 79 days later. At the peak the company was fielding up to 8,000 calls a day on its consumer hotline and received around 40,000 letters.
Two flaws in the research explain the gap.
The tests were sip tests. A small sample in a controlled setting favors a sweeter product, which is not the same as drinking a full can. The measurement captured immediate preference and missed consumption.
The second flaw is the one that should worry anyone running a survey. The company never asked how people would feel if the new formula replaced the original. The tests compared liquids. The decision was about removing something from people’s lives, and no question in the instrument addressed that.
Then the part that makes this a story about looking rational rather than a story about bad research. Focus groups did surface the anger. Participants told researchers that removing the original would upset them. Those signals were reconciled against surveys that said otherwise, and the surveys won, because a survey with 190,000 responses looks more like evidence than a room full of people expressing feelings.
Donald Keough, then president, said afterward that all the time and money and skill poured into the research could not measure or reveal the deep and abiding emotional attachment people felt for the original. He added a line that has aged well: some cynics say the whole thing was planned, and the truth is the company was neither that dumb nor that smart.
Jeff Bezos gave the operating rule for that situation at a leadership forum in 2018. He calls himself a fan of anecdotes in business and notes that Amazon monitors an enormous number of metrics. Then: “when the anecdotes and the data disagree, the anecdotes are usually right.” His reason matters more than the line. The disagreement means something is wrong with the way you are measuring. You still need the data, he argues, and you need to check it against your instincts.
Coca-Cola had that exact disagreement in front of them. The focus groups were the anecdotes, the 190,000 taste tests were the data, and the anecdotes were right, because the instrument was measuring sips instead of attachment. The disagreement was the signal, and it was resolved in favor of the larger sample.
What he means by magic
Sutherland’s alchemist is not a person who understands universal laws. The alchemist is a person who spots the many cases where those laws do not apply. The title means that, and not much more.
The value he is describing gets created by changing perception rather than physical reality, and it is usually available at a fraction of the cost of the engineering alternative. He calls these psychological moonshots. An engineering moonshot spends heavily to make something ten times better. A psychological moonshot produces a comparable result by changing how the thing is experienced.
His example is the Uber map. It does not reduce the wait for a car by a single second. It reduces the uncertainty of waiting, which is the part people actually mind, and by his estimate makes waiting about 90 percent less frustrating. No engineering budget would have funded that, since the map does not move a car any faster.
The Eurostar case is the one he is best known for. Britain spent around six billion pounds on high-speed rail to cut the London to Paris journey by roughly 40 minutes. Sutherland’s alternative: for about a tenth of that money, hire the world’s top supermodels to serve free Château Pétrus to every passenger for the whole journey. You keep five billion in change, and passengers ask for the trains to be slowed down.
The response to that proposal matters more than the proposal. It was dismissed as absurd on the grounds that any improvement would be psychological, and therefore somehow not real. Sutherland’s reply is that in human behavior psychological improvements are the only ones that matter, since nobody makes decisions objectively.
Which produces the third failure mode, and the most expensive of the three.
The first two are errors. An average that fits nobody is a wrong number. A taste test that never asks the relevant question is a wrong instrument. Both show up eventually, because something breaks and someone investigates.
The third one never shows up at all. It is the good idea that was never proposed, because the person who had it could not justify it in advance and knew what would happen in the meeting. Magic, in Sutherland’s sense, resists pre-justification by definition. Its value is psychological, so it cannot be modeled in advance, so it dies at the point where someone asks for the projected return. That cost appears on no ledger. Nobody audits the ideas that were never brought to the room.
Sutherland’s line for this: when you demand logic, you pay a hidden price, and the price is magic.
Where Sutherland overreaches
The honest note: Alchemy is a book of excellent observations attached to a loose theory, and it has been fairly criticized on that basis. Reviewers have pointed out that many of his examples amount to saying a model would be better with more variables in it. That is an ordinary and entirely logical claim rather than evidence that logic has failed. Calling the improvement psycho-logic rather than better modeling does not add much.
The criticism is fair about the theory. It does not touch the chapter this article draws on, where the evidence is Daniels’ and the reasoning is arithmetic. Sutherland also scopes his own claim more carefully than his critics allow. He has said he does not want a conceptual artist running air traffic control, and that his argument applies where success depends on human perception. What survives is the three claims at the top of this piece, and Sutherland is better read as a catalogue of cases than as a theory of mind.
What the book leaves you with
Four things I took away and now use.
Ask how many real people the average describes. This is Daniels’ question and it takes five minutes. Pull the underlying distribution and count how many individuals sit near the mean on the dimensions that matter. If the answer is few, the average is an artifact rather than a description.
Report the spread alongside the mean, always. A velocity of 34 ranging from 12 to 51 is a different planning input than 34 ranging from 31 to 36. The second supports a commitment and the first does not, and the mean alone cannot tell them apart.
Check your hiring rubric for the composite candidate. If the scoring rewards acceptable performance on every dimension, it is selecting for evenness, and evenness is rare among people who are exceptional at anything. Ask what a candidate would need to be strong at for a weakness elsewhere to be worth absorbing, and write that down before the loop starts.
Watch for the number that arrives late. When a team reaches for a metric after a decision is already made, that number is there to survive the review rather than to improve the choice. Naming it in the room, gently, beats the alternative, which is a company full of decisions nobody will own.
Worth reading, with a caveat
Alchemy is a 384-page book and this article has covered roughly one chapter of it. The rest ranges across signaling, satisficing, why Red Bull succeeded by tasting bad, and a great deal of advertising history. Some of it is excellent. Some of it is a man who is very good company telling stories that do not quite add up to the thesis he claims for them.
Read it for the chapter on maths, and for the habit of mind it builds. Sutherland’s real contribution is a persistent, well-armed suspicion of any number presented as settling an argument, rather than a theory of decision-making. That suspicion is useful to have in a planning meeting.
Buy it, skim the parts that do not apply to you, and read page 99 twice.
Reader discussion