Ask Your Dashboard What It Reads When Nothing Changed
Work metrics move when you do work. Impact metrics move when the work mattered. AI made the first kind free to move, and most dashboards…
Ask Your Dashboard What It Reads When Nothing Changed
Work metrics move when you do work. Impact metrics move when the work mattered. AI made the first kind free to move, and most dashboards can’t tell them apart.

Baltimore
The number that only went up
I shipped a dashboard this spring with one headline metric: tokens saved. I was building a prompt compressor, a tool that shortens what you send to an AI model so each request costs less. The metric counted what the tool removed, and it kept going up. Big, purple, satisfying, and trivially cheated. Set the compressor to its most aggressive rate and the metric hits an all-time high while the product does its worst possible work. Nothing else on the dashboard could register the damage, because nothing on the dashboard measured whether the answers coming back were still any good.
What gets measured gets done, the saying goes, and the compressor proved it. Deleting tokens was what got measured, so deleting tokens was what got done. Goodhart’s law says a measure that becomes a target stops being a good measure. Knowing the law changed nothing on Monday morning. Two questions did.
Two questions
Does this measure the work or the impact? Work metrics count what the system did: tickets closed, pull requests merged, lines of code, emails sent, tokens saved. Impact metrics measure what happened to the thing you cared about: revenue, latency, retention, whether the answer got worse. Work metrics have a property that makes them irresistible. They always move. Do more work and the number goes up, whether or not anything improved.
There’s a test, and it runs in one morning. Arrange the system so that no impact is possible and watch the number. Close ten duplicate tickets and tickets-closed climbs. Merge a refactor that changes no behaviour and PR count climbs. Tell my compressor to keep one word in ten and the tokens-saved number has its best day of the quarter. A number that moves when nothing improves is a work metric, whatever slot it occupies on the dashboard. The headline slot implies impact. Nobody demotes a work metric from it, because work metrics always have good news.
What does it read when nothing changed? This one is for the impact metrics, and it’s the question this article exists to ask. An impact metric arrives through an instrument: a survey, a log pipeline, a grader, someone’s memory of the quarter. Instruments wobble. Until you’ve taken a reading with the input held still, you don’t know how much of any movement is real. The reading doesn’t require a lab. Split a week’s traffic in half at random and grade the halves against each other. Compare this quarter to the same quarter last year, when nothing shipped. Hold one region out of the rollout. Anything that runs a known nothing through the instrument counts.
The judge gave nothing a 93.8%
I learned the second question by flunking it. The compressor’s whole risk is that shorter input produces worse answers, so I needed a quality metric to sit beside tokens saved. I hired a second AI as a judge. It reads two answers to the same question, one produced from the full text and one from the shortened text, and scores how much quality survived the cut.
Before trusting the judge I borrowed a trick from drug trials. Every trial has a control arm, the group that takes the sugar pill, because you can’t know what the drug did until you know what happens to people who took nothing. My control arm took nothing. The same full-length text went through twice, produced two answers, and the judge compared them. Nothing was shortened, so the true score is 100 percent: no quality can be lost where nothing was changed.
The judge returned 93.8%.
Six percent of quality loss appeared from nowhere, on text nobody touched. The judge was punishing ordinary variation: the model words its answers a little differently every time, and the judge itself scores a little differently every time. My tolerance for real loss was one percent, which means I had been trying to detect a one percent effect with an instrument that wobbles six. The result I was ten minutes from publishing, text a third shorter at 95.7% quality, fell entirely inside the wobble. Averaged over three runs it came back as a 9.5% loss. I hadn’t found a good trade. I had watched a coin land well, once.
The rows that never move
Dashboard hygiene says to delete the rows that never move. Reviews are long, attention is scarce, cut the flatlines. The control arm inverts that advice. A metric that reads flat when nothing changed is the error bar for every other row on the page, the one line telling you the instrument works. My control arm will read roughly 94% forever, and the day it doesn’t is the most important day on that dashboard.
Watch instead for a different pairing: a work metric climbing next to a flat impact metric. That gap is the most honest sentence a dashboard can produce. Effort went in, nothing came out, and now you know.
The same logic settles an old argument about whether teams should report metrics they can’t control. Report them, but file them as instrument readings: context, denominators, noise floors. The trouble starts when an uncontrollable number migrates into a performance review and someone collects credit for something nobody in the room can move. A team can no more move the noise floor than it can move the exchange rate, and a review that rewards either is measuring luck with decimal places.
Picking the number
Both questions are audits, and audits assume the hard choice already got made. Out of everything measurable, which number actually moves when the business goal moves? That’s the hardest problem in the exercise. The second hardest is instrumentation: finding a system that can read that number reliably enough to act on. Most dashboards fail the first problem by default, filling up with whatever the tooling emits for free, and the tooling emits work metrics because the tooling can only see the system. The goal lives outside it.
Quality is my leading candidate, and the case doesn’t start with software. In the 1920s the Bell System had a manufacturing problem: the phone network ran on millions of identical parts, strung across the country and buried under streets, and every inconsistent part was a future field repair someone had to dig for. A physicist named Walter Shewhart, working at Western Electric, the Bell System’s manufacturing arm, sent his boss a one-page memo in May 1924. It contained a sketch that became the control chart, and the idea inside carried further than the sketch. Every process varies on its own, even when nothing is wrong, so the first job is to chart the process and learn its natural band. Movement inside the band is noise; react to it and you make the process worse. Movement outside the band is a signal worth chasing. A phone company worked out what a reading means when nothing changed, fifty years before anyone owned a dashboard.
W. Edwards Deming, a statistician who learned the method from Shewhart himself, carried it to Japan in 1950 and lectured to engineers and executives rebuilding Japan’s industrial base. They listened. Japanese manufacturers spent the next thirty years beating American firms on quality, and Deming stayed largely unknown in the United States until a 1980 NBC documentary, “If Japan Can… Why Can’t We?”, profiled his work and American industry started paying attention. My control arm does what Shewhart’s chart did, a century late and pointed at a language model.
Quality earns the candidacy because it reads at every point where the work meets someone who didn’t do it. A defect rate reads it at the loading dock. Hold time reads it at the phone line. My judge reads it in the answer coming back from the model. Find the touchpoint your business goal runs through, and the quality reading there is usually the impact metric you were hunting for.
Fast past the point of feeling
Operational metrics survive the work-or-impact test better than most. Latency and uptime trace to a customer outcome without much argument; no customer has ever asked for slower, flakier, or more expensive. But they carry their own trap: diminishing returns. Every impact metric has a threshold past which the customer stops being able to feel the improvement, and past that threshold the number quietly converts into a work metric. It still moves. The experience it stood in for stopped moving a quarter ago, and nothing on the dashboard says which side of the threshold you’re spending on.
Sometimes the margin that matters isn’t even the one being measured. In 2012 Alex Stone opened a New York Times piece called “Why Waiting in Line Is Torture” with a story about a Houston airport. Passengers complained about waits at baggage claim, so the airport added handlers and drove the average wait down to eight minutes, inside industry benchmarks. The complaints kept coming. A closer look found passengers walked one minute from the gate and then stood seven at the carousel. So the airport moved the arrival gates farther away and routed bags to the outermost carousel. Passengers walked six times longer, their bags were waiting when they arrived, and complaints fell to near zero. The measured wait barely changed. The felt wait collapsed, because occupied time reads shorter than unoccupied time.
Streaming gives software the same lever. Total completion time for a long LLM answer might run twenty seconds and resist every optimization you can afford. Time to first token is a different margin. If the first token arrives in half a second, the customer spends the other nineteen and a half reading, watching the answer assemble, getting a window into the reasoning as it forms. That’s occupied time, and for the less generous read, at least the feeling that something is getting done. A latency dashboard that tracks total completion and not time to first token is measuring the bags and ignoring the walk.
Ten minutes of hold music
In December 2023, on the Lex Fridman podcast, Jeff Bezos told a story from Amazon’s early years about this exact failure. The metric in the weekly business review said customers calling the 1–800 support line waited under sixty seconds. Complaints kept insisting otherwise. During one review, with the head of customer service defending the number, Bezos picked up a phone and dialed. The room sat with the hold music for more than ten minutes.
His summary became its own aphorism: when the data and the anecdotes disagree, the anecdotes are usually right. He was precise about the meaning, and the precision is the useful part. The point is not that anecdotes trump data. It’s that the data is usually measuring the wrong thing rather than being miscollected, so a pile of complaints your metrics say shouldn’t exist is a reason to doubt the metrics.
Read the story as instrumentation and it sharpens further. The complaints were a second instrument, independently sampled from the same reality. Two instruments disagreed. Disagreement between instruments never tells you which one is right; it tells you to calibrate. The phone call was the calibration: one manual sample, taken outside both pipelines, in front of everyone whose credibility depended on the number. Bezos ran a control arm on speakerphone, and it cost the ten most expensive minutes of executive time that meeting had.
What AI changed
Nothing so far requires AI to exist. The principles date to Goodhart in 1975 and to hold music in the 1990s. What changed is the price of moving a work metric.
Moving one used to cost human effort. Lines of code, commits, pull requests were always work metrics, but a person had to type them, so they correlated loosely with something real. That correlation made them tolerable proxies for decades. AI removed the effort, and the correlation went with it. Faros AI’s July 2025 telemetry report across more than ten thousand developers found teams with high AI adoption merging 98% more pull requests while review time rose 91% and DORA delivery metrics stayed flat. Work metrics at all-time highs, impact metrics unmoved. The climbing-work, flat-impact pairing has stopped being a warning sign to hunt for. On an engineering dashboard in 2026 it’s the default reading. The impact side still exists, it just isn’t where the AI numbers accumulate. Deployment frequency, lead time, change failure rate, defects reaching production: the rows that stayed flat while PR volume doubled are the rows to read.
The instruments degraded on the same schedule, starting with the one every retro and every earnings call runs on: how fast people feel. METR, a research nonprofit, ran a randomized controlled trial in early 2025 with sixteen experienced open-source developers completing 246 real tasks on mature repositories, each task randomly assigned to allow or forbid AI tools. Going in, the developers predicted AI would make them 24% faster. Coming out, having done the work, they estimated it had made them 20% faster. The measured result was 19% slower. Perception missed reality by 39 points, in the confident direction, after the fact.
The sequel teaches more than the headline. In February 2026 METR announced it was redesigning its follow-up experiment. The raw numbers hinted the slowdown had reversed, but developers who benefited most from AI wouldn’t join a study that might forbid it, and the selection effect swamped the estimate. Their revised public position on whether AI speeds developers up amounts to: we currently can’t say. An organization whose entire job is measurement took a reading of its own instrument, found the wobble, and declined to print a number it couldn’t defend. Most dashboards never extend themselves that courtesy. Most dashboards print the 93.8% and call it quality.
What it costs to know
I still run the compressor. Tokens saved is still on the dashboard, one row down, under a quality gauge with an error bar the control arm paid for. The questions travel to any metric you own. Ask what the number does when the work happens and nothing improves. Ask what it reads when nothing changed at all.
By Joshua McDonald on July 16, 2026.
Exported from Medium on August 26, 2026.
Reader discussion