A Third of Feedback Makes Performance Worse

Kluger and DeNisi found feedback backfiring in 38 percent of cases. The failures have two things in common: no specific action and distance…


A Third of Feedback Makes Performance Worse

Kluger and DeNisi found feedback backfiring in 38 percent of cases. The failures have two things in common: no specific action and distance from the work.

Austin, Texas

Bottom line. Feedback works when it is tied to a specific action at the moment of correction and becomes unreliable as it drifts from that. Across 607 effect sizes and 23,663 observations, the average effect was positive, while 38 percent of feedback interventions made performance worse. The pattern behind the failures is usable. Feedback helps when it gets someone thinking about the work and hurts when it gets them thinking about themselves. Name a specific action while the work is fresh, and they stay on the work. That finding is a problem for the annual review. A rating tied to compensation and a permanent record make the conversation about the person, while the work being discussed is months gone. Feedback meant to change behavior belongs in the moment, attached to the work, close enough to the event that both people remember the details. Evaluation is a separate job and belongs in a separate meeting.

The 1996 meta-analysis

Avraham Kluger and Angelo DeNisi published their meta-analysis in Psychological Bulletin in 1996. They pulled 607 effect sizes from 131 papers, covering 12,652 participants and 23,663 observations. Then they asked a question the field had assumed the answer to.

Average feedback effect: d = 0.41, a moderate improvement. That figure travels widely.

The rest of the result travels less: over a third of the effects were negative. In 38 percent of cases, giving feedback left performance worse than giving none. Kluger and DeNisi checked whether sampling error explained it, or whether negative feedback accounted for all the damage while praise stayed safe. Neither held. The effect was real, and praise carried the same risk as criticism.

They also documented something uncomfortable about the field itself. Negative results had been showing up since the beginning of the century and had been largely ignored because feedback that improved performance was treated as too obvious to test.

Where attention goes

Kluger and DeNisi built a theory to explain the failures, and it is the practical part of the paper.

They proposed that feedback works by moving a person’s attention, and that attention sits at one of three levels. At the bottom is task learning, the details of how the thing gets done. In the middle is task motivation, the effort put into doing it. At the top are meta-task processes, which means the self: how competent am I, what does this say about me, what do they think of me.

The finding: feedback effectiveness drops as attention moves up that hierarchy, away from the task and toward the self. Feedback aimed at the task improves performance. Feedback aimed at the self degrades it, and the damage happens whether the self-directed message is flattering or harsh.

A concrete pair makes the difference visible. “This query does a full table scan on a five-million-row table, so it will slow down as the data grows,” points out the work. “You need to think more carefully about performance,” points at the person. Both can be said kindly.

What a performance review is made of

A review assigns a rating and compares the person to others, sometimes openly through ranking sessions. The rating attaches to compensation and enters a permanent record that follows the person through promotion decisions. All of it happens in a scheduled meeting whose subject is the person’s worth to the organization.

Every one of those properties pushes the conversation to the top of the hierarchy, where feedback stops working. A manager delivering careful, task-level coaching inside a review meeting is fighting the container. The employee hears “your design docs would be stronger with more detail on failure modes.” They are processing what it means for their rating. The advice reaches a mind that has already moved up a level.

Delay compounds the problem. Feedback about an incident from eight months ago arrives when neither person remembers the specifics, which forces the conversation into generalities about character and pattern. Generalities about a person are self-directed by definition. The review format converts task feedback into self-feedback through delay alone.

The timing evidence is split

Research on feedback timing is genuinely contested, and the split runs along a line worth understanding.

In laboratory studies of memory and retention, delayed feedback sometimes outperforms immediate feedback. Delaying corrective information can serve as spaced practice and improve long-term retention, as several studies support. Some of that literature is confounded, since delayed feedback often ends up being administered closer to the final test, inflating its apparent benefit through recency.

The workplace research found the opposite. Research on feedback timing in multi-step tasks finds a clear winner. Feedback delivered right after a decision is implemented best promotes learning and future performance. That moment carries the lowest cost of learning. Feedback delivered before implementation discourages learning. Feedback delivered after long delays increases learning costs and leads to worse outcomes.

What reconciles the two is the kind of learning involved. Memorizing vocabulary and improving how someone runs a design review are different tasks. For work, where the lesson is embedded in a specific situation with specific details, the value decays as the details fade from both memories.

Two jobs, one meeting

The review is usually asked to do two things that pull in opposite directions.

Job one is evaluation. Rating performance, deciding compensation, and documenting a record. This job is inherently self-directed, and no amount of framing changes that. It also requires care, documentation, and fairness across a group.

Job two is improvement. Helping someone do the work better. This job requires task-level attention, specific detail, and proximity to the event.

Running both in the same meeting means job one wins. Money and rating consequences crowd out the coaching. The coaching gets delivered and absorbed by none of it, and both people leave believing feedback happened.

Separating them costs nothing structurally. The evaluation conversation stays where it is and stops pretending to be coaching. The improvement conversations move into the weeks where the work happens, where they can be specific and brief.

The rating depends on the coaching

The two jobs stay connected even after you separate them.

A manager who has given feedback all year has a record to rate from. A manager who has not is working from memory, and a year of it comes back weighted toward the last few weeks. Recency bias is just how recall works. Nobody is being lazy. The result is a rating that measures the last six weeks of the year.

So, the continuous practice is the prerequisite for the annual one. The coaching conversations supply the evidence the evaluation runs on.

Make sure that someone writes them down. Fifty useful conversations that live in nobody’s notes leave the manager rating from memory anyway, back at recency with extra steps. A line per person per week is enough. Date, artifact, what changed.

The employee has the same problem. Someone who got steady feedback all year and had a bad final month will hear the rating as a verdict on that month. Naming the timeline in the review conversation helps: what the year looked like, where the hard stretch sat, and how it weighed. The notes make that possible. Without them, the manager is arguing from memory against an employee who is also arguing from memory, and the recent stretch wins both arguments.

What immediate feedback looks like

The practice that works is smaller than it sounds.

It happens within a day or two of the thing it refers to, while both people can recall the specifics. It refers to an artifact: this pull request, this incident review, this customer call. Good feedback hands over a next action the person could take within the week, and runs one or two sentences most of the time rather than a scheduled sit-down. The ceremony of a scheduled session is itself a signal that the subject is the person, not the work.

And it happens often enough that no single instance carries weight. A manager who gives task-level feedback twice a week has instances that cost nothing emotionally. A manager who gives it twice a year has created an event, and events point at the employee, not the work.

The failure mode to avoid is manufacturing feedback to hit a frequency target. Saved-up observations delivered on a schedule repeat the review problem on a smaller scale.

What to do differently

Give feedback while the work is still on the screen. Two days is the outer limit for anything specific.

Name the artifact rather than the person every time. If the sentence cannot be written about a piece of work, it belongs in a different conversation. Role fit and expectations are evaluation topics.

Stop saving observations for the review. The review should contain nothing the person has not already heard, and a review that surprises someone is a report on the manager’s communication over the preceding year.

Write the feedback down as you give it. One line per person per week. You will rate those notes in eleven months. They are the only defense against rating the last six weeks.

Then check the ratio. A manager whose feedback arrives mostly in scheduled meetings about the person is operating in the 38 percent. The fix is venue and timing, not wording.

By Joshua McDonald on August 25, 2026.

Canonical link

Exported from Medium on August 26, 2026.