Why do people optimize the reward instead of the stated goal?
Before blaming motivation or culture, inspect the operative reward system from the recipient's point of view. A visible proxy can make A rational even while leaders announce B. Kerr does not claim that formal rewards determine every action, or that every apparent mismatch is a design error. Some cases reveal that leaders prefer A, while others reflect a real choice to prioritize equity or morality over efficiency.
Steven Kerr · Academy of Management Journal, 18(4), 769-783 · December 1, 1975 prototype 5 min read Explained by Superalignment Research
The 60-second answer
AI labs, evaluation programs, and governance teams often combine stated safety goals with visible delivery metrics. Kerr provides a compact way to inspect the incentive channel before treating unwanted behavior as a character flaw.
Kerr argues that stated goals do not by themselves explain behavior. People look for the actions that actually bring approval, pay, promotion, safety, or status. In a manufacturer, lower-level employees perceived conformity and risk avoidance as more acceptable than senior managers said they wanted. In an insurer, complaint counts, fast claim handling, narrow merit increments, and a strict attendance rule pulled behavior away from accurate claims work. Kerr links these mismatches to visible metrics, supposedly objective criteria, concealed preferences, and competing values.
- People respond to the consequences they expect, which can differ sharply from the behavior leaders publicly praise.
- Kerr separates proxy and visibility failures from hidden preferences and legitimate choices to value equity or morality over efficiency.
- The first diagnostic is to ask members what behavior actually earns approval, not to infer the reward system from policy language.
Written for: Technical generalists who design metrics, incentives, teams, or governance. Useful prerequisites: The distinction between a goal and a proxy, Basic incentive reasoning.
- The question
- Why do people rationally pursue behavior an organization says it does not want?
- What the authors did
- Kerr develops an organizational argument through examples from public life, medicine, education, business, and sport. He then reports interviews and a companywide approval-expectation survey in a manufacturer, describes reward practices in an insurance claims division, groups four reasons for reward-goal mismatch, and compares three possible remedies.
- The source
- On the Folly of Rewarding A, While Hoping for B
Which behavior does the operative reward select?
Policy view selected. The diagram assumes the stated goal and operative reward agree.
| View | What is inspected | Interpretive limit |
|---|---|---|
| Policy | Stated goal, praise, and formal rules. | May miss the consequences members actually expect. |
| Perceived payoff | Approval, pay, promotion, status, blame, and safety attached to concrete actions. | A mismatch may be a proxy error, a hidden preference, or a competing value. |
Walk through the argument
Each step below points to the source location that carries it. The source map follows the walkthrough for source-level checking.
Read the operative reward, not the poster
An organization may announce that it values careful judgment, teamwork, or long-term quality. Its members still have to decide what to do on Monday morning. They look at which actions bring approval, money, promotion, status, or protection from blame.
Kerr's A-versus-B pattern appears when those consequences favor A while leaders say they hope for B. The behavior can be personally rational even when it is bad for the stated organizational goal.
Source: Journal pages 769 to 775, opening argument and societal, organizational, and business examples
Ask what members expect to be approved
In the manufacturer, Kerr did not infer rewards from the formal policy. Interviews and a companywide questionnaire asked employees how much approval or disapproval they expected for concrete actions. The survey was anonymous and administered without company staff handling the forms.
Senior managers complained about conformity and risk avoidance. Lower-level workers, especially in one division, were more likely to report that those same behaviors brought approval. The data capture perceived consequences, not a randomized causal effect, but they expose a disagreement that policy language hid.
Source: Journal pages 775 to 778, A Manufacturing Organization and Table 1
Map every signal in the reward stack
The insurance claims division tracked returned checks and complaints as accuracy signals. Underpayment produced complaints, overpayment often did not, and requesting clarification threatened a separate two-day speed target. The local rule became to overpay when uncertain.
A small difference between merit raises weakened the performance signal, while losing the entire raise after three absence or lateness events made attendance highly salient. Calling this one reward system hides several signals with different strength and visibility.
Do not collapse four causes into one
Kerr identifies two direct design problems. Leaders can become fascinated with an objective-looking criterion, or reward only what is easy to observe. Both cases let a visible measure displace a less visible goal.
The other two cases are different. Leaders may secretly prefer the rewarded behavior, or they may openly prioritize morality or equity over efficiency. Only the first two are reward systems that truly pay for something the rewarder does not want.
Source: Journal pages 779 to 781, Causes
Audit before changing the incentive
Kerr is skeptical that selection will reliably find people whose motives match management, and skeptical that training will reliably rewrite those motives. He therefore emphasizes changing the reward system, beginning with a study of what members believe it rewards now.
That does not mean attaching money to every desired action. A new metric can become another A. The useful procedure is to map stated goal, operative criterion, perceived payoff, likely behavior, and the values that any redesign would trade away.
Source: Journal pages 779 to 781, Causes, Journal pages 781 to 783, Conclusions
Keep formal rewards in their place
Kerr explicitly says formal rewards and punishments do not determine all behavior. People can act from care, duty, identity, or professional standards even when the organization does not reinforce them.
The narrower management claim is causal responsibility. If the desired behavior appears despite an opposing reward system, the organization is a fortunate bystander. It should not assume the current design will keep producing that behavior under greater pressure.
Source map
These are the source locations that carry the argument. Use them to check this explanation against the original rather than trusting the summary alone.
| Locus | Why it matters | Source |
|---|---|---|
| Journal pages 769 to 775, opening argument and societal, organizational, and business examples | Introduces the reward-goal mismatch and shows how visible or operative consequences can make apparently unwanted behavior rational. | Open source → |
| Journal pages 775 to 778, A Manufacturing Organization and Table 1 | Describes the interviews, companywide Expect Approval survey, response conditions, and differences in perceived approval for conformity and risk avoidance. | Open source → |
| Journal pages 778 to 779, An Insurance Firm | Shows how complaint counts, a two-day processing target, small merit differences, and an attendance rule created several competing reward signals. | Open source → |
| Journal pages 779 to 781, Causes | Separates objective-criterion and visibility problems from hypocrisy and legitimate emphasis on morality or equity. | Open source → |
| Journal pages 781 to 783, Conclusions | Compares selection, training, and reward-system change, proposes auditing perceived rewards, and limits the claim about formal reinforcement. | Open source → |
The Assumption Switch
One result. One assumption exposed. Turn it and see what changes.
Assumption under test
The organization's stated goal is a reliable description of which behavior its members expect to be rewarded.
- Held in the source
- Leaders infer that announcing B and praising B means the operative reward system supports B.
- Turn it
- Ask members what actually brings approval, pay, promotion, safety, or status, and map those consequences to behavior A or B.
- What changes
- A can become the rational response even when leaders say they hope for B. The mismatch may reflect a poorly chosen proxy, a hidden preference for A, or a legitimate competing value rather than one universal cause.
The common misreading
The slogan is often read as a universal claim that people do exactly what formal incentives reward. Kerr says formal rewards do not determine all organizational behavior and notes that patriotism, professional concern, or care can survive without them. His narrower point is that leaders should not treat desired behavior as caused by the organization when its reward system points elsewhere.
Outside the ML frame
AI research management
What behavior does an AI lab reward when it says safety is a priority?
A lab can praise careful evaluation while promotion, publication, and release decisions reward speed, benchmark wins, or visible launches. Kerr's method suggests asking researchers which actions they expect to pay off and comparing that answer with the stated safety goal. This is an organizational diagnostic for AI work, not evidence that any named lab has the mismatch.
Where the result stops
The paper is an illustrative organizational essay, not an estimate of how often reward-goal mismatch occurs. Its manufacturer survey covers one company, has no external benchmark for the approval scale, and supports perception differences rather than a causal effect of rewards on behavior. The insurance account is descriptive and does not report a sampling protocol. Several societal examples rely on simplified assumptions, and Kerr explicitly narrows the claim for hypocrisy and competing-values cases.
What remains open
- How can an organization measure members' expected rewards before a mismatch becomes costly behavior?
- Which reward changes align behavior without creating a new narrow proxy or suppressing professional judgment?
- How can leaders distinguish a mistaken incentive from an honest tradeoff among efficiency, equity, legality, and care?
- When do informal status rewards dominate formal pay, promotion, or performance systems?
What this bears on
Superalignment maintains a public register of the claims it makes and the evidence that would change its mind. The entries below connect this source to the exact public claims it bears on.
- C4. Behavioral evaluation cannot carry a deployment decision alone. This record bears on it, suggestively. Kerr shows how a visible performance criterion can direct behavior away from a stated objective and how members' perceived rewards can differ from policy. The paper does not study AI deployment decisions. See the claim and what would change our mind →
Source audit and review status
| Field | Result | Checked | By | Against |
|---|---|---|---|---|
title | exact | 2026-08-17 | codex-primary-source-review | source |
authors | exact | 2026-08-17 | codex-primary-source-review | source |
date | exact | 2026-08-17 | codex-primary-source-review | source |
venue | exact | 2026-08-17 | codex-primary-source-review | source |
full_text | exact | 2026-08-17 | codex-primary-source-review | source |
- Review status prototype.
- Program collection seminal v1; theory; wave 3, release slot unassigned.
- Source access public full text, PDF. Open the reading copy →
- Provider-authored safety claim no.
- Explained by Superalignment Research.
- Reviewed by No named human reviewer yet.
- Explainer dates created 2026-08-17; updated 2026-08-17.
- AI assistance AI assisted with primary-source retrieval, full-text extraction, manifestation checking, locus mapping, first-pass prose, figure design, and implementation. The page remains a prototype until a named human review is recorded.
- Rights and access The article is publisher-controlled. A complete publisher-produced PDF is publicly readable from an MIT course archive, and a second institutional reading copy is hosted by the U.S. Air Force. Neither copy carries an open-content license. Link to them rather than redistributing their text or pages.
- Corrections Read the correction policy or report an error.
Provenance
- First seen 2026-08-16, via seed_library.py, special_docs shard of the Stampy snapshot.
- Work id
work:kerr-rewarding-a-hoping-b, which groups manifestations of the same intellectual work. - Record id
url:journals.aom.org/c934f11f2a, the natural key for this catalog manifestation. - 2026-08-16 seeded from the special_docs shard
- 2026-08-17 publisher identity verified, kind and venue corrected, full paper read, and implementation-ready Explained prototype prepared
Full audit data, including this record under id
url:journals.aom.org/c934f11f2a:
/library/records.jsonl.
Compact browser index:
/library/corpus.json.
Catalog method and counts:
/library/index.json.