Superalignment

The 60-second answer

AI benchmarks increasingly decide funding, access, releases, and reputations. Campbell offers a way to ask how those decisions alter the evidence itself before treating a leaderboard as a neutral report of capability or safety.

Campbell argues that evaluation must be built for a political and administrative world, not an ideal laboratory. Quantitative measures can omit context, qualitative accounts can be selectively persuasive, implementations drift, and records change with the institutions producing them. In the best-known section, he proposes a conditional pressure mechanism: the more an indicator is used for consequential decisions, the more incentives arise to corrupt the measure and distort the activity being measured. He presents that claim as pessimistic, largely anecdotal, and especially grounded in the U.S. setting of his examples.

  • Consequential use can send pressure backward into both an indicator and the activity that produces it.
  • Campbell presents a conditional institutional risk supported mainly by anecdotes, not a universal theorem about every metric.
  • Independent criticism, contextual evidence, and auditable designs matter more than simply adding another score.

Written for: Technical generalists who use benchmarks, dashboards, audits, or performance metrics. Useful prerequisites: Basic causal reasoning, The distinction between a measure and the goal it represents.

The question
How can social programs be evaluated when measurement, politics, implementation, and the stakes attached to indicators all change what gets observed?
What the authors did
Campbell develops a methodological argument from evaluation design, social-science examples, and predominantly anecdotal cases. He compares quantitative and qualitative evidence, reviews time-series, experimental, regression, and quasi-experimental designs, then examines how decision use can pressure an indicator and the process it represents.
The source
Assessing the Impact of Planned Social Change

When an indicator becomes a decision rule

How decision pressure can change an indicator and its underlying process A service process produces records, an indicator, and a decision. In descriptive mode the decision mainly reads the indicator. In consequential mode feedback paths from the decision reach record production and service behavior, exposing two ways score and objective may diverge. The ordinary reporting path Service process the objective in practice Records what gets counted Indicator reported value Decision low-stakes use When the reported value controls stakes Pressure on what enters the record Pressure on how the service itself operates Record path Selection, reclassification, omission, or presentation can change the reported score. Process path Work can shift toward what raises the score while leaving the underlying objective behind. Descriptive mode: feedback pressure is possible but not strongly activated by the decision rule.

Descriptive indicator selected. The figure keeps the decision feedback paths faint.

Mechanism and evidentiary boundary
AssumptionMechanism exposedWhat the source supports
Low-stakes descriptive useThe indicator mainly reports on the process.Measurement can still omit context or contain error.
Consequential decision usePressure can feed back into records and operations.A corruption risk supported mainly by institutional examples, not inevitability or a measured rate.
The switch changes a low-stakes descriptive indicator into a consequential decision target. In Campbell's account, decision pressure can feed back into both record production and the service process. The diagram presents a possible mechanism, not a measured frequency or an inevitable result. Arrow thickness, position, and color do not encode measured effect size, prevalence, or certainty.

Walk through the argument

Each step below points to the source location that carries it. The source map follows the walkthrough for source-level checking.

A dashboard is also an intervention

A dashboard looks like a window onto an organization until pay, status, or permission depends on what it shows. Then the people and processes behind the window have reasons to alter the view, sometimes by improving the work and sometimes by improving only what is visible.

Campbell places this problem inside a much larger account of program evaluation. Measures are produced through administrative routines, political choices, implementation histories, and selective records. Evaluation therefore studies a changing social system, not a fixed object waiting to be counted.

Source: Authorized reprint pages 3 to 6, introduction, Authorized reprint pages 7 to 10, quantitative and qualitative knowing

Separate the score from the process

A school can raise a reported score by teaching more effectively, coaching only tested material, excluding inconvenient cases, or changing what gets recorded. Those moves do not have the same relation to the educational objective even when the dashboard moves in the same direction.

Campbell's mechanism has two paths. The indicator itself can become less trustworthy, and the social process can be distorted to maximize what the indicator rewards. The distinction matters because a cleaner database does not automatically repair a damaged service, while process reform does not guarantee an honest record.

Source: Authorized reprint pages 34 to 36, Corrupting Effect of Quantitative Indicators

Read the law as conditional pressure

A speedometer does not corrupt driving merely by displaying speed. The institutional switch occurs when one number becomes a consequential decision rule and actors can influence either the number or the activity beneath it. Greater stakes create greater pressure, not a guarantee of successful gaming.

Campbell calls his formulations pessimistic, anchors them especially in U.S. examples, and describes the evidence as predominantly anecdotal. A careful explainer should preserve those qualifiers. The paper offers a mechanism and warning signs, not a measured corruption rate or a theorem without exceptions.

Source: Authorized reprint pages 3 to 6, introduction, Authorized reprint pages 34 to 36, Corrupting Effect of Quantitative Indicators

Use an evidence portfolio

A single photograph can be precise and still omit everything outside its frame. Campbell treats quantitative measures similarly: abstraction is useful, but a result can conflict with participant observation, narrative history, or implementation detail that reveals what the number left out.

His answer is not to replace numbers with stories. Qualitative accounts also admit selective attention and persuasion. The practical response is criticism across methods, explicit alternative explanations, replicated administrative experiments where possible, and room for minority reports that challenge the official account.

Source: Authorized reprint pages 7 to 10, quantitative and qualitative knowing, Authorized reprint pages 14 to 32, evaluation designs

More metrics are not an automatic cure

Adding gauges to a cockpit helps only if they expose different failure paths and cannot all be manipulated through the same lever. A bundle of correlated indicators may create the appearance of triangulation while preserving one shared blind spot.

Campbell considers multiple indicators and outside watchdogs, but he does not declare either sufficient. Independence, access to underlying records, and the ability to investigate the process remain design questions. The strongest lesson is to make criticism operational rather than to search for an ungameable number.

Source: Authorized reprint pages 36 to 37, watchdogs, multiple indicators, and summary

Carry the mechanism into AI evaluation

An AI benchmark begins as a probe. Once model access, release approval, investment, or public standing depends on it, developers can optimize training data, prompts, exclusions, and reporting around the probe. Some adaptation is real progress, while some narrows the distance between the test and the training target.

Campbell does not establish that current AI evaluations are corrupt. He supplies a disciplined question: how does the decision use alter data production and behavior, and what independent evidence could reveal divergence? That question turns Goodhart-style rhetoric into an inspectable institutional mechanism.

Source: Authorized reprint pages 34 to 36, Corrupting Effect of Quantitative Indicators, Authorized reprint pages 36 to 37, watchdogs, multiple indicators, and summary

Source map

These are the source locations that carry the argument. Use them to check this explanation against the original rather than trusting the summary alone.

Where the argument lives
LocusWhy it mattersSource
Authorized reprint pages 3 to 6, introductionFrames evaluation as both methodological and political, and warns that the argument may not be universal across social systems.Open source →
Authorized reprint pages 7 to 10, quantitative and qualitative knowingExplains why numerical abstractions can conflict with contextual knowledge and why neither quantitative nor qualitative evidence is infallible.Open source →
Authorized reprint pages 14 to 32, evaluation designsReviews time-series, randomized, regression, and quasi-experimental approaches together with threats from changing records, attrition, timing, and alternative explanations.Open source →
Authorized reprint pages 34 to 36, Corrupting Effect of Quantitative IndicatorsStates the conditional indicator-pressure claim, labels the supporting evidence predominantly anecdotal, and works through institutional examples.Open source →
Authorized reprint pages 36 to 37, watchdogs, multiple indicators, and summaryConsiders outside evaluators and multiple indicators, preserves doubts about easy fixes, and notes that evaluation success stories exist.Open source →

The Assumption Switch

One result. One assumption exposed. Turn it and see what changes.

Assumption under test

The indicator is mainly descriptive and carries limited consequences for the people producing it.

Held in the source
Records can still be incomplete or biased, but the indicator does not strongly reshape the work whose performance it is meant to summarize.
Turn it
Budgets, status, sanctions, or rewards become tightly coupled to the reported value while the record remains open to strategic influence.
What changes
Pressure now runs backward from the decision rule into data production and operational behavior. Score and objective can diverge, but the source supports a risk mechanism rather than an inevitable universal law.

The common misreading

Campbell's law is often compressed into the claim that any target metric must be corrupted. Campbell instead describes a pressure that grows with consequential decision use and a risk of corrupting both indicator and process. He also calls the evidence anecdotal, limits the setting, discusses successful evaluations, and does not present multiple metrics as an automatic cure.

Outside the ML frame

Machine learning evaluation

What changes when a benchmark stops describing a model and starts deciding access, funding, or deployment?

A benchmark can become part of the training and governance environment that model developers adapt to. Campbell's mechanism suggests examining who can influence the data, protocol, exclusions, and presentation once a score carries stakes. This is an institutional interpretation for AI evaluation, not evidence that every benchmark is already gamed.

Where the result stops

The famous indicator claim is not estimated from a defined sample, and Campbell describes its evidence as predominantly anecdotal. He warns that the politico-methodological argument may not generalize across all social and political systems. The full text checked here is an authorized 2011 reprint of the December 1976 Dartmouth occasional paper. The publisher says the canonical 1979 journal article contains minor revisions and additions to an earlier version, so page loci below use the openly readable reprint and should not be silently converted to 1979 pagination.

What remains open

  • Which features of an indicator and its surrounding institution predict corruption pressure before visible gaming appears?
  • When do independent audits and qualitative evidence detect score-objective divergence without creating another targetable score?
  • How can evaluators distinguish legitimate process improvement from adaptation that preserves the number while degrading the objective?
  • Which parts of Campbell's argument travel across political systems, professions, and machine learning benchmark ecosystems?

What this bears on

Superalignment maintains a public register of the claims it makes and the evidence that would change its mind. The entries below connect this source to the exact public claims it bears on.

  • C4. Behavioral evaluation cannot carry a deployment decision alone. This record bears on it, suggestively. Campbell shows why a behavioral score can be altered by consequential use, administrative records, and implementation context. His evidence is a methodological warning, not direct AI deployment evidence. See the claim and what would change our mind →

Source audit and review status

What was checked, and against what
FieldResultCheckedByAgainst
titleexact2026-08-17codex-primary-source-reviewsource
authorsexact2026-08-17codex-primary-source-reviewsource
dateexact2026-08-17codex-primary-source-reviewsource
venueexact2026-08-17codex-primary-source-reviewsource
full_textminor variant2026-08-17codex-primary-source-reviewsource
  • Review status prototype.
  • Program collection seminal v1; theory; wave 2, release slot unassigned.
  • Source access public full text, PDF. Open the reading copy →
  • Provider-authored safety claim no.
  • Explained by Superalignment Research.
  • Reviewed by No named human reviewer yet.
  • Explainer dates created 2026-08-17; updated 2026-08-17.
  • AI assistance AI assisted with primary-source retrieval, full-text extraction, manifestation checking, locus mapping, first-pass prose, figure design, and implementation. The page remains a prototype until a named human review is recorded.
  • Rights and access The canonical 1979 journal article is publisher-controlled. The complete December 1976 Dartmouth occasional paper is available as a 2011 reprint explicitly published with permission, but it carries no open-content license. Link to the authorized reprint rather than redistributing its text or pages.
  • Corrections Read the correction policy or report an error.

Provenance

  • First seen 2026-08-17, via cross-disciplinary seminal-source survey and full-source review.
  • Work id work:campbell-assessing-planned-social-change, which groups manifestations of the same intellectual work.
  • Record id doi:10.1016/0149-7189(79)90048-x, the natural key for this catalog manifestation.
  • 2026-08-17 full authorized source read and implementation-ready Explained prototype prepared with scope and manifestation caveats

Full audit data, including this record under id doi:10.1016/0149-7189(79)90048-x: /library/records.jsonl. Compact browser index: /library/corpus.json. Catalog method and counts: /library/index.json.