Superalignment

The 60-second answer

AI organizations often reward what they can count while relying on the same people for hard-to-measure safety, judgment, and maintenance. This paper explains why a better metric does not by itself solve the allocation problem.

The paper adds attention allocation to incentive design. When a measured task competes with an important unmeasured task, paying harder for the visible output can pull effort away from what the principal also values. Under the paper's strongest substitute-effort cases, a fixed wage can be optimal even when the measured output is accurate and the agent responds to incentives. Ownership, restrictions, and job boundaries then become part of the same incentive system because they change the opportunity cost of each task.

  • An incentive directs attention among tasks as well as increasing total effort, so its effect depends on the agent's whole job.
  • When an important unmeasured task competes with a measured task, the optimal visible-task incentive can be muted or even zero.
  • Compensation, asset ownership, activity restrictions, and job boundaries are connected instruments rather than separate design choices.

Written for: Technical generalists comfortable with optimization, incentives, and model assumptions. Useful prerequisites: Principal and agent, Opportunity cost, Noisy performance measurement.

The question
When can a stronger incentive on a useful performance measure make the principal worse off?
What the authors did
Holmstrom and Milgrom analyze a formal linear principal-agent model in which one agent allocates effort across several tasks. Compensation can depend on noisy performance signals, the agent is risk averse, and task costs can interact. Specialized models derive propositions about missing incentive clauses, fixed wages, asset ownership, limits on outside activities, unity of responsibility, and grouping tasks by measurability.
The source
Multitask Principal-Agent Analyses: Incentive Contracts, Asset Ownership, and Job Design

Do the measured and unmeasured tasks compete?

How task substitution changes a measured-task incentive One agent performs a measured task and an unmeasured task. A control compares separable costs with competing attention. When attention competes, the visible reward pulls effort toward the measured task and away from the unmeasured task. One agent, two valuable tasks Agent allocates effort across the whole job Measured task visible output signal reward can depend on it Unmeasured task quality, care, judgment still valuable to principal Visible reward pay tied to signal useful in isolation Illustrative allocation of attention measured task can rise on its own unmeasured task need not fall higher opportunity cost Separable tasks: the visible reward need not displace unmeasured work. Bar lengths are qualitative. They are not a solved allocation or reported data.

Separable task costs selected. The measured-task reward does not automatically pull effort from the other task.

Task interaction changes the incentive problem
AssumptionEffect of stronger measured-task rewardDesign implication in the model
Task costs are separableMeasured effort can rise without directly raising the cost of unmeasured effort.Set the incentive mainly from the measured task's value, noise, and responsiveness.
Tasks compete for one attention poolMeasured effort becomes more attractive and unmeasured effort more costly.Mute the reward, protect the other task, or redesign the job when displacement is costly.
The switch changes task interaction. With separable effort costs, a reward for the measured task need not pull effort from the unmeasured task. With substitutable attention, the same reward raises the opportunity cost of unmeasured work, which can justify a weaker incentive. The schematic does not solve for an effect size. Arrow thickness, bar length, box size, color, and position do not encode an optimal contract, empirical prevalence, or quantitative effect size.

Walk through the argument

Each step below points to the source location that carries it. The source map follows the walkthrough for source-level checking.

An incentive directs attention

A one-task model asks how much effort a reward buys. A multitask model also asks where that effort comes from. A teacher can spend time on tested basics or harder-to-measure reasoning. A production worker can increase output or protect quality and equipment.

The same commission can therefore raise measured output and lower another valuable activity. Its net value depends on the full task portfolio, not only on whether the measured task is useful.

Source: Journal pages 24 to 29, Introduction

Read the linear contract as a control input

The agent chooses a vector of efforts. Those efforts create benefits for the principal and noisy performance signals for a linear wage. Stronger coefficients motivate work but expose a risk-averse agent to more noise, creating the standard incentive-versus-risk tradeoff.

The multitask step adds cross-effects in the agent's cost. When tasks are separable, incentives can be chosen more independently. When they are substitutes, a reward on one task raises the opportunity cost of the other.

Source: Journal pages 29 to 33, Section 2, Equations 1 to 7

Turn on competition for attention

Suppose output is measured accurately but quality is not measured. If output and quality draw on one pool of attention, an output reward makes quality more expensive for the agent to supply. The missing quality clause now changes the right output incentive.

In the paper's strongest case, the unmeasured task is essential and effort is perfectly substitutable. Proposition 1 then makes a fixed wage optimal even for a risk-neutral agent. The result follows from protecting allocation, not from claiming workers ignore incentives.

Source: Journal pages 29 to 33, Section 2, Equations 1 to 7, Journal pages 33 to 35, Sections 3.1 and 3.2, Proposition 1

Move beyond the compensation lever

Who owns an asset changes which hard-to-measure return the agent already internalizes. Under the model, an employee whose firm owns the asset receives muted production incentives to avoid neglecting asset value, while an owner-contractor can receive stronger production incentives.

Restrictions work through the same opportunity-cost channel. When performance is hard to measure and direct rewards are weak, limiting competing outside activities can preserve attention for the principal's task. Stronger measured responsibility can support more discretion.

Source: Journal pages 35 to 38, Section 3.3, Proposition 2, Journal pages 38 to 43, Section 4, Figure 1, Propositions 3 and 4

Use job boundaries to separate conflicts

The two-agent model derives sole responsibility for each small task and groups the hardest-to-measure tasks in one job, with easier-to-measure tasks in another. Separation lets strong incentives reach measurable work without pulling the same person away from unmeasured work.

That result is suggestive rather than a universal organization chart. Real tasks can be large, inseparable, correlated, complementary, and learned through rotation. The authors list these omissions and describe the model as a first pass.

Source: Journal pages 44 to 50, Section 5, Propositions 5 to 7 and Caveats

Audit the incentive system as a whole

Before increasing a benchmark reward, list the other tasks the same people perform, how effort moves among them, what signals exist, who owns the long-run result, which activities are restricted, and whether the job can be redesigned.

The paper's durable contribution is this systems view. A local incentive that looks efficient in isolation can be globally costly. The exact remedy still depends on assumptions that should be tested rather than inherited from the model.

Source: Journal pages 50 to 52, Conclusion and references, Journal pages 44 to 50, Section 5, Propositions 5 to 7 and Caveats

Source map

These are the source locations that carry the argument. Use them to check this explanation against the original rather than trusting the summary alone.

Where the argument lives
LocusWhy it mattersSource
Journal pages 24 to 29, IntroductionMotivates multitask incentives through teaching, production, asset care, outside activities, and job design, and summarizes the paper's main conditional results.Open source →
Journal pages 29 to 33, Section 2, Equations 1 to 7Defines the linear principal-agent model, risk term, incentive constraints, task interactions, and the benchmark contrast between separable and substitute activities.Open source →
Journal pages 33 to 35, Sections 3.1 and 3.2, Proposition 1Derives the fixed-wage result when an important unmeasured activity competes for perfectly substitutable attention with a measured activity.Open source →
Journal pages 35 to 38, Section 3.3, Proposition 2Connects muted employee incentives and stronger contractor incentives to who owns hard-to-measure asset returns and to measurement and risk parameters.Open source →
Journal pages 38 to 43, Section 4, Figure 1, Propositions 3 and 4Shows how restrictions on outside activities can substitute for performance incentives and predicts more discretion when measured responsibility is stronger.Open source →
Journal pages 44 to 50, Section 5, Propositions 5 to 7 and CaveatsDerives sole responsibility and grouping by measurability in a simplified two-agent model, then states the assumptions and omitted effects that limit those conclusions.Open source →
Journal pages 50 to 52, Conclusion and referencesStates the system-level lesson that compensation, ownership, restrictions, and job design must be analyzed together when performance measures are incomplete.Open source →

The Assumption Switch

One result. One assumption exposed. Turn it and see what changes.

Assumption under test

The measured and unmeasured tasks draw on effort that can be shifted from one task to the other.

Held in the source
If task costs are separable, the incentive for the measured task can be set largely from that task's value, noise, and responsiveness.
Turn it
If the tasks compete for one pool of attention, raising the measured task's reward also raises the opportunity cost of the unmeasured task.
What changes
The optimal measured-task incentive can weaken or fall to zero even when that task is valuable and its performance signal is accurate. The conclusion follows from substitution, not from measurement imperfection alone.

The common misreading

The paper is sometimes compressed into the claim that incentives are bad whenever metrics are incomplete. Its result is comparative and conditional. Incentives can be strong when relevant performance is well measured or tasks can be separated. Muted incentives become attractive when rewarding a measured task raises the opportunity cost of another valuable task that cannot be measured or protected well.

Outside the ML frame

AI evaluation and research organization

What happens when benchmark progress and unmeasured safety work compete for the same researchers?

A benchmark incentive can improve the measured result while redirecting attention from threat modeling, documentation, negative results, or maintenance that is harder to score. The model suggests changing compensation, ownership, restrictions, or job design as a system. This is a theoretical transfer, not evidence that a particular AI benchmark has caused effort substitution.

Where the result stops

This is a theory paper, not an empirical estimate. The core model uses linear performance pay, exponential utility, normal noise, and a risk-neutral principal. Several results specialize to effort that is perfectly substitutable across tasks. The job-design section assumes small tasks, flexible grouping, identical agents, and independent measurement errors; the authors call it a first pass and list omitted task size, correlation, complementarity, and rotation effects. Examples illustrate the mechanism but do not validate it.

What remains open

  • How can an organization estimate whether two tasks are substitutes, complements, or largely separable before changing incentives?
  • Which empirical designs can identify attention reallocation rather than only changes in measured output?
  • When does separating measured and unmeasured tasks improve incentives, and when does it destroy useful integration or shared context?
  • How do repeated interaction, professional norms, intrinsic motivation, and nonlinear rewards change the fixed-wage result?
  • Which ownership or authority changes protect hard-to-measure safety work without creating rigid bureaucracy?

What this bears on

Superalignment maintains a public register of the claims it makes and the evidence that would change its mind. The entries below connect this source to the exact public claims it bears on.

  • C4. Behavioral evaluation cannot carry a deployment decision alone. This record bears on it, suggestively. The model shows that a valid behavioral measure can still redirect effort away from unmeasured objectives, so a score cannot be evaluated without the surrounding task and incentive system. It is theory, not AI deployment evidence. See the claim and what would change our mind →

Source audit and review status

What was checked, and against what
FieldResultCheckedByAgainst
titleminor variant2026-08-17codex-primary-source-reviewsource
authorsexact2026-08-17codex-primary-source-reviewsource
dateexact2026-08-17codex-primary-source-reviewsource
venueexact2026-08-17codex-primary-source-reviewsource
full_textexact2026-08-17codex-primary-source-reviewsource
  • Review status prototype.
  • Program collection seminal v1; theory; wave 3, release slot unassigned.
  • Source access public full text, PDF. Open the reading copy →
  • Provider-authored safety claim no.
  • Explained by Superalignment Research.
  • Reviewed by No named human reviewer yet.
  • Explainer dates created 2026-08-17; updated 2026-08-17.
  • AI assistance AI assisted with primary-source retrieval, full-text extraction, equation and proposition mapping, manifestation checking, first-pass prose, figure design, and implementation. The page remains a prototype until a named human review is recorded.
  • Rights and access Oxford University Press controls the version of record and currently requires access on the article page. Paul Milgrom publicly hosts a complete author copy from his Stanford site, but it carries no open-content license. Link to the source rather than redistributing its text or pages.
  • Corrections Read the correction policy or report an error.

Provenance

  • First seen 2026-08-17, via cross-disciplinary seminal-source survey and full-source review.
  • Work id work:holmstrom-milgrom-multitask-principal-agent, which groups manifestations of the same intellectual work.
  • Record id doi:10.1093/jleo/7.special_issue.24, the natural key for this catalog manifestation.
  • 2026-08-17 full author copy read, equations and propositions mapped, and implementation-ready Explained prototype prepared with model-assumption caveats

Full audit data, including this record under id doi:10.1093/jleo/7.special_issue.24: /library/records.jsonl. Compact browser index: /library/corpus.json. Catalog method and counts: /library/index.json.