Superalignment

The 60-second answer

Organizations often count evaluators or decision makers as independent even when all of them inherit the same model, benchmark or blind spot.

Kleinberg and Raghavan identify a correlation externality in shared algorithmic rankings. Under two stated conditions, for any accuracy of the independent human rankings there exists a slightly more accurate shared algorithm that strictly dominates the human ranking for each firm. Both firms therefore choose the algorithm. Yet the two independent rankings produce higher system welfare because their errors and discoveries are less correlated. The paper also gives a Plackett-Luce ranking family where this monoculture effect is zero.

  • The theorem compares independent lower-accuracy rankings with one shared ranking that is slightly more accurate for each user.
  • Each firm can rationally adopt the shared algorithm while the pair of firms selects lower total candidate quality.
  • The missing quantity is an independence dividend: separate errors and discoveries can improve coverage across the system.
  • The result is conditional, and the paper's Plackett-Luce family produces no monoculture welfare effect.

Written for: Technical generalists comfortable with expected value and simple game theory. Useful prerequisites: Expected value, A ranking as an ordered list, The difference between private payoff and system welfare.

The question
Can a shared ranking algorithm that is more accurate for every decision maker still make the full system worse?
What the authors did
The paper models two firms choosing between independent, lower-accuracy rankings and one shared, slightly more accurate ranking. Each firm selects one candidate from the top of its ranking. The authors compare each firm's expected payoff with system welfare, defined by the sum of the qualities of the two selected candidates, and prove an existence result under two conditions on the ranking model.
The source
Algorithmic Monoculture and Social Welfare

When private accuracy and system welfare point in opposite directions

Private incentives and system welfare under shared ranking algorithms A two by two qualitative strategy table. Under the theorem conditions, each firm has a private incentive to choose the shared, slightly more accurate algorithm, while the system has higher welfare when both use independent rankings. A button changes to the paper's Plackett-Luce countercase, where the monoculture effect is zero. The choice each firm sees The theorem's two ranking conditions hold Private dominance points to A / A, while total welfare can be higher at H / H.

Theorem conditions selected. Private incentives and system welfare can point in opposite directions.

What changes when the assumption turns
Ranking familyPrivate choiceSystem result
Paper's two theorem conditionsA can strictly dominate H for each firmThere exists a slightly more accurate A for which welfare at H / H is higher than A / A
Plackett-Luce countercaseThe best available ranking remains optimalThe paper derives no monoculture welfare effect
The theorem state shows only the paper's qualitative inequalities. Under its two ranking conditions, the shared algorithm can strictly dominate for each firm while the pair of independent rankings has higher system welfare. The switch activates the paper's Plackett-Luce countercase, where the monoculture effect is zero. No payoff values are invented. The table shows proved preference directions and equilibrium structure only. Cell positions and colors do not encode payoff magnitudes or empirical frequencies.

Walk through the argument

Each step below points to the source location that carries it. The source map follows the walkthrough for source-level checking.

Start with two objectives that look similar

Two firms each choose one candidate from the top of a ranking. A firm's private payoff is the expected quality of its own choice. System welfare is the sum of the qualities of both selected candidates. The same ranking can improve the first quantity and reduce the second.

The alternative rankings are independent but less accurate. The shared algorithm is slightly more accurate, yet both firms receive the same ordering from it. Accuracy and correlation therefore move together.

Source: Algorithmic Hiring as a Case Study, Modeling Ranking and Modeling Selection

Independent mistakes can improve system coverage

If two imperfect rankings make different mistakes, one can surface a strong candidate that the other misses. Their combined selections can cover more candidate quality than two choices driven by one ordering. The paper calls this a preference for independence.

That benefit is system-level. A firm deciding alone sees only whether the shared algorithm improves its own expected pick, not the candidate quality another firm loses when both rankings become correlated.

Source: A Preference for Independence, Algorithmic Hiring as a Case Study, Modeling Ranking and Modeling Selection

The theorem is an existence result with two conditions

Under the paper's two conditions on the ranking distribution, any given accuracy for the independent rankings admits a slightly more accurate shared algorithm. Each firm strictly prefers that algorithm, so shared adoption is privately rational.

At the same time, the independent rankings can yield strictly higher total welfare. The theorem proves that this reversal can happen. It does not estimate how often it happens in hiring or any other deployed system.

Source: Stating the Main Result, Theorem 1, Proving Theorem 1 and Proof of Theorem 1

The equilibrium ignores a correlation externality

Each firm captures its accuracy gain and pushes part of the correlation cost onto the system. Neither firm's private objective pays for the lost chance that an independent ranking would discover a different strong candidate.

This is why adding evaluators can fail to add assurance. If they share a model family, data lineage or ontology, their agreement may be one correlated signal rather than several independent checks. That application is an institutional inference, not a tested claim in the paper.

Source: Stating the Main Result, Theorem 1, A Preference for Independence

A countercase shows what the theorem does not say

In the paper's Plackett-Luce ranking family, the relevant independence dividend disappears and the monoculture effect is zero. Shared rankings are therefore not harmful by definition.

The practical audit question is narrower: does this decision process contain valuable independent errors or discoveries, and does adoption of one shared system erase them? Without evidence about that correlation structure, the theorem supplies a mechanism, not a verdict.

Source: Instantiating with Ranking Models, RUMs

Source map

These are the source locations that carry the argument. Use them to check this explanation against the original rather than trusting the summary alone.

Where the argument lives
LocusWhy it mattersSource
Algorithmic Hiring as a Case Study, Modeling Ranking and Modeling SelectionDefines two firms, candidate rankings, independent human rankings and the shared algorithmic ranking.Open source →
Stating the Main Result, Theorem 1States the two conditions and the existence of a shared algorithm that each firm prefers even when system welfare falls.Open source →
A Preference for IndependenceExplains why independent rankings can cover more high-quality candidates than correlated rankings.Open source →
Proving Theorem 1 and Proof of Theorem 1Constructs the accuracy interval where private adoption and higher welfare under independence coexist.Open source →
Instantiating with Ranking Models, RUMsShows that the Plackett-Luce family has no monoculture effect, bounding the theorem's reach.Open source →

The Assumption Switch

One result. One assumption exposed. Turn it and see what changes.

Assumption under test

The ranking distribution satisfies the paper's two conditions that create a value for independent errors and discoveries.

Held in the source
Under those conditions, each firm can prefer the same slightly more accurate algorithm even when independent rankings have higher total welfare.
Turn it
Under the paper's Plackett-Luce ranking family, the relevant independence dividend disappears and the monoculture welfare effect is zero.
What changes
The result is conditional and existential, not a universal indictment of shared models. The institutional question is whether the deployed ranking process creates correlated blind spots that its users do not bear privately.

The common misreading

The paper does not show that humans are generally better than algorithms, that algorithm sharing is always harmful or that the paper's constructed four percent welfare loss is an empirical threshold. It shows that correlation can reverse a welfare comparison under a stated model even when the shared ranking is slightly more accurate for each user.

Outside the ML frame

Institutional design

Who pays for lost independence when every actor chooses the privately better tool?

The paper describes a correlation externality. Each firm captures the private gain from a more accurate ranking but does not price the social loss from making its choice more correlated with another firm's choice. For AI assurance, this suggests that adding more evaluators is not enough when they share a model, training lineage, benchmark or ontology. That application is our interpretation, not a result tested in the paper.

Where the result stops

This is a conditional existence theorem in a stylized ranking model, not evidence that shared algorithms generally reduce welfare. Welfare is the sum of the qualities of the two selected candidates, not a full account of fairness, diversity or downstream outcomes. The result requires two conditions on the ranking distribution. The paper's Plackett-Luce countercase has no monoculture effect, and its numerical examples are constructions rather than field estimates.

What remains open

  • Which empirical decision systems have enough shared error to create a material independence dividend?
  • How should procurement or audit rules reward error diversity without preserving avoidable inaccuracy?
  • What changes when firms train related but nonidentical models rather than adopting one shared ranking?
  • Can system-level evaluation measure correlation costs before a deployment concentrates decisions?

What this bears on

Superalignment maintains a public register of the claims it makes and the evidence that would change its mind. The entries below connect this source to the exact public claims it bears on.

  • C4. Behavioral evaluation cannot carry a deployment decision alone. This record bears on it, indirectly. The theorem shows that evaluating a ranking only by each user's accuracy can miss a system-level welfare loss caused by correlated decisions. It does not by itself establish a deployment rule. See the claim and what would change our mind →

Source audit and review status

What was checked, and against what
FieldResultCheckedByAgainst
titleexact2026-08-17codex-primary-source-reviewsource
authorsexact2026-08-17codex-primary-source-reviewsource
dateexact2026-08-17codex-primary-source-reviewsource
venueexact2026-08-17codex-primary-source-reviewsource
full_textexact2026-08-17codex-primary-source-reviewsource
  • Review status prototype.
  • Program collection related prototype; theory; wave 1, release slot unassigned.
  • Source access public full text, HTML. Open the reading copy →
  • Provider-authored safety claim no.
  • Explained by Superalignment Research.
  • Reviewed by No named human reviewer yet.
  • Explainer dates created 2026-08-17; updated 2026-08-17.
  • AI assistance AI assisted with source discovery, full-text extraction, theorem-scope checking, first-pass prose and implementation. The page remains a prototype until a named human review is recorded.
  • Rights and access The author manuscript is available on arXiv and the published full text is available through PubMed Central.
  • Corrections Read the correction policy or report an error.

Provenance

  • First seen 2026-08-17, via cross-disciplinary citation-closure and anti-monoculture review.
  • Work id work:algorithmic-monoculture, which groups manifestations of the same intellectual work.
  • Record id arxiv:2101.05853, the natural key for this catalog manifestation.
  • 2026-08-17 full paper read and added as an Explained v2 prototype with a conditional Assumption Switch

Full audit data, including this record under id arxiv:2101.05853: /library/records.jsonl. Compact browser index: /library/corpus.json. Catalog method and counts: /library/index.json.