Superalignment

The 60-second answer

AI oversight and institutional reporting often rely on informed parties describing states that outsiders cannot inspect directly. The paper shows why a channel can sound candid and remain systematically too coarse for the decision at hand.

Costless communication need not be either fully revealing or useless. When sender and receiver preferences are partly aligned, equilibrium messages can identify intervals of the hidden state while withholding distinctions inside each interval. Greater preference divergence generally supports fewer informative intervals under the paper's assumptions. The receiver's inability to commit matters because each message must induce an action the receiver prefers after hearing it.

  • Costless strategic messages can reveal broad intervals of a hidden state without revealing distinctions inside each interval.
  • Preference divergence can coarsen the finest sustainable partition because the receiver chooses its preferred action after each message.
  • The result depends on a spare one-sender model and does not show that all biased communication is uninformative or deceptive.

Written for: Technical generalists working on oversight, incentives, governance, or information systems. Useful prerequisites: Bayesian updating, Nash equilibrium, The difference between a message and verifiable evidence.

The question
How much information can a better-informed sender transmit when the receiver's preferred action is close to, but not the same as, the sender's?
What the authors did
Crawford and Sobel build a Bayesian cheap-talk game. A privately informed sender sends a costless message, an uninformed receiver chooses an action without precommitment, and their ideal actions differ. They characterize equilibria as finite partitions of the sender's information and derive comparative results under additional assumptions, including a quadratic-uniform example.
The source
Strategic Information Transmission

How preference bias coarsens a message

Preference divergence and the granularity of cheap talk A hidden-state interval is divided into four, two, or one message regions. Buttons change the sender and receiver preference divergence. More divergence makes the displayed partition coarser. One hidden state, several possible message bins The sender sees the position. The receiver hears only the bin label. Hidden state low high message A message B message C message D one receiver action one receiver action one receiver action one receiver action Nearby preferences: several coarse but informative messages can be sustained. Receiver learns: one of four illustrative state regions. Schematic reconstruction of the partition mechanism, not calculated equilibrium boundaries.

Nearby preferences selected. Four illustrative message bins are visible.

Assumption switch and information granularity
Preference relationDisplayed partitionReceiver learnsScope
NearbyFour illustrative binsA relatively fine region, not the exact stateQualitative partition mechanism
More divergentTwo illustrative binsA broader regionCoarsening in the quadratic-uniform example
Large divergenceOne binNothing from the messageThe only equilibrium beyond the example's stated threshold
The switch reconstructs the paper's partition mechanism. Nearby preferences can sustain several messages for different state intervals. Greater divergence supports fewer distinctions in the quadratic-uniform example, and sufficiently large bias leaves only uninformative communication there. This is a qualitative schematic, not a plot of equilibrium values or an empirical effect. Interval widths, number of displayed regions, colors, and spacing are illustrative. They do not encode the paper's calculated equilibrium boundaries, welfare, frequency, or empirical magnitude.

Walk through the argument

Each step below points to the source location that carries it. The source map follows the walkthrough for source-level checking.

Give one side the hidden state

The sender observes a state that the receiver cannot see. The sender then chooses a message, and the receiver chooses an action. Messages have no direct cost and do not carry proof.

Both parties care about the action and the state, but their favorite actions do not coincide. The receiver also cannot promise in advance how it will react. Its response must be optimal after interpreting the message.

Source: Journal pages 1431 to 1433, abstract and Section 1, Journal pages 1433 to 1435, Section 2 and Equations 1 to 2

Replace full revelation with a partition

An informative equilibrium groups neighboring states into intervals. The same message is sent throughout an interval, and the receiver chooses one action for that whole region. The message is informative because it identifies a region, but it is coarse because it hides position inside the region.

At every boundary, the sender must be indifferent between the actions induced by adjacent messages. Those incentive constraints determine which partitions can persist.

Source: Journal pages 1435 to 1440, Section 3, Lemmas 1 to 2, Theorem 1, and Equations 3 to 19

Increase the preference bias

In the quadratic-uniform example, the sender always wants an action shifted by a fixed amount from the receiver's ideal. As that shift grows, fewer interval boundaries satisfy the sender's incentive constraints.

The figure's assumption switch reconstructs this mechanism qualitatively. It does not plot the paper's equilibrium boundaries or estimate an effect in real communication systems.

Source: Journal pages 1440 to 1444, Section 4, Equations 20 to 25, and Figure 1, Journal pages 1444 to 1450, Section 5 and Theorems 2 to 5

Keep multiple equilibria in view

For a given bias, the model can support partitions with different numbers of intervals, including uninformative communication. The paper's example gives both parties an ex ante preference for the most informative available equilibrium, but the game does not itself select it.

A receiver therefore cannot infer full informativeness merely because a more revealing equilibrium exists. Coordination, conventions, and institutional design still matter.

Source: Journal pages 1435 to 1440, Section 3, Lemmas 1 to 2, Theorem 1, and Equations 3 to 19, Journal pages 1440 to 1444, Section 4, Equations 20 to 25, and Figure 1, Journal pages 1444 to 1450, Section 5 and Theorems 2 to 5

Translate the model into design questions

The model suggests three levers for oversight: reduce preference conflict, let the receiver commit to responses, or add evidence that is not under the sender's control. Each changes a premise rather than asking rhetoric alone to solve strategic disclosure.

It does not establish how an AI system communicates. A real system may have many messages, repeated interactions, uncertain preferences, external checks, and limitations unrelated to strategy.

Source: Journal pages 1433 to 1435, Section 2 and Equations 1 to 2, Journal page 1450, conclusion

Read the result as theory

Theorems show what follows inside the specified game. They do not measure how often strategic coarsening occurs or identify the model that best fits a particular organization.

The useful empirical question is whether changing incentives, commitment, or verification changes the granularity of information that reaches a decision maker.

Source: Journal pages 1444 to 1450, Section 5 and Theorems 2 to 5, Journal page 1450, conclusion

Source map

These are the source locations that carry the argument. Use them to check this explanation against the original rather than trusting the summary alone.

Where the argument lives
LocusWhy it mattersSource
Journal pages 1431 to 1433, abstract and Section 1Introduces strategic information transmission, partial revelation, applications, and the contrast with models driven by signaling costs.Open source →
Journal pages 1433 to 1435, Section 2 and Equations 1 to 2Defines the sender, receiver, private state, costless message, receiver action, Bayesian Nash equilibrium, and absence of receiver precommitment.Open source →
Journal pages 1435 to 1440, Section 3, Lemmas 1 to 2, Theorem 1, and Equations 3 to 19Shows why informative equilibria take a finite partition form and establishes an upper bound on the number of induced actions under the stated assumptions.Open source →
Journal pages 1440 to 1444, Section 4, Equations 20 to 25, and Figure 1Works through quadratic preferences with a uniform state, links greater bias to coarser partitions, and compares the multiple equilibria in that example.Open source →
Journal pages 1444 to 1450, Section 5 and Theorems 2 to 5States sufficient conditions for comparative results about bias, the number of partition elements, and sender and receiver welfare.Open source →
Journal page 1450, conclusionMarks unresolved issues around lying, credibility, equilibrium selection, verification, and richer communication settings.Open source →

The Assumption Switch

One result. One assumption exposed. Turn it and see what changes.

Assumption under test

The sender and receiver prefer nearby actions for each hidden state.

Held in the source
Several state intervals can induce distinct receiver actions that both parties are willing to sustain in equilibrium.
Turn it
The sender's preferred action moves farther from the receiver's for every state while messages remain costless and unverifiable.
What changes
The finest sustainable partition becomes coarser in the paper's example, and beyond its stated threshold only an uninformative equilibrium remains.

The common misreading

Cheap talk is often paraphrased as either honest revelation or meaningless babble. Crawford and Sobel's central result is the middle case: strategic messages can be informative only up to a partition. It is also wrong to read the quadratic example's bias threshold as a general empirical cutoff for organizations or AI systems.

Outside the ML frame

AI oversight and organizational communication

What can a monitor learn from an informed system whose preferred intervention differs from the monitor's?

The model suggests that an oversight channel can carry genuine but strategically coarse information. It directs attention to preference divergence, receiver commitment, and independent verification instead of treating fluent disclosure as full revelation. This is a theoretical transfer, not evidence about model internals or AI behavior.

Where the result stops

The model has one sender, one receiver, common knowledge of preferences, a scalar state and action, costless messages, no exogenous reputation or verification, and a receiver who cannot commit before the message. The strongest characterization relies on assumptions including single-peaked preferences and monotonicity. The paper leaves equilibrium selection open and says its framework gives an incomplete operational account of lying and credibility.

What remains open

  • How do verifiable evidence, repeated interaction, reputation, or penalties for false statements change the sustainable information partition?
  • What happens when several senders have correlated information and different conflicts with the receiver?
  • Can a receiver design commitment or audit mechanisms that recover finer information without making honest participation unattractive?
  • How can empirical studies distinguish strategic coarsening from limited knowledge, ambiguity, or ordinary compression?

What this bears on

Superalignment maintains a public register of the claims it makes and the evidence that would change its mind. The entries below connect this source to the exact public claims it bears on.

  • C5. The binding constraint on overseeing stronger workers is conditions, not capability. This record bears on it, indirectly. The model shows that oversight quality can depend on preference conflict, receiver commitment, and verification conditions even when the informed party can communicate. It does not study stronger AI workers or identify the binding practical constraint. See the claim and what would change our mind →

Source audit and review status

What was checked, and against what
FieldResultCheckedByAgainst
titleexact2026-08-17codex-primary-source-reviewsource
authorsexact2026-08-17codex-primary-source-reviewsource
dateexact2026-08-17codex-primary-source-reviewsource
venueexact2026-08-17codex-primary-source-reviewsource
full_textexact2026-08-17codex-primary-source-reviewsource
  • Review status prototype.
  • Program collection seminal v1; theory; wave 4, release slot unassigned.
  • Source access public full text, PDF. Open the reading copy →
  • Provider-authored safety claim no.
  • Explained by Superalignment Research.
  • Reviewed by No named human reviewer yet.
  • Explainer dates created 2026-08-17; updated 2026-08-17.
  • AI assistance AI assisted with primary-source retrieval, full-text extraction, manifestation checking, locus mapping, first-pass prose, figure design, and implementation. The page remains a prototype until a named human review is recorded.
  • Rights and access The authors' UC San Diego copy is public for reading. Vincent Crawford's publication page permits downloading, printing, and reproduction for personal or classroom use, not commercial redistribution. Link to the source rather than republishing its pages.
  • Corrections Read the correction policy or report an error.

Provenance

  • First seen 2026-08-17, via cross-disciplinary seminal-source survey and full-source review.
  • Work id work:crawford-sobel-strategic-information-transmission, which groups manifestations of the same intellectual work.
  • Record id doi:10.2307/1913390, the natural key for this catalog manifestation.
  • 2026-08-17 full source read and implementation-ready Explained prototype prepared with equilibrium-selection and scope caveats

Full audit data, including this record under id doi:10.2307/1913390: /library/records.jsonl. Compact browser index: /library/corpus.json. Catalog method and counts: /library/index.json.