Superalignment
Convergence Programming

The paper, in seven figures.

From a plausible first app to evidence for its next action. Explore the argument through the paper’s seven visual ideas, rebuilt as interactions you can inspect.

Kevin Daniel Pantasdo & Peter Wu · The paper is being prepared for release. These are conceptual adaptations and scripted illustrations, not measured research results.

A trajectory connects three perspectives.

Adapted from paper Figure 1

The request, the generated interpretation, and the observed behavior constrain one another. Follow a double booking from plausible proposal to a bounded, retested repair.

Human intent, as surfaced, constrains an AI interpretation. Executing two overlapping reservations reveals a double booking. The owner confirms exclusivity; an atomic reservation repair is proposed, then retested. A later successful scenario supports that requirement within the tested conditions, not universal readiness.

The same loop can end five ways.

Adapted from paper Figure 2

Discovering an unmet requirement changes what is known. Repairing it changes behavior. Explore true convergence, false convergence, plateau, regression, and premature action without confusing these changes.

Four explicit requirements define this synthetic world. Known unmet counts can rise when discovery occurs. The evaluator-only record, hidden until requested, distinguishes missing knowledge from satisfied behavior. Five constructed outcomes are possible patterns, not performance measurements.

What capability alone cannot establish.

Adapted from paper Figure 3

The paper states these limits under explicit assumptions. The four interactions below separate missing information, indistinguishable conditions, critical blockers, and the scope of continued action.

Specification debt

The initial request states opening hours and session duration. Workflow observation and owner clarification reveal preparation and cancellation rules. This supplies new information; it does not repair the generated behavior. The four requirements are a synthetic finite example.

No free readiness

Two studios give the same request, preview, and accepted booking. Observing and confirming the local preparation rule distinguishes 15 minutes from 30 minutes. The same 10:15 offer fits one rule and conflicts with the other.

Critical blockers do not average

Adding routine passing checks raises a naive average from 7/8 to 99/100 while a critical privacy failure remains. Repair and rechecking change the supported action. A changed email service reopens the evidence question.

Autonomy spends evidence

Under an illustrative per-action risk upper bound of 0.2%, a union bound reaches a chosen 1% window budget after five actions. A sixth would exceed it. Changing the service requires new evidence; the old record does not vanish. No real risk estimate or permission is provided.

Make discoveries survive the next change.

Adapted from paper Figure 4

Select a function to see what it contributes and which failure it helps address. This is one reference architecture; equivalent support can come from other tools and human workflows.

The generator proposes executable behavior. Other functions seek grounding, track tentative and confirmed requirements, retain evidence, check behavior, preserve earlier obligations, handle critical blockers, and bound continued action. No arrangement of boxes by itself guarantees readiness.

Change the question an evaluation asks.

Adapted from paper Figure 5

A completed task and justified permission are different evaluation targets. Explore those targets under stated and initially latent constraints. The map describes questions, not a ranking of benchmarks.

A qualitative map crosses stated versus initially latent constraints with task completion versus evidence for permission. Benchmarks can cover several regions. The proposed evaluation focus includes discovering hidden constraints and grounding action decisions. The map has no numerical scores or measured benchmark placements.

A repair can erase earlier progress.

Adapted from paper Figure 6

Follow the paper’s clinic example through discovery, a privacy repair that breaks duplicate handling, and a localized correction. The requirement record shows what is known, what holds, and what remains to check.

Eight local clinic requirements form a stipulated illustration: intake fields, nurse review, queue ordering, private SMS, duplicate merging, missing-insurance follow-up, access control, and audit logging. Targeted observations reveal missing requirements. Removing information to repair SMS privacy breaks duplicate handling; a localized repair preserves both. Counts are derived from discrete states, not the paper’s unreconciled numerical trace. Clinical rules are fictional workflow assumptions.

Different worlds. The same oversight problem.

Adapted from paper Figure 7

Choose a proposed task world, inspect the local requirement behind its visible request, and see how repair changes the evidence. These are illustrated evaluation designs, not released environments or measured agent runs.

The paper proposes eight illustrative worlds: clinic intake, cold-chain logistics, restaurant compliance, school pickup, support escalation, finance approval, government permit, and field-service dispatch. Each has a visible task, a local operating constraint, a discoverable scenario, and a bounded action decision. No live environment or empirical evaluation runs here.