Superalignment

The 60-second answer

AI safety work depends on people surfacing weak signals, failed tests, and uncomfortable disagreement before the evidence is polished for a decision. A technically strong team can still lose those signals if disclosure is socially costly.

Edmondson defines team psychological safety as a shared belief that a team is safe for interpersonal risk taking. Across 51 teams, psychological safety was associated with reported learning behaviors such as asking for help, discussing errors, seeking feedback, testing assumptions, and experimenting. Learning behavior was associated with observer-rated performance and statistically mediated the safety-performance relationship in the paper's analysis. Context support and leader coaching were associated with safety, while team efficacy added less once safety was considered.

  • Psychological safety is a shared expectation about the interpersonal cost of speaking up, not a synonym for comfort or agreement.
  • In the 51-team field study, safety was associated with learning behaviors that bring hidden errors, questions, and feedback into the team.
  • The evidence is multimethod but cross-sectional and drawn from one company, so the causal and general claims remain provisional.

Written for: Technical generalists leading research, engineering, evaluation, or safety teams. Useful prerequisites: Correlation versus causation, The difference between individual skill and team process.

The question
Why do capable teams sometimes hide errors and questions instead of using them to learn?
What the authors did
Edmondson ran a three-phase multimethod field study at one office-furniture manufacturer. Preliminary interviews and observations developed the constructs, surveys measured 51 teams through members and outside observers, structured interviews supplied an additional view of team design, and follow-up fieldwork compared selected high- and low-learning teams. Group-level regressions, mediation tests, and GLM analyses tested eight hypotheses.
The source
Psychological Safety and Learning Behavior in Work Teams

Will the error become a learning signal?

How interpersonal cost can change whether an error becomes a learning signal A team member observes the same error in both states. A control changes the expected team response. Under a low interpersonal cost, the member speaks up and the team can learn. Under expected punishment, the member stays silent and the signal remains hidden. The technical signal stays fixed Error or concern known to one member same signal in both states Low interpersonal cost expected team response help, interest, no rejection Speak up ask, admit, dissent Learning signal enters team process What the team can do with a surfaced signal Seek feedback bring in another view Discuss error make causes collective Test assumption compare belief with data Experiment change and observe The signal remains private silence protects the member but blocks team learning expected punishment makes silence individually safer Low interpersonal cost: a hidden signal can become learning behavior. Safety enables a process. It does not replace skill, resources, task design, or accountability.

Low interpersonal cost selected. Speaking up connects the signal to learning behavior.

Same signal, different expected response
Expected responseIndividually safer actionTeam consequence proposed by the model
Help, interest, and no rejectionSpeak up, ask, admit, or dissent.The signal can enter feedback, reflection, error discussion, and experimentation.
Blame, embarrassment, rejection, or punishmentStay silent or move the concern off-line.The member avoids interpersonal cost while the team loses information needed for learning.
The switch holds an error signal fixed and changes the expected interpersonal response. In the low-cost state, disclosure can feed questions, feedback, and learning behavior. In the punitive state, silence protects the individual and deprives the team of the signal. This schematic presents the paper's mechanism, not a measured treatment effect. Arrow thickness, box size, color, and position do not encode causal strength, prevalence, performance gain, or statistical effect size.

Walk through the argument

Each step below points to the source location that carries it. The source map follows the walkthrough for source-level checking.

Treat an error report as a social risk

Admitting an error can help a team while making the speaker look incompetent. Asking for help can improve the work while exposing uncertainty. A team member therefore weighs an organizational benefit against an immediate risk to image, status, or relationships.

Edmondson's mechanism begins with that asymmetry. Silence can be individually rational even when it deprives the group of information. The question is not only whether members know something, but whether the local climate makes it safe to reveal.

Source: Journal pages 350 to 357, introduction, model, hypotheses, and Figure 1

Define psychological safety narrowly

Team psychological safety is a shared belief that the team is safe for interpersonal risk taking. It concerns expected reactions to speaking up, making a mistake, asking for help, or stating a different view.

The construct is not group cohesion. A cohesive team can suppress disagreement. It is not permissiveness or constant positive feeling either. The relevant confidence is that a well-intentioned contribution will not bring embarrassment, rejection, or punishment.

Source: Journal pages 350 to 357, introduction, model, hypotheses, and Figure 1, Journal pages 382 to 383, Appendix survey scales

Follow the behavior between climate and outcome

Edmondson does not treat safety as a direct performance input. It enables learning behavior: seeking feedback, sharing information, asking for help, discussing errors, testing assumptions, reflecting, and experimenting.

That middle step matters. A team can feel safe and still perform poorly if it lacks skill, resources, or a useful task. The model predicts value when safety changes whether relevant information becomes collective action.

Source: Journal pages 350 to 357, introduction, model, hypotheses, and Figure 1, Journal pages 382 to 383, Appendix survey scales

Read what the field study actually observed

The study began with interviews and meeting observations, then surveyed 496 members across 53 recruited teams. It received responses from 427 members in 51 teams and from 135 outside observers. A separate researcher interviewed managers about team design, and later fieldwork compared selected high- and low-learning teams.

At the team level, psychological safety was consistently associated with member- and observer-rated learning behavior. Learning behavior predicted observer-rated performance, and the paper's mediation analysis was consistent with safety affecting performance through learning. Team efficacy was less robust once safety entered the models.

Source: Journal pages 358 to 365, Methods, Tables 1 to 3, and measurement notes, Journal pages 365 to 369, Results and Tables 4 to 8

Hold the error fixed and switch the response

In the high-learning production case, members described criticism as information intended to improve the product. They acknowledged mistakes, sought second opinions, and tested changes. The interpersonal interpretation made the signal usable.

In a low-learning publications team, members described tension, weak support, and reluctance to hear bad news. Questions and concerns stayed private. The case contrast illustrates the proposed mechanism, but it does not isolate safety as the only causal difference.

Source: Journal pages 369 to 377, high- and low-learning team comparisons and Discussion

Keep the causal direction open

The survey is a snapshot. High-performing teams may become safer, safe teams may learn more, good leaders may produce both, and repeated experiences may create feedback loops in every direction. The design cannot separate those paths over time.

Edmondson presents the work as a first step in establishing a construct. The useful conclusion is conditional: interpersonal consequences can shape whether teams expose learning signals. The size, direction, and intervention strategy require stronger evidence in each setting.

Source: Journal pages 377 to 380, Study Limitations and Model Applicability and Conclusion

Source map

These are the source locations that carry the argument. Use them to check this explanation against the original rather than trusting the summary alone.

Where the argument lives
LocusWhy it mattersSource
Journal pages 350 to 357, introduction, model, hypotheses, and Figure 1Defines learning behavior and team psychological safety, distinguishes safety from cohesion and trust, and states the proposed mediation model.Open source →
Journal pages 358 to 365, Methods, Tables 1 to 3, and measurement notesDescribes the site, team types, three research phases, member and observer samples, scales, construct checks, and aggregation to the team level.Open source →
Journal pages 365 to 369, Results and Tables 4 to 8Reports associations among safety, efficacy, learning behavior, performance, coaching, and context support, including the mediation analyses.Open source →
Journal pages 369 to 377, high- and low-learning team comparisons and DiscussionUses field cases to show how responses to errors and feedback differed across teams and how design conditions interacted with shared beliefs.Open source →
Journal pages 377 to 380, Study Limitations and Model Applicability and ConclusionStates construct, common-method, cross-sectional, sample-size, single-company, task-applicability, and causal limitations.Open source →
Journal pages 382 to 383, Appendix survey scalesLists the exact team psychological safety, learning behavior, performance, coaching, efficacy, and observer items used in the study.Open source →

The Assumption Switch

One result. One assumption exposed. Turn it and see what changes.

Assumption under test

A team member expects that raising an error, question, or dissenting view will not trigger rejection or punishment.

Held in the source
The interpersonal cost is low enough that the member can surface the signal, seek help, test an assumption, or invite feedback.
Turn it
The same action is expected to damage image, status, relationships, or career prospects inside the team.
What changes
Silence becomes individually safer even when disclosure could help the team. Information needed for learning stays private, and capability alone does not recover it.

The common misreading

Psychological safety is often treated as comfort, niceness, consensus, or freedom from standards. Edmondson defines it more narrowly as confidence that a team will not embarrass, reject, or punish someone for an interpersonal risk such as admitting an error or asking for help. She distinguishes it from cohesion and does not model it as a direct substitute for task design, competence, or performance.

Outside the ML frame

AI incident reporting and evaluation

Will researchers surface a model failure that also exposes their own mistake?

An evaluation system can collect logs and still miss what team members are afraid to say. Edmondson's mechanism suggests treating error disclosure, requests for help, and dissent as socially risky actions whose local consequences affect the evidence an organization receives. This is a transfer to AI governance, not a study of AI labs.

Where the result stops

The survey is cross-sectional, so it cannot establish causal direction or the self-reinforcing dynamics the paper proposes. The 51-team sample comes from one company, participation was voluntary, and the sample was chosen for variance rather than representativeness. Antecedent analyses rely partly on the same survey, team efficacy and context support had low internal-consistency estimates, and the new construct was not conclusively separated from trust. The paper also warns that learning behavior may add less value for tightly constrained routine tasks.

The numbers, with their measurands

Each value below states its measurand, evidence type, source location and evidence base when the source reports one. These fields distinguish self reported results from independent measurements.

  • 427 members from 51 teams. team-member survey responses included in the group-level analysis. Reported as self reported, Methods, phase 2, journal pages 361 to 362. Evidence base: 496 members across 53 recruited teams were administered the survey. Check it →
  • 135 of 150 observers. outside-observer surveys returned for team learning and performance ratings. Reported as self reported, Methods, phase 2, journal page 362. Evidence base: two or three identified recipients of each team's work. Check it →
  • adjusted R-squared .63 and .35. variance accounted for by the single-predictor psychological-safety models of member-rated and observer-rated team learning behavior. Reported as self reported, Table 5, journal page 367. Evidence base: 51 teams. Check it →

What remains open

  • Which leader behaviors causally build psychological safety without weakening accountability or technical standards?
  • How quickly does safety change after a punished disclosure, a leadership transition, or a public failure?
  • When do anonymous channels improve error discovery, and when do they prevent the team from learning together?
  • How does psychological safety interact with power, status, professional identity, and incentives across organizations?
  • Which team tasks benefit most from learning behavior, and which are constrained enough that other mechanisms dominate?

What this bears on

Superalignment maintains a public register of the claims it makes and the evidence that would change its mind. The entries below connect this source to the exact public claims it bears on.

  • C5. The binding constraint on overseeing stronger workers is conditions, not capability. This record bears on it, indirectly. The study finds that interpersonal conditions predict whether team members surface errors, questions, and feedback, while team efficacy adds less once safety is considered. It does not study stronger-than-human workers or establish the binding AI oversight constraint. See the claim and what would change our mind →

Source audit and review status

What was checked, and against what
FieldResultCheckedByAgainst
titleexact2026-08-17codex-primary-source-reviewsource
authorsexact2026-08-17codex-primary-source-reviewsource
dateexact2026-08-17codex-primary-source-reviewsource
venueexact2026-08-17codex-primary-source-reviewsource
full_textexact2026-08-17codex-primary-source-reviewsource
  • Review status prototype.
  • Program collection seminal v1; field study; wave 3, release slot unassigned.
  • Source access public full text, PDF. Open the reading copy →
  • Provider-authored safety claim no.
  • Explained by Superalignment Research.
  • Reviewed by No named human reviewer yet.
  • Explainer dates created 2026-08-17; updated 2026-08-17.
  • AI assistance AI assisted with primary-source retrieval, full-text extraction, manifestation checking, locus mapping, first-pass prose, figure design, and implementation. The page remains a prototype until a named human review is recorded.
  • Rights and access The article is publisher-controlled and the current publisher page retains Cornell's 1999 copyright notice. A complete version-of-record reading copy is publicly hosted by MIT, but it carries JSTOR use terms rather than an open-content license. Link to the sources rather than redistributing their text or pages.
  • Corrections Read the correction policy or report an error.

Provenance

  • First seen 2026-08-17, via cross-disciplinary seminal-source survey and full-source review.
  • Work id work:edmondson-psychological-safety-team-learning, which groups manifestations of the same intellectual work.
  • Record id doi:10.2307/2666999, the natural key for this catalog manifestation.
  • 2026-08-17 full version-of-record paper read and implementation-ready Explained prototype prepared with construct, causal, and generalization caveats

Full audit data, including this record under id doi:10.2307/2666999: /library/records.jsonl. Compact browser index: /library/corpus.json. Catalog method and counts: /library/index.json.