top of page
Vanguard Voices logo with swirl

How to measure psychological safety in the workplace

8 hours ago
11 min read

For organizational measurement, it helps to separate three kinds of evidence. Validated surveys measure how people experience interpersonal risk. Psychosocial risk assessment examines working conditions that can affect psychological health. Operating evidence shows what the organization did when someone raised a concern or challenged a decision. One score cannot answer all three questions.


Contents



What did the score actually measure?


An engagement survey closes. Six weeks later a slide appears in a leadership meeting with a psychological safety score on it, colored green, and the meeting moves on.


The score may be useful. But before drawing conclusions from it, there is a more basic question: what did it actually measure?


Psychological safety is often discussed as though one number can describe the whole subject. In practice, organizations can be trying to understand several things at once: whether people believe they can raise a problem, admit a mistake or disagree with someone senior; what conditions at work are shaping that experience; and what the organization actually does when somebody takes the risk.


Those questions sit beside one another, but they require different evidence.

Question

Evidence

What it shows

Experience

Validated psychological safety surveys

Perceived interpersonal climate

Conditions

Psychosocial risk evidence

Workplace conditions and controls

Response

Operational records

What happened in practice


Vanguard Voices diagram showing three ways to examine psychological safety: employee experience, psychosocial conditions, and organizational response.
Psychological safety can be examined through three kinds of evidence: employee experience, psychosocial conditions, and organizational response.

The trouble begins when evidence designed to answer one of those questions is used as proof of all three.


How do you measure the experience?


If the question is whether people experience a team as psychologically safe, ask the people on that team.


Psychological safety concerns what people expect will happen when they take an interpersonal risk. Can I admit I made a mistake? Ask for help? Question a decision? Raise a difficult issue without being embarrassed, rejected or punished?


Amy Edmondson's 1999 research defined team psychological safety as a shared belief that a team is safe for interpersonal risk-taking and introduced a seven-item scale for measuring it.


Google later drew on that research in Project Aristotle, its study of team effectiveness. Psychological safety emerged as an important team dynamic, and Google used questions about mistakes, difficult issues, asking for help, risk-taking and whether people's abilities were valued as part of that work.


The important point is that these measures are about experience.


CIPD's 2024 evidence review recommends using Edmondson's established scale and cautions against casually rewriting it. Changing a validated instrument can reduce confidence that it is still measuring the same construct consistently.


That matters because companies often rewrite survey questions to make them sound more familiar or more aligned with internal language. The resulting number may still be interesting, but it becomes harder to compare with the underlying research or sometimes even with an earlier version of the organization's own survey.


Used properly, a psychological safety survey tells an organization something worth knowing: this is how people describe the interpersonal climate of their team.


The result becomes more useful when organizations look below the company-wide average. Two teams inside the same organization can have very different experiences, and an aggregate score can smooth those differences away.


What does a survey score fail to establish?


A survey has boundaries.


Imagine an employee who answers positively to questions about asking for help, admitting mistakes and raising difficult issues. That tells you something real about how she experiences her team.


It does not establish whether an investigation six months earlier was sufficiently independent. It does not show whether a concern stayed traceable when it moved between a manager and HR, who was responsible for responding, or what happened afterward.


It cannot tell you whether someone facing possible retaliation received appropriate protection. It cannot reconstruct the criteria behind a later performance, promotion, role or exit decision.


Those questions sit in operational evidence.


This is where psychological safety measurement becomes muddled. A perception score begins doing work it was never designed to do because it is the most visible number available.


The reverse problem matters too. An organization can have policies, investigation procedures, reporting lines and detailed records while employees still believe challenging authority is costly.


A complete process does not invalidate what people say about their experience. A positive survey does not prove that every organizational mechanism works.


The two forms of evidence tell you different things.


How do you measure the conditions around psychological safety?


Between an individual's experience and the handling of a specific concern sits another layer: the conditions people work inside every day.


Workload, role clarity, organizational change, leadership behavior, communication, conflicting demands, relationships at work and participation in decisions can all affect psychological health.


This is where ISO 45003 is particularly relevant.


ISO 45003:2021 provides guidance for managing psychosocial risk within an occupational health and safety management system based on ISO 45001. It covers identifying conditions, circumstances and workplace demands that can affect psychological health, then assessing and managing the risks associated with them.


That requires a broader evidence base than a psychological safety questionnaire.

Depending on the organization and the risks being examined, relevant information can include psychosocial hazard assessments, worker consultation, workload information, absence and turnover data, incidents, organizational changes and other management-system records.


The measurement question has changed.


A psychological safety scale asks about people's experience of interpersonal risk within a team.


Psychosocial risk management examines workplace conditions with the potential to affect psychological health and what the organization is doing to identify and control those risks.


There is overlap, but the two should not quietly become substitutes for one another. ISO 45003 is a management-system standard for psychosocial risk. It does not prescribe a single psychological safety score.


How do you measure what happens after someone speaks up?


This is the part of the measurement picture that receives much less attention.


Someone challenges a decision in front of the person who made it. An employee raises a concern that will be difficult to deal with. A team repeats an issue it has raised before.


Someone says they are experiencing consequences after speaking up. A person asks for a material employment decision to be reviewed.


At that point something has happened inside the organization, and it can leave evidence.


Follow the concern through the system and a different set of questions appears.

  • Was it received and recorded? A concern can enter through a reporting line, a manager, HR, an ethics function, a grievance route or another channel. The first measurement question is whether relevant concerns remain visible once they enter the organization.

  • Who owned the response? A record should make it possible to see who was responsible for the next action and what obligation applied.

  • What happened when responsibility moved? Concerns often involve more than one organizational function. The evidence needs to survive those handoffs rather than becoming fragmented into unrelated records.

  • Did the person receive a response? That does not mean the organization has to agree with the person who raised the issue. It means there should be a way to determine what was considered, what happened and, where appropriate, what route remained available for review or escalation.

  • Were relevant safeguards considered? Where raising an issue creates a credible risk of retaliation or other harm, the evidence should show whether that risk was recognized and how it was handled.

  • Can later material people decisions be reconstructed? If a subsequent decision affects someone's formal performance status, pay, advancement, authority, role or continued employment, an assessor needs to be able to examine the criteria and evidence behind that decision rather than rely only on a retrospective explanation.

  • What happens when failures repeat? Repeated problems should be visible as a pattern rather than remaining isolated inside separate channels.


None of this requires an organization to create a second company underneath the first one.


Much of the evidence already exists in case-management systems, HR processes, grievance files, employee-listening mechanisms, leadership records and people-decision documentation. The challenge is whether those records form a sufficiently complete and testable picture.


This changes the measurement question from how do people experience the system? to what did the system actually do?


Vanguard Voices is developing its Standard around that distinction. Employee perception remains important evidence. The proposed requirements add tests of organizational performance across leadership accountability, workplace safeguards, voice and response, and material people decisions, with a separate certification-accountability layer governing independent assessment.


The survey is not being replaced.


It is being asked to answer the question it is actually capable of answering.


Why are complaint numbers so hard to read?


Organizations often reach for another apparently objective metric: the number of concerns or complaints raised.


The number is easy to count and difficult to interpret.


A low number can mean there are few problems. It can also reflect reluctance to report them.


A higher number can indicate serious underlying issues. It can also reflect greater confidence in reporting routes or better visibility of problems that previously stayed outside the formal system.


That is why complaint volume on its own tells you very little about whether an organization is psychologically safe.


The number becomes more useful when it is connected to what happened around it.

  • Were relevant concerns captured across different routes, or only those submitted through one formal channel?

  • How long did unresolved actions remain open?

  • Did the organization identify repeated themes across cases?

  • Were safeguards considered where retaliation risk existed?

  • Did recurring failures eventually trigger a different organizational response?


This is a much richer measurement problem than whether complaints went up or down this year.


It is also why rewarding managers simply for reducing the number of reported concerns creates a dangerous incentive. The metric improves when the organization knows less.


Should all of it become one psychological safety score?


A single number is attractive for understandable reasons.


A dashboard tile reading Psychological safety: 82% is easy to communicate. It can be compared with last year, placed beside engagement and turnover, and shown to a board without much explanation.


That simplicity becomes a problem when different types of evidence are compressed into one result.


Imagine an organization with strong team survey scores and mature psychosocial risk processes, but an independent assessment discovers that serious employee concerns were systematically missing from the population being examined.


Now take another organization where employee perceptions fall sharply during a difficult restructuring, but the organization can demonstrate that concerns were captured, review routes worked, people decisions were documented and leaders responded to difficult feedback.


Those findings need to remain visible.


A weighted average can make a serious failure disappear inside strong performance somewhere else.


That is why the Vanguard Voices Standard is not being designed around a single overall psychological safety percentage. A serious failure in one area should not disappear inside a strong average in another. Requirements are assessed against the evidence relevant to them, with critical failures treated separately rather than averaged away.


The purpose of measurement is not to produce the most attractive number.


It is to make important information harder to lose.


Can psychological safety be audited?


The experience itself cannot be audited in the same way as a financial transaction.


People's perceptions are real, and two people sitting in the same meeting can experience the environment differently. Surveys and qualitative evidence remain important because they tell us about that experience.


Organizational systems leave a different kind of trail.


A commitment can be recorded. A concern can remain traceable. A response can be examined against a defined requirement. A review can show who made a decision and whether relevant conflicts were managed. A material people decision can have contemporaneous criteria and evidence behind it, or lack them.


Those things can be assessed.


That is the distinction behind psychological safety accountability: employee experience remains one source of evidence, while independent assessment examines the organizational infrastructure around it.


Sometimes the most useful finding will be the disagreement between those sources.


Leadership may believe the reporting system works while employees consistently describe raising difficult issues as costly. An organization may have poor survey results during a turbulent period while operational evidence shows that difficult concerns are nevertheless being handled consistently.


Neither finding automatically cancels the other.


The gap between them may be exactly what needs attention.


Where does ISO 45003 end and the Vanguard Voices Standard begin?


ISO 45003 is an important foundation for the Vanguard Voices work.


It gives organizations guidance for managing psychosocial risks as part of an occupational health and safety management system based on ISO 45001. Its scope includes identifying workplace conditions that can affect psychological health and taking action to prevent injury and ill health and promote well-being.


Vanguard Voices is developing a narrower accountability layer around what organizations can demonstrate when voice, challenge and organizational power meet.


The proposed Standard examines areas including leadership accountability, workplace safeguards, voice and response, and material people decisions.


Certification against a future Vanguard Voices Standard would not be certification to ISO 45003.


The intention is also not to make organizations build parallel systems simply for an assessment. Where an existing OH&S process, HR system, reporting mechanism or case-management system already produces appropriate evidence, that evidence should be usable.


The distinction is in what is being tested.


ISO 45003 provides guidance for managing psychosocial risk.


The Vanguard Voices Standard is being designed to test whether defined organizational obligations can be demonstrated through evidence and independent assessment.


How should an organization measure psychological safety?


Nothing here requires abandoning the tools organizations already use. It requires being more precise about what each tool can establish.


Measure employee experience with a validated psychological safety instrument. Ask whether people believe they can raise problems, admit mistakes, ask for help and take interpersonal risks within their teams. Keep the integrity of the instrument and examine variation between teams rather than relying only on one company-wide average.


Assess psychosocial conditions through the organization's risk-management system. Identify workplace conditions that can affect psychological health, determine what controls are needed and examine whether those controls are operating. ISO 45003 provides an established framework for doing this within occupational health and safety management.


Examine operating evidence. Where the organization says a response process, safeguard, escalation route or review mechanism exists, examine what happened when people actually used it. The question moves from whether the process exists to whether there is evidence that it operated.


Read the evidence together without pretending it is interchangeable. A sharp fall in one team's perception data should prompt questions about conditions and operational evidence. Repeated failures in case handling should make employee experience in that area worth examining more closely. Neither source should simply overrule the other.


A useful psychological safety dashboard therefore needs more precision than one percentage.


It should make clear which question each measure answers and what remains outside it.


What better measurement changes


A psychological safety survey can tell an organization something important about how people experience a team. Psychosocial risk evidence can show what conditions are shaping that experience. Operating records can show whether the organization's own mechanisms worked when someone actually used them.


Keeping those forms of evidence separate does not make measurement more complicated for the sake of it. It makes the conclusion harder to manipulate.


A high survey score cannot erase a failed safeguard. A complete case file cannot overrule a team telling you that challenging authority feels costly. A low complaint count cannot prove that nothing is wrong.


The useful information is often in the relationship between the evidence.


That is also what makes independent assessment possible. An assessor does not need to certify a feeling. They can examine whether relevant evidence exists, whether the population being assessed is sufficiently complete, whether sampling is independently controlled and whether defined requirements were met.


Vanguard Voices is a Swiss not-for-profit developing an independent, evidence-based public standard for psychological safety accountability. The work is still under development, and no organization has been certified.


The measurement question is therefore bigger than What is our psychological safety score?


It is whether we are measuring the right thing for the question we are trying to answer, and whether important failures remain visible when the evidence is put together.



Interested in helping shape the Standard?


Sources


  • Edmondson, A. C. (1999). “Psychological Safety and Learning Behavior in Work Teams.” Administrative Science Quarterly, 44(2), 350–383.

  • Google re:Work. Understand team effectiveness - Project Aristotle.

  • CIPD (2024). Trust and psychological safety: an evidence review.

  • ISO 45003:2021. Occupational health and safety management - Psychological health and safety at work - Guidelines for managing psychosocial risks. ISO describes it as guidance for managing psychosocial risk within an OH&S management system based on ISO 45001.

  • ISO 45001:2018. Occupational health and safety management systems - Requirements with guidance for use.



bottom of page