Have you ever been in a meeting where everyone nods—"Sounds good," "Agreed," "No concerns here"—and then the moment people hit the hallway, the whispering starts: "That plan won't work."

It's not an isolated problem. Google ran a research initiative called Project Aristotle, spending two years analyzing 180 teams to figure out what makes teams effective. The result surprised everyone: the number-one predictor of team performance wasn't member skill, resources, or leadership style. It was psychological safety.

But "psychological safety" is an abstract phrase. Tell an engineering lead, "Your team's psychological safety needs improvement," and you'll likely get a blank stare—"Safety? Nobody's in physical danger here."

The problem is fundamental: you can't manage what you can't measure. For psychological safety to be discussed and improved, it first has to be visible.

What Psychological Safety Actually Means

Let's clear up a common misconception first: psychological safety ≠ "everyone being nice to each other."

Amy Edmondson, a Harvard Business School professor, introduced the concept in 1999. Her definition: "A shared belief held by members of a team that the team is safe for interpersonal risk-taking—that members can bring their full selves to the team, ask questions, raise concerns, admit mistakes, and disagree with the group without fear of negative consequences."

Important distinction: this is not the same as "comfort." Teams with high psychological safety actually have more disagreement—because people feel safe enough to speak honestly. Teams with low psychological safety appear harmonious on the surface, but the real problems are all underwater.

Edmondson also uncovered a counterintuitive pattern: teams with high psychological safety report more errors and problems. Not because they make more mistakes, but because they're willing to surface them. Low-safety teams make plenty of mistakes too—everyone just hides them, and management only ever sees the tip of the iceberg.

The Edmondson 7-Item Scale

Edmondson developed a 7-item scale to measure psychological safety. It's the most widely used and validated instrument in academic research today.

Code Item (Rate your true experience on this team)
PS1 If you make a mistake on this team, it is often held against you
PS2 Members of this team are able to bring up problems and tough issues
PS3 Members of this team sometimes disagree with each other, but they can do so openly
PS4 It is safe to take a risk and propose ideas that differ from the majority on this team
PS5 It is safe to ask other members of this team for help
PS6 No one on this team would deliberately undermine my efforts
PS7 My unique skills and talents are valued and utilized on this team

Score on a 5-point Likert scale (1 = strongly disagree, 5 = strongly agree). Note that PS1 is reverse-scored—higher scores indicate lower psychological safety. Reverse it before calculating the team total.

Seven items, roughly 2 minutes to complete. Short enough that almost no one skips it, but enough to produce a psychological safety "health profile" for the team.

Adaptation note: If you're translating this scale into another language, don't just do a literal translation—the wording needs to sound natural in the target workplace context. Before launching, have 3-5 target respondents read through the items and confirm there's no ambiguity. One awkward phrasing can skew an entire dimension.

How to Interpret the Results

Once you have the data, look at three things.

1. Look at the Average

If the team average is above 4.0, psychological safety is in healthy territory. 3.0-4.0 means "yellow flags"—the surface looks fine, but certain areas are already unsafe. Below 3.0 is "minefield territory"—team members have stopped speaking up.

But the average is just the entry point. Two teams with the same 3.5 average can have completely different problems.

2. Look at Item-Level Scores

Which item scored lowest? What behavior does that item correspond to?

I've worked with many teams where PS2 (raising tough issues) and PS4 (proposing differing ideas) scored particularly low, while everything else was fine. What does that tell you? Team members aren't afraid of admitting mistakes (PS1 is okay) and they're willing to help each other (PS5 is fine)—but "challenging the status quo" feels unsafe. These teams execute well but struggle with innovation. Everyone follows the process; nobody dares to ask "is this even the right approach?"

3. Look at the Standard Deviation

This is the most overlooked—and most valuable—metric.

If PS4 averages 3.5 but the standard deviation is large (say, 1.2), what does that mean? Team members perceive this issue very differently—some rated it 5 ("totally safe") while others rated it 1 ("completely unsafe").

This situation is more alarming than a low average. It means two different realities coexist within the same team. Some people experience safety that others don't. This usually correlates with role, tenure, or a specific past incident. Wherever the standard deviation is large, that's where you need to dig deeper.

"Psychological safety isn't a 'team attribute'—it's each member's individual experience. The average tells you 'how things are overall.' The standard deviation tells you 'who's getting left behind.' The latter is where improvement starts."

Using It in a Workshop

Psychological safety assessment has two typical workshop applications.

Application 1: Pre-assessment for team development workshops. Send the assessment link 3 days before the workshop. At the opening, display the aggregate results—but never individual data. Show the 7 items as a bar chart so the team can "see" their psychological safety profile.

After displaying, facilitate a discussion. Don't ask "Why is PS4 so low?"—that phrasing makes people defensive. Instead ask: "Looking at these results, what's your reaction? Is there an item that feels particularly true?"

Focus the discussion on real scenarios behind low-scoring items. If PS4 is low, prompt the group: "In what situations would you feel unsafe raising a different idea?" Start from concrete situations, not abstract concepts.

Application 2: Self-awareness tool in leadership workshops. Have managers rate their team using the 7-item scale, then rate "where I want my team to be on each item." The gap between the two is the manager's improvement direction.

This is far more effective than telling someone "you should improve psychological safety." The manager sees the gap themselves—nobody has to lecture them.

What to Do When Scores Are Low

Assessment is diagnosis. After diagnosis, you need a prescription.

Building psychological safety isn't something one workshop can solve, but a workshop can be the starting point. Here are some practical interventions:

Building psychological safety is a long-term project. But the first step—turning it from "hard to articulate" to "visible"—starts with a 7-item assessment.

🛠️ Run Psychological Safety Assessments in FormLM

Psychological safety assessment has specific tool requirements—anonymity is non-negotiable, data visualization is essential, and regular tracking delivers long-term value. FormLM covers all of these:

  • Scale fields: 7 items on a 5-point Likert scale, smooth mobile experience, 2 minutes to complete
  • Anonymous collection: No login or name required—reduces social desirability bias (achievable via no-name share links)
  • Statistical overview: Automatically calculates per-item averages and standard deviations—instantly highlights which dimensions need attention
  • Radar chart: 7-dimension radar chart visualization—perfect for displaying live during a workshop
  • Insight pre-post comparison: Same questionnaire for quarterly re-tests; system auto-compares both rounds to quantify improvement
Create a Psychological Safety Assessment →

📌 Key Takeaways

  • Psychological safety is the #1 predictor of team performance (Google Project Aristotle)
  • The Edmondson 7-item scale is the most validated measurement tool—adapt wording carefully for your context
  • Interpret data at three levels: average score, item-level scores, and standard deviation
  • Large standard deviations are more alarming than low averages—they signal "two realities" within one team
  • Anonymous collection is the baseline; regular re-testing turns a snapshot into a trend
← Previous: Pre-Workshop Assessment Next: Workshop Summary Report Automation →