Position Paper
Why More Is Not Better
Team Diagnostics Beyond the Maturity Model
The safest team in the room
Imagine a leadership team that scores near the top of every team-culture survey it takes. Members report that they feel safe to speak. Meetings are warm. Feedback is described as frequent and kind. Conflict, when the survey asks about it, is rated as “handled respectfully.” By every instrument the team has ever completed, it is a model of psychological safety.
It is also, quietly, failing. Weak ideas pass unchallenged because challenging them would feel unkind. Feedback is frequent precisely because it has become weightless, a social lubricant rather than a performance instrument. The hardest conversations, the ones about a colleague who is no longer delivering and a strategy that is no longer working, have not happened in two years. The team has become what one member privately calls “a mutual appreciation society.” Its safety is real. Its safety is also the problem.
No maturity-model survey can see this team. Every item on such a survey points one way, toward more voice, more feedback, more openness, more collaboration, and this team has more of everything. It scores green across the board. The instrument congratulates the pathology.
This paper is about why that happens, why it is a design flaw rather than a data problem, and what a team diagnostic has to look like instead.
How a good idea became a ceiling
Psychological safety deserves its reputation. Kahn (1990) first described it as the condition of being able to show oneself “without fear of negative consequences” to self-image, status or career. Edmondson’s field research (1999, 2012) then demonstrated, with unusual rigour for the field, that teams which feel safe to take interpersonal risks learn faster and perform better. These findings are among the most robust in organisational psychology, and nothing in this paper disputes them.
But it is worth noticing what the founders of the concept actually said. Kahn framed safety as one of three psychological conditions of engagement, not a sufficient one. Edmondson has spent much of the last decade insisting that safety must be paired with accountability and ambition, that a safe team with low standards is a comfort zone, not a learning zone. Hawkins (2020, 2021) has argued that the whole “high-performing team” framing needs to be transcended, not perfected. The scholars were always describing a condition within a tension. It was the market that flattened the tension into a target.
The flattening was commercially understandable. A target can be scored, benchmarked, colour-coded and sold. “Make your team psychologically safe” fits on a slide in a way that “hold safety and challenge in productive tension, recalibrating continuously as your context changes” does not. So an entire generation of team diagnostics was built on a single silent assumption: that for every quality worth measuring, more is better.
That assumption is false, and the evidence against it is not marginal. It is one of the best-documented patterns in management science.
The evidence for too much
Grant and Schwartz (2011) surveyed the psychological literature and found the inverted-U everywhere: virtually every positively valued trait and practice, generosity, optimism, conscientiousness, empathy, even happiness, confers benefits up to a point and costs beyond it. Pierce and Aguinis (2013) did the same for management research and named it the “too much of a good thing” effect, showing that ordinarily beneficial factors such as experience, assertiveness, organisational identification and empowerment reliably turn harmful at high levels. Their conclusion was blunt: the assumption of monotonic benefit is one of the most consequential and least examined errors in how organisations measure themselves.
Team life supplies the examples daily. Trust is essential; unconditional trust removes the scrutiny that catches bad decisions. Candour is essential; unfiltered candour delivered without care silences the very voices it claims to liberate. Cohesion is essential; cohesion past a threshold is the raw material of groupthink, a finding as old as Janis (1972). Experimentation drives learning; perpetual experimentation means nothing is ever consolidated and quality never stabilises. Even the willingness to challenge the leader, a hallmark of safety in most surveys, has a documented far end, where authority erodes and the team loses coherent direction altogether.
None of this is exotic. Every experienced team coach has sat with a team that had too much of something good. That’s not the strange part. The strange part is that almost no team diagnostic can detect it.
The measurement trap
It is a structural problem, and the solution needs a new structure. A Likert item is monotonic by construction. When a survey asks a team to rate “we give each other feedback regularly and directly” from 1 to 5, it has built a ruler with good at one end and bad at the other. The only failure such a ruler can register is deficit. A team drowning in casual, weightless feedback, feedback so normalised it can no longer carry a serious message, scores a 5 and is told it is a role model.
This single design decision produces a cascade of familiar diagnostic failures.
Every inverted-U documented by Pierce and Aguinis becomes invisible: an over-rotation blind spot. The teams most in need of recalibration, the over-safe, the over-candid, the over-cohesive, are precisely the teams the instrument scores highest.
Then there’s the ceiling problem. Once a team scores in the 4s, a maturity survey has nothing left to say except “stay there.” The diagnostic exhausts itself exactly when a good team’s real questions begin. Mature teams do not stop having tensions; as their season changes, their tensions shift too.
Some instruments try to add depth by asking teams to rate both their current state and their needed future state, with their priorities to grow being guided by the sizes of the gaps discovered. The instinct is sound; the execution fails predictably, and it fails in the same way every time. On a more-is-better scale, no one rates their needed future below “consistently well.” Social desirability compresses every aspiration toward the top of the scale. When the future column is a near-constant, the gap is arithmetically just the inverse of the current score, and the prioritisation adds no information. Worse, gap logic rewards low ambition: the complacent team that aspires modestly is coloured green everywhere. It’s a kind of aspirational collapse.
Worst of all is what happens next: a monotonic instrument does not merely fail to see over-rotation, it actively prescribes it. Tell a conflict-avoidant team to raise its “challenge” score and it may genuinely improve. Tell a team already bruised by combative debate to raise the same score, because the same item flagged amber for a different reason, and the instrument has just recommended the disease as the cure.
These are not implementation defects that better item-writing can fix. They follow from the geometry of the scale. If the instrument is a ruler, it cannot measure a see-saw.
Tensions, not traits
Organisational scholarship arrived at the alternative some time ago; diagnostics have been slow to follow. Smith and Lewis (2011) synthesised two decades of research into a theory of paradox: the observation that organisations are constituted by contradictory-yet-interdependent demands, stability and change, control and autonomy, care and challenge, that persist over time and cannot be resolved, only navigated. Johnson (1992) had already built the practitioner version: polarity management, the discipline of treating chronic tensions as polarities to be managed rather than problems to be solved, each pole carrying genuine upsides and a genuine shadow.
Applied to teams, the shift is fundamental. The qualities a maturity model treats as traits to maximise are better understood as positions on a tension to calibrate. Safety is not a score; it is one pole of several distinct tensions, safety and standards, voice and discretion, vulnerability and confidence, challenge and deference, each with two legitimate positions and two genuine failure modes.
A diagnostic built on this understanding looks structurally different from a survey. In the iTeam framework, every measured dimension is a continuum between two named poles, and the design imposes four non-negotiable rules. Both poles must have genuine merit: neither is the correct answer, and a team may rightly stand anywhere between them depending on its context and season. Both extremes must be genuinely negative: the far end of even the most virtuous-sounding pole is explicitly described as a failure mode, so that “high trust” carries its shadow (blind trust, scrutiny abandoned) as visibly as “low trust” carries its own. The poles must be true opposites, creating real tension rather than restating one idea twice. And every continuum must name a distinct paradox, so the instrument maps the team’s actual tensions rather than sampling one construct twelve ways.
The practical consequence is that the diagnostic can finally see in both directions. The over-safe team from our opening page does not score green on a continuum instrument. Its markers cluster hard against one pole, and the instrument’s own text, in the extreme clause attached to that pole, describes what the team is living: an environment where every idea is welcomed and nothing is truly tested, where the quality of thinking quietly atrophies because no one will apply the pressure rigorous work requires. The team does not receive a score. It receives a mirror, with the consequences written on it.
The average is not the team
There is a second silent assumption in conventional team measurement, and it deserves its own indictment: the assumption that a team can be summarised by its mean.
Report a dimension as “3.2, range 1–5” and you have told the reader almost nothing that matters. A team averaging 3.2 with responses spanning the full scale is very plausibly not a moderately safe team at all. It is two teams wearing one name: half its members experiencing genuine safety, half experiencing none. In team-coaching terms, that split is the diagnosis. It usually maps directly onto the real fault line: an in-group and an out-group, a founding cohort and the newcomers, the people in the room and the people on the video call. Averaging the two halves together does not simplify the finding. It deletes it.
A continuum instrument reports differently by design. It shows the full distribution of individual placements along each tension, so a split team looks like a split, two visible clusters, the moment the coach opens the report. The platform has anonymity built into its design. It records whether a member completed, never what they answered.
In this case, absolute anonymity can also be too much of a good thing, as having the option to unpack named roles versus team experience can, conditionally, be very valuable. When a named role’s placement is shown — the leader’s, a co-leader’s, the chair’s or the coach’s — it appears only through an explicit reveal control. This design and functionality enable something no calculated average can: it makes the team’s genuine disagreement or misalignment about its own reality discussable.
Engineering is not magic, and its limits must also be named. In a very small team, or where a member writes a comment unmistakably their own, a reader may still infer who said what. That is a matter for the coach’s framing and for judgement about team size, and it is better stated as a caveat than ignored.
Some teams are ready and need to see the gaps between themselves and their leaders, some are not. Regardless, the most productive systemic team conversations rarely begin with “our score is 3.2.” They begin with “what is going on that we are standing in two different places?”
Calibration, not maximisation
The deepest practical difference between the two paradigms shows up in the question each one hands the team at the end.
A maturity model asks: how do we raise our score? The question has a fixed answer, more, and it is the same answer for every team, in every context, forever.
A continuum instrument asks: where does this team need to stand on this tension, for the season we are in? That question has a real answer, and the answer moves. A team integrating three new members may deliberately stand closer to inclusive contribution than meritocratic voice for a quarter, accepting slower debate as the price of building the base. A team entering a crisis may consciously shift toward decisive deference and away from open-ended challenge, not because challenge stopped mattering, but because this season demands speed, and the team is choosing its trade-off with open eyes. Six months later, the right position may be different, and the team returns to the continuum not to check whether its number went up but to ask whether its position still fits its context.
It is here, too, that the aspiration question, sound in instinct, broken in execution on a monotonic scale, comes back to life. Ask each member not how much better should we be? but where does this team need to stand on this tension, for the season it is in, given its mandate and stakeholders? and the answer is a position with a reason, not a number pressed against a ceiling. And because the answers are individual, their spread is itself a finding: a team that agrees on where it stands but disagrees on where it needs to stand has an alignment question before it has a development question, a distinction no averaged instrument can draw.
This is also why a continuum instrument does not exhaust itself. There is no ceiling to reach, because there is no top of the scale, only a position, a context, and the question of fit between them. The best teams do not graduate from the instrument. They get better at the conversation it structures. That conversation still needs a skilled coach to hold it: the instrument does not replace the facilitator, it arms the room. Calibration is a permanent discipline in a way maximisation never was. Maximisation ends at the ceiling; calibration ends never.
Four tests for any team diagnostic
For coaches and leaders evaluating instruments, including ours, the argument compresses into four questions. Ask them before you commission anything.
Can it see too much as well as too little? Ask the vendor to show you how the instrument would flag a team that has over-rotated into its virtues: too harmonious, too candid, too cohesive. If the honest answer is that such a team scores well, the instrument has a blind spot exactly where experienced coaches know the subtle pathologies live.
Does it show you the spread or sell you the average? Ask to see how a genuinely divided team appears in the report. If the answer is a mean with a range in brackets, the instrument will flatten the most diagnostic finding a team can produce.
Does it have anything to say to a good team? Ask what the report looks like for a team scoring at the top of the scale. If the answer amounts to congratulations, the instrument is a deficit-finder, and its useful life with any given team is one engagement.
Does its advice change with the team’s position? Picture two teams that both flag on the same dimension, say, challenge. One flags because nobody ever pushes back; the other flags because debate has turned combative and people are getting hurt. These teams need opposite interventions: the first needs more challenge, the second needs less. Ask the vendor what guidance each team would receive. If the instrument gives both teams the same advice, because its only direction of travel is up, then for one of those teams it is prescribing the disease as the cure.
These are not rhetorical questions. Each one can be answered in a demo and checked against a sample report, and a diagnostic built on paradox principles should be able to pass all four. Ours included.
Conclusion
Return, finally, to the team this paper opened with, the safest team in the room. Nothing was wrong with its people, and nothing was wrong with its intentions. What failed it was its instruments: every survey it ever took could see in only one direction, told it that more was better, and congratulated it as it drifted to the far end of a tension it did not know it was in. Seen through a continuum, that team is not a mystery. It is a position: markers pressed hard against one pole, the extreme’s description reading like a diary of its last two years, and a calibration question waiting that nobody had ever thought to ask it.
Psychological safety was never the destination, and the people who founded the concept never claimed it was. Nor is it a single dial to turn up. What the term names is a whole territory of team life, spanned by a family of distinct tensions, safety and standards, voice and discretion, vulnerability and confidence, challenge and deference, which combine, in turn, with every other tension a team holds. Each pole in that family is real and essential, and each is capable of harm in excess, as every virtue can be. The research on too-much-of-a-good-thing effects is extensive; the theory of organisational paradox is mature; the practice of polarity management is thirty years old. The only part of the field still behaving as if more is always better is the measurement layer, and that is a design choice, not a necessity.
Teams do not need a better ruler. They need a mirror that shows both edges of every tension, the honest spread of where their members actually stand, and a way to keep asking, season after season, whether their position still fits their context.
The question is not how safe is this team? It is not even how mature is this team? It is: what is this team’s own best thinking about where it needs to stand on each of its tensions, in service of its stakeholders and its mandate, for the season it is in?
No instrument can answer that question for a team. The right instrument makes it answerable. That is what iTeam was built to do.
Bibliography
Edmondson, A. (1999). Psychological safety and learning behavior in work teams. Administrative Science Quarterly, 44(2), 350–383.
Edmondson, A. (2012). Teaming: How Organizations Learn, Innovate, and Compete in the Knowledge Economy. Jossey-Bass.
Grant, A. M., & Schwartz, B. (2011). Too much of a good thing: The challenge and opportunity of the inverted U. Perspectives on Psychological Science, 6(1), 61–76.
Hawkins, P. (2020, 23 June). We need to move beyond ‘High Performing Teams’ [Blog post]. Renewal Associates. renewalassociates.co.uk/we-need-to-move-beyond-high-performing-teams
Hawkins, P. (2021). Leadership Team Coaching: Developing Collective Transformational Leadership (4th ed.). Kogan Page.
Janis, I. L. (1972). Victims of Groupthink. Houghton Mifflin.
Johnson, B. (1992). Polarity Management: Identifying and Managing Unsolvable Problems. HRD Press.
Kahn, W. A. (1990). Psychological conditions of personal engagement and disengagement at work. Academy of Management Journal, 33(4), 692–724.
Pierce, J. R., & Aguinis, H. (2013). The too-much-of-a-good-thing effect in management. Journal of Management, 39(2), 313–338.
Smith, W. K., & Lewis, M. W. (2011). Toward a theory of paradox: A dynamic equilibrium model of organizing. Academy of Management Review, 36(2), 381–403.