"Diversity trumps ability" is one of the most quoted — and most contested — claims in decision science. The strong slogan outruns the evidence, but the weaker claim that survives scrutiny still matters enormously: on complex problems, a team's performance depends less on its smartest member than on whether its members think differently and whether those differences actually reach the decision.
- The c factor: group intelligence correlates only weakly with member IQ; process (turn-taking, social sensitivity) predicts more — supported by a 1,356-group meta-analysis, still debated
- Hong-Page: diverse groups can beat "best-member" groups on hard problems — under model conditions, and the theorem's math has been seriously criticized
- The solid core: the diversity prediction theorem — collective error = average error − prediction diversity — is an identity, not an opinion
- Practice: use diversity for complex decisions, expertise for routine ones, and structure the process so divergent views are captured independently and judged on merit
The most seductive claim in management science
Every leader has felt the pull of a simple staffing theory: hire the smartest people you can find, put them in a room, and the best decisions will follow. It has the virtue of being obvious. It has the drawback of being, at best, half true. Over the past two decades, a body of research on collective intelligence has converged on an uncomfortable finding: the intelligence of a team is not the sum — or even the maximum — of the intelligence of its members. Something else is doing a large share of the work, and that something looks a lot like difference: different information, different heuristics, different ways of being wrong.
That finding hardened into a slogan — “diversity trumps ability” — which then did what slogans do: it detached from its conditions and floated free into keynote decks. The backlash was equally sweeping, dismissing the whole line of research as politicized mathematics. Both extremes are wrong, and the truth in the middle is genuinely useful. This article walks through the two research programs behind the claim — the collective intelligence “c factor” and the Hong-Page theorem — including the critiques that are usually left out, and ends with what a careful reader can actually take into a hiring or team-design decision.
The c factor: teams have an intelligence of their own
In 2010, Anita Williams Woolley, Christopher Chabris, Alex Pentland, Nada Hashmi, and Thomas Malone published a study in Science that asked a deceptively simple question: if individuals have a general intelligence factor (the psychometric g), do groups have one too? They tested 699 people working in groups of two to five on a wide battery of tasks — brainstorming, moral reasoning, negotiation, puzzles — and then checked whether a group’s performance on one kind of task predicted its performance on the others.
It did. A single statistical factor — which they called c, for collective intelligence — explained roughly 43% of the variance in group performance across tasks. The surprise was what c did not track: it correlated only weakly with both the average intelligence of group members and the intelligence of the smartest member. What predicted c instead were three properties of the group itself: the members’ average social sensitivity (measured with the Reading the Mind in the Eyes test), the equality of conversational turn-taking — groups dominated by one or two voices scored lower — and the proportion of women in the group, an effect statistically mediated by social sensitivity.
Read carefully, this is not a claim that ability doesn’t matter. It is a claim that at the team level, ability is not the bottleneck. The bottleneck is whether the group can extract and combine what its members know — which is a property of process, not of any individual skull.
The replication fight you don’t hear about in keynotes
The c factor has been through a genuine scientific fight, and honesty requires reporting it. In 2017, Marcus Credé and Garett Howardson reanalyzed data from six previously published samples and argued that the empirical support for a general collective intelligence factor is weak — that a single factor explains little variance on many group tasks, and that the construct may be partly an artifact of scoring choices. Woolley, Yeonjeong Kim, and Malone replied in 2018, arguing that the critique rested on a scoring procedure and simulation assumptions that did not match the tasks groups actually performed.
The most substantial evidence since then is a 2021 meta-analysis in PNAS by Christoph Riedl, Young Ji Kim, Pranav Gupta, Thomas Malone, and Anita Woolley, covering 22 studies and 1,356 groups — university students, military personnel, online workers, and gamers. It found consistent evidence for a c factor across these populations, with group collaboration processes and member skills both contributing. Where does that leave a practitioner? The construct is supported by a large meta-analysis and still debated in its structure — roughly where g itself stood for decades. What is not seriously debated is the underlying pattern: member IQ is a surprisingly poor predictor of team performance, and interaction process is a surprisingly good one.
“Diversity trumps ability”: what Hong and Page actually proved
The second pillar of the diversity argument comes from economics and complexity science. In 2004, Lu Hong and Scott Page published a paper in PNAS with the memorable title “Groups of diverse problem solvers can outperform groups of high-ability problem solvers.” They built computational models in which agents equipped with different problem-solving heuristics searched a rugged solution landscape. Each agent climbs until stuck; then another agent takes over from that point with different moves.
The result: under the model’s conditions, a randomly selected group of agents — random selection guaranteeing heuristic diversity — outperformed the group composed of the individually best-performing agents. The mechanism is intuitive once seen: the best individual performers tend to have learned similar heuristics, so they get stuck at the same local optima. A team of clones, however talented, searches like one mind. A diverse team’s members rescue each other from their respective dead ends.
But the conditions matter, and Page himself has been clearer about them than his popularizers: the problem must be genuinely hard (no individual can solve it alone), the agents must all be reasonably capable (diversity among the clueless does not help), the pool from which the team is drawn must be large, and the setting is iterative problem-solving where agents build on one another’s positions. “Any diverse group beats any expert group” is not what was proved, and treating it that way sets the claim up to fail.
The counterattack: a theorem under fire
In 2014, mathematician Abigail Thompson published a pointed critique in the Notices of the American Mathematical Society (vol. 61, no. 9, pp. 1024–1030) titled “Does Diversity Trump Ability? An Example of the Misuse of Mathematics in the Social Sciences.” Her argument had two prongs. First, that the formal theorem in the Hong-Page paper, once its assumptions are unpacked, is close to trivial — in her reading, the conditions effectively build the conclusion in, and the mathematical content is thin for the social weight placed on it. Second, that the simulations backing the broader claim are sensitive to parameter choices and do not license the sweeping conclusions drawn from them.
The critique drew responses of its own. A 2017 note in Critical Review by Daniel Kuehn argued that Thompson’s mathematical objections were variously incorrect, misleading, or beside the point of how formal models function in social science — where the value of a model is the mechanism it isolates, not the depth of its proof. Philosophers of science working on epistemic communities, including Daniel Singer, have developed related models in which diversity’s value survives in modified forms.
Our reading, for what a practitioner needs: the theorem is weaker than its title, the mechanism is real and observable, and the sweeping slogan should be retired in favor of the conditional claim. When you face a hard, multidimensional problem and your candidate pool is deep, a team selected for complementary ways of thinking will often search the solution space better than a team selected purely on individual scores — because individual scores cluster people who think alike. That claim needs no overreaching theorem, and the next section shows why.
The math nobody disputes: the diversity prediction theorem
Strip away the contested models and one piece of mathematics remains standing that no one argues with, because it is an identity — algebra, not empirics. Scott Page calls it the diversity prediction theorem (in The Difference, 2007). For any group making numerical estimates, with error measured as squared distance from the truth:
Collective error = Average individual error − Prediction diversity
where prediction diversity is the variance of the individual estimates around the group’s average estimate
Two consequences fall straight out. First, the group’s aggregate estimate is always at least as accurate as its average member — the crowd can’t be worse than the typical individual, though it can be worse than its best one. Second, holding individual accuracy fixed, every increase in genuine disagreement mechanically reduces collective error. Diversity is not a feel-good bonus in this equation; it is one of the two terms. A group of accurate clones and a group of moderately accurate contrarians can land on identical collective error by different routes.
The catch — and it is the same catch that runs through all of this research — is the word genuine. The theorem rewards diversity of predictions, which requires diversity of information and models. If everyone reads the same three sources, their predictions huddle together, the diversity term collapses, and the crowd degrades into an echo. This is precisely the failure mode documented in the wisdom of crowds literature: correlated errors don’t cancel, they compound.
When diversity helps
The problem is complex and multidimensional
Strategy, product design, policy, diagnosis of novel situations — problems where no single perspective sees the whole landscape. Different heuristics get stuck at different points, so one member’s dead end is another’s starting position.
Errors are uncorrelated
The mathematical engine of collective accuracy is error cancellation. If members draw on different information sources and different mental models, their mistakes point in different directions and partially cancel in aggregate.
The diversity is genuinely cognitive
Different training, different industries, different local knowledge, different toolkits. What matters for prediction and problem-solving is diversity in how people think — not diversity on paper that leaves everyone reasoning from identical sources.
The process lets diverse views surface
Woolley’s research points at process: equal turn-taking and social sensitivity predicted group performance. A diverse team dominated by one loud voice performs like a team of one.
When it doesn’t — the part the slogan leaves out
An honest account has to include the costs, because they are real and they are the reason diverse teams sometimes underperform homogeneous ones in practice even when the theory favors them.
Coordination costs exceed the benefit
Diverse groups take longer to build shared language and trust. For routine, well-understood execution tasks, that overhead buys little — there is no hidden perspective to find.
There is no psychological safety
Divergent views only help if they get spoken. In teams where dissent feels risky, diversity produces silence plus friction — the costs without the information.
The problem has a known best method
When one person has unambiguous, verified expertise and the task is inside their domain, adding voices adds noise. Expertise should carry decisions it genuinely covers.
The pattern across the organizational literature is consistent with the models: diversity’s benefits are contingent on task complexity and on process quality, and its costs are front-loaded (coordination, conflict, slower norm formation) while its benefits are back-loaded (better search, fewer blind spots). Teams that quit early capture the costs and miss the payoff.
Cognitive diversity is not demographic diversity
A distinction that keeps this debate honest: the research above is about cognitive (or epistemic) diversity — differences in information, training, heuristics, and mental models. Demographic diversity is a different variable. The two are correlated in the real world, because life experience shapes information and perspective, but they are not interchangeable, and the strongest research effects attach to the cognitive kind. A demographically diverse team drawing on identical sources and identical training can be cognitively uniform; a demographically similar team can contain an engineer, a lawyer, and a field operator who disagree about everything that matters. Note also that Woolley’s c-factor correlates — social sensitivity and turn-taking — are about process, not composition of either kind. For a deeper treatment of the distinction and the research behind it, see our guide to epistemic diversity.
What to actually do: composing and running the team
Translating the defensible core of this research into practice yields advice that is less quotable than the slogan but far more robust:
Sort decisions before staffing them
Complex, consequential, multidimensional decisions get the diverse group. Routine decisions inside one person’s verified expertise get the expert. Applying diversity indiscriminately buys coordination costs with no information payoff.
Recruit for complementary heuristics
When hiring the fourth analyst, the question isn’t “is she the smartest applicant?” but “does she think in a way the first three don’t?” Marginal value comes from covering unsearched territory.
Protect independence before discussion
Collect estimates and arguments individually, in writing, before the group converges. The diversity term in the prediction theorem only exists if opinions form before social influence flattens them.
Engineer the process, not just the roster
Woolley’s correlates are process variables: balanced turn-taking, attention to each other’s signals. A diverse roster run as a monologue performs like a homogeneous one. Structure participation so every perspective is actually extracted.
Making diverse input count: structure as the multiplier
Notice that every practical recommendation above is about capturing diversity, not just assembling it. That is where structured methods earn their keep. In an argument map, each stakeholder contributes their reasoning independently — before the meeting converges on a frontrunner — and every argument stands or falls on its rated merit, not on the seniority of whoever offered it. This severs the status-influence link that lets one confident voice homogenize a diverse room.
Argumentree implements exactly this pattern for collaborative decision making: multi-stakeholder input collected asynchronously and independently, arguments organized into a shared pro/con tree, and multi-dimensional ratings (helpfulness, clarity, accuracy, completeness) aggregated into consensus scores. The diversity your hiring produced becomes information your decision actually uses — with a full audit trail showing which perspectives shaped the outcome. And because contributions can be made in 66 languages, the pool of perspectives isn’t limited to the ones fluent in the meeting’s language.
For the longer arc of the research this article draws on — from Condorcet’s jury theorem to today’s human-AI teams — see our companion piece, Collective Intelligence: 240 Years of Research.
Frequently Asked Questions
Does diversity really beat ability in teams?
Sometimes, under conditions — and the honest answer is that the strong slogan outruns the evidence. Hong and Page’s 2004 model showed that for hard problems, with capable agents drawn from a large pool, a cognitively diverse group can outperform a group of the individually best performers, because the top performers tend to think alike and get stuck in the same places. But the result depends on those model conditions, and its mathematical formulation has been seriously criticized. What survives scrutiny is weaker but still important: on complex problems, diverse perspectives with uncorrelated errors add real value that individual ability alone does not capture.
What is the c factor in collective intelligence research?
The c factor is a proposed general collective-intelligence factor for teams, analogous to the g factor for individuals. Woolley and colleagues (Science, 2010) tested 699 people in small groups across many tasks and found a single statistical factor explaining roughly 43% of the variance in group performance. Strikingly, c correlated only weakly with the average and maximum intelligence of members — and more strongly with average social sensitivity, equality of conversational turn-taking, and the proportion of women in the group.
Has the c-factor research been replicated?
It is a live scientific debate. Credé and Howardson (2017) reanalyzed six samples and argued the evidence for a general factor is weak. Woolley, Kim, and Malone (2018) replied that the critique rested on scoring and simulation assumptions that do not match the tasks. A 2021 meta-analysis by Riedl and colleagues in PNAS, covering 22 studies and 1,356 groups, found support for a c factor across student, military, gamer, and online-worker populations. Supported by a large meta-analysis, still contested in its structure — that is the accurate summary.
What is the Hong-Page "diversity trumps ability" theorem?
In a 2004 PNAS paper, Lu Hong and Scott Page built computational models of agents solving hard problems with different heuristics. Under specified conditions — a difficult problem, individually capable agents, and a large pool to select from — a randomly selected (and therefore cognitively diverse) group outperformed the group of individually best performers, because the best performers used similar heuristics and got stuck at the same local optima. It is a model result about problem-solving diversity, not a blanket law about hiring.
What is the criticism of the Hong-Page theorem?
In 2014, mathematician Abigail Thompson published a critique in the Notices of the American Mathematical Society arguing that the paper’s mathematical theorem is essentially trivial once its assumptions are unpacked, that it is mislabeled as being about diversity, and that the accompanying simulations do not support the sweeping social claims made on the theorem’s behalf. Defenders — including a 2017 response in Critical Review — countered that the objections misread how formal models are used in social science. Scott Page himself scopes the result carefully: it holds for hard problems, capable agents, and large pools.
What is the diversity prediction theorem?
It is a mathematical identity, popularized by Scott Page in The Difference (2007), that decomposes group prediction error: the collective’s squared error equals the average individual squared error minus the diversity of predictions. Nothing about it is controversial — it is algebra. Its practical reading: a crowd is more accurate than its average member whenever predictions genuinely disagree, and holding individual accuracy constant, more prediction diversity always reduces collective error.
How can teams actually benefit from cognitive diversity?
Match diversity to the problem: use diverse input for complex, consequential, multidimensional decisions and let verified expertise carry routine ones. Collect judgments independently before group discussion so perspectives stay uncorrelated. Create conditions where dissent is safe and expected. And evaluate contributions on their merits rather than their source — structured methods such as argument mapping, where every argument is rated on evidence and reasoning rather than the rank of whoever offered it, convert diversity from friction into information.
Argumentree Team
Decision Science
The Argumentree team explores the science of better decisions—from 18th-century mathematics to modern AI.
Turn diverse perspectives into better decisions.
Argumentree collects independent input from every stakeholder and rates arguments on merit — so the diversity you hired actually reaches the decision.
Start Free 14-Day Trial →Related Articles
Wisdom of Crowds: Galton's Ox — The 1906 Experiment That Proved Crowds Beat Experts
The Condorcet Jury Theorem: The 240-Year-Old Math Behind Better Decisions
When Crowds Go Wrong: Bubbles, Cascades, and the Madness of Herds
Across the ecosystemLLM Ensembles
The diversity-trumps-ability theorem is why an ensemble of different LLMs beats a single stronger one.

