The idea that groups can be smarter than their smartest member is not a Silicon Valley slogan — it is a 240-year research program. This is the full timeline, from Condorcet’s 1785 theorem to the 2026 science of human-AI superminds.
- 1785: Condorcet proves majorities of independent, better-than-chance voters approach certainty
- 1907: Galton’s crowd guesses an ox’s weight within 1% — better than the experts
- 2010: Woolley et al. measure a group “c-factor” driven by social process, not member IQ
- 2021: A 22-study, 1,356-group meta-analysis supports the c-factor after a real scientific fight
- 2024–2026: Human-AI combinations underperform the best of either alone unless the collaboration is designed — the lesson the whole history was building toward: structure decides whether groups get smarter
In 1794, a French mathematician died in a revolutionary prison cell, nine years after publishing a theorem almost nobody read. In 1906, an 84-year-old Victorian statistician wandered through a livestock fair collecting used raffle cards. In 2026, researchers at MIT are running hundreds of thousands of negotiations between AI agents to learn how machines should join human teams. These three scenes belong to a single, continuous research story — the study of collective intelligence, the capacity of groups to act in ways that seem intelligent. What follows is that story, era by era, with the actual findings and the actual fights.
The Enlightenment era
Condorcet proves mathematically that a group of imperfect judges can be almost perfectly right — if they judge independently.
The statistical era
Galton measures crowd wisdom at a county fair; Hayek shows prices aggregate dispersed knowledge; RAND invents the Delphi method.
Cybernetics and organizations
Game theory formalizes strategic interaction; Engelbart reframes computers as intellect augmenters; groupware emerges.
The swarm era
Termites, ants, and simulated birds show that intelligence can emerge from simple local rules — with no leader at all.
The internet era
Wikipedia, open source, and prediction markets put collective intelligence into daily practice; MIT starts measuring it.
The debate decade
The c-factor faces its critics and a 1,356-group meta-analysis; the field grows up through replication and reply.
AI-human superminds
LLMs join the group. Meta-analyses show human-AI combinations only beat both alone when the collaboration is designed well.
1785: The Enlightenment wager — Condorcet’s theorem
Marie Jean Antoine Nicolas de Caritat, Marquis de Condorcet, was the kind of Enlightenment figure who makes modern polymaths look narrow: a mathematician elected to the Académie des sciences in his twenties, an early and vocal advocate of public education, women’s suffrage, and the abolition of slavery. In 1785 he published the Essai sur l’application de l’analyse à la probabilité des décisions rendues à la pluralité des voix — an essay on applying probability theory to majority voting. Buried in it is the result we now call the Condorcet Jury Theorem.
The theorem says something startling. Take a yes/no question with a correct answer. Suppose each voter is right with a probability even slightly better than a coin flip — say 55% — and suppose the voters judge independently of one another. Then as the group grows, the probability that the majority is right climbs toward certainty. Mediocre individual judgment, aggregated properly, becomes near-perfect collective judgment. The theorem has a dark twin, too: if each voter is slightly worse than a coin flip, the majority becomes almost certainly wrong as the group grows. Aggregation amplifies whatever competence — or incompetence — you feed it. We tell the full story of the theorem, its assumptions, and its modern uses in our companion piece on the Condorcet Jury Theorem.
The same essay contains a second discovery that would take 160 years to detonate: the Condorcet paradox. With three or more options, majority preferences can cycle — the group prefers A to B, B to C, and C to A — so “what the majority wants” may not even be well defined. That observation is the seed of social choice theory, the field Kenneth Arrow would later transform with his impossibility theorem: no ranked voting system can satisfy a short list of reasonable fairness conditions at once. Voting systems, ranked-choice methods, and the mathematics of preference aggregation are a story of their own — a different kind of aggregation from the belief-averaging this article follows — and one we will return to in a future series.
Condorcet did not live to see any of it matter. Outlawed during the Terror for opposing the Jacobin constitution, he hid for months — writing, remarkably, his optimistic Sketch for a Historical Picture of the Progress of the Human Mind — was arrested in March 1794, and was found dead in his cell at Bourg-la-Reine two days later. The mathematics of collective wisdom began as the work of a man destroyed by a collective gone mad. The irony has hung over the field ever since: groups can be brilliant, and groups can be monstrous, and the difference is not the people — it is the conditions.
1906–1950s: The statistical era — Galton, Hayek, and Delphi
For its first empirical test, collective intelligence had to wait 121 years — and it arrived at a livestock fair. In the autumn of 1906, Francis Galton, then 84, attended the West of England Fat Stock and Poultry Exhibition in Plymouth, where a weight-judging contest invited visitors to guess the dressed weight of a live ox for sixpence a ticket. Galton — a founder of statistics, and no democrat by temperament — expected the entries to prove the ignorance of the average voter. He obtained the used cards, discarded 13 illegible ones, and analyzed the remaining 787 guesses. The middlemost estimate was 1,207 lb; the ox’s actual dressed weight was 1,198 lb — an error under 1%, better than the cattle experts present. He published the result in Nature on 7 March 1907 under the title “Vox Populi,” and in follow-up correspondence acknowledged the arithmetic mean was more accurate still: 1,197 lb, one pound off. The full story — including what Galton got wrong and the century of replications since — is in our deep dive on Galton’s ox, and the general phenomenon in our guide to the wisdom of crowds.
Why does the trick work? Each guess can be split into signal plus error. If the errors are independent and scattered on both sides of the truth, aggregation cancels them while the shared signal survives. That “if” is Condorcet’s independence assumption wearing statistical clothes, and it is the hinge on which every crowd-wisdom success and failure turns. Kenneth Wallis’s 2014 re-examination of Galton’s competition confirmed how well the fair’s conditions — private written entries, an incentive to be accurate, no visible running tally — happened to protect it.
The next leap came from economics. In “The Use of Knowledge in Society” (American Economic Review, 1945), Friedrich Hayek argued that the knowledge relevant to economic decisions never exists in concentrated form — it is dispersed across millions of people as local, partial, often tacit fragments. No central planner can collect it; but a price system aggregates it automatically, compressing scattered knowledge into a single number that tells everyone how to act. Hayek had, in effect, identified the first large-scale aggregation mechanism running in production. Half a century later, prediction markets would take his insight literally and build markets whose only product is the information in the price.
The 1950s added the first deliberately engineered collective intelligence process. At RAND Corporation, researchers developing long-range forecasts faced a familiar problem: put experts in a room and the loudest or most senior voice anchors everyone else. Their answer, the Delphi method, collected expert judgments anonymously and in iterated rounds with controlled feedback — protecting independence from rank and social pressure while still letting information circulate. Delphi is the direct ancestor of every modern structured-input process, from planning poker to the independent-first contribution flow used in collaborative decision making platforms.
1940s–1990s: Cybernetics, game theory, and augmented intellect
While statisticians studied aggregation, another lineage studied interaction. John von Neumann and Oskar Morgenstern’s Theory of Games and Economic Behavior (1947 edition) gave social science a mathematics of strategic interdependence — what happens when the outcome of my choice depends on yours. Game theory reframed groups not as bags of independent estimators but as systems of interacting agents, a view that would matter enormously once researchers started asking why rational individuals produce irrational collectives (the subject of information cascades research decades later).
The second thread was technological. In his 1962 SRI report Augmenting Human Intellect: A Conceptual Framework, Douglas Engelbart proposed that computing’s purpose was not to replace human thinking but to augment it — including the collective thinking of groups tackling problems too complex for any individual. Engelbart’s framing seeded the research field of computer-supported cooperative work (CSCW) that took shape in the 1980s, and the groupware wave that followed. The augmentation-versus-replacement question he posed in 1962 is, almost verbatim, the question the 2024–2026 human-AI literature is now answering with data.
This middle era contributed something easy to underrate: it turned collective intelligence from a property you observe into a system you build. Delphi was a designed protocol; groupware was designed software; Engelbart’s framework was explicitly an engineering agenda for group cognition. The statistical era had asked “are crowds accurate?” The cybernetic era asked the question every modern collaboration platform inherits: what structure makes a group smarter than it would otherwise be? Structured argumentation itself has a parallel intellectual lineage here — Stephen Toulmin’s 1958 model of argument and the later argument-mapping tradition gave the deliberative side of collective intelligence its data structure, the argument map, just as prices had become the data structure of market aggregation.
1959–1999: The swarm era — intelligence without anyone in charge
Meanwhile, biology was quietly dismantling the assumption that intelligence requires a brain in charge. In 1959, the French zoologist Pierre-Paul Grassé, studying termite nest construction, coined the term stigmergy: coordination through traces left in the environment. No termite holds a blueprint; each responds to the local state of the mound, and the mound itself coordinates the colony. It was the first precise mechanism for leaderless coordination — and it turned out to generalize far beyond insects.
In 1987, computer graphics researcher Craig Reynolds showed how little individual sophistication emergence actually needs. His SIGGRAPH paper “Flocks, Herds, and Schools: A Distributed Behavioral Model” animated flocks of simulated birds — boids — governed by just three local rules: separation, alignment, and cohesion. No leader, no script, and yet the flock wheels and splits and reforms like the real thing. Two years later, Gerardo Beni and Jing Wang coined the term swarm intelligence in the context of cellular robotic systems, giving the phenomenon its name.
The engineering payoff followed fast. Marco Dorigo’s 1992 PhD thesis at Politecnico di Milano, Optimization, Learning and Natural Algorithms, translated ant foraging — pheromone trails reinforcing shorter paths, as revealed by the double-bridge experiments with Argentine ants — into Ant Colony Optimization, a family of algorithms for routing and scheduling problems. James Kennedy and Russell Eberhart’s Particle Swarm Optimization (1995) did the same for flocking. By 1999, Bonabeau, Dorigo, and Theraulaz’s Swarm Intelligence: From Natural to Artificial Systems (Oxford) had codified the field, and ant-based algorithms were routing telephone calls (Schoonderwoerd and colleagues at BT Labs, 1996) and delivery trucks. Later, robotics made the point physically: in 2014, Rubenstein, Cornejo, and Nagpal self-assembled shapes from 1,024 Kilobots (Science), and Thomas Seeley’s Honeybee Democracy (2010) showed honeybee swarms choosing nest sites through scouting, waggle-dance advocacy, and quorum sensing — a natural, decentralized approximation of a jury theorem in action. The full account is in our guide to swarm intelligence.
The swarm era’s lesson for human groups is double-edged. It proved that decentralization works — local knowledge plus simple interaction rules can outperform central planning, exactly as Hayek argued. But it also marked a boundary: swarms coordinate actions, not reasons. Ants do not deliberate. Human collective intelligence adds a layer no swarm has — the exchange and evaluation of arguments — and that layer needs different machinery.
1999–2015: The internet era — collective intelligence goes into production
Then the internet made everyone a potential contributor, and collective intelligence stopped being a laboratory curiosity. Open-source software demonstrated that thousands of loosely coordinated volunteers could build systems as complex as an operating system — Eric Raymond’s 1999 essay The Cathedral and the Bazaar supplied the era’s manifesto. Wikipedia, launched in 2001, scaled the model to human knowledge itself: an encyclopedia with no editor-in-chief, coordinated stigmergically through the artifact being edited — termite logic, applied to text.
In 2004, two publications gave the era its theory. James Surowiecki’s The Wisdom of Crowds assembled a century of evidence and distilled the four conditions under which groups are wise: diversity of opinion, independence, decentralization, and aggregation — remove any one, and the crowd gets dumber, not smarter. The same year, Lu Hong and Scott Page published their “diversity trumps ability” result in PNAS: in their computational model, a random, cognitively diverse group of problem solvers outperformed a group composed of the individually best performers. Honesty requires the footnote: in 2014, mathematician Abigail Thompson published a sharp critique in the Notices of the American Mathematical Society, arguing the underlying theorem was trivial and mislabeled and the simulations overinterpreted; defenders including Page himself replied that the result holds under stated conditions — hard problems, capable agents, large pools — and was never a claim that any diverse group beats any expert team. The debate, and what survives it, is examined in our piece on diversity versus ability and our guide to epistemic diversity.
Hayek’s price mechanism got its controlled experiment in this era too. The Iowa Electronic Markets, a real-money research market founded at the University of Iowa in 1988, accumulated enough history for a definitive comparison: Berg, Nelson, and Rietz (International Journal of Forecasting, 2008) matched market prices against 964 national polls across the 1988–2004 US presidential elections and found the market closer to the eventual outcome 74% of the time, with its edge strongest at horizons beyond 100 days. Robin Hanson’s logarithmic market scoring rule (2003) meanwhile solved the thin-market problem algorithmically, letting even small groups run internal prediction markets.
Expert judgment, by contrast, took a beating. Philip Tetlock’s Expert Political Judgment (2005) reported a roughly two-decade study of 284 experts and about 28,000 probability judgments; the average expert barely outperformed simple extrapolation — the finding behind the famous “dart-throwing chimpanzee” caricature. But Tetlock’s follow-up work in the IARPA forecasting tournaments (2011–2015) found the constructive half of the story: a small cadre of “superforecasters,” distinguished not by credentials but by probabilistic thinking and relentless belief updating, reportedly outperformed intelligence analysts with access to classified information. Forecasting skill and calibration deserve — and will get — their own deep treatment in a future series.
The era’s institutional landmark came in 2006, when Thomas Malone founded the MIT Center for Collective Intelligence around a deceptively simple question: how can people and computers be connected so that, collectively, they act more intelligently than any person, group, or computer has ever done before? Early outputs mapped the design space — the “Collective Intelligence Genome” (Malone, Laubacher, and Dellarocas, MIT Sloan Management Review, 2010) catalogued the building blocks of CI systems along four questions: who is performing the task (crowd or hierarchy), why (money, love, or glory), what (create or decide), and how (independently or dependently). By 2015, Malone and Bernstein’s Handbook of Collective Intelligence (MIT Press) could define the field’s object plainly: “groups of individuals acting collectively in ways that seem intelligent.”
And in 2010, the field acquired its most provocative measurement. Anita Williams Woolley, Christopher Chabris, Alex Pentland, Nada Hashmi, and Thomas Malone (Science, 330: 686–688) tested 699 people working in groups of two to five on a wide battery of tasks. Group performance across tasks was correlated enough to yield a single statistical factor — a collective intelligence factor, c — explaining about 43% of the variance, the group-level analogue of individual IQ’s g-factor. The shock was what predicted it. Average and maximum member intelligence correlated only weakly with c. What correlated strongly: the group’s average social sensitivity (measured by the Reading the Mind in the Eyes test), equality of conversational turn-taking, and the proportion of women in the group (itself largely mediated by social sensitivity). How smart a group is, the data said, depends less on who is in it than on how it interacts.
The shadow history: when collectives fail
Running underneath the whole optimistic timeline is a second, darker literature — the study of how groups of intelligent, well-meaning individuals produce collective stupidity. It matters here because it is not a separate subject: every documented failure mode is the violation of a condition the success stories depend on.
The psychological branch came first. Irving Janis’s Victims of Groupthink (1972) dissected how President Kennedy’s exceptionally capable advisers talked themselves into the 1961 Bay of Pigs invasion. Janis catalogued eight symptoms in three clusters — overestimation of the group (illusion of invulnerability, belief in inherent morality), closed-mindedness (collective rationalization, stereotyped views of outsiders), and pressure toward uniformity (self-censorship, an illusion of unanimity, direct pressure on dissenters, and self-appointed “mindguards” who filter disturbing information). Two years later, Jerry Harvey’s “Abilene Paradox” (Organizational Dynamics, 1974) described the stranger cousin: groups agreeing on a course of action that no individual member wanted, because everyone mistook everyone else’s silence for preference. Both are failures of the same input: honest, independent judgment never enters the pool.
The economic branch arrived in 1992, and it was more unsettling because it required no psychology at all. Sushil Bikhchandani, David Hirshleifer, and Ivo Welch (“A Theory of Fads, Fashion, Custom, and Cultural Change as Informational Cascades,” Journal of Political Economy, 100: 992–1026) and, independently, Abhijit Banerjee (“A Simple Model of Herd Behavior,” Quarterly Journal of Economics, 1992) showed that perfectly rational agents deciding in sequence will, after only a few observations, rationally ignore their own private information and copy their predecessors. The resulting information cascade is individually sensible and collectively fragile — it aggregates almost none of the group’s dispersed knowledge, and it can lock a whole population onto the wrong answer that a handful of early movers happened to pick. Lisa Anderson and Charles Holt reproduced cascades on demand in the laboratory (American Economic Review, 1997). The mechanism illuminates everything from asset bubbles to viral misinformation: Vosoughi, Roy, and Aral’s analysis of roughly 126,000 story cascades on Twitter (Science, 2018) found false news spread significantly farther, faster, and deeper than truth, with falsehoods about 70% more likely to be retweeted. The full failure catalogue — from tulip mania to 2008 — is in When Crowds Go Wrong.
Notice what both branches share. Groupthink destroys independence socially; cascades destroy it informationally. Either way, the crowd stops being many judgments and becomes one judgment wearing many faces — exactly the condition under which Condorcet’s theorem flips from a guarantee of wisdom into a guarantee of confident error. The shadow history is why the field’s practical wing became obsessed with process: collect judgments before exposing people to each other’s views, protect dissent structurally, and make the aggregation explicit.
2015–2023: The debate decade — replication, critique, and reply
Any measurement that surprising was going to be attacked, and the attack came in 2017. Marcus Credé and Garett Howardson (Journal of Applied Psychology) reanalyzed six samples and argued the empirical support for a general c-factor was weak — that group performance might be better described by task-specific abilities than by one underlying factor. Woolley, Kim, and Malone replied in 2018, disputing the critique’s scoring procedures and arguing its simulation assumptions did not match the tasks groups actually performed. This is what a healthy field looks like: the fight was conducted in data and reanalysis, not press releases.
The closest thing to a verdict arrived in 2021. Riedl, Kim, Gupta, Malone, and Woolley published a meta-analysis in PNAS spanning 22 studies and 1,356 groups — students, military personnel, online workers, gamers — and found consistent evidence for a collective intelligence factor across all of them, with group collaboration processes and member social perceptiveness among its strongest antecedents. The construct’s structure and boundary conditions remain debated, as they should be; but “group intelligence is measurable and is not reducible to member IQ” now rests on a much larger evidence base than the original 2010 paper.
Two complementary findings rounded out the decade. Malone’s Superminds (2018) supplied a taxonomy of the collective minds civilization already runs on — hierarchies, democracies, markets, communities, and ecosystems — each a different answer to the aggregation question, each with characteristic strengths and failure modes. And Google’s Project Aristotle (2012–2015), an internal study of 180 teams, found that psychological safety — the shared belief that a team is safe for interpersonal risk-taking — was the strongest predictor of team effectiveness, ahead of composition and seniority. Woolley’s turn-taking equality and Google’s psychological safety are two measurements of the same underlying truth: the bottleneck of group intelligence is whether information held by members actually enters the group. Group dynamics — safety, deliberation, and their pathologies like groupthink — are a cluster of their own, and we treat their failure modes in When Crowds Go Wrong.
2024–2026: AI-human superminds — the current frontier
The current era began when large language models joined the group. Burton and 27 co-authors across disciplines mapped the stakes in “How large language models can reshape collective intelligence” (Nature Human Behaviour, 8: 1643–1655, 2024): LLMs can lower barriers to participation, summarize sprawling deliberations, and aggregate dispersed knowledge at unprecedented scale — and they can also homogenize viewpoints, manufacture false consensus, and thin out the diversity of the information ecosystem that crowd wisdom feeds on. Every condition on Surowiecki’s list is simultaneously strengthened and threatened by the same technology.
The synergy question — Engelbart’s 1962 question — finally got its meta-analysis. Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone analyzed 106 experiments and 370 effect sizes (“When combinations of humans and AI are useful,” Nature Human Behaviour, 8: 2293–2303, 2024) and found that, on average, human-AI combinations performed worse than the best of the human or the AI alone. The losses concentrated in decision-making tasks; the genuine gains appeared in content-creation tasks. Simply putting a “human in the loop” does not improve outcomes — often the human overrides a better machine judgment, or defers to a worse one. Synergy is a design achievement, not a default. The emerging science of human-AI systems — automation bias, trust calibration, when to let which party decide — is another future chapter of this series.
That design agenda is now the explicit program of the MIT Center for Collective Intelligence, whose research is organized in 2026 under three umbrellas: Generative AI and Collective Intelligence (AI to augment human creativity), Supermind Design (designing innovative combinations of people and computers), and Designing Human-AI Teams (designing and allocating tasks in human-machine teams — pursued with the Singapore-MIT M3S program as a systematic science of task allocation: which subtasks go to people versus AI, who supervises, and whether the AI acts as tool, assistant, peer, or manager). The earlier flagship projects — the Climate CoLab and CoLabs platform, the CI Design Lab, the Measuring Collective Intelligence program — are now formally legacy work, absorbed into these AI-era questions.
The concrete outputs show what “designed synergy” looks like. Supermind Ideator, an LLM-based ideation tool that scaffolds users through structured creative problem-solving moves, produced significantly more innovative ideas in an experimental study than either ChatGPT alone or people working unaided (ACM Collective Intelligence Conference 2024; Collective Intelligence, 2024); beta testers have generated more than 10,000 ideas with it, and a sibling tool, DesignAID, applies the approach to design work. The scaffolding implements Supermind Design’s six moves — Zoom Out, Zoom In, Analogize, Groupify, Cognify, Technify — a checklist for reimagining who (or what) does each part of a task. The lesson mirrors the Vaccaro meta-analysis exactly: the value is not in the model, it is in the structure wrapped around it. On the measurement side, AGI-Elo (Sun and colleagues, NeurIPS 2025; arXiv 2505.12844) — with CCI-affiliated co-authors including Malone — rates AI models and individual test cases against each other like chess players, so a task’s difficulty and a model’s competency land on one comparable scale: a prerequisite for deciding which parts of a job to hand to the machine.
Two more 2026 results sketch where this is heading. A PNAS paper, “Advancing AI Negotiations: A Large-Scale Autonomous Negotiations Competition,” reported an international competition in which participant-designed AI agents conducted more than 180,000 negotiations with one another. Human negotiation theory survived the species change: agents exhibiting warmth — positivity, gratitude, question-asking — consistently achieved better outcomes on deals and value, while dominance-flavored marathon exchanges predicted impasses; the agents also evolved AI-native tactics, from chain-of-thought reasoning to attempted prompt injection. And a deep ontology of work (Cai and colleagues, arXiv 2603.20619, March 2026) reorganized roughly 20,000 O*NET work activities and mapped 13,275 AI applications plus a worldwide tally of 20.8 million robotic systems onto them — finding AI market value extraordinarily concentrated, with the top 1.6% of activities accounting for over 60% of it. Superminds of humans and machines are no longer a metaphor; they are an installed base being inventoried.
What 240 years teach us
Lay the eras side by side and one through-line emerges: every success in this history is a success of structure, and every failure is a failure of the same few conditions. Condorcet’s theorem runs on independence and better-than-chance competence. Galton’s fair happened to enforce independent, incentivized, written judgments. Hayek’s prices and Hanson’s market makers are aggregation mechanisms; Delphi is an independence-protection mechanism; Surowiecki’s four conditions are simply the recurring requirements named. The c-factor literature found that interaction process — turn-taking, social sensitivity, safety to speak — governs whether a group can use the intelligence it contains. And the AI-era meta-analyses found that human-AI teams win precisely when the collaboration is deliberately designed. When the conditions break — when judgments stop being independent and people start copying the visible behavior of others — you get information cascades, bubbles, and groupthink: individually rational, collectively wrong.
This is also, not coincidentally, a map of where collective intelligence sits in the life of a decision. Deciding well means attending to the right problem, framing it, generating options, investigating them, aggregating dispersed judgments, deliberating over reasons, and only then choosing, implementing, and monitoring. The 240-year research program is overwhelmingly about those middle stages — aggregation and deliberation — because that is where individual knowledge either becomes collective knowledge or dies in silence.
Argumentree is built as a direct application of these findings. Independent, asynchronous argument submission protects Condorcet’s and Delphi’s independence condition before the group converges. Structured pro/con argument trees give every perspective — including dissent — a first-class place, the condition Woolley’s turn-taking equality and Google’s psychological safety point to. Multi-dimensional rating (helpfulness, clarity, accuracy, completeness), aggregated into consensus scores, is an explicit aggregation mechanism in the Galton-Hayek lineage — except that unlike a market price, it preserves the reasoning. And the full audit trail turns each decision into an organizational record, so the group can learn from itself — the monitoring stage the lifecycle ends with. AI extraction from meeting transcripts follows the 2024–2026 lesson: use AI where it demonstrably helps (structuring and content processing) while humans retain the judgment. The history’s verdict is consistent from 1785 to 2026: groups do not get smart by accident. They get smart by structure — and the structure is now something you can adopt rather than improvise. Start with our guide to collective intelligence, or see how the principles run inside a real decision process in collaborative decision making.
Frequently Asked Questions
What is collective intelligence?
Collective intelligence is the capacity of groups to act in ways that seem intelligent — often more intelligent than any individual member. Malone and Bernstein’s Handbook of Collective Intelligence (MIT Press, 2015) defines it as “groups of individuals acting collectively in ways that seem intelligent.” It spans phenomena as different as a crowd accurately guessing an ox’s weight (Galton, 1907), Wikipedia’s collaborative editing, prediction markets aggregating beliefs into prices, and modern human-AI teams. The common thread is that individual judgments, knowledge, or actions are combined through some aggregation mechanism into a collective result.
Who invented the concept of collective intelligence?
No single person invented it, but the mathematical foundation is usually traced to the Marquis de Condorcet, whose 1785 jury theorem proved that a majority of independent, better-than-chance voters becomes almost certainly correct as the group grows. Francis Galton provided the first famous empirical demonstration in 1906-1907 with his “Vox Populi” analysis of 787 weight guesses at an English fair. The modern research field crystallized when Thomas Malone founded the MIT Center for Collective Intelligence in 2006, and James Surowiecki’s 2004 book The Wisdom of Crowds popularized the idea for a general audience.
What is the c-factor in collective intelligence research?
The c-factor is a measurable general collective intelligence factor for groups, analogous to the g-factor of individual IQ. Woolley, Chabris, Pentland, Hashmi, and Malone (Science, 2010) tested 699 people in small groups and found a single statistical factor explaining about 43% of the variance in group performance across diverse tasks. Strikingly, c correlated only weakly with members’ average or maximum individual intelligence — and strongly with average social sensitivity, equality of conversational turn-taking, and the proportion of women in the group.
Is the c-factor replicated, or is it contested?
Both. Credé and Howardson (Journal of Applied Psychology, 2017) reanalyzed six samples and argued the evidence for a general factor was weak; Woolley, Kim, and Malone replied in 2018, disputing the scoring and simulation assumptions. The largest evidence to date is a meta-analysis by Riedl and colleagues (PNAS, 2021) covering 22 studies and 1,356 groups, which found support for a c-factor across populations from students to military teams and online workers. The honest summary: a well-supported but still actively debated construct.
What are Surowiecki’s four conditions for crowd wisdom?
In The Wisdom of Crowds (2004), James Surowiecki argued that groups are collectively wise only when four conditions hold: diversity of opinion (each person brings private information or a different interpretation), independence (opinions are not dictated by others), decentralization (people draw on local knowledge), and aggregation (a mechanism turns private judgments into a collective decision). Remove any one and the crowd tends to get dumber, not smarter — which is why information cascades, echo chambers, and groupthink are so destructive.
What is the MIT Center for Collective Intelligence?
The MIT Center for Collective Intelligence (cci.mit.edu), founded at MIT Sloan and directed by Thomas Malone, studies how groups of people — and increasingly people plus AI — can be organized to act more intelligently than any individual, systems it calls superminds. Its crowdsourcing-era projects such as the Climate CoLab are now legacy work; as of 2026 its research runs under three umbrellas — Generative AI and Collective Intelligence, Supermind Design, and Designing Human-AI Teams — with tools like Supermind Ideator and DesignAID, the AGI-Elo rating system, and, with the Singapore-MIT M3S program, a systematic science of human-AI task allocation.
How is AI changing collective intelligence?
Two ways, in tension. Large language models can broaden participation, summarize huge deliberations, and aggregate dispersed knowledge — but Burton and colleagues (Nature Human Behaviour, 2024) warn they can also homogenize viewpoints and manufacture false consensus. And synergy is not automatic: a meta-analysis of 106 experiments by Vaccaro, Almaatouq, and Malone (Nature Human Behaviour, 2024) found that human-AI combinations on average performed worse than the best of human or AI alone, with losses concentrated in decision-making tasks. The design of the collaboration — who does what, and how judgments are combined — determines whether AI helps.
What is the difference between collective intelligence and the wisdom of crowds?
The wisdom of crowds is one specific form of collective intelligence: the statistical accuracy of aggregated independent judgments, as in Galton’s ox experiment or a prediction market. Collective intelligence is the broader umbrella — it also covers collaborative creation (Wikipedia, open source), deliberation and argumentation, swarm coordination in nature and robotics, and human-AI teams. Crowd wisdom relies on keeping judgments independent and then averaging them; other forms of collective intelligence rely on structured interaction, specialization, and deliberate aggregation mechanisms.
Argumentree Team
Decision Science
The Argumentree team explores the science of better decisions—from 18th-century mathematics to modern AI.
Put 240 years of research to work.
Argumentree structures your group’s reasoning the way the science says wise crowds work: independent input, explicit aggregation, and a permanent record of the arguments.
Start Free 14-Day Trial →Related Articles
Wisdom of Crowds: Galton's Ox — The 1906 Experiment That Proved Crowds Beat Experts
Diversity Beats Ability: Why Your Smartest Hire Isn't Your Best Decision-Maker
Collaborative Decision Making: 240 Years of Proof That Groups Beat Individuals
Across the ecosystemCollective AI Intelligence
The 240-year arc from Condorcet to today now extends to machines: many models reasoning together outperform one.

