Researchers asked people in India and the United States to write short passages about their favorite food, favorite festival, favorite public figure, and a request for time away from work. Half wrote alone. The others wrote with an autocomplete system powered by GPT-4o.
Our buddy, ChattyG, had preferences.
When the prompt asked about food, its most common openings involved pizza and sushi. When it asked about a festival, it repeatedly proposed Christmas. For the Indian participants, the first food suggestion was always pizza or sushi, and the first festival suggestion was always Christmas. Many writers ignored the words hovering in front of them. The suggestions still arrived first.
But let’s look at something seemingly more innocuous. Indian participants who wrote about biryani without assistance mentioned Malabar preparation, nutmeg, raita, lemon pickle, date chutney, and the dish’s Mughal history. With ChattyG, biryani retained its name but became rich, aromatic, and filled with tender meat. Unaided descriptions of Diwali included worship, lamps, cows, firecrackers, and several days of ritual. Assisted descriptions leaned toward sweets, gifts, family gatherings, happiness, and warmth.1
The machine did not necessarily make the writing false. It did, however, make the truth more interchangeable. Biryani remained biryani. Diwali remained Diwali. The culturally specific object survived while the language around it moved toward a description that could be attached to almost any beloved meal or major holiday. The food was flavorful. The festival brought people together. Everyone was happy to be there.
Not a single thing in those sentences is objectionable. That is what gives them power.
We have become good at spotting the visible style of generated prose. Certain rhythms now travel through offices and schools with the efficiency of an invasive plant: the patient preamble, the tidy distinction, the paragraph that raises a question, answers it moderately, and closes by observing that the matter remains complex. There are also lists where every item is grammatically identical and emotionally denuded.
These tells are useful until they are noticed. Models change, users edit, and the obvious habits become unfashionable. The deeper aesthetic is harder to remove because it is the reason the prose works. A plausible answer includes the expected considerations, acknowledges uncertainty in familiar places, and reaches a conclusion proportionate to the evidence it has chosen to describe. Its movement is familiar enough that the reader never has to decide what kind of thing has arrived. The prose gives no trouble, creates no friction, draws no blood.
This is commonly described as blandness, an aesthetic party-foul. The culture will become more boring, everyone will write in one smooth professional dialect, and the novel will die beneath a deluge of competent paragraphs. Civilizations have survived a great deal of bad prose. The larger problem begins when the same systems help many people decide what matters, which objections belong, and what conclusion a reasonable person should reach.
Agreement contains information when separate observers approach a question through sufficiently different evidence, experience, and methods. If ten people independently reach the same conclusion, their convergence gives us some reason to believe they found something in the world rather than in one another. That word, “independently,” is doing nearly all the work.
Ten people repeating one source do not provide ten confirmations. Ten analysts applying the same flawed method do not create ten checks. Ten thermometers built with the same defect do not become more accurate when placed in a row. Generative AI can make judgments less independent without making them look copied.
The outputs appear under separate names in different documents and meetings. Each user supplies a different prompt, keeps different examples, and adjusts the tone. The finished recommendations may contain substantial human work and reflect real differences among their authors. The danger begins when the visible number of judgments exceeds the number of meaningfully separate searches that produced them. Nobody receives instructions from a central office. Everyone simply produces something that sounds about right.
You may have seen this phenomenon before the advent of AI.
The research is young and dispersed. The studies use writing experiments, creativity tasks, and model benchmarks, often with semantic distance as an imperfect proxy. None establishes a society-wide loss of independent thought. Together, they expose a distinction that ordinary performance measures miss: improving individual output and preserving collective independence are different achievements.
In a 2024 experiment, Anil Doshi and Oliver Hauser asked participants to write short stories with no AI assistance, one AI-generated idea, or a choice among several AI-generated ideas. Access to AI improved outside ratings of novelty, usefulness, writing quality, and enjoyment, with the largest gains among participants who scored lower on an initial creativity measure. The assisted stories were also more similar to one another.2
The individual writer “wrote” a better story. Writ large, the population received a narrower range of stories. No participant could experience the second result as a loss in the moment. The writer saw a blank page become manageable, and the reader got a more polished story. The reduction in variety appeared only when the outputs were placed beside one another.
A smaller study by Barrett Anderson, Jash Shah, and Max Kreminski found a related pattern. ChatGPT users produced more detailed ideas than users of a non-AI creativity aid, but ideas from different users were less semantically distinct. They also felt less responsible for what they produced.3 The machine can make each room larger and move the rooms closer together.
That sounds paradoxical because we usually treat creativity as a property of a person. Someone produces five ideas instead of four and develops them in greater detail, so creativity has increased. Collective creativity asks how far apart the searches conducted by different people were, and whether one person explored an area another missed.
The distinction matters outside art. Institutions rely on multiple people because people notice different things. One reviewer sees the methodological flaw. Another recognizes the historical analogy. A third has worked with the population being described and knows that the clean administrative category obscures reality. Analysts begin with different private inventories, including different experiences of failure and different suspicions about what a tidy dataset conceals. Even their mistakes can be useful when the mistakes are not shared. Agreement carries weight because the routes were capable of diverging.
The trade is difficult because the improvement is real and morally relevant. AI can give people access to forms of expression blocked by language conventions, disability, confidence, or time. Asking a weaker writer to refuse useful assistance so everyone else can enjoy a more varied culture would be a peculiar form of impressment into the human. A person does not owe the culture an awkward sentence merely because the awkwardness is distinctive.
The evidence also shows that convergence is not inevitable. A 2026 study comparing 102 people and 22 language models found that the models performed around the human level on individual originality while producing substantially more similar responses.4 Raising temperature increased variation, but at the highest setting much of the output deteriorated into unusable language. Five people shouting unrelated nonsense are unlikely to improve a meeting, though I have attended meetings where it might have helped.
Better-designed interventions produced better results. In one product-development experiment, ordinary pools of GPT-4 ideas were less diverse than ideas generated by human groups, while careful prompting increased dispersion.5 Another experiment involving more than eight hundred participants from over forty countries found that heavy exposure to AI-generated examples increased collective idea diversity, although it did not improve individual creativity on the study’s measure.6 Homogenization is not a substance that leaks automatically from a model into everyone who touches it. The effect depends on how the tool enters the work.
One confident completion can direct a writer toward a mode. A deliberately dispersed field of examples can interrupt an existing human convergence. Someone who forms a view and asks the model to attack it is having a different encounter from someone who begins tabula rasa and requests the best answer. Current interfaces often conceal the difference. A blank field invites a request, and the response arrives as a unit. The user experiences abundance because the machine can continue indefinitely, but endless alternatives can still orbit one big idea.
This is why sounding “about right” matters. A plausible answer supplies more than language. It offers a selection of what counts as relevant and an account of what completeness should look like. It decides which concern deserves a paragraph, which objection must be acknowledged, and what level of confidence is socially appropriate. Those choices are easy to mistake for neutral features of a well-formed answer because professional writing has trained readers to expect them.
In institutional analysis, the shape of completeness is often recognizable before the substance is tested. A proposal should contain benefits, risks, and implementation concerns. A failure should be translated into causes, lessons, and corrective actions. A contested policy should acknowledge tradeoffs and end with a calibrated recommendation. These are often sensible demands. They can also become a quiet theory of relevance, ruling observations in or out before the writer has decided what the subject requires.
Two answers can differ in wording while inheriting the same selection. One says the proposal requires a balanced approach. Another says the tradeoffs demand careful consideration. A third says success will depend on thoughtful implementation. Nobody has copied anyone. The room may still have learned from a common source which thoughts belong in the room.
The influence is difficult to see because the output contains the user. The examples may come from her work. She may delete sections, add an objection, correct facts, and replace the ending. The finished document can be genuinely hers while containing assumptions she never consciously selected because they arrived disguised as the ordinary contents of a complete answer. The model’s defaults become candidates for the room’s common sense.
The opening experiment makes this visible because cultural specificity is easy to name after it disappears. Malabar style becomes rich flavor. Ritual becomes celebration. The same reduction occurs less visibly in analysis. A policy choice becomes a balance between innovation and safeguards. An institutional failure becomes a need for clearer communication. A moral conflict becomes a tension among legitimate perspectives. These formulations may be accurate. They may also translate a resistant object into language the institution already knows how to process.
Biryani keeps its name while losing the details that made it this biryani. A policy problem can undergo the same treatment. Its population, history, and conflict remain named while their specific resistance is translated into the familiar language of tradeoffs and implementation. The subject survives as a noun. The judgment around it becomes portable.
A meeting can therefore contain ten intelligent people and less independent judgment than the head count suggests. Similar writing does not prove shared reasoning, and common model use does not erase differences in human knowledge or revision. Dependence is a matter of degree. The question is how much additional evidence each new judgment contributes after its common inputs are considered.
Research on correlated model errors shows why this matters. Elliot Kim, Avi Garg, Kenny Peng, and Nikhil Garg evaluated more than 350 language models using benchmark datasets and a resume-screening task. On one benchmark, when two models were both wrong, they selected the same wrong answer about 60 percent of the time. Larger and more accurate models continued to show highly correlated errors, including across distinct providers and architectures.7
Suppose five advisers each make a mistake ten percent of the time. If their errors are independent, disagreement can reveal danger. If all five tend to fail on the same cases, adding advisers creates confidence faster than protection. Their average accuracy can be excellent while the group retains a shared blind spot.
Consistency can be merciful. A shared system may reduce arbitrary treatment, prevent one manager’s private eccentricity from deciding a career, and give similar cases similar outcomes. Brian Hedden and Manish Raghavan have argued that many objections to algorithmic monoculture overstate the value of difference. Some forms of consistency reduce bias and noise, and diversity has no automatic value simply because it is human.8 Ten prejudiced managers do not become a wisdom-of-crowds mechanism.
The question is what survives after consistency improves. An institution still needs more than one way to detect the system’s remaining errors, especially when the accepted method encounters an unusual case or a changed environment. A consensus is worth more when its members could have been wrong in different ways.
The same tension appears in systems that teach. A study of roughly 52,000 users of an online chess platform found that AI feedback widened the skill gap and made players’ decision strategies less diverse.9 The players were learning in related directions. Chess would be ridiculous if players preserved losing openings to maintain intellectual biodiversity. The concern begins when the strongest teacher has blind spots, the environment changes, or continued exploration would have produced knowledge the dominant method cannot.
The principle is familiar in science. A replication using the same data, assumptions, and analytical pipeline provides less confirmation than one conducted through a method with different vulnerabilities. If every analyst asks a related model to summarize evidence, identify risks, and recommend an approach, their agreement may be sincere and the documents may still inherit a common map of relevance. The institution sees ten recommendations. It may have received fewer than ten independent searches.
A shared factual error can be checked, a fabricated citation can be opened, and a numerical mistake can be recalculated. A shared omission is harder to notice. Nobody sees the factor the model did not surface or asks the question the generated structure did not make room for. The reports agree because they have all been made complete according to a related idea of completeness.
The missing thing leaves no red underline. It is absent professionally.
Plausibility helps the omission pass because the answer contains what the reader expects. The risks are present. The language is careful. The conclusion follows. An unusual idea often begins by violating one of those expectations, perhaps by assigning importance to a detail everyone else treated as incidental or refusing a moderate conclusion because the underlying facts are not moderate.
Surprise does not make an idea true. Most strange claims deserve their obscurity. The world contains many people courageously resisting consensus because the consensus has asked them to stop emailing. A functioning culture needs filters. It also needs some way for a true observation to remain alive while it still sounds wrong.
Individual performance is easy to count. Collective independence requires examining outputs together, tracing shared inputs, and asking how many distinct searches the apparent agreement represents. An organization can improve every employee’s work and weaken its own ability to discover that everyone is wrong. No individual dashboard will show the loss.
Using several models or prompting for neglected evidence can recover meaningful variation. Those practices are useful, but ChattyG can play the dissenter while remaining the author of the disagreement. Real independence carries the possibility that another source rejects the framing or asks why the institution is solving this problem instead of another one.
That kind of difference is inefficient. It slows meetings and produces recommendations that are difficult to merge into one clean answer. The machine-assisted answer arrives ready to circulate; the independent one arrives needing explanation. Under time pressure, institutions will prefer the first for understandable reasons. Repeated preference teaches people which answers travel and makes friction look increasingly like low quality.
Plausibility is also a social signal. Generative AI can democratize it, which is one of its real gifts. People disadvantaged by language conventions, disability, or educational history can produce work institutions will read. The answer cannot be to take the uniform away from newcomers while established people retain the authority of cultivated prose. The challenge is to widen access to the signal without mistaking it for independent judgment.
Some tasks benefit from convergence. We want accurate arithmetic and consistent eligibility decisions. Other tasks require search. Scientific hypotheses, strategic forecasts, institutional diagnoses, and unfamiliar crises depend on people exploring different possibilities. A routine problem becomes novel when the environment changes, and the institution may suddenly need whatever differences it previously treated as noise.
Design can help without pretending that independence can be manufactured through a clever prompt. People can form an initial view before using the model, teams can preserve at least one separate evidence route, and reviewers can be told when apparent agreement rests on common tools. The goal is modest: do not mistake the number of documents for the number of judgments.
I have sat in versions of this meeting for years. A senior official asks several analysts to examine the same question, and the work returns in separate memos. The analysts have different expertise and different ideas about what will survive contact with the institution. When their recommendations converge, that convergence has traditionally carried information because it suggests that separate people searched the problem and found roughly the same landscape.
That inference was never automatic. Analysts have always shared sources, professional norms, and institutional habits. What has changed is that part of the common route can now pass through private exchanges the organization never sees. Each analyst may consult an AI system, reject some suggestions, add personal experience, and produce a thoughtful memo that is genuinely their own. The documents can differ in tone while sharing an earlier decision about which facts matter and what a complete answer should contain.
The official has no easy way to tell how much independence remains. The subject stays specific, just as biryani remained biryani, but the judgment surrounding it may have moved toward a common account of what matters. The memos carry ten names and an unknown number of independent searches. They may all be right. What has changed is how much their agreement proves.
* * *
The Second Order is a series within The Slow Panic. View the section alone at https://slowpanic.substack.com/s/the-second-order, or subscribe to the full publication.
Notes
1. Dhruv Agarwal, Mor Naaman, and Aditya Vashistha, “AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural Nuances,” in Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (New York: Association for Computing Machinery, 2025), 1–21, https://doi.org/10.1145/3706598.3713564.
2. Anil R. Doshi and Oliver P. Hauser, “Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content,” Science Advances 10, no. 28 (2024): eadn5290, https://doi.org/10.1126/sciadv.adn5290.
3. Barrett R. Anderson, Jash Hemant Shah, and Max Kreminski, “Homogenization Effects of Large Language Models on Human Creative Ideation,” in Proceedings of the 16th ACM Conference on Creativity & Cognition (New York: Association for Computing Machinery, 2024), 413–425, https://doi.org/10.1145/3635636.3656204.
4. Emily Wenger and Yoed N. Kenett, “Large Language Models Are Homogeneously Creative,” PNAS Nexus 5, no. 3 (2026): pgag042, https://doi.org/10.1093/pnasnexus/pgag042.
5. Lennart Meincke, Ethan R. Mollick, and Christian Terwiesch, “Prompting Diverse Ideas: Increasing AI Idea Variance,” arXiv:2402.01727, January 27, 2024, https://doi.org/10.48550/arXiv.2402.01727.
6. Joshua Ashkinaze, Julia Mendelsohn, Li Qiwei, Ceren Budak, and Eric Gilbert, “How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment,” Proceedings of the ACM Collective Intelligence Conference (2025), https://doi.org/10.1145/3715928.3737481.
7. Elliot Kim, Avi Garg, Kenny Peng, and Nikhil Garg, “Correlated Errors in Large Language Models,” paper presented at the 42nd International Conference on Machine Learning, 2025, arXiv:2506.07962, https://doi.org/10.48550/arXiv.2506.07962.
8. Brian Hedden and Manish Raghavan, “Algorithmic Monoculture and Its Critics,” arXiv:2604.06047, April 7, 2026, https://doi.org/10.48550/arXiv.2604.06047.
9. Christoph Riedl and Eric Bogert, “Who Benefits from AI? Self-Selection, Skill Gap, and the Hidden Costs of AI Feedback,” arXiv:2409.18660, rev. April 20, 2026, https://doi.org/10.48550/arXiv.2409.18660.


