The most-quoted number from the largest study yet of AI's effect on science is 4.84: the multiple by which researchers who use AI tools out-cite those who do not. The number that should interest science policy is 4.63. That is the percentage by which the collective range of topics under study contracts as AI adoption spreads. One is a private return, the other a public cost, and the same Nature paper reports both. Qianyue Hao, Fengli Xu and Yong Li of Tsinghua University, with James Evans of the University of Chicago, published the analysis in January. It has been reorganizing the argument about AI and science ever since, including in Nature's own pages this week, where Evans resurfaced with a prescription.
The study is enormous and unusually blunt about its own machinery. The team trained a BERT-based classifier to read the text of 41.3 million natural-science papers indexed in OpenAlex and flag those whose methods were augmented by AI. Roughly 311,000 cleared the bar; validated against expert-labelled data, the classifier scored an F1 of 0.875. The individual-level results are lopsided. Scientists who adopt AI publish 3.02 times more papers a year and collect 4.84 times more citations. Junior researchers who adopt it become project leaders 1.37 years sooner, 7.33 years versus 8.70. AI-augmented work is 18.60% more likely to land in a first-quartile journal. Teams get smaller: 1.33 fewer people on average, with the junior ranks thinning fastest, from 2.89 members to 1.99.
None of that is a causal estimate, and the authors say so. There is no instrument here, no natural experiment, no difference-in-differences; the headline ratios come from group comparisons and t-tests across 27 million papers with intact reference records. The team does try to close the obvious hole, comparing scientists matched on early-career position and finding the gap survives, and it replicates on Web of Science and within the generative-AI era. But the limitations paragraph is candid: "despite consistently suggestive evidence, we cannot fully identify the causal linkage between AI adoption and scientific impact." Two further cracks are worth naming. The classifier reads what papers say they did, so quiet or unmentioned AI use contaminates the comparison group. And the career finding is thinner than the abstract implies: the 13.64% edge in a junior scientist's odds of reaching project lead held in only four of six disciplines, at p<0.2.
The citation multiple deserves the same scrutiny, and the paper supplies the ammunition. AI-augmented work clusters in data-rich, crowded topics inside high-impact journals, which is exactly where the citing population is largest. Reading 4.84x as 4.84 times the scientific value assumes citations track contribution rather than traffic. The authors' own diagnosis undercuts that reading. Embedding papers into a 768-dimensional space, they find AI-augmented work occupies a measurably smaller and lower-entropy region of knowledge in every field examined, while engagement between scientists falls 22.00%. They call the result "lonely crowds": popular topics drawing concentrated attention even as the papers citing the same work stop engaging with each other. The mechanism they propose is "collective hill-climbing," everyone scaling the same mountain by the same route, and research done "under the lamp post" of data-rich phenomena. AI, they conclude, "appears to drive problem solution over generation."
That framing has propagated fast. In a January News and Views, Georgia State's Veda Storey called it a paradox: AI adoption "expands scientists' impact but narrows the set of domains that research is carried out in." In June, Xizhe Zhang of Nanjing Medical University asked in Nature whether AI would bring a renaissance or "a diffuse monoculture," an outcome he argued depends less on model capability than on whether researchers, reviewers and funders reward originality over speed. In February, Cecilie Steenbuch Traberg of Copenhagen Business School, with Jon Roozenbeek and Sander van der Linden of Cambridge, mapped a self-reinforcing loop of topical, methodological and linguistic convergence they call epistemic monocropping, closing with the sharpest line in the literature: "In warning that AI might make us think alike, we may have begun to think alike about AI." On Monday, reviewing Alexander Krauss's The Engine of Scientific Discovery in Nature, Evans endorsed the claim that roughly a quarter of scientific fields are defined by the instruments that made them possible, and argued that designing new ones is where AI would earn its keep.
Why It Matters
The narrowing figure is small, and that is the point. A 4.63% contraction in topical breadth is nothing next to a threefold productivity gain, which is why no individual scientist will act on it. The incentives run entirely one way: more papers, more citations, faster promotion, smaller teams. The costs land on a commons nobody is scored against. That asymmetry makes monoculture a structural problem rather than a moral one, and every proposed fix is institutional: ring-fenced funding for data-poor questions, deliberate methodological rotation, reviewer pools drawn from outside the computational mainstream. A second-order worry sits in the team-size data. AI-augmented teams carry 31% fewer junior scientists, so a field that lets AI substitute for the graduate students who used to run the analysis trades near-term output for a thinner pipeline of people trained to ask questions a model cannot yet frame.
What To Watch
Three things. Whether anyone replicates the narrowing outside the natural sciences, which the authors explicitly exclude, and where Traberg and colleagues argue convergence is already further advanced. Whether the effect survives the generative-AI wave; the post-ChatGPT period is handled as a robustness check, not the main event, so the LLM-native cohort is barely measured. And whether funders move. Hao and colleagues end with a demand rather than a finding: "To preserve collective exploration in an era of AI use, we will need to reimagine AI systems that expand not only cognitive capacity but also sensory and experimental capacity," enabling scientists to gather new kinds of data rather than optimizing analysis of what already exists. That is a research agenda for AI labs, not a caution for scientists. Nobody has built the instrument it describes.
“In warning that AI might make us think alike, we may have begun to think alike about AI.”— Traberg, Roozenbeek and van der Linden, Copenhagen Business School and University of Cambridge