الحوسبة والذكاء الاصطناعينسخة أوليةتحليل بياناتمدة القراءة: 3 د

لم يُترجَم بعد: النص الأصلي بالإنجليزية.

IS SCIENCE STARTING TO REPEAT ITSELF?

The scientific literature keeps growing — by roughly 4% a year, doubling about every 17 years, according to the figures cited in the paper. A larger literature could mean a wider frontier of ideas. It could also mean more papers that repeat what was already written. Ilan Doron-Arad and Elchanan Mossel, mathematicians at MIT, set out to measure one side of that question.

Originality has many dimensions and no agreed measure. But one of them can now be measured at scale: textual distinctiveness, how much of a paper reads like recent work. Modern language models can turn a passage into a list of numbers that captures its meaning, so that two passages saying the same thing in different words end up close together.

The authors are careful about what this does and does not capture. A genuinely new finding can be written with familiar sentences; an unusual phrase is not a new idea.

Holding the comparison fixed

A naive comparison would be misleading: as the literature grows, any passage has more chances to resemble something. So the authors kept the comparison pool constant. For each of eight fields — astrophysics, bioinformatics, chemistry, climate change, machine learning, medicine, psychology and quantum computing — and each year from 2015 to 2025, they took 300 papers from the open-access database CORE and compared them with 3,000 papers from the five previous years. In total, about 26,000 full texts.

Each paper was cut into overlapping passages of 160 words. A match required a high similarity of meaning, at least six shared specific terms, no identical run of five words — so plain copying is excluded — no shared author, and not two versions of the same paper. On a test set built for the purpose, the detector made almost no false matches (99.6% precision) while missing some true ones.

Stable, then a sharp rise

Until 2020, the share of a paper’s text matching recent work hovered around 0.43%. From 2021 it climbed, reaching 1.34% in 2025 — 3.1 times the earlier level. A statistical test places the start of the rise in 2021.

Two line charts: the share of text matched and the share of papers with a match, both flat until 2020 then rising steeply to 2025.

Left: mean share of each paper’s text resembling recent work. Right: share of papers with at least one such passage. — Figure 1a–b, Doron-Arad & Mossel (2026), arXiv:2609.32963.

The rise did not come from a few heavy borrowers. The share of papers containing at least one echoing passage reached 25% in 2025. All eight fields increased, at different rates.

Echoes without citations

Did authors simply cite the work they echo? Rarely. A citation link could be found for only 6.4% of matched pairs of papers, and a manual check of 176 pairs gave 7%. Between 2015 and 2025, matched pairs grew from 449 to 1,831 — but those with a citation stayed at 59 both years.

Bar chart of matched paper pairs per year, mostly without a detected citation, growing sharply after 2021.

Number of matched pairs of papers each year, by citation status: nearly all the growth is in pairs with no detected citation. — Figure 3a, Doron-Arad & Mossel (2026), arXiv:2609.32963.

The team also asked whether the ideas repeat, not just the wording. AI models from Anthropic summarised each passage in a few words, then judged, without seeing the original text, whether two summaries expressed essentially the same scientific claim. Using a strict definition, such pairs rose from an average of about 26 per year in 2015–2020 to 139 in 2025.

The pattern held under many checks — different thresholds, passage lengths, removal of methods sections and of the most-copied sources — and it replicated in a second, independent database, Europe PMC, where the share of echoing text doubled.

A suspect, not a verdict

The largest increase coincides with the rapid adoption of AI tools for writing science, and the authors see them as a plausible contributor, since they tend to spread the most common phrasing. But the study does not establish the cause. Shared disciplinary language and the narrowing of attention to the same few papers may also play a part; on this view, AI would be the latest accelerator of a convergence already under way. The authors also stress that semantic similarity does not imply misconduct.

They suggest tools that make originality visible to authors, reviewers, editors and funders — with one warning: never make a distinctiveness score a target, since it would reward unusual wording over real insight, and would be easy to game.

Legal notice