کمپیوٹنگ تے مصنوعی ذہانتپری پرنٹتجربہپڑھن لئی ۳ منٹ

حالے ترجمہ نئیں ہویا: اصل انگریزی لکھت۔

NEARLY A THIRD OF THE WEB IS NOW WRITTEN BY AI

Large language models learn mostly from text gathered on the web. More and more of that text is itself written by AI. The authors call it “wild” AI text: written by many different models for human readers, published online, and swept into training data mixed with human text and unlabelled. It is neither “model collapse”, where a model is retrained on its own output, nor curated synthetic data produced on purpose to help training.

Jenna Russell, Mohit Iyyer and colleagues at the University of Maryland and at Pangram Labs, a company that sells an AI-text detector, set out to measure how much of this text there is and what it does to a model.

How much AI is on the web?

Using the FineWeb pipeline, which filters the huge Common Crawl archive of the web, the team rebuilt the corpus up to June 2026: 96 million documents, 83 billion tokens, the word fragments a model reads. Each document was labelled with the company’s detector, Pangram, whose reported false-positive rate is 0.05%; on 60,000 documents from 2021, before ChatGPT’s release in late 2022, it flagged only 0.062% as AI.

Share of AI-written tokens among web text that passes FineWeb’s quality filters:

  • under 0.1% in June 2021;
  • 10.1% in June 2024, 16.1% in June 2025;
  • 27.5% in June 2026, 31.1% in August 2026;
  • forecast: 42.3% at the end of 2027, 50.7% at the end of 2028.

The rise is uneven. In 2026 about half of tutorial and “knowledge article” tokens are AI-labelled, against 9% of news and 5% of personal blogs. Worse, the quality filters used to build training sets favour AI text: FineWeb keeps AI documents 2.3 times as often as human ones, and the DCLM filter 9.8 times as often.

800 models put to the test

The team pretrained 800 small language models, from 19.9 million to 973 million parameters, adding from zero to 64 AI tokens per human token to a fixed body of human text, and compared each with the same model given the same number of fresh human tokens instead.

  • AI text helped only data-starved models, below about 10 human tokens per parameter.
  • From the usual compute-optimal budget of about 20 tokens per parameter, adding AI text raised the error on human text, while extra human text kept lowering it.
  • The best share of AI text for predicting human text fell from 37% at 5 tokens per parameter to 1% at 20. For predicting AI text, it stayed above 90%.
  • Repeating the same human text beat adding fresh AI text: 2.3% lower error after one extra pass, about 6.3% after eight.
  • Models trained on AI text started writing like AI: phrases such as “delve” and “tapestry” rose from 0.16 to 0.54 per 1,000 words.

Existing “scaling laws” — formulas that predict a model’s error from its size and data — treat an AI token like a human one and get this wrong. The authors propose a new law with a benefit that saturates and a harm that grows logarithmically, so the value of an AI token can turn negative. Fitted on models up to 268 million parameters, it predicts models 3.6 times larger with 41% less error than the best existing law, although the appendix notes its lead is not significant on two test sets at low AI shares.

The bill for not filtering

According to this law, training on unfiltered web text at August 2026’s AI share costs 1.6 times the computing power of training on its human part alone — 3.0 times at the share forecast for 2028. And the damage is easy to miss: AI text is easier to predict, so a validation set with 22.3% AI text reported 95.5% of harmful runs as improvements. Standard benchmark scores also rose about as much with AI text as with human text.

The authors recommend filtering AI text when the goal is human text, repeating human data before adding AI data, and reporting errors on human and AI text separately. The limits are real: models no larger than 973 million parameters, English only, pretraining only, and a single detector, made by the authors’ company — whose business these conclusions favour. To allow independent checks, they release the corpus, all 800 models and the code.

Legal notice