Artificial intelligence now appears to have a substantial role in writing about one in 10 English-language webpages, according to new research.
The Pew Research Center published the findings on Aug 20 after examining nearly 490,000 webpages collected from the Common Crawl web archive. Researchers analyzed pages published between January 2021 and July 2026.
Additionally, Pew found much higher rates among pages published after ChatGPT launched in November 2022. About 35 per cent of those pages showed significant signs of AI authorship.
Researchers used Open Pangram, an AI detection model developed by Pangram Labs, to analyze the collected text. The software looks for statistical patterns across large amounts of writing rather than individual words or phrases.
However, researchers found substantial differences among different types of websites.
Pages using .com domains showed signs of AI writing at about 10 times the rate of .edu and .gov pages. Both institutional domain categories recorded rates near 1 per cent.
Meanwhile, approximately 4.6 per cent of pages on .org domains showed significant signs of AI involvement. The four domain categories had recorded similarly low rates in 2021.
Pew also tracked changes in writing patterns commonly associated with generative AI systems. Em dashes now appear about twice as frequently as they did before 2023.
In addition, use of the Oxford comma has increased by 63 per cent. Words including “delve,” “interplay” and “testament” have more than doubled in frequency.
Researchers also found that negative parallelism has nearly tripled since 2023. The construction contrasts two ideas using phrases such as “not just X, but Y.”
However, Pew said the writing pattern remains relatively uncommon despite its rapid increase.
Read more: Electricity, water concerns drive US data center backlash
Read more: Bedrock Robotics pushes the envelope with autonomous excavators
Differences among domains reflect organizational norms
The researchers cautioned that AI detection software can incorrectly classify both human and machine-generated writing. Consequently, a detection result does not prove that AI generated an entire webpage.
Some flagged material may instead involve human writers using AI tools to edit or assist their work.
Pangram Labs’ technology has also appeared in separate research examining AI use by American newspapers. That research detected AI-generated text in approximately 9 per cent of U.S. newspaper articles examined this year.
Further, AI detection technology itself could change as major developers explore more reliable identification methods.
Anthropic has reportedly worked on model-level text fingerprinting for output from its Claude AI models. Such systems could potentially identify AI-generated writing without relying entirely on linguistic patterns.
The differences among domains may partly reflect how organizations publish information. Academic and government websites commonly use editorial review, institutional approval and slower publishing schedules.
Conversely, commercial domains include news outlets, marketing operations and automated content businesses publishing material at much greater scale.
The change has accelerated rapidly. AI-associated writing on .com pages rose from roughly 1 per cent in January 2021 to 9.35 per cent by January 2026.
The findings could also complicate how developers train future AI models. Common Crawl says its archive provides an estimated 70 to 90 per cent of the training-data tokens used by nearly all major large language models.
Researchers have warned that repeatedly training AI systems on machine-generated material can degrade their performance, a phenomenon called model collapse.
University of Toronto computer engineering professor Nicolas Papernot compared the process to repeatedly photocopying a photocopy.
“You lose some of the information,” Papernot said. His research has also found that using synthetic data can reinforce existing errors, biases and unfairness.