35% of Web Pages Published Since ChatGPT Launch Show AI Signs
A Pew Research Center study reveals a massive shift toward synthetic content, particularly within commercial domains.
The public internet is undergoing a fundamental transformation in how content is produced, with machine-generated text now appearing in a significant portion of new web pages. A study released August 20, 2026, by the Pew Research Center found that 35% of English-language web pages published since the launch of ChatGPT show significant signs of AI authorship or substantial editing.
To reach this conclusion, researchers analyzed nearly 500,000 English-language pages sourced from the Common Crawl archive. The study employed Open Pangram detection technology to identify linguistic patterns typical of large language models. The findings reveal a stark divide in adoption based on the nature of the website: AI authorship is most prevalent in .com domains, where approximately 10% of 2026 samples showed AI signs. In contrast, .edu and .gov domains remained largely human-authored, with AI prevalence hovering around only 1%.
The Linguistic Fingerprint of AI
The shift toward synthetic content is visible through specific "tells" that have surged since 2023. According to the Pew Research Center, the frequency of em dashes has doubled, and the use of Oxford commas has increased by 63%. More tellingly, the use of specific vocabulary favored by AI models has skyrocketed; words such as "delve" and "tapestry" have more than doubled in frequency across the analyzed pages.
This data arrives nearly four years after ChatGPT's November 2022 debut, a period during which AI integration has moved from novelty to utility. The Pew study quantifies the result of this behavioral shift, showing that the tools used by individuals are now scaling into the infrastructure of the web itself.
Implications for Digital Trust
The concentration of AI content in commercial .com domains suggests that businesses are automating information production far more aggressively than academic or government institutions. This divergence creates a tiered internet where official and educational records remain human-centric, while the commercial web becomes increasingly synthetic.
This trend raises urgent questions regarding authenticity and the long-term health of the information ecosystem. Industry experts warn of a potential "feedback loop," where future AI models are trained on a web increasingly populated by synthetic data rather than human thought. If the majority of new content is machine-generated, the diversity and accuracy of the training sets for the next generation of AI could be compromised.
What Remains Unconfirmed
While the study provides a broad snapshot of the current landscape, the exact ratio of fully AI-generated pages versus those that were merely AI-edited remains difficult to parse. Furthermore, as detection technology evolves, it remains to be seen if AI models will adapt their linguistic patterns to avoid the "tells" identified by Open Pangram, potentially masking the true scale of synthetic authorship in the years to come.