One-Third of New Web Content Shows Signs of AI Authorship
Analysis of nearly 500,000 webpages reveals how generative AI tools have reshaped online writing since ChatGPT's 2022 debut.

One-Third of New Web Content Shows Signs of AI Authorship
More than one-third of webpages published after ChatGPT's November 2022 release show significant signs of AI authorship, according to new research analyzing nearly half a million English-language sites.
Pew Research Center examined webpages from the past five years using the Common Crawl web archive, running the text through an AI detection tool called Open Pangram. The analysis reveals a sharp upward trend that began precisely when OpenAI made ChatGPT publicly available and accelerated as competing tools like Claude and Gemini entered the market.
Looking at a July 2026 snapshot of the entire web—which includes both old and newly published material—10% of all pages displayed markers of AI authorship. But when researchers filtered for only content published after ChatGPT's launch, that figure jumped to over one-third.
Commercial Sites Lead AI Adoption
The distribution of AI-generated content varies dramatically across different corners of the internet. Commercial sites with .com domains show the highest concentration, with roughly one in ten pages displaying AI authorship signals in 2026 samples—double the rate found on .org domains at 4.6%.
Educational and government sites (.edu and .gov) show minimal AI adoption, with only about 1% of pages flagged. When ChatGPT first launched, AI linguistic patterns appeared at similar rates across all these domain types, suggesting the divergence reflects deliberate choices about where organizations deploy generative AI tools.
The Linguistic Fingerprints of AI
While AI detection models can misclassify individual documents, large-scale analysis reveals consistent patterns. AI models favor certain punctuation marks and phrasing conventions at rates that differ measurably from human writing.
Comparing 2026 web snapshots to 2023 baselines, researchers found em dashes now appear about twice as frequently across webpages. Oxford comma usage increased 63%. Certain words favored by language models—including "delve," "interplay," and "testament"—more than doubled in frequency.
AI-generated text also shows a preference for listing items in threes and using negative parallelism structures ("it's not just X, it's Y"), though the latter remains relatively rare overall. The detection model used in this research goes beyond these surface signals, identifying subtle statistical patterns in word choice and sentence structure that distinguish machine-generated prose.
Why It Matters
The rapid proliferation of AI-authored content fundamentally changes assumptions about online information. Business leaders making decisions based on market research, competitive intelligence, or customer sentiment increasingly encounter synthetic text that may reflect training data patterns rather than genuine human perspective. Organizations need strategies to assess source authenticity, particularly when AI-generated content can influence everything from SEO rankings to public discourse. The concentration of AI text on commercial domains suggests competitive pressure is driving adoption faster than quality considerations in some sectors.
This analysis was conducted by Pew Research Center's Data Labs and first reported by the organization in August 2026. The research was led by senior data scientist Samuel Bestvater.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call