Pew puts a number on it: 35% of pages published after ChatGPT show AI authorship
Pew Research ran nearly half a million pages from Common Crawl through a detector. Among pages published after November 2022, more than a third came back with significant signs of AI writing, and .com domains showed it at ten times the rate of .edu and .gov.
Source
How Much of the Internet Is Written With AI?Pew Research Center · Original published August 20, 2026
Pew Research put a number on something the web has felt for two years. Screening pages collected from the Common Crawl archive, the study found significant signs of AI authorship in 35% of pages published after ChatGPT's release in November 2022. In a random July 2026 sample of 10,000 pages, which necessarily includes pages written before AI writing tools existed, the figure was around 10%.
How the number was produced
Pew collected close to half a million English-language pages spanning roughly five years, starting a couple of years before ChatGPT shipped, and ran them through detection technology from Open Pangram. The two figures answer two different questions: 10% is how much of a random slice of the web looks machine-written today, and 35% is how much of the web written since November 2022 does.
The split by domain is the part worth keeping. Pages on .com showed signs of AI authorship at roughly ten times the rate of .edu and .gov, which both landed near 1%. Pages on .org came in at 4.6%. Commercial publishing is where the volume is, which is exactly where you would expect it.
- Random sample, July 2026: about 10% of pages show significant signs.
- Pages published after November 2022: 35%.
- .edu and .gov: about 1% each. .org: 4.6%. .com: roughly ten times the .edu rate.
- Detection ran on nearly 500,000 English-language pages from Common Crawl.
What the number cannot tell you
Detectors misclassify. Pangram and every tool like it will call some human pages machine-written, and the error is not evenly spread: heavily edited prose, technical writing, and non-native English all read as more machine-like to a classifier. Pew says so directly, and the honest reading is that the shape of the trend is solid while any single page verdict is not.
The study also tracked stylistic tells that have grown more common over the years, including em dashes, Oxford commas, and the it-is-not-X-it-is-Y construction. Those are correlations with machine drafting, not proof of it, and treating them as proof punishes people who simply write that way.
Why it matters for anyone publishing
The context is a web where Cloudflare has already reported bot traffic overtaking human traffic. Put the two findings side by side and much of the web is machine-written pages being read by machines, with the human audience a smaller share of both ends than the raw page count suggests.
For a publisher the practical consequence is not to avoid the tools. It is to be able to say what you did. A page that states how it was written, cites what it is based on, and gets corrected when it is wrong survives a filter that a page with none of that does not.
„Much of it is bots reading web pages written by other bots.“
Sources
Related
How to label AI content under Article 50, and which part of it is not your job
Article 50 of the EU AI Act has applied since August 2, 2026, and it binds anyone serving people in the Union, wherever the server is. Most of the panic is about the machine-readable marking requirement, which for a site owner who calls somebody else's API is somebody else's obligation. Here is what is actually yours: a chatbot that says what it is, published text that either carries a name or carries a label, and a deepfake that admits it.
Evropska komisijaverified

Gemma passes a billion downloads and 100,000 derived models
On August 20, Google announced that the Gemma family of open models has passed a billion downloads. In two years the community has published more than 100,000 derived models, and the most recent Kaggle competition drew more than 1,600 entries. An Awesome Gemma repository launched with the announcement, meant as an index of vetted projects, fine-tuned models, tutorials, and tools. The figure matters to someone building sites too, because a model you run on your own server stops being exotic, which makes text processing without sending data to a third party workable.
Googleverified

Google gives sites a button readers use to name them a preferred source
On August 20, Google published a button a publisher places on its own page, which a reader clicks once to add that site to their preferred sources in Search. The consequence is that the site's pages appear more prominently in Top Stories, in AI Overviews, and in AI Mode. Google says users have chosen more than 600,000 different sources so far. Two reader-side changes arrived with it: tuning topics in the Discover feed through a three-dot menu, and a customizable audio briefing in Google News for Android.
Googleverified