Skip to content
Industrystrong signalverified

Pew puts a number on it: 35% of pages published after ChatGPT show AI authorship

Pew Research ran nearly half a million pages from Common Crawl through a detector. Among pages published after November 2022, more than a third came back with significant signs of AI writing, and .com domains showed it at ten times the rate of .edu and .gov.

By Redakcija WebAiRadarPublished 2 min readwritten by a model

Source

How Much of the Internet Is Written With AI?

Pew Research Center · Original published August 20, 2026

Pew Research put a number on something the web has felt for two years. Screening pages collected from the Common Crawl archive, the study found significant signs of AI authorship in 35% of pages published after ChatGPT's release in November 2022. In a random July 2026 sample of 10,000 pages, which necessarily includes pages written before AI writing tools existed, the figure was around 10%.

How the number was produced

Pew collected close to half a million English-language pages spanning roughly five years, starting a couple of years before ChatGPT shipped, and ran them through detection technology from Open Pangram. The two figures answer two different questions: 10% is how much of a random slice of the web looks machine-written today, and 35% is how much of the web written since November 2022 does.

The split by domain is the part worth keeping. Pages on .com showed signs of AI authorship at roughly ten times the rate of .edu and .gov, which both landed near 1%. Pages on .org came in at 4.6%. Commercial publishing is where the volume is, which is exactly where you would expect it.

  • Random sample, July 2026: about 10% of pages show significant signs.
  • Pages published after November 2022: 35%.
  • .edu and .gov: about 1% each. .org: 4.6%. .com: roughly ten times the .edu rate.
  • Detection ran on nearly 500,000 English-language pages from Common Crawl.

What the number cannot tell you

Detectors misclassify. Pangram and every tool like it will call some human pages machine-written, and the error is not evenly spread: heavily edited prose, technical writing, and non-native English all read as more machine-like to a classifier. Pew says so directly, and the honest reading is that the shape of the trend is solid while any single page verdict is not.

The study also tracked stylistic tells that have grown more common over the years, including em dashes, Oxford commas, and the it-is-not-X-it-is-Y construction. Those are correlations with machine drafting, not proof of it, and treating them as proof punishes people who simply write that way.

Why it matters for anyone publishing

The context is a web where Cloudflare has already reported bot traffic overtaking human traffic. Put the two findings side by side and much of the web is machine-written pages being read by machines, with the human audience a smaller share of both ends than the raw page count suggests.

For a publisher the practical consequence is not to avoid the tools. It is to be able to say what you did. A page that states how it was written, cites what it is based on, and gets corrected when it is wrong survives a filter that a page with none of that does not.

Much of it is bots reading web pages written by other bots.
TechCrunch, August 20, 2026

Sources

BrandsChatGPT

Related

EU AI ACT50the article in force since August 2, 2026
Industrystrong signal

How to label AI content under Article 50, and which part of it is not your job

Article 50 of the EU AI Act has applied since August 2, 2026, and it binds anyone serving people in the Union, wherever the server is. Most of the panic is about the machine-readable marking requirement, which for a site owner who calls somebody else's API is somebody else's obligation. Here is what is actually yours: a chatbot that says what it is, published text that either carries a name or carries a label, and a deepfake that admits it.

Evropska komisijaverified

The Gemma wordmark over a starfield, with the line "1 billion downloads" below it.
Industrymedium signal

Gemma passes a billion downloads and 100,000 derived models

On August 20, Google announced that the Gemma family of open models has passed a billion downloads. In two years the community has published more than 100,000 derived models, and the most recent Kaggle competition drew more than 1,600 entries. An Awesome Gemma repository launched with the announcement, meant as an index of vetted projects, fine-tuned models, tutorials, and tools. The figure matters to someone building sites too, because a model you run on your own server stops being exotic, which makes text processing without sending data to a third party workable.

Googleverified

A phone showing a news article with a black "Add to Preferred Sources" button; behind it a woman holding a coffee cup reads her own phone.
Industrymedium signal

Google gives sites a button readers use to name them a preferred source

On August 20, Google published a button a publisher places on its own page, which a reader clicks once to add that site to their preferred sources in Search. The consequence is that the site's pages appear more prominently in Top Stories, in AI Overviews, and in AI Mode. Google says users have chosen more than 600,000 different sources so far. Two reader-side changes arrived with it: tuning topics in the Discover feed through a three-dot menu, and a customizable audio briefing in Google News for Android.

Googleverified