Common Crawl: The 2007 Nonprofit Archive Behind Most AI Models, Now Facing Publisher Revolt Over Copyright

September 8, 2026 AI Angst avatar — a robot head with a distressed expression. JBS

A vast, glowing archive of stacked server racks shaped like an open book, with faint strands of light flowing out of it into logos representing several AI companies, set against a dim data-center background.AI Label

Almost nobody outside of AI research has heard of Common Crawl, yet a huge share of the AI models you've used were trained in part on its archive. This small nonprofit has quietly supplied the raw substrate behind some of the biggest names in AI for nearly two decades.

Since 2007, this small nonprofit has been quietly crawling the public web and giving away the archive for free. That obscurity ended in 2026, when some of the largest news publishers in the world decided Common Crawl wasn't a research project anymore. They called it a copyright problem, and they want it fixed.


What Common Crawl Actually Is

Common Crawl is a 501(c)(3) nonprofit, founded in 2007, that has crawled the public web for close to two decades and gives the resulting archive away for free.

Detail What It Is
Founded 2007, by Gil Elbaz, as a 501(c)(3) nonprofit
Mission Democratize access to web-scale data that was previously available only to large search engine companies
What it publishes A free, petabyte-scale archive of crawled web pages, indexes, and extracted text, hosted on Amazon S3
Crawler name CCBot, built on Apache Nutch
Who runs it Executive Director Rich Skrenta, previously founder of the Blekko search engine and the Open Directory Project
Annual operating cost Millions of dollars, per Skrenta, funded in part by major AI and tech companies

Why This Obscure Archive Actually Matters

Common Crawl isn't itself an AI company, and it doesn't train models, but it supplies the raw material that a huge share of the AI industry builds on. That role makes it one of the most consequential pieces of infrastructure in AI that almost nobody outside the field has heard of:

  • OpenAI's own GPT-3 research paper reported that Common Crawl-derived data made up the majority of that model's training tokens

  • Widely used cleaned datasets like FineWeb, C4, and OSCAR, staples of open-source model training, are all built directly on top of Common Crawl's raw archive

  • Its funders and downstream users include OpenAI, Google, Anthropic, Nvidia, Meta, and Amazon

That's what makes Common Crawl different from any single AI lab's crawler. Blocking GPTBot only affects OpenAI. Blocking, or not blocking, Common Crawl affects a meaningful share of the entire open-model ecosystem at once, since so many training pipelines start from the same archive.


How The Archive Is Actually Built

Common Crawl's crawler, CCBot, works from a frontier of more than 1 trillion known URLs but only samples a small fraction of it in any given cycle. Each monthly release typically captures somewhere between 2 and 5 billion pages, a mix of previously seen pages revisited to track how they've changed and newly discovered ones, rather than any attempt to exhaustively crawl the entire web at once.

  • The full historical archive now exceeds 300 billion captured pages, spanning more than a dozen petabytes of raw data collected since crawling began in 2008

  • To decide which pages are worth prioritizing out of that trillion-URL frontier, Common Crawl uses Harmonic Centrality, a graph-based measure of how close a page sits to the structural core of the web, alongside PageRank

  • The crawler focuses on text-based formats, primarily HTML and PDFs, rather than image or video files, which keeps the dataset's size manageable relative to the web it's sampling

  • CCBot publishes the IP ranges it crawls from, identifies itself by user-agent, and says it honors robots.txt directives and rate-limits requests to avoid overloading any single server

That last point is the organization's core pitch for why it should exist at all: a single, disciplined crawler respecting server limits is a lighter load on the web than dozens of AI companies each running their own uncoordinated scrapers against the same sites.


How Common Crawl Handles Sensitive And Illegal Content

Common Crawl says it does not run aggressive automatic filters across its raw archive, in order to preserve the dataset's research value, but it does act on specific reports of illegal or dangerous material. That reactive process, rather than proactive scanning of billions of pages, is also part of why complete removal of any single publisher's flagged content has proven slower than a one-time request might suggest.

  • The organization has described working with outside researchers to identify and remove child sexual abuse material and non-consensual intimate imagery when flagged

  • It has also said it removes exposed personal data covered by regulations like the GDPR, along with accidentally leaked cryptographic keys and credentials, once such material is brought to its attention

  • Publishers who prefer not to be archived at all can add their domains to Common Crawl's public Opt-Out Ledger, formerly called the Opt-Out Registry, so that downstream model developers can identify and exclude those sites from their own training runs

That ledger is real, and it's grown to include entries from the BBC, The Guardian, the Financial Times, The Washington Post, Reuters, and hundreds of other outlets. What it doesn't do, according to reporting cited later in this piece, is come with any enforcement mechanism forcing AI companies who already downloaded a dataset to actually delete a publisher's content from it.


The Publisher Revolt, Step By Step

What had been a slow-burning standards dispute turned into a formal legal confrontation over the course of about seven months in 2025 and 2026.

When What Happened
November 4, 2025 The Atlantic publishes an investigation questioning whether Common Crawl actually removes content after publishers request it. Common Crawl publishes a same-day rebuttal calling the characterization untrue.
April 29, 2026 The News/Media Alliance sends a formal letter to Rich Skrenta on behalf of publishers including NBCUniversal, CNN, McClatchy, Vox Media, Ziff Davis, and USA Today, citing removal requests that had gone unanswered for over two and a half years.
June 3-4, 2026 Digital Content Next sends a cease-and-desist on behalf of the Associated Press, the New York Times, NBCUniversal, Bloomberg, NPR, and Fox, arguing that copyright is not an opt-out regime and Common Crawl should need permission before including content at all.

That last point is the argument that changes the stakes. An opt-out system puts the burden on publishers to notice they've been scraped and ask to be removed. Digital Content Next's position flips that: Common Crawl should have to ask first. If that framing gains legal traction, it doesn't just affect Common Crawl, it questions the default assumption most web archiving and AI training has operated under for two decades.


Common Crawl's Defense

Rich Skrenta has pushed back directly and repeatedly against the accusation that Common Crawl has misled publishers about removing their content. In his November 2025 response to The Atlantic, he said the organization communicates honestly with publishers who contact it and has always operated in good faith and in public view, in line with its mission to serve the public good. He's rejected the idea that Common Crawl has become an extension of AI industry interests, pointing out that no donor, corporate or otherwise, controls what the archive collects, publishes, or removes.

His deeper argument is about the archive itself, not any single publisher's dispute: pulling historical material out of a shared, decades-spanning web archive doesn't just resolve one company's complaint, it starts dismantling one of the only comprehensive public records of what the web actually looked like at a given point in time. Once that precedent is set for one publisher, he's warned, it's set for everyone who wants their history rewritten too.

Both sides are defending something real. Publishers are right that their journalism trained systems that now compete with them for attention and revenue, without consent or payment. Common Crawl is right that a web archive which permanently deletes its own history on request stops being an archive. Neither of those truths cancels the other, which is exactly why this keeps ending in letters instead of resolutions.

The Part Neither Side Fully Controls

Even where everyone agrees on what should happen, the technical and structural reality is messier than a letter can fix. The New York Times and Danish publishers represented by the Danish Rights Alliance both secured agreements to have their content removed from Common Crawl's archives, and The Atlantic later reported that content from both was still turning up in the dataset months afterward. That's not necessarily bad faith, decades of crawl data spread across enormous archive files is genuinely difficult to scrub completely, but it means a removal agreement and an actual removal are currently two different things.

Digital Content Next's lawyers have gone further, saying they're reviewing whether Common Crawl's past statements about complying with removal requests, later followed by claims that technical costs and delays prevented full removal, may have been inaccurate or misleading. And the Opt-Out Ledger itself has drawn criticism for being easy to sign and hard to enforce: it's listed as one of dozens of subsections on Common Crawl's site, carries no directive to the developers actually downloading the data, and does nothing to un-scrape content that was captured before a publisher ever opted out.

Blocking future crawling is more straightforward, disallowing CCBot in robots.txt does stop new scraping, but it comes with its own cost. Research published in early 2026 found publishers who blocked AI crawlers saw roughly a 23% monthly traffic decline, without a matching drop in how often they were still cited by AI tools. Blocking protects future content. It doesn't recover past content, and it isn't free.


The Counter-Argument: Why Some Say Brands Should Stay Crawlable

Not everyone in this debate is arguing to be removed. A separate line of argument, mostly from marketers and AI-visibility consultants rather than publishers, says opting out of Common Crawl carries its own cost: invisibility to the conversational AI tools increasingly replacing traditional search.

The logic runs like this: as more people ask AI assistants for recommendations instead of typing queries into a search bar, a brand's visibility depends on whether it exists inside the datasets those assistants were trained on in the first place. A company that blocks CCBot doesn't just skip out on AI training, in this view, it risks having AI systems answer questions about its industry with no knowledge that the company exists at all.

It's worth naming as a real trade-off, and worth being clear about who's making the argument: companies selling "AI visibility" and "AI SEO" services have an obvious commercial interest in persuading brands to stay maximally crawlable. The traffic-decline data earlier in this piece cuts the other way for content creators specifically, so neither position is a free lunch, and the honest answer is that publishers and brands are currently choosing between two different kinds of exposure, not between exposure and safety.


What This Means For The Open Web Going Forward

Common Crawl was built on a premise that made sense in 2007, and that premise is what's now being tested. Crawling the web once and sharing the result saves everyone else from hammering the same servers with redundant scrapers, and keeps web-scale data out of the hands of only the largest search companies. That premise didn't anticipate a world where the same shared archive would become the default training substrate for an entire industry worth trillions of dollars, built on content its original authors never priced for that use.

Whatever gets resolved between Common Crawl and this wave of publisher demands will likely set the template other archives, libraries, and open datasets get judged against next. The open web's future partly depends on whether "open" can still mean "freely available to read" without also meaning "freely available to train billion-dollar AI models on," and right now, nobody involved has fully answered that question.


Common Crawl And The Open Web: FAQ

Common Crawl is a 501(c)(3) nonprofit founded in 2007 by Gil Elbaz that crawls the public web and freely publishes the resulting archive for anyone to download. The archive, currently more than 10 petabytes across over 300 billion pages, is hosted on Amazon S3. It was built to democratize access to web-scale data that was previously available only to large search engine companies. Its crawler identifies itself with the user-agent CCBot.

Common Crawl is the raw substrate underneath many of the datasets used to train large language models. OpenAI's GPT-3 paper reported that Common Crawl-derived data made up the majority of its training tokens. Widely used cleaned derivatives like FineWeb, C4, and OSCAR are all built on top of Common Crawl's raw archive, making it a foundational, if largely invisible, piece of AI infrastructure.

Publishers are asking Common Crawl to stop scraping their content, delete what it already scraped, and make its terms explicitly prohibit AI training use. The News Media Alliance's April 2026 letter and Digital Content Next's June 2026 cease-and-desist, sent on behalf of outlets including the New York Times, Associated Press, NBCUniversal, Bloomberg, NPR, and Fox, also ask Common Crawl to add enforceable protections to its Opt-Out Ledger.

Executive Director Rich Skrenta has disputed characterizations that Common Crawl misled publishers about content removal, saying the organization operates in good faith and in accordance with its mission to serve the public good. He has separately warned that removing already-archived material from a shared resource like Common Crawl threatens the open web's function as a historical record, even while acknowledging that running the organization costs millions of dollars a year and that some removal requests have taken years to fully process.

Only partially, and only going forward. Disallowing CCBot in robots.txt can stop future crawling, but it does not remove content already captured in past archives, and some publishers have reported that removal requests for historical content went unresolved for years. Separately, research has found that sites blocking AI crawlers saw meaningful traffic declines without a corresponding drop in how often they were cited by AI tools, since not all AI systems rely on fresh crawling to reference a site.

Common Crawl's total archive exceeds 300 billion captured pages across more than 10 petabytes of data, built from monthly crawls that each typically add 2 to 5 billion pages. The crawler works from a known frontier of over 1 trillion URLs but only samples a fraction each cycle, prioritizing pages using Harmonic Centrality and PageRank, and focuses on text-based formats like HTML and PDFs rather than images or video.

Common Crawl says it responds to specific reports of illegal or sensitive material rather than running blanket automated filters across its archive. The organization has described working with outside researchers to identify and remove child sexual abuse material and non-consensual intimate imagery when flagged, along with exposed personal data and leaked credentials brought to its attention, though this reactive approach means removal isn't instantaneous or guaranteed to be complete.


Jans Bock-Schroeder, AI Expert and Founder of AI Angst

Jans Bock-Schroeder

Publisher & Founder of AI Angst

Coming from the world of art, photography, and the luxury market, Jans launched AI Angst in 2025 to explore the cultural, ethical, and psychological impacts of artificial intelligence. His work bridges creative vision with critical technology analysis, offering clarity in an era of rapid technological change.


Sources and Citations

This article is based on the following sources:

  1. Common Crawl Foundation: "Setting the Record Straight" and organizational "About" page
    Primary source for Common Crawl's founding, mission, funding, and Rich Skrenta's response to publisher allegations.
    https://commoncrawl.org/blog/setting-the-record-straight-common-crawls-commitment-to-transparency-fair-use-and-the-public-good
  2. PPC Land: "News publishers target Common Crawl, the AI training data backdoor"
    Source for the News Media Alliance's April 29, 2026 letter and its list of publisher signatories.
    https://ppc.land/news-publishers-target-common-crawl-the-ai-training-data-backdoor/
  3. Digital Applied: "Publishers vs Common Crawl: AI Training-Data Showdown"
    Source for the Digital Content Next cease-and-desist letter of June 2026 and its "copyright is not an opt-out regime" argument.
    https://www.digitalapplied.com/blog/publishers-common-crawl-ai-training-data-showdown-2026
  4. Press Gazette: "US publishers tell Common Crawl to stop scraping and delete archive"
    Source for the New York Times and Danish Rights Alliance removal requests and The Atlantic's follow-up reporting.
    https://pressgazette.co.uk/media_law/common-crawl-ai-news-publishers-scraping-cease-and-desist-letter/
  5. Playwire: "Common Crawl Is an AI Training Pipeline. Publishers Are Done Pretending Otherwise."
    Source for the traffic-decline data on publishers that blocked AI crawlers and the critique of the Opt-Out Ledger's enforceability.
    https://www.playwire.com/blog/common-crawl-is-an-ai-training-pipeline-publishers-are-done-pretending-otherwise
  6. Common Crawl Foundation: "About" page and "Common Crawl Foundation Opt-Out Registry" blog post
    Primary source for crawl volume, total archive size, the 1-trillion-URL frontier, Harmonic Centrality prioritization, and the Opt-Out Ledger's scope.
    https://commoncrawl.org/about
  7. Search Engine Land: "Publishers push Common Crawl to stop collecting content for AI training"
    Source for Digital Content Next's concerns about the accuracy of Common Crawl's past compliance statements.
    https://searchengineland.com/publishers-common-crawl-content-ai-training-479831
  8. Editor and Publisher: "News publishers demand accountability from Common Crawl over unauthorized use of content"
    Source for the specific list of requested actions in the News Media Alliance's letter and comments from NMA president Danielle Coffey.
    https://www.editorandpublisher.com/stories/news-publishers-demand-accountability-from-common-crawl-over-unauthorized-use-of-content,261410

Published: September 8, 2026. Sources verified at time of publication. All external links open in a new tab.

A packed conference hall in Mumbai with delegates seated at rows of tables, laptops open, a large screen at the front showing a network topology map glowing over a stylized map of the Asia Pacific region.

APNIC 62: Over 600 Internet Experts Meet in Mumbai as AI Traffic Strains Global Infrastructure.


A glowing star-shaped AI core suspended above a dark control room, with faint chain-like lines of light around it fraying and dissolving into static on one side.

OpenAI Calls Astra The Start Of The AGI Era. Its Own CEO Calls "AGI" A Marketing Term.


Nvidia's green geometric logo mark merging into Hugging Face's yellow emoji-style logo at the center of the frame, set against a dark server-rack background with faint circuit-board lines.

Nvidia Just Bought The Place Where Open-Source AI Lives, For $12.9 Billion.