Blocking Internet Archive Won't Stop AI, but Will Erase Web's Historical Record

https://www.eff.org/deeplinks/2026/03/blocking-internet-archive-wont-stop-ai-it-will-erase-webs-historical-record

56•pabs3•4h ago

Comments

xnx•1h ago

Does Internet Archive have a distributed residential IP crawler program? I would enthusiastically contribute to that.

There must be some mechanism to prevent tampering in such a setup.

progval•24m ago

The Internet Archive does not, but Archive Team does: https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior

SlinkyOnStairs•21m ago

Devil's advocate: Anyone seeking to limit AI scraping doesn't have much of a choice in also blocking archivists.

And it's genuinely not that weird for news organisations to want to stop AI scraping. This is just a repeat of their fight with social media embedding.

Sure. The back catalogue should be as close to public domain as possible, libraries keeping those records is incredibly important for research.

But with current news, that becomes complicated as taking the articles and not paying the subscription (or viewing their ads) directly takes away the revenue streams that newsrooms rely on to produce the news. Hence the "Newspaper trying to ban linking" mess, which was never about the links themselves but about social media sites embedding the headline and a snippet, which in turn made all the users stop clicking through and "paying" for the article.

Social media relies on those newsrooms (same with really, most other kinds of websites) to provide a lot of their content. And AI relies on them for all of the training data (remember: "Synthetic data" does not appear ex nihilo) & to provide the news that the AI users request. We can't just let the newsrooms die. The newsroom hasn't been replaced itself, it's revenue has been destroyed.

---

And so, the question of archives pops up. Because yes, you can with some difficulty block out the AI bots, even the social media bots. A paywall suffices.

But this kills archiving. Yet if you whitelist the archives in some way, the AI scrapers will just pull their data out of the archive instead and the newsrooms still die. (Which also makes the archiving moot)

A compromise solution might be for archives to accept/publish things on a delay, keep the AI companies from taking the current news without paying up, but still granting everyone access to stuff from decades ago.

There's just major disagreement about what a reasonable delay is. Most major news orgs and other such IP-holders are pretty upset about AI firm's "steal first, ask permission later" approach. Several AI firms setting the standard that training data is to be paid for doesn't help here either. In paying for training data they've created a significant market for archives, and significant incentive to not make them publicly freely accessible.

Why would The Times ever hand over their catalogue to the Internet Archive if Amazon will pay them a significant sum of money for it? The greater good of all humanity? Good luck getting that from a dying industry.

---

Tangent: Another annoying wrinkle in the financial incentives here is that not all archiving organisations are engaging in fair play, which yet further pushes people to obstruct their work.

To cite a HN-relevant example: Source code archivist "Software Heritage" has long engaged in holding a copy of all the sourcecode they can get their hands on, regardless of it's license. If it's ever been on github, odds are they're distributing it. Even when licenses explicitly forbid that. (This is, of course, perfectly legal in the case of actual research and other fair use. But:)

They were notable involved in HuggingFace's "The Stack" project by sharing a their archives ... and received money from HuggingFace. While the latter is nominally a donation, this is in effect a sale.

---

I find it quite displeasing that the EFF fails to identify the incentives at play here. Simply trying to nag everyone into "doing the thing for the greater good!" is loathsome and doesn't work. Unless we change this incentive structure, the outcome won't change.

user_7832•19m ago

> But in recent months The New York Times began blocking the Archive from crawling its website, using technical measures that go beyond the web’s traditional robots.txt rules. That risks cutting off a record that historians and journalists have relied on for decades. Other newspapers, including The Guardian, seem to be following suit.

I'm a bit surprised I never read about this till now, though while disappointing it is unfortunately not surprising.

> The Times says the move is driven by concerns about AI companies scraping news content. Publishers seek control over how their work is used, and several—including the Times—are now suing AI companies over whether training models on copyrighted material violates the law. There’s a strong case that such training is fair use.

I suspect part of it might be these corps not wanting people to skip a paywall (whether or not someone would pay even if they had no access is a different story). But this argument makes no sense for the Guardian.

user_7832•16m ago

I went to Guardian's website to cross check their motto (getting confused with WaPo's motto) and got served this (hilarious? sad?) banner. As if blocking cross website tracking is somehow bad.

> Rejection hurts … You’ve chosen to reject third-party cookies while browsing our site. Not being able to use third party cookies means we make less from selling adverts to fund our journalism.

We believe that access to trustworthy, factual information is in the public good, which is why we keep our website open to all, without a paywall.

If you don’t want to receive personalised ads but would still like to help the Guardian produce great journalism 24/7, please support us today. It only takes a minute. Thank you.

OpenCode – Open source AI coding agent

Mamba-3

Atuin v18.13 – better search, a PTY proxy, and AI for your shell

FFmpeg 101 (2024)

A Japanese glossary of chopsticks faux pas (2022)

Molly Guard

Ghostling

Fujifilm X RAW STUDIO webapp clone

Linux Applications Programming by Example: The Fundamental APIs (2nd Edition)

Padel Chess – tactical simulator for padel

We rewrote our Rust WASM parser in TypeScript and it got faster

The Los Angeles Aqueduct Is Wild

We give every user SQL access to a shared ClickHouse cluster

Blocking Internet Archive Won't Stop AI, but Will Erase Web's Historical Record

Cryptography in Home Entertainment (2004)

Attention Residuals

The worst volume control UI in the world (2017)

The Ugliest Airplane: An Appreciation

An industrial piping contractor on Claude Code [video]

Show HN: We built a terminal-only Bluesky / AT Proto client written in Fortran

Turing Award Honors Bennett and Brassard for Quantum Information Science

France's aircraft carrier located in real time by Le Monde through fitness app

VisiCalc Reconstructed

The Story of Marina Abramovic and Ulay (2020)

Lent and Lisp

Why One Key Shouldn't Rule Them All: Threshold Signatures for the Rest of Us

Our commitment to Windows quality

ArXiv declares independence from Cornell

Entso-E final report on Iberian 2025 blackout

Delve – Fake Compliance as a Service

Blocking Internet Archive Won't Stop AI, but Will Erase Web's Historical Record

Comments

OpenCode – Open source AI coding agent

Mamba-3

Atuin v18.13 – better search, a PTY proxy, and AI for your shell

FFmpeg 101 (2024)

A Japanese glossary of chopsticks faux pas (2022)

Molly Guard

Ghostling

Fujifilm X RAW STUDIO webapp clone

Linux Applications Programming by Example: The Fundamental APIs (2nd Edition)

Padel Chess – tactical simulator for padel

We rewrote our Rust WASM parser in TypeScript and it got faster

The Los Angeles Aqueduct Is Wild

We give every user SQL access to a shared ClickHouse cluster

Blocking Internet Archive Won't Stop AI, but Will Erase Web's Historical Record

Cryptography in Home Entertainment (2004)

Attention Residuals

The worst volume control UI in the world (2017)

The Ugliest Airplane: An Appreciation

An industrial piping contractor on Claude Code [video]

Show HN: We built a terminal-only Bluesky / AT Proto client written in Fortran

Turing Award Honors Bennett and Brassard for Quantum Information Science

France's aircraft carrier located in real time by Le Monde through fitness app

VisiCalc Reconstructed

The Story of Marina Abramovic and Ulay (2020)

Lent and Lisp

Why One Key Shouldn't Rule Them All: Threshold Signatures for the Rest of Us

Our commitment to Windows quality

ArXiv declares independence from Cornell

Entso-E final report on Iberian 2025 blackout

Delve – Fake Compliance as a Service