frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

AI companies are shredding rare books

https://xcancel.com/HedgieMarkets/status/2081534588485296565
124•anon373839•46m ago

Comments

sherr•36m ago
I see mentions of Bradbury's "Fahrenheit 451" in that thread but what this really seems to be mostly like is Vernor Vinge's "shred and scan" factory in his novel "Rainbows End".
vessenes•26m ago
Perhaps the last great near-term predictor. I often wish he'd written more. To remind us all, he predicted shred and scan would be a short stop over done by villains on the way to nondestructive scanning.

That said, supporting Anna's archive is one of the best things you could do for humanity long term in my opinion.

orthoxerox•24m ago
Were they the villains? I remember the rogue three-letter-agency executive being the BBEG.
clickety_clack•23m ago
That’s digital though, so it requires the continued survival of readers for the data that is stored. The best thing you could for the long term is probably to buy a few hundred physical books to keep in a bookcase in your home.
thechao•33m ago
Which book that was rare was destroyed? I'm interested to know a few titles.
glimshe•27m ago
No evidence. But someone said it on the Internet so it must be true.
Cynddl•24m ago
The 404media article mentions notably https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi... which says:

> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).

infinite_spin•18m ago
> Barrett's Traditional Fairy Tales (2021)

How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.

ACCount37•1m ago
Niche text. It's not impossible that there was only ever under a thousand of them printed and released into circulation.

A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.

est31•26m ago
> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.

Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.

Scanning books you own should be legal from a copyright point of view, and not require shredding.

Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.

croes•20m ago
Books that are shredded can’t be scanned by competitors.
sethops1•18m ago
> Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.

https://www.404media.co/ai-companies-are-buying-tons-of-old-...

ACCount37•14m ago
Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.

stuartjohnson12•26m ago
I think people are, on the whole, too precious about old things. In the case of books produced after major commercial printing began, I don't believe it is the paper that imbues the book with historical value.

Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content!

There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who.

For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

Cynddl•22m ago
> For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

What makes you think they will? What would be the incentives for these companies to do so?

stuartjohnson12•8m ago
Well, you probably weren't going to go and find any rare, non-digitised books to physically go and read (unless you were going to, in which case, rock on), so we can start by benchmarking relative probability there.

1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.

2. Availability via Google Books or similar.

3. Availability via AI model reference.

4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.

I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.

ACCount37•25m ago
The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd.

Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.

What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.

SirFatty•23m ago
I see.. so the various AI companies are in the right on this?
master_crab•13m ago
It can be the case that everyone in “a fight” is wrong.
infinite_spin•9m ago
I think they are in the legal sense of right, and I think they only discarded the remains of these dissected books because previous rulings (e.g. archive.org's lending practices of digital copies of books they physically owned) gave rise to a situation where destruction bore less legal risk. As for the moral case, I don't have much to say on that, we all have our own lines in that sand.
tencentshill•7m ago
So they're not valuable... except to AI companies. They should pay a fair amount.
vessenes•24m ago
I don't see any proof of shredding here. Most book scanners I'm aware of are from Google's scanning days, and those had cameras plus page turning.

If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun.

Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.

lousken•23m ago
That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.
Incipient•17m ago
Publishers don't care if rare books get shredded?
azan_•13m ago
Yeah, why would it be bad for publishers? If anything they'd most likely encourage more book shredding!
the-grump•12m ago
And, regrettably, The Archive lent books regardless of physical possession.

Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.

I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.

It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.

infinite_spin•6m ago
of the rare books, which was the rarest of them all? What year was it published?
kingstnap•9m ago
The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction.

But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.

Good4boothee•19m ago
> A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time.

Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.

fmaccomber•18m ago
That's not an accurate characterization of the ruling
retinaros•16m ago
if they do this you can foresee what else they can do
azan_•9m ago
What else they can do based on this?
qsera•1m ago
If they are destroying old books, then it shows where their values lie...
timcobb•12m ago
This kind of reads like a blood libel. My guess is they're buying all those books that University libraries are throwing away these days (to see other HM threads for that), kinda sad but probably aren't "rare books" in the way people are thinking
azan_•11m ago
The comments there are absolutely unhinged. There are some good reasons for being anti-AI, but why dilute it with this kind of bullshit:

> It is equivalent to book burning in the past. A form of thought control

johnxianren•11m ago
I have zero proof for this, but just a what if: what if Anthropic's strict anti-China stance actually means the Chinese training corpus is way more valuable than people realize?
skybrian•7m ago
It’s unclear whether ISBNdb will scan books without ISBN’s, which were invented in the late 1960’s. Customers appear to be ordering books to be scanned by ISBN? Here is one book seller’s experience:

> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.

Article is paywalled, but I saved a few quotes here:

https://skybrian-links.exe.xyz/post/1026

hackernudes•7m ago
Also discussed here https://news.ycombinator.com/item?id=44381838 from June 2025.
enaaem•7m ago
This is the destruction of Western civilisation.
Springtime•6m ago
It seems a key contention of theirs is the possibility that rare books are being destroyed this way, yet the things they cite don't seem to suggest this (based on their paraphrasing), they just throw the following at the end to make it seem like it's occurring to irreplaceable books:

> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.

Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)

pu_pe•3m ago
From what I understand, the rare books in question are not some historically relevant medieval manuscripts, but rather some relatively recent books (still under copyright) for which there are few print copies available for purchase.

I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.

lejalv•1m ago
"To preserve them digitally"

For whom? is the relevant question

musha68k•2m ago
[delayed]
qsera•2m ago
I will make a robot scanner for books. I will then scan all the books in my state libraries and make a digital copy of them before these things come for them.

I wish...

mc32•6m ago
Old rare books where there are single digit copies should enjoy some sort of patrimonial protection just like museum pieces. You can own them but have the state have the option to buy it if you’re about to significantly deface it or destroy it.
croes•19m ago
So if the Mona Lisa is part of a model we can burn it?
infinite_spin•17m ago
If you purchased the Mona Lisa (or some rare book), in this hypothetical, you can burn it.
Invictus0•8m ago
> For instance, a painter may insist on proper attribution of their painting, and in some instances may sue the owner of the physical painting for destroying the painting even if the owner of the painting lawfully owned it.[1]

https://en.wikipedia.org/wiki/Visual_Artists_Rights_Act

infinite_spin•2m ago
The Mona Lisa's painter isn't alive, they can't sue, and this act doesn't apply to printed books.
stuartjohnson12•4m ago
I pre-empted this - my argument does not apply to texts where the physical object is a major part of the historical value of the thing. No, I'm not saying to destroy one of the four remaining Magna Cartas that were meticulously copied by hand. But even if I was, we're only dealing with texts here that are irrelevant enough to have never been digitised already - we tend to digitise most things of value and so the Mona Lisa and Magna Carta would never have been part of this discussion in the first place.

I am however OK with destroying one of the remaining 50 children's books of which only 300 copies were ever printed in a small town in Ohio in the 1970s as a test run for a failed book which was subsequently never commercialised.

chii•3m ago
> They should pay a fair amount.

they should pay the marginal value that the next buyer would buy.

Do you also think that a person dying of thirst ought to pay the maximum price they could possibly pay for water?

ACCount37•2m ago
They are paying a fair amount. In the ballpark of $5 per book.

You know, piracy online is nice and simple - but it's kind of hard to get physical media without paying what the previous owner considers "a fair amount" to part with it.

Kimi-K3 Releases on HuggingFace 7/27

https://huggingface.co/moonshotai/Kimi-K3
552•nateb2022•7h ago•254 comments

How is the Bun Rewrite in Rust going?

https://lockwood.dev/ai/2026/07/27/how-is-the-bun-rewrite-in-rust-going.html
159•tomlockwood•2h ago•99 comments

AI companies are shredding rare books

https://xcancel.com/HedgieMarkets/status/2081534588485296565
124•anon373839•46m ago•51 comments

Elevated errors on Claude Opus 5

https://status.claude.com/incidents/mfdtrknpxghq
31•croemer•1h ago•25 comments

PGSimCity - How PostgreSQL Works

https://nikolays.github.io/PGSimCity/
724•jonbaer•12h ago•69 comments

Magnolias Are So Old That They're Pollinated by Beetles, Not Bees

https://mymodernmet.com/magnolia-ancient-flowers-beetles/
123•speckx•4d ago•48 comments

Libsm64: Mario 64 as a library for use in external game engines

https://github.com/libsm64/libsm64
43•klaussilveira•3h ago•8 comments

Removing React.js from the codebase and adapting Htmx for UI interactivity

https://misago-project.org/t/removing-reactjs-from-the-codebase-and-adapting-htmx-for-ui-interact...
37•Ralfp•3h ago•16 comments

Building a Fast Lock-Free Queue in Modern C++ from Scratch

https://blog.jaysmito.dev/blog/04-fast-lockfree-queues/
26•ibobev•4d ago•6 comments

Show HN: Physically accurate black hole you can put in your room

https://blackhole.plav.in
381•aplavin•3d ago•129 comments

The Birth of the American 12-string Guitar

https://www.harpguitars.net/history/grunewald/12-string.htm
31•bilegeek•2h ago•15 comments

Towards a Theory of Bugs: The Ruliology of the Unexpected

https://writings.stephenwolfram.com/2026/07/towards-a-theory-of-bugs-the-ruliology-of-the-unexpec...
11•nsoonhui•3d ago•0 comments

Shay Locomotives

https://www.shaylocomotives.com/
23•Rygian•3h ago•5 comments

VLC for Unity now supported on Linux

https://code.videolan.org/videolan/vlc-unity
53•martz•4h ago•14 comments

The Proof Machine (2016)

https://incredible.pm/
5•BenoitP•49m ago•0 comments

Worse on Purpose

https://ledger.worseonpurpose.com/brands
25•bookofjoe•49m ago•4 comments

Modern email can be built from borrowed parts

https://en.andros.dev/blog/d7ed8b07/modern-email-can-be-built-from-borrowed-parts/
42•andros•4h ago•11 comments

Scriptc by Vercel: TypeScript-to-Native compiler, no JavaScript engine in binary

https://github.com/vercel-labs/scriptc
221•maxloh•14h ago•110 comments

Decker, a platform that builds on the legacy of Hypercard and classic macOS

https://beyondloom.com/decker/
337•tosh•18h ago•76 comments

US citizen charged after GrapheneOS phone wipes during airport search

https://www.techspot.com/news/113236-us-prosecutors-charge-atlanta-man-after-grapheneos-phone.html
949•eecc•14h ago•731 comments

Elevated errors on Claude Opus 5

https://status.claude.com/incidents/lhqp09kxq7pb
31•flyaway123•4h ago•16 comments

We have proof automation now

https://www.imperialviolet.org/2026/07/26/zstd-lean.html
195•zdw•16h ago•78 comments

I wanted a clock that never needed setting. Things escalated

https://arstechnica.com/gadgets/2026/07/i-wanted-a-clock-that-never-needed-setting-things-escalated/
139•lee_ars•4d ago•118 comments

Chinese chipmaker shares surge 470%

https://www.bbc.com/news/articles/c9q9w3x9qn2o
117•pingou•4h ago•103 comments

Introduction to Data-Oriented Design [pdf]

https://www.gamedevs.org/uploads/introduction-to-data-oriented-design.pdf
197•tosh•19h ago•53 comments

Google Chrome Arrives on ARM64 Linux, Widevine DRM Included

https://www.omgubuntu.co.uk/2026/07/chrome-arm64-linux-available
16•twapi•1h ago•5 comments

Measuring developer productivity with the DX Core 4

https://getdx.com/research/measuring-developer-productivity-with-the-dx-core-4/
31•saikatsg•3d ago•25 comments

How Unix spell ran in 64 kB of RAM

https://blog.codingconfessions.com/p/how-unix-spell-ran-in-64kb-ram
60•donw•4h ago•4 comments

Simulate cassette tape audio profiles using FFmpeg

https://github.com/AARomanov1985/Audio-Cassette-Simulation
150•xterminal•17h ago•65 comments

8086 Emulator Inside Scratch

https://turbowarp.org/1248315967?size=640x400
15•rickcarlino•4d ago•1 comments