frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

AI companies destroy physical books – let's scan rare books before it's too late

https://annas-archive.gl/blog/physical-destruction.html
44•Cider9986•1h ago

Comments

ezfe•16m ago
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.

Instead, they enforce the copyright and force AI companies to shred books they want to ingest.

edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.

breezybottom•12m ago
They don't "force" anything. Trillion dollar AI companies and their owners have as much agency as book publishers.
tptacek•8m ago
To do what?
zmmmmm•11m ago
It is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights.

I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.

alightsoul•10m ago
Ai companies don't use ebooks, because they are more expensive than second hand books
breezybottom•6m ago
They absolutely do. Meta torrented 81 terabytes of ebooks. They just have no incentive to pay when the law looks the other way.
jacobo37•11m ago
this is plainly stupid ... many of these books are likely to have no current publisher nor any way to "reprint" the book. "ai" companies are simply burning our cultural context ...
HedonicEscal8r•8m ago
If only this complaint was being posted by an organization ideologically opposed to copyright itself!
RajT88•7m ago
The articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed.

These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.

imperio59•16m ago
Getting 10 million people to do anything is really, really hard. Getting 10 million people to spend hours scanning a book (which takes a really long time with a home scanner) sounds impossible :(
alightsoul•12m ago
No it is not. https://reddit.com/r/Annas_Archive/comments/1vrvt9e/athome_s...
xvxvx•9m ago
Pretty funny that they just took Anna’s archive and ingested it.

As for the story: they make it sound like AI companies are buying up all existing copies of rare books and stealing the knowledge, which isn’t the case, as far as I know.

HedonicEscal8r•8m ago
The piracy organizations are playing 4D chess while everyone else is playing checkers. The irony of this entire situation - AI companies being legally required to shred books due to kafkaesque copyright laws, then used as a marketing tactic by Anna's Archive - is a work of art.

I support Anna's Archive, by the way. Information wants to be free.

shakna•7m ago
[delayed]
landgenoot•4m ago
Isn't this a matter of regulation? I'm not sure about US, but in EU you have old houses/buildings that are protected. Sure, you can buy them, but you can't modify or destroy them (being cultural heritage).

AI companies destroy physical books – let's scan rare books before it's too late

https://annas-archive.gl/blog/physical-destruction.html
53•Cider9986•1h ago•21 comments

The August 17 outage

https://github.blog/news-insights/company-news/the-august-17-outage-and-the-work-ahead/
375•0xedb•8h ago•420 comments

I like 'em thick: an apology to my English teachers

https://www.experimental-history.com/p/i-like-em-thick
602•Ariarule•2d ago•266 comments

There's no such thing as a small software team anymore

https://jacob.gold/posts/theres-no-such-thing-as-a-small-software-team/
36•mooreslaw•3h ago•69 comments

HTML Can Do That

https://chrisburnell.com/html-can-do-that/
617•encyclopedism•1d ago•168 comments

Malicious Rust crate Arrayref runs a build-time payload

https://safedep.io/arrayref-proc-macro1-rust-build-time-malware/
415•abhisek•14h ago•374 comments

I should have loved biology (2020)

https://jsomers.net/i-should-have-loved-biology/
206•tyre•9h ago•78 comments

Make a 6-Tesla-class high-temperature superconducting dipole magnet at 4.2 K

https://journals.aps.org/prab/abstract/10.1103/4nhs-bkwh
17•supermagnet•6d ago•1 comments

Why aren't smart people happier? (2022)

https://www.experimental-history.com/p/why-arent-smart-people-happier
107•rafaelc•8h ago•159 comments

CIA funding helped keep NeXT afloat in the 80s

https://www.wsj.com/tech/steve-jobs-apple-next-cia-161b65f9?st=NWWds1&reflink=desktopwebshare_per...
361•EwanG•1d ago•220 comments

Show HN: I trained a 125M model to autocomplete piano on-device

https://simedw.com/2026/08/20/midi-autocomplete/
530•simedw•15h ago•110 comments

Vomit: Clean up Claude 5's token output with a separate LLM

https://github.com/zachahn/vomit
199•Bluestein•12h ago•215 comments

Linux 7.2

https://www.igalia.com/2026/08/19/Linux-72-Released.html
208•mariuz•11h ago•73 comments

Show HN: Huzzah – a novel approach to coding with AI

https://www.danielvaughn.dev/posts/huzzah/
236•danielvaughn•8h ago•137 comments

Detecting scraper bots through scroll behaviour

https://niki.cat/detecting-scraper-bots-through-scroll-behaviour
29•theanonymousone•4h ago•10 comments

Speeding Up (Small) Ruby Hashes

https://byroot.github.io/ruby/performance/2026/08/13/speeding-up-ruby-hashes.html
28•arto•6d ago•0 comments

AliExpress runs silent WebAudio fingerprinting that breaks Bluetooth multipoint

https://blog.laserphile.com/2026/08/aliexpress-webpage-keeping-multipoint.html
919•emctech•17h ago•295 comments

Consumer Rights Wiki

https://consumerrights.wiki/w/Main_Page
251•gregsadetsky•9h ago•34 comments

Mojo is now open source

https://www.modular.com/blog/mojo-open-source
354•visheshdembla•2d ago•81 comments

SpacetimeDB: A Short Technical Review

https://strn.cat/posts/spacetime/
63•hurrrr•8h ago•15 comments

How to compromise your system with a job interview

https://www.codedge.de/posts/how-to-compromise-your-system-with-a-job-interview
127•codedge•11h ago•106 comments

Anti-AI fonts are useless and harmful

https://blog.yaros.ae/anti-ai-fonts-are-useless-and-harmful/
126•speckx•12h ago•86 comments

Git at any scale

https://cursor.com/blog/git-at-any-scale
292•meetpateltech•2d ago•95 comments

Artificial Intelligence Policy

https://www.law.berkeley.edu/academics/registrar/academic-rules/artificial-intelligence-policy/
23•hackerBanana•2h ago•13 comments

Sixtyfour (YC P25) Is Hiring

https://www.ycombinator.com/companies/sixtyfour/jobs/39SkSrA-software-engineering-intern
1•HPMOR•10h ago

Captain Zilog

https://www.zilog.com/captain_zilog/
6•rbanffy•3d ago•0 comments

Aaron Swartz was prosecuted for scraping, while Meta does it without consequence

https://blog.curiousquail.com/im-upset-again-about-a-co-creator-of-rss-being-prosecuted-for-somet...
1154•speckx•7h ago•256 comments

Copyright does not protect AI-generated content in EU

https://mathstodon.xyz/@maxpool/117128107757895678
149•u1hcw9nx•3h ago•142 comments

DiffusionGemma Technical Report

https://arxiv.org/abs/2608.00146
139•gmays•14h ago•34 comments

Every Model Cheats

https://dreadnode.io/research/every-model-cheats-prompt-level-mitigation-of-cheating-on-offensive...
83•vga805•13h ago•70 comments