frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

AirTag reveals Amazon is trashing rare books to train AI

https://arstechnica.com/tech-policy/2026/08/hidden-airtag-reveals-amazon-is-trashing-rare-books-to-train-ai/
73•jefurii•1h ago

Comments

rohan_•45m ago
you do realize "rare books" in this context most likely means random technical manuals nobody cares about and not collectors items, right?
wahnfrieden•44m ago
If most but not all, what number are ones that people and collectors do care about?
ButlerianJihad•37m ago
Any bookseller worth their salt will ensure that a truly valuable book will not languish on their shelf for 1 minute longer than it takes to find a buyer willing to pay a fair price.

If AI-scanners are somehow bid-sniping bona fide collectors and wealthy aficionados, there may be cause for concern. But that is most certainly not happening here.

wahnfrieden•33m ago
How do you have confidence in that? These buyers are reportedly exceptionally non price conscious.

Personally when seeking out rare books with few copies in existance and extremely scarce availability, I have often found listings that have lasted for quite some time. Not all rare books go to auction. You might be thinking only of some extreme of notable works and not a wider spectrum of desired but scarce publications that does indeed exist contrary to your confident assertion.

timmmmmmay•25m ago
because that ended up being the case the last couple times this topic came up and now we've gotten wolf-cried

the part where the article avoids mentioning what book this was is a tell

cortesoft•22m ago
If collectors cares about them, the sellers wouldn’t have sold it for pennies.
wahnfrieden•20m ago
These buyers are reportedly exceptionally non price conscious. What evidence is there that all books are sold for pennies, not just most?
cortesoft•15m ago
Because they want quantity, and they don’t need specific titles, so they have no reason to pay more than the cheapest bulk prices.
jasonvorhe•24m ago
Unless we know this to be the case it's reasonable to assume it might be more.
nh23423fefe•21m ago
It's reasonable to assume the clickbait article with no information is actually important?
k33n•19m ago
Why would that be a reasonable assumption? This is so blown out of proportion. The imagery invoked by the narrative is one of huge corporations destroying the final copies of literary treasures. That's just not happening.
a2ff6eeb0•14m ago
Oh, that's really interesting, I'd love to see the list of books that they've digitized too. Where did you find it?
anygivnthursday•43m ago
Another discussion on this https://news.ycombinator.com/item?id=49335216
tacticalturtle•17m ago
Which itself is a dupe of the original reporting on this by 404 media:

https://news.ycombinator.com/item?id=49330742

zamadatix•24m ago
I'm less concerned with how these rare books didn't rot in large lots of unused books and more concerned with whether or not this helps preserve the content longer.
wahnfrieden•21m ago
The content is never made available. It would be illegal to.
progval•10m ago
It will be legal to share in some decades. Not that Amazon will bother, though.
nullorempty•12m ago
The content is preserved obviously, and I hope that sometime in the future it will be made available in its original form.

What worries me about the trends is inevitable sanitization of content or straight out falsification.

jeffbee•23m ago
I remember when hackers believed in the doctrine of first sale. You can do whatever with the stuff you own.
onion2k•19m ago
In a strictly libertarian sense, sure.

In more liberal sense, because you might be destroying something unique that later generations might actually like to see, the life rule that states "Don't be a dick" probably applies trumps even first sale doctrine.

radial_symmetry•19m ago
They are rare because nobody cares about them otherwise. Why is everyone acting as if they are trashing Gutenberg Bibles or first edition LOTR copies?
joshstrange•19m ago
I really have trouble getting worked up about this. "Rare books" is thrown around regularly but my gut feeling is that's not the case. These are used (often? always?) books and while I'm sure there is waste, in general they just want 1 of every book.

While I wish there was a repository of every book that was already digitized (it pains me this is the best solution), there isn't one and so I think this is not a real problem.

It'd be a different story if they had furnaces that ran only on rare books that they had to continually feed books to but that's not what's happening here. And that 1 destroyed copy will "live on" in a way that it otherwise might not.

nemomarx•8m ago
It'll live on if they publish those scans or contribute them to a national archives or something. Proprietary data has a habit of being lost over time though.
hinkley•7m ago
Rare books tend to be out of copyright, true.
Nicook•17m ago
I remember when google was scanning a bunch of rare books, I mean they might still be doing that? Either way, that was cool.

I have a few "rare books" and have read many, you'd be suprised at what is publicly available on google books since like ~2010ish.

jmyeet•12m ago
Years ago I saw an article on this. For Google Books, they had two processes.

The first was destructive. This was for mainstream books currently being published so they had no value. It's (I believe) where you cut off the spine and scan the pages.

For rarer books, there was a non-destructive process. Basically the book was opened to each page and scanned. This was slower but didn't destroy the book.

I don't understand why these companies haven't just licensed the scans Google has already done. Why is each company doing this rather than just scanning the books once and sharing the scans?

nkurz•6m ago
> I don't understand why these companies haven't just licensed the scans Google has already done.

Because for the vast majority of the books, Google doesn't have the legal right to license those scans. It would be legal if the books were out of copyright, but despite the connotation of "rare books" those generally aren't the books we're talking about. Further, in the cases where Google didn't destroy the original of the book they scanned, their scanned copy may be considered infringing under the new standard, so Google doesn't want call undue attention to what they have.

xiphias2•14m ago
The scariest thing for me about companies not caring even minimally about conservation is that when AI gets more powerful than humanity (which is clearly a when not an if, even if there's a lot of uncertainty and differences in opinion about the time horizon here), I want to hope that AI will care more about conservation of human people.

So far from how I see how powerful organizations work, I'm not as certain as I would like to be.

pstuart•14m ago
A couple thoughts:

  * They're only buying a single copy of that book as they only need one to scan

  * If the book was public domain (or should be), then there should be an effort to "democratize" that data into a public commons of intellectual property?
Would that an enhancement to the Library of Congress or such?
atleastoptimal•10m ago
They are only destroying the books because they are required to by copyright law. They obviously wouldn't do so if they were allowed to merely copy the book and preserve the original.
criddell•9m ago
If it was $1 cheaper to destroy the books even if they didn't have to, they probably would. Storing books is expensive and selling it on means it could be picked up and scanned by a competitor.
huslage•5m ago
There's no copyright law that requires the owner of a copy of a book to destroy it. What are you talking about?
pavon•5m ago
Not true. To the extent that it is fair use to digitize a work for various purposes, it is also fair use to keep the original. It is only if they wanted to resell the original that they would have to delete their digitized copy. The reason they are destroying them because removing the binding is the most efficient way to scan them, they have no use for the originals after they have been scanned, and don't want to spend money storing them.
hinkley•8m ago
Am I too old now, expecting someone to make a Rainbows End reference? Vernor Vinge predicted this 20 years ago.

(Also the person who coined Singularity, though Ray Kurzweil really wanted everyone to think it was his idea.)

olalonde•7m ago
This feels like a manufactured controversy. What difference does it make to me what someone does with a book after they buy it? It's effectively unavailable to me regardless of what they do. If people are really concerned about these "rare" books, they should lobby the copyright owners to release them online or print more copies.
declan_roberts•5m ago
It's better value to me that someone is digitizing and training with books then they just languish somewhere in perpetuity.
pohl•5m ago
If there’s only 3 copies in existence, and everyone on the frontier wants it in their corpus, what do you think will happen?

AI;DR (AI; Didn't Read)

https://www.rickmanelius.com/p/aidr-ai-didnt-read
100•mooreds•27m ago•23 comments

A Preview of DuckDB v2.0

https://duckdb.org/2026/08/17/duckdb-20-highlights
420•ibotty•6h ago•66 comments

GPU Offload in Rust: Portable, Safe, and Fast

https://arxiv.org/abs/2608.13759
48•linggen•2h ago•9 comments

AI-Generated GitHub Copilot “Autofix” Allowed Compromise of Snowflake's Jira

https://www.wiz.io/blog/red-agent-snowflake-copilot-cicd-bug
249•galnagli•5h ago•111 comments

Incident with Github.com

https://www.githubstatus.com/incidents/zkxwbgr0cnmx
386•SpyCoder77•6h ago•789 comments

Sun Clock

https://sunclock.net/
81•Gecko4072•3h ago•32 comments

How to disable or avoid intrusive AI

https://www.librarian.net/notoai/
191•ColinWright•6h ago•93 comments

Will you have spent more of your life with computers than your family?

https://beachfront.bearblog.dev/will-you-have-spent-more-of-your-life-with-computers-than-your-fa...
6•Gecko4072•18m ago•0 comments

GPT 5.6 Sol is the best "vision" model OpenAI ever released

https://blog.roboflow.com/openai-gpt-5-6/
254•plurby•8h ago•132 comments

Launch HN: Speko (YC S26) – OpenRouter for Voice AI

https://speko.ai/
75•abdik•4h ago•46 comments

Judge sets framework for Nine PBS to retrieve archival data

https://current.org/2026/08/judge-sets-framework-for-nine-pbs-to-retrieve-archival-data/
67•qingcharles•4h ago•23 comments

AirTag reveals Amazon is trashing rare books to train AI

https://arstechnica.com/tech-policy/2026/08/hidden-airtag-reveals-amazon-is-trashing-rare-books-t...
78•jefurii•1h ago•46 comments

India built the biggest digital payments miracle: Now comes the bill

https://www.bbc.com/news/articles/c8xnwqe00v1o
7•monkey_monkey•49m ago•1 comments

Olo (Color)

https://en.wikipedia.org/wiki/Olo_(color)
192•inigyou•5d ago•46 comments

Roboflow Playground: Try and Compare 30 Computer Vision Models

https://blog.roboflow.com/roboflow-playground/
13•Bluestein•1h ago•1 comments

How I Over-Engineered My Book

https://ben.balter.com/2026/08/17/how-i-over-engineered-my-book/
33•benbalter•1h ago•24 comments

The Oldest Bar in Every US State

https://www.businessinsider.com/oldest-bar-every-state
29•NaOH•4d ago•7 comments

Qwen3.8 27B scores 52 on Artificial Analysis

https://artificialanalysis.ai/models/qwen3-8-27b
185•anana_•2h ago•95 comments

Ask HN: Alternatives to GitHub

395•dhruv3006•6h ago•253 comments

The Lonely Men Who Work in Patagonia, at the End of the World

https://www.newyorker.com/culture/photo-booth/the-lonely-men-at-the-end-of-the-world
58•bookofjoe•1h ago•22 comments

Marketers are Addicted to Bad Data (2020)

https://www.jacquescorbytuech.com/writing/marketers-addicted-bad-data
12•zbentley•4d ago•14 comments

How to put 170 atoms in an atom

https://signoregalilei.com/2026/08/02/how-to-put-170-atoms-in-an-atom/
79•surprisetalk•5h ago•16 comments

Anthropic's ‘watermark’ text adulteration in Claude is a perversion of writing

https://daringfireball.net/2026/08/anthropics_watermark_text_adulteration_in_claude_is_a_perversi...
737•ropbear•22h ago•641 comments

Amazon, which started off selling books, is destroying rare texts to train AI

https://techcrunch.com/2026/08/17/amazon-once-an-online-bookseller-is-destroying-rare-books-to-tr...
70•rzk•2h ago•34 comments

A particle made of force: physicists say they've found mysterious 'glueball'

https://www.nature.com/articles/d41586-026-02498-1
49•Brajeshwar•5d ago•2 comments

How I developed an Am29000 C compiler and web browser

https://nanochess.org/am29000_c_compiler_web_browser.html
62•nanochess•23h ago•9 comments

We Are Forking dotenvy into dotenv-ng

https://secretspec.dev/blog/we-are-forking-dotenvy-into-dotenv-ng/
17•linggen•2h ago•16 comments

Show HN: Sokoban AI Solver

https://mkornreich.me/projects/sokoban/
59•enjoyyourlife•7h ago•33 comments

On AI regulation and messaging

https://twitter.com/DarioAmodei/status/2088758816376807762
218•jacquesm•18h ago•458 comments

Show HN: Learn Flags Quiz

https://flagquizzes.com/
33•artiomyak•6h ago•18 comments