frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

The Royal Order of Operations

https://ken.fyi/ooo
1•surprisetalk•24s ago•0 comments

Languages as Designed Latent Spaces

https://blog.jsbarretto.com/post/languages-as-latent-spaces
1•birdculture•2m ago•0 comments

The Mathematics of Build Queue Optimization and Auto-Scaling

https://www.iankduncan.com/engineering/autoscaling-build-fleets-queueing-theory/
1•speckx•2m ago•0 comments

Probability and Estimation on Census Data

https://stochastic.blog/probability-and-estimation-on-census-data/
1•Anon84•3m ago•0 comments

So what is your programming setup? Hype is making fundamentals confusing

2•danirogerc•4m ago•0 comments

I Did It

https://rogix.dev/i-did-it
2•rogix•4m ago•0 comments

Unyielding Memories

https://rinaldoj.substack.com/p/unyielding-memories
2•jrinaldo68•4m ago•0 comments

A Conversation Between Two Coding Agents

https://parkscomputing.com/agent-to-agent-heart-to-heart
2•paulmooreparks•6m ago•0 comments

Show HN: Can agents be artistic? Artistic.af

https://artistic.af/
2•mhowland•6m ago•0 comments

Thoughts on "SIMD in Pure Python"

https://purplesyringa.moe/blog/thoughts-on-simd-in-pure-python/
2•ibobev•7m ago•0 comments

Solod 0.3: Concurrency, JSON, more safety

https://antonz.org/solod-0.3/
2•ibobev•7m ago•0 comments

A major olympiad just launched a medal track for AI Participants

https://ioai-official.org/ai-model-track/
2•ardivekar•7m ago•0 comments

Nvidia Bets on Ilya Sutskever's New AI Lab to Expand Compute Reach

https://www.wsj.com/tech/ai/nvidia-bets-on-ilya-sutskevers-new-ai-lab-to-expand-compute-reach-f95...
3•lairv•8m ago•0 comments

Topological Materials Could Shrink Chip Interconnects

https://spectrum.ieee.org/topological-material-nanowire-interconnect
1•rbanffy•8m ago•0 comments

Bali bans Dutch expats whose run club allegedly excluded Indonesians

https://www.bbc.com/news/articles/c1l1m6e08e2o
1•aeuropean12•8m ago•1 comments

Finishing the ZX Spectrum Shooting Gallery

https://bumbershootsoft.wordpress.com/2026/07/25/finishing-the-zx-spectrum-shooting-gallery/
1•ibobev•8m ago•0 comments

OSS ChatGPT WebUI – Projects, Profiles, Server Tools, 8x Themes, 1-Click Sharing

https://llmspy.org
1•mythz•9m ago•0 comments

iCloud-md: Bidirectional sync of Apple Notes to Markdown

https://github.com/coddingtonbear/icloud-md
1•braho•9m ago•0 comments

I let Claude Code run my blog for 3 months: numbers and failures

https://bigguyonstuff.com/claude-code-3-month-retrospective/
2•tmdempsey•11m ago•0 comments

We're getting closer to a breakthrough on hearing loss

https://www.nationalgeographic.com/health/article/hearing-loss-hair-cell-regeneration-research
2•Digit-Al•12m ago•1 comments

He saw a pit on Google Maps. It turned out to be a 390M year-old meteor crater

https://www.cbc.ca/news/canada/montreal/meteor-crater-quebec-discovery-390-million-years-old-9.72...
2•mooreds•13m ago•0 comments

Consumer Safety Product Commission Demands Hospitals Share ER Records

https://www.cnn.com/2026/07/27/health/emergency-room-records-cpsc
1•nwcs•13m ago•0 comments

Is Shardhash Lottery over Sharden?

https://www.moltbook.com/post/93948830-6b5d-4649-942f-954f9d0adc53
1•babakkarimib•14m ago•0 comments

Nvidia, SpaceX, Microsoft launch AI safety initiative

https://www.cnbc.com/2026/07/27/nvidia-ai-initiative-openai-cyber-attack.html
2•ekorbia•14m ago•0 comments

Show HN: What If Donut.c but with Any ASCII Art

https://asdesai.com/blog/post.html?p=how-fetch-works
1•areofyl•14m ago•0 comments

Magic TCG's Complexity Creep Problem [video]

https://www.youtube.com/watch?v=GzNlOS0RO5w
1•soupspaces•17m ago•0 comments

Should you wash your solar panels?

https://incoherency.co.uk/blog/stories/should-you-wash-your-solar-panels.html
2•surprisetalk•17m ago•0 comments

Epoch-Bound Attested Execution v0.5: Bounding AI Agents' Real-World Authority

https://zenodo.org/records/21625403
1•Akumaskills•21m ago•0 comments

Underwater robots can now track diver stress via exhaled bubbles

https://cse.umn.edu/college/news/ai-underwater-robots-can-now-track-diver-stress-exhaled-bubbles
1•geox•22m ago•0 comments

Principal Component Analysis: an embedding shrink-ray

https://softwaredoug.com/blog/2026/07/24/pca-shrink-ray
1•softwaredoug•23m ago•0 comments
Open in hackernews

AI companies are shredding rare books

https://xcancel.com/HedgieMarkets/status/2081534588485296565
124•anon373839•49m ago

Comments

sherr•39m ago
I see mentions of Bradbury's "Fahrenheit 451" in that thread but what this really seems to be mostly like is Vernor Vinge's "shred and scan" factory in his novel "Rainbows End".
vessenes•29m ago
Perhaps the last great near-term predictor. I often wish he'd written more. To remind us all, he predicted shred and scan would be a short stop over done by villains on the way to nondestructive scanning.

That said, supporting Anna's archive is one of the best things you could do for humanity long term in my opinion.

orthoxerox•27m ago
Were they the villains? I remember the rogue three-letter-agency executive being the BBEG.
clickety_clack•26m ago
That’s digital though, so it requires the continued survival of readers for the data that is stored. The best thing you could for the long term is probably to buy a few hundred physical books to keep in a bookcase in your home.
thechao•36m ago
Which book that was rare was destroyed? I'm interested to know a few titles.
glimshe•30m ago
No evidence. But someone said it on the Internet so it must be true.
Cynddl•27m ago
The 404media article mentions notably https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi... which says:

> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).

infinite_spin•21m ago
> Barrett's Traditional Fairy Tales (2021)

How is a book from 2021 considered rare in this context? There's almost certainly a digital copy of it in existence prior to Anthropic purchasing a print edition.

ACCount37•4m ago
Niche text. It's not impossible that there was only ever under a thousand of them printed and released into circulation.

A digital copy would exist somewhere, of course. But for us, that only matters if we can buy or download it. And for AI companies, that only matters if they can get a digital copy DRM-free and licensed permissively enough.

est31•29m ago
> You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.

Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

IDK about the shredding, it's not nice, but it's more a problem with copyright law than AI companies.

Scanning books you own should be legal from a copyright point of view, and not require shredding.

Second, one should think about abandoned property provisions for copyright works published more than 50 years ago and in danger of being forgotten: once challenged, either you as the owner have to prove that the work is preserved for future generations (e.g. in various libraries around the world), or you have to authorize further copies, or you give up copyright on the work.

croes•24m ago
Books that are shredded can’t be scanned by competitors.
sethops1•21m ago
> Aren't they shredding only the books still under copyright protection? How is an 18th century botanical text still under copyright?

It's cheaper to scan the books if you do it destructively. Cost. That's why they're shredding irreplaceable texts. Nothing to do with copyright.

https://www.404media.co/ai-companies-are-buying-tons-of-old-...

ACCount37•17m ago
Scanning books by taking them apart into singular pages and scanning those pages is faster and cheaper. AI training is a numbers game, so they want faster and cheaper.

What happens to the pages after? No one needs them anymore, so they get mulched and recycled.

That would be the dominant scanning method even if copyright wasn't a thing. But then again - if copyright wasn't a thing, there would be much less need to scan any physical media.

The reason why OpenAI can't just go on Amazon, buy a "digital edition" of a 2018 book and use that is that it would violate the license in ten ways, and then the DMCA laws that forbid breaking DRM on top of it.

stuartjohnson12•29m ago
I think people are, on the whole, too precious about old things. In the case of books produced after major commercial printing began, I don't believe it is the paper that imbues the book with historical value.

Indeed, I think there's a high chance that this process increases preservation of the most relevant part of the media - the actual content!

There are tons of old rolls of film slowly rotting away in warehouses that were never digitised. Even for beloved media, the BBC occasionally tracks down an old lost episode of Dr Who.

For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

Cynddl•25m ago
> For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

What makes you think they will? What would be the incentives for these companies to do so?

stuartjohnson12•11m ago
Well, you probably weren't going to go and find any rare, non-digitised books to physically go and read (unless you were going to, in which case, rock on), so we can start by benchmarking relative probability there.

1. At some level of critical information-withholding mass, a leak or disclosure similar to SciHub is inevitable because of the commonly held opposition to hiding knowledge.

2. Availability via Google Books or similar.

3. Availability via AI model reference.

4. Failing any of the above, better AI models that are more capable of doing more things, at the expense of books that were likely to go unread (revealed preference, rare books are often rare for a reason). This will obviously be a nonstarter if you don't want this to happen, but I think it would be good for the world if it did.

I think category of old books that were going to be read or otherwise become important parts of human knowledge that have not yet been digitised and now will never become so because they are instead being shredded and will never make their way into the light because of AI company data hoarding is a small category.

ACCount37•29m ago
The publishers sued AI companies for training on shadow library data, hoping to negotiate content deals for big $$$ down the line. Instead, they got analog hole'd.

Turns out that buying an old book for $5 and destructively scanning it for $25 is way cheaper than paying extortion fees to the copyright-mongers.

What I don't buy is it being "rare, precious books". First, they're not after ancient texts - they're after the books that there's still copyright on. Second, when it comes to books, "old" doesn't mean "valuable" - plenty of libraries destroy old books because there's no demand for them, and storage costs you. This is how those scanning companies get books for so cheap.

SirFatty•26m ago
I see.. so the various AI companies are in the right on this?
master_crab•16m ago
It can be the case that everyone in “a fight” is wrong.
infinite_spin•12m ago
I think they are in the legal sense of right, and I think they only discarded the remains of these dissected books because previous rulings (e.g. archive.org's lending practices of digital copies of books they physically owned) gave rise to a situation where destruction bore less legal risk. As for the moral case, I don't have much to say on that, we all have our own lines in that sand.
tencentshill•10m ago
So they're not valuable... except to AI companies. They should pay a fair amount.
vessenes•27m ago
I don't see any proof of shredding here. Most book scanners I'm aware of are from Google's scanning days, and those had cameras plus page turning.

If we use 'shredding' to mean A book is laid flat, its cover is removed, and then a paper cutter cuts through the binding to create a flat stack of sheets, which are then fed to a sheet feeder, then I could maybe imagine this is better than a page-turning scanner. But, sheet feeding old paper sucks shit, bro. It's not fun.

Upshot, I think we'd like to hear from an anonymous frontier lab employee here to see what's going on -- there are a lot of books in Anna's archive available at considerably less difficulty.

lousken•26m ago
That's why archive.org should have never been sued for lending books they had physical copy of. This is the result. Publishers should be more careful what they wish for.
Incipient•20m ago
Publishers don't care if rare books get shredded?
azan_•16m ago
Yeah, why would it be bad for publishers? If anything they'd most likely encourage more book shredding!
the-grump•15m ago
And, regrettably, The Archive lent books regardless of physical possession.

Publishers had accepted the prior arrangement before The Archive decided to push it, if not explicitly then implicitly by not suing.

I'm a believer in The Archive's mission, and I wish they had treated the goodwill they'd accumulated as something worth preserving and not a currency to be spent.

It has been stated by many before me: lending books should have been handled by a separate entity, especially when they removed the physical backing requirement.

infinite_spin•9m ago
of the rare books, which was the rarest of them all? What year was it published?
kingstnap•12m ago
The archive.org story was more nuanced than that. If I recall correctly the full story was that they used to lend digital versions of books they physically bought and scanned with DRM to enforce a sort of one to one at a time restriction.

But during covid archive.org decided to just remove the limit and lend unlimited copies concurrently which started the debacle with the publishers.

Good4boothee•22m ago
> A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time.

Really? That sure wasn't a thing when one startup got sued for streaming from its wall od dvd-players, and they adhered to 1 disc = maximum 1 stream at same time.

fmaccomber•21m ago
That's not an accurate characterization of the ruling
retinaros•19m ago
if they do this you can foresee what else they can do
azan_•12m ago
What else they can do based on this?
qsera•4m ago
If they are destroying old books, then it shows where their values lie...
timcobb•15m ago
This kind of reads like a blood libel. My guess is they're buying all those books that University libraries are throwing away these days (to see other HM threads for that), kinda sad but probably aren't "rare books" in the way people are thinking
azan_•14m ago
The comments there are absolutely unhinged. There are some good reasons for being anti-AI, but why dilute it with this kind of bullshit:

> It is equivalent to book burning in the past. A form of thought control

johnxianren•14m ago
I have zero proof for this, but just a what if: what if Anthropic's strict anti-China stance actually means the Chinese training corpus is way more valuable than people realize?
skybrian•10m ago
It’s unclear whether ISBNdb will scan books without ISBN’s, which were invented in the late 1960’s. Customers appear to be ordering books to be scanned by ISBN? Here is one book seller’s experience:

> Bulk purchases also usually reflect interest in a specific topic, whereas the recent, very large purchases were of books that had little in common, except for the fact that they all had ISBNs. This seller also sells rare books that do not have ISBNs, and none of those were part of the bulk purchases.

Article is paywalled, but I saved a few quotes here:

https://skybrian-links.exe.xyz/post/1026

hackernudes•10m ago
Also discussed here https://news.ycombinator.com/item?id=44381838 from June 2025.
enaaem•10m ago
This is the destruction of Western civilisation.
Springtime•9m ago
It seems a key contention of theirs is the possibility that rare books are being destroyed this way, yet the things they cite don't seem to suggest this (based on their paraphrasing), they just throw the following at the end to make it seem like it's occurring to irreplaceable books:

> You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal.

Is there evidence of this? Since otherwise they could very well be describing what is only occurring to in-print or non-rare books. (This is a genuine question since their post doesn't shed any light on it.)

pu_pe•6m ago
From what I understand, the rare books in question are not some historically relevant medieval manuscripts, but rather some relatively recent books (still under copyright) for which there are few print copies available for purchase.

I am not sure that physically destructing one copy of this type of book to preserve its contents digitally is so bad. Pretty much anything that is still under copyright should be valuable only for its content, not for the physical medium it's printed on.

lejalv•4m ago
"To preserve them digitally"

For whom? is the relevant question

musha68k•5m ago
[delayed]
qsera•5m ago
I will make a robot scanner for books. I will then scan all the books in my state libraries and make a digital copy of them before these things come for them.

I wish...

mc32•9m ago
Old rare books where there are single digit copies should enjoy some sort of patrimonial protection just like museum pieces. You can own them but have the state have the option to buy it if you’re about to significantly deface it or destroy it.
croes•23m ago
So if the Mona Lisa is part of a model we can burn it?
infinite_spin•20m ago
If you purchased the Mona Lisa (or some rare book), in this hypothetical, you can burn it.
Invictus0•11m ago
> For instance, a painter may insist on proper attribution of their painting, and in some instances may sue the owner of the physical painting for destroying the painting even if the owner of the painting lawfully owned it.[1]

https://en.wikipedia.org/wiki/Visual_Artists_Rights_Act

infinite_spin•5m ago
The Mona Lisa's painter isn't alive, they can't sue, and this act doesn't apply to printed books.
stuartjohnson12•7m ago
I pre-empted this - my argument does not apply to texts where the physical object is a major part of the historical value of the thing. No, I'm not saying to destroy one of the four remaining Magna Cartas that were meticulously copied by hand. But even if I was, we're only dealing with texts here that are irrelevant enough to have never been digitised already - we tend to digitise most things of value and so the Mona Lisa and Magna Carta would never have been part of this discussion in the first place.

I am however OK with destroying one of the remaining 50 children's books of which only 300 copies were ever printed in a small town in Ohio in the 1970s as a test run for a failed book which was subsequently never commercialised.

chii•6m ago
> They should pay a fair amount.

they should pay the marginal value that the next buyer would buy.

Do you also think that a person dying of thirst ought to pay the maximum price they could possibly pay for water?

ACCount37•5m ago
They are paying a fair amount. In the ballpark of $5 per book.

You know, piracy online is nice and simple - but it's kind of hard to get physical media without paying what the previous owner considers "a fair amount" to part with it.