frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

A 'bananas' order for 5000 obscure book titles fuels suspicion

https://www.irishtimes.com/world/europe/2026/08/10/a-mysterious-buying-spree-is-unsettling-europes-booksellers/
48•ohaodha•1h ago

Comments

jvanderbot•48m ago
An instant and obvious good these companies could do is to digitize these books ala google books.
discreteevent•41m ago
Then they would have to deal with copyright. AI is the washing machine that leaves everything spotless.
conartist6•30m ago
Remember kids, if you say "I was doing it for AI", it's legal. If you say "I was doing it for the good of humanity," well, we still remember Aaron Swartz.

But of course, unlike humanity, AI is really worth it!

jvanderbot•23m ago
There will be an Open AI / Anthropic leak eventually. One can only hope.
mapt•37m ago
Preservation rather than ingesting them into the gaping maw of an Intelligence Engine that preserves the soul of the book in crystalized neural network form? That would be a copyright violation.
wongarsu•30m ago
And get sued even harder by publishers?

Publishing any out-of-copyright books would be awesome thing to do. But it seems none of the books in this order would fall under that, and it's unclear how many old books they are actually ingesting. Current reports are mostly talking about relatively recent books

Cthulhu_•46m ago
This is a good thing to a point, right? They sell heaps of stock, some old / outdated stock (like "Pass Your Driving Test, 2018 Edition" as per the article), in bulk, at full price, no questions asked. Retailer's dreams come true.

(I don't believe them being digitized for consumption by AI will impact book sales much, as ebooks, google books, project gutenberg, etc didn't either)

femto•37m ago
As long as the AI devouring a book isn't making it harder for a person to get access to a copy of the book. There might not be many copies of some out-of-print books.
PunchyHamster•36m ago
Would be nice to at least force them to open their scanned books data, after all they stole everything else, they can contribute some back
sobiolite•11m ago
"Open" it how, exactly? Make it available for free? As in violate copyright?
abanana•23m ago
It's great for the retailer, although I think that example was picked out because it's an extreme example of a book that'll never sell. But the volume of these orders shows another way the AI companies are pissing away their investors' money.
phoghed•19m ago
BLKNSLVR•41m ago
Is this a second form of AI psychosis, as in AI training corpus psychosis?

The need for all the content, Moar!! Feed me. Is this the Paperclip Maximiser in the form of the Training Material Maximiser?

Will it be that, in the end, the lack of "Pass Your Driving Test, 2018 Edition" was the cause of driving rule hallucinations in all previous models? Will this finally get OpenAI back in front of Anthropic? Hurrah! We found it!

luxuryballs•37m ago
I was thinking the inverse the psychosis is that someone who sells books is alerted that they are selling more books, as if someone buying books needs to give an explanation. If I want to seed software with books why is that newsworthy? Print more books, maintain profitable margins on books, enjoy success as a book seller?

If a book is rare and you wanted to preserve it don't sell it so easily, increase the price, or put it into a contract system like Ferrari does with cars they sell.

PunchyHamster•34m ago
Till it hits books that are out of print, we already had news about AI scourge companies destroying rare books coz scanning them in destructive ways was slightly cheaper.

And as is happens in history, any time some entity mass destroys the books, they turn out to be bad guys

luxuryballs•32m ago
But anyone can mistreat or fail to care for a book, if it's rare and you don't want to risk it don't sell it.
cwnyth•29m ago
There's a very simple solution to this, but the people who rabidly argue for a copyright length of 70 years + the lifetime of the author won't want to hear it.
vintermann•41m ago
> with a focus only on books with an ISBN number, the identification system introduced in 1966

So it's not likely to be rare and precious books. It's things which are already in the Library of Congress or its many equivalents.

If they're forced by laws to destroy the results of the scans after training on them, as some have implied, that's bad, on the chance there is some actual lost media in there. But if they keep the scans, or even the transcripts, that's probably an improvement on the status quo, to be honest.

I've had family members in the used books business, and trust me the fate of the vast majority of these books was always to be pulped.

iamacyborg•37m ago
There are plenty of rare and precious books with ISBNs.
abanana•27m ago
There are, but, quoting from the article:

> Anthropic said: "None of our data acquisition programmes buy and destroy rare or antiquarian books."

This statement was clearly carefully worded, using the words "rare or antiquarian" to mean "pre-1966, pre-ISBN". Their statement is answering a question that wasn't asked, so that (if they need to) they can claim in the future that they never said they weren't destroying books with ISBNs.

bloak•6m ago
Rare, certainly: there are plenty of vanity publications that were chucked into the bin by almost everyone who was unlucky enough to be given a copy. Precious? Well, with the help of an electronic friend I found ISBN 978-3-8365-7349-8. Apparently a copy of that book is worth about £40k. Can anyone beat that?
toofy
ck2•36m ago
if they are going to just violate copyright blatantly and price is no object

why not just every textbook for every subject in high school and college?

paweladamczuk•29m ago
> “Under German law scanning books – regardless of the purpose – would not be permissible and constitute a clear violation of copyright law,”

Excuse me, what?

ZenDroid•12m ago
https://law.stackexchange.com/a/93882

"Individual reproductions of a work by a natural person for private use on any medium are permitted, provided they are not used directly or indirectly for commercial purposes …"

wasmitnetzen•11m ago
German copyright law has no Fair Use. There are copyright exemptions for private or academic use, but commercial use has no such exemption. Books have special protection even according to § 53 (4) b).

And yes, it's the duplication which is illegal, not the publication of them.

[1]: https://www.gesetze-im-internet.de/urhg/__53.html

lluisantoni•27m ago
I didn't think about this happening. It reminds me of the cross-border challenges with crypto. Each country has laws to control its author rights or money supply, but those are hard to enforce in the international setup.
khalic•27m ago
Not exactly top journalism...
stockresearcher•22m ago
In the case of English-language books, they’re just working to legitimize the stuff they already obtained via, um, dubious methods.

Public domain books - no need to bother, they’ve scraped PG and any other source and can do with it what they like.

Used booksellers should raise the price of used books that are of no interest to any real reader.

criddell•16m ago
When I hear these stories of AI companies buying all the books, I think back to Kevin Kelly in 2008 or 2009. He talked about how books are less expensive than at any other time in history and easier to buy and that could change so it makes sense to buy a lot of books.

> [...] I was near to the point of actually digitizing and getting rid of all my paper books.

> I was that close about five years ago, but then I had an epiphany. I went to private library, and I realized that books were never as cheap as they are today. They never will be as cheap, and that there's some power about having these things in paper always available, no batteries, never obsolete, and that if you made a library now, you would never be able to make some of these libraries in 50 years, so I decided to keep and to cultivate this paper library as something that was going to be powerful in the future.

I remember seeing photos of his library but can't find them anymore. This is the only thing I could find:

https://colossus.com/article/flounder-mode/

breezybottom•10m ago
Am I supposed to know who that is?
Download content to train? Immoral and illegal.

Buy it? You’re wasting investor money.

carlosjobim•10m ago
Your local school throws away books by the dumpster load once a year or every other year.
abanana•21m ago
> The need for all the content, Moar!! Feed me

"MOAR input!!!"

- Johnny 5, Short Circuit. That was a great film. Seems it was unexpectedly prescient too.

•
37m ago
> But if they keep the scans, or even the transcripts, that's probably an improvement on the status quo, to be honest.

only if they share the scans/transcripts and don’t hide it away.

DougBTX•27m ago
Yeah, we’re in a funny position. By all accounts it is fair use (at least in the US) to train models (and build search indexes, e.g. Google Books), but sharing the books dataset itself is absolutely forbidden (clear non-transformative copying).

Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!

FartyMcFarter•4m ago
> Anyone that wants to train a model needs to procure and destroy their own physical copy of each book!

Why? Couldn't they resell or give away the books after scanning them?

trescenzi•3m ago
It's less work if you use a destructive method. They rip the spine off to get individual flatter sheets instead of scanning a bound book.
Kim_Bruning•4m ago
How is this in any way, shape, or form a promotion of the arts and sciences anymore?

This could very easily be turned into a preservation and archiving operation with just a tweak of the laws, or a carve-out.

And make the bank once, and make it legal to train on? How people in the bank get compensated is a different question, but -while almost impossible to settle on an individual basis- could be settled reasonably in bulk by some form of mandatory implied statutory contract?

woodrowbarlow•33m ago
letting companies train on books for only the price of a paperback and then forcing them to destroy their scans would truly be the worst of all worlds. (ps: the scanning process typically destroys the book too; they chop off the spines.)
troupo•13m ago
> on the chance there is some actual lost media in there. But if they keep the scans, or even the transcripts, that's probably an improvement on the status quo

Where can I see the scans and the transcripts in this "better status quo"?

vintermann•6m ago
It's not enough of an improvement that you might see them. It's enough of an improvement that they will likely still exist somewhere.
sixtyj•11m ago
SBN started in 1967 in UK (W.H.Smith), official international book numbering (ISBN) is after 1970.

Source: https://www.isbn.org/ISBN_history

Meta Muse Glimmer – open weights 30B local coding model

https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model
287•riordan•2h ago•111 comments

Because It's Not Fun Enough: why languages fail

https://bytecode.news/posts/2026/08/because-it-s-not-fun-enough
48•jottinger•1h ago•24 comments

Docker Sandboxes – Disposable, isolated sandboxes for AI agents

https://www.docker.com/products/docker-sandboxes/
323•etoxin•6h ago•187 comments

Tail-call optimization in C is relatively recent

https://lwn.net/Articles/1034703/
29•prakashqwerty•1h ago•11 comments

Squeak/Smalltalk 6.1 Release Notes

https://squeak.org/release_notes/6.1/
11•fniephaus•46m ago•1 comments

50k Boat Names

https://www.beautifulpublicdata.com/boat-names/
3•jonathanmkeegan•3m ago•0 comments

Parametron: 50s Japanese computer that uses neither transistors nor vacuum tubes

https://ethw.org/Milestones:Parametron,_1954
20•xeonmc•2h ago•1 comments

What Happened to HackerOne?

https://blog.teknogeek.io/posts/what-happened-to-hackerone/
282•hipparchus•10h ago•142 comments

Run Android ARM64 VR APKs on Apple Vision Pro

https://github.com/shinyquagsire23/Klepton
114•LorenDB•9h ago•16 comments

Tail-Call Interpreters in Rust – Jimmy Ostler

https://lordgoati.us/blog/tail-call/
45•amatheus•3d ago•15 comments

An Interesting Fourier Transform – 1/F Noise

https://www.dsprelated.com/showarticle/40.php
69•q7m•3d ago•15 comments

Show HN: Voice driven murder mystery, Interview AI suspects with your voice

https://www.whodunnitai.com/
125•MrRowTheBoat•9h ago•52 comments

How Blackwing Pencils are Made [video]

https://www.youtube.com/watch?v=fow-LsdaH2E
30•NaOH•4d ago•12 comments

How I use LLMs to learn complex topics

https://laurentiugabriel.github.io/blog/articles/how-i-use-llms-to-learn/
705•laurentiurad•17h ago•446 comments

Taxi drivers rarely die of Alzheimer's

https://theconversation.com/taxi-drivers-rarely-die-of-alzheimers-how-complex-mental-maps-and-spa...
322•jader201•21h ago•229 comments

DeepSeek costs OpenCode Go user $1.14/day; dual DGX breaks even in 24 years

https://twitter.com/thdxr/status/2086599224674681242
11•delduca•37m ago•3 comments

Findphone: Locate a nearby Bluetooth device by signal strength

https://github.com/ben-z/findphone
5•helsinkiandrew•5d ago•1 comments

How We Pushed CDC into Postgres

https://www.snowflake.com/en/blog/engineering/postgres-to-snowflake-replication-mirroring/
108•craigkerstiens•12h ago•20 comments

Ask HN: What are you working on? (August 2026)

270•david927•19h ago•934 comments

A 'bananas' order for 5000 obscure book titles fuels suspicion

https://www.irishtimes.com/world/europe/2026/08/10/a-mysterious-buying-spree-is-unsettling-europe...
48•ohaodha•1h ago•42 comments

Cool URIs Don't Change (1998)

https://www.w3.org/Provider/Style/URI
255•Klaster_1•22h ago•62 comments

ATProto for Distributed Systems Engineers

https://atproto.com/articles/atproto-for-distsys-engineers
101•LelouBil•3d ago•18 comments

Over 181,000 AI meeting recordings left wide open in note taking app

https://bobdahacker.com/blog/tldv-hack
7•colesantiago•35m ago•1 comments

An alias-based formulation of the borrow checker (2018)

https://smallcultfollowing.com/babysteps/blog/2018/04/27/an-alias-based-formulation-of-the-borrow...
17•parksb•2d ago•1 comments

Picophysics: Single file physics for games on platforms like N64, PSX, DC

https://gitlab.com/Kazade/picophysics
68•klaussilveira•5d ago•22 comments

Tuxedo No. 2 – Cocktail recipes

https://tuxedono2.com
109•smartmic•16h ago•35 comments

Nearest Pint

https://knowwhereconsulting.co.uk/maps/pubs/
41•bookofjoe•5d ago•23 comments

Everything you do is being recorded

https://www.theatlantic.com/technology/2026/05/ai-wearable-surveillance-countermeasures/687203/
345•ike_usawa•1d ago•296 comments

I made tinnitus my friend, then it disappeared [video]

https://mynoise.net/vlog.php?ep=20260803
176•gregsadetsky•18h ago•138 comments

Meta's new open-weight model targets local agentic AI

https://twitter.com/finkd/status/2086754845218726027
16•bakigul•2h ago•2 comments