Jokes aside, I don't know what I would do in this situation either. Maybe find a way to mark "rare" or "out of print" books and sell/auction the "pages"? Or make the scanned book available via some mechanism?
Roughly nobody would have had access to that copy. Roughly everyone can now benefit from its content.
You mean, roughly everyone who uses that particular model provider’s products. For truly rare books, this has an anticompetitive flavor, since it ensures others can’t train models from the same knowledge.
Confuses "(not all that) old books" with "rare books".
Ragebait, nothing more.
"This is not a novel to be tossed aside lightly. It should be thrown with great force."
More troubling to me is the focus on pre-2022 data. This indicates that the companies are seriously worried about model collapse, where the future models become worse becuase they are being fed by generated data instead of real human data.
This is actually my worse case scenario for AI: the models become good enough to disrupt entire industries in the next 10 years, but then stay frozen at that level. Model drift then kicks in because the world will keep changing but the models do not, and in a few decades we actually regress instead of progress becuase they won't be enough skilled humans to drive advances.
But changing how non-AI people write, that's an interesting angle. Because where do we go from here? In 100 years will we still be overvaluing pre-AI sources? That doesn't make sense. Of course a lot can (and will) change in 100 years.
But i'm reminded of pre-atomic steel, which is steel made before the first atomic bombs were detonated and thus have really low background radiation. This is necessary for making MRIs and such. People will go and find it from shipwrecks and such (ironically, many of which are from WW2). It's also a finite resource. What happens when we run out?
Now pre-2022 texts aren't consumed (other than destructively scanning books of course) but it is also finite. We can't make more of it.
This story has a lot of details left out and is a magnet for illinformed commentators: the reason the originals are destroyed is because it is legal to do so for the purposes of format conversion, but not legal to retain the originals under copyright law.
What's blowing my mind lately though is the people complaining about copyright today are some of the same people I saw 30 years ago screaming "Information wants to be free" and running around with the DeCSS code on their T-Shirts. But suddenly, now that they hate AI, copyright is important and the information should be locked away. I can't wait for this all to settle down.
Define excessive. Who makes that determination and how?
https://edconway.substack.com/p/the-eerie-story-of-low-backg...
"Rare and out-of-print" is fabricating a lot of aura here. It's technically correct (the best kind of correct) but I've yet seen evidence that these are culturally significant copies being destroyed for scanning.
From https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...:
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
The oldest is from 1999! If that's the best they can actually enumerate to bolster this outrage farming cycle you could just wonder how irrelevant the rest are.
Really, please, kill this news cycle. There's a lot of issues deserving proper attention right now and this one is a straight-up nothingburger.
Oh no, a lump of cellulose is gone. Will no one think of the fibers?
But it ... won't be? The companies have no motivation to do so. The copyright on lots of these books are surely already expired, so if they wanted to they could be putting these up now. I'm sure internally this is viewed as a corpus of knowledge they have that their competitors do not, so they will not release them unless something forces them to.
You mean every Anthropic?
Ancient Rome stopped creating aqueducts because they had all the ones they needed. They failed to pass that knowledge on to the next generation and so they just forgot how to create aqueducts.
Once AI starts making the majority of content, people will simply forget how to make content. Before we know it we're all fat slobs in floating chairs like in Wall-E.
Best case scenario I think is similar to what we see in Ian M. Banks Culture books where the machines basically take care of us out of the goodness of their hearts and we just kinda fuck off into obscurity.
They aren't. Point me to the links to the digital versions of the originals.
On the contrary, they're mangling the originals: "Character.ai, for example, offers a Books feature that permits users to rewrite public domain works such as “Pride and Prejudice” and “Frankenstein” by changing endings, settings, and inserting themselves into the narrative."
Anthropic: "We use a ‘soft codename’ for it because we don’t want it to be known that we are working on this." If destroying old books were a public service, Anthropic wouldn't need to hide it.
inigyou•54m ago
lookingdesk•51m ago
inigyou•48m ago
netsharc•42m ago
Will I get sued?
pfdietz•37m ago
themaninthedark•22m ago
Edit(seriously): >Transformativeness is a characteristic of such derivative works that makes them transcend, or place in a new light, the underlying works on which they are based. In computer- and Internet-related works, the transformative characteristic of the later work is often that it provides the public with a benefit not previously available to it, which would otherwise remain unavailable.(Wikipedia)
I am sure Napster or the like argued that what they were doing was transformative.
Not sure how AI companies are arguing about PUBLIC benefit if you have to pay...
pfdietz•33m ago
madaxe_again•46m ago
Don’t hate the player, hate the game.
themaninthedark•19m ago