frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

F-Droid 2.0: A New Chapter for Android Freedom

https://f-droid.org/2026/09/24/f-droid-2.0-a-new-chapter-for-android-freedom.html
230•daveoc64•1h ago•60 comments

GitHub has not removed malicious imitation software after 3 weeks

https://successfulsoftware.net/2026/09/24/github-has-not-removed-malicious-imitation-software-aft...
40•hermitcrab•54m ago•13 comments

Dynamic Abliteration: Non-Destructive Refusal Suppression via Engram Steering

https://blog.madhukaraphatak.in/non-destructive-refusal-supression-using-engram
85•phatak-dev•2h ago•29 comments

The science of Monkey Island: can grog dissolve a metal mug that fast?

https://jgeekstudies.org/2026/09/23/the-science-of-monkey-island-can-grog-actually-dissolve-a-met...
74•zdw•1d ago•15 comments

Enjoy Every Sandwich

https://bradmontague.substack.com/p/enjoy-every-sandwich
127•NaOH•1d ago•53 comments

Nokia Design Archive (2025)

https://nokiadesignarchive.aalto.fi/index.html
167•pillars•6h ago•86 comments

Linux support is coming to Snapdragon X2 Series

https://www.qualcomm.com/news/onq/2026/09/snapdragon-summit-agentic-ai-pcs-linux
566•aaronday•18h ago•236 comments

I Have a Confession: I Built This Site with AI – Please Forgive Me

https://dynamicallytyped.org/blog/i-have-a-confession-i-built-this-site-with-ai
17•joelberger•50m ago•14 comments

QR Codes That Route to the Appropriate App Store

https://matthuggins.com/blog/posts/qr-codes-that-route-to-the-appropriate-app-store
9•matthuggins•54m ago•4 comments

LinkedIn wins court order blocking mass scraping of user data

https://therecord.media/linkedin-wins-court-order-blocking-mass-scraping
19•ilamont•42m ago•0 comments

Ideas on modernizing the open-source desktop

https://lwn.net/SubscriberLink/1095425/2d9f411252325784/
314•signa11•13h ago•375 comments

RAM: the forgotten history (2024)

https://blog.coredump.cx/p/memory-the-forgotten-history
91•Luc•2d ago•2 comments

WaveDigger: Dig into wireless signals to discover their physical locations

https://github.com/christianrowlands/wavedigger
12•882542F3884314B•1d ago•1 comments

ArXiv receives multiyear commitments to support it as an independent nonprofit

https://blog.arxiv.org/2026/09/23/arxiv-receives-multiyear-investment/
275•JohnHammersley•17h ago•37 comments

Disney+ and Hulu raise prices by up to 13 percent after doubling profits

https://arstechnica.com/gadgets/2026/09/disney-and-hulu-raise-prices-by-up-to-13-percent-after-do...
106•Brajeshwar•1h ago•123 comments

What Is RLCD? The Secret Behind Jev

https://di-zhang-llm.github.io/blog/what-is-rlcd-the-secret-behind-jev/
34•tnspacetime•4h ago•3 comments

When the Debugger Lies

https://danielmangum.com/posts/when-the-debugger-lies/
48•hasheddan•2d ago•15 comments

The newest ESP32 can run Linux and it's getting close to a Raspberry Pi

https://www.xda-developers.com/newest-esp32-run-linux-close-to-raspberry-pi/
149•adunk•5h ago•66 comments

VSCode's SSH Agent Is Bananas (2025)

https://fly.io/blog/vscode-ssh-wtf/
290•Rapzid•19h ago•185 comments

Meta takes down a critical video about meta AI Glasses after filming at Meta

https://www.reddit.com/r/facebook/comments/1wotwrk/meta_takes_down_a_critical_video_about_meta_ai/
545•pieterr•8h ago•320 comments

Contrastive Language Models

https://contrastive-lm.notion.site/
137•erichocean•12h ago•39 comments

The Year of Internal Tools

https://www.geocod.io/code-and-coordinates/2026-09-23-the-year-of-internal-tools
49•thecodemonkey•9h ago•11 comments

The "Windows XP Box" (2003)

https://www.mini-itx.com/projects/windowsxpbox/
207•doubletwoyou•2d ago•43 comments

Coulomb's law remains tricky to test at home

https://chillphysicsenjoyer.substack.com/p/coulombs-law-remains-tricky-to-test
13•surprisetalk•2d ago•9 comments

Fixing the Portobello Police Station Clock

https://pointinthecloud.com/2026-04-11-211700.html
502•avidly•1d ago•112 comments

Hackers influence ChatGPT and Gemini to direct users to scam centers

https://medium.com/@arielsimon/dark-sourcery-how-hackers-manipulate-ai-to-scam-you-88df434d2073
94•ArielSimon•4h ago•30 comments

Virtio-nvgpu: Near-native Nvidia GPU access inside a KVM guest

https://github.com/nestrilabs/virtio-nvgpu
143•WanjohiRyan•15h ago•58 comments

Best LLM for every budget, updated daily

https://bestmodelforyourbudget.terrydjony.com/
127•terryds•2h ago•80 comments

Owners mourn spoiled food after firmware update bricks Samsung smart fridges

https://arstechnica.com/gadgets/2026/09/owners-mourn-spoiled-food-after-firmware-update-bricks-sa...
194•nonfamous•3h ago•198 comments

Why 'What's Opera, Doc?' looks like that

https://animationobsessive.substack.com/p/why-whats-opera-doc-looks-like-that
118•CharlesW•1d ago•22 comments
Open in hackernews

Japanese used bookstores see 5x sales surge as books are being bought by the ton

https://www.tomshardware.com/tech-industry/artificial-intelligence/japanese-used-bookstores-see-5x-sales-surge-as-books-are-being-bought-by-the-ton-one-50-ton-order-sent-to-the-us-for-ai-scanning-and-destruction-multitude-of-suspicious-bulk-buys-thought-to-end-up-in-foreign-ai-scan-and-shred-facilities
57•speckx•1h ago

Comments

duchanjo•1h ago
Could this be for AI training data?
ErneX•1h ago
It’s on the headline if you visit the link.
panny•1h ago
A dark age will come. AI shredders destroy all the books then hallucinate what they once contained.
manarth•54m ago
Once upon a time, Hansel and Gretel were walking through the woods when they met a Sleeping Beauty called Snow White. As they tried to wake Beauty, a naked Emperor walked in screaming "Off with his head" before a Big Bad Wolf started huffing and puffing.
aprilthird2021•29m ago
I'm the opposite of an AI doomer but this is actually scary to me. Once knowledge is hollowed out like this how can we get it back?
MarkusQ•22m ago
We could go out in the world, have experiences, cogitate upon them, learn to write well, and then do so? It ain't easy, but that's the way it used to be done.
mistercheph•7m ago
This is unironically part of their business plan: don't just regurgitate the world's information but also destroy all other sources of information.
t1234s•1h ago
The value in these AI companies will be more in their proprietary training data than the models.
RobotToaster•1h ago
The sad part is, imagine how positive this could be if the scans were made available to the public.
IrishTechie•58m ago
Might add insult to injury for the publishers/authors though?
doublerabbit•55m ago
A book, song that has been out in public-domain for more than 10 years, should be downloadable. Even if you were to rebuy the same CD again the amount that the artist received would be a pointless pittance. Why not just let it be free to be enjoyed by all?

OCR Scanned for training, then tossed away or burnt. Great for nature.

RobotToaster•54m ago
Maybe for current in print books, but the concern is about rare books being pulped in this process. If a book is rare then it isn't in print, so nobody is making money out of it.
nightpool•31m ago
Unfortunately copyright law does not have a squatters-rights exception
ChickeNES•17m ago
I've believed for years that it really should have one, or at least a "you aren't selling this to the public at a reasonable market-rate price (or via subscription, I care about access far FAR more than ownership), you lose all rights to it" regime.
genxy•58m ago
It doesn't matter what language the tokens are in, now Eye of Sauron seeks to consume all knowledge.
pfdietz•44m ago
How does one "consume knowledge"?
wccrawford•43m ago
Well, in this case, I imagine they mean by shredding books after scanning them. Since that's what's happening.
pfdietz•37m ago
That's not consuming knowledge, that's consuming cellulose and ink. If it were, then printing another copy of a book would be "producing knowledge".

BRB, going to do "programming" by copying source files to another directory. Look how productive I can be.

rjh29•54m ago
Good memories of visiting used bookshops near Kyoto University with stacks upon stacks of obscure literary works and research material. Lots of interesting books about the Japanese language that were never digitized and I always left with 2-3 new books. So I'm not super happy about AI companies hoovering this all up and not making the scans available.
zirkonit•50m ago
Our family volunteers at a nonprofit that moves a huge number of books. We take donations and run massive charity sales, clearing tens of thousands of books a month. Pricing works like a ladder: you try to sell a book for a couple of bucks, then for a dollar, then by the $5 bag, then for free, and you still end up with thousands of books nobody wants even at no cost. These used to go straight to pulp. Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

I'm sure the Anthropics of the world have shredded the last known copies of plenty of titles. But nobody talks about the last known copies that get pulped every day simply because there are so many books nobody wants and storing them is expensive.

pixl97•37m ago
Yea, there are a ton of people that seem they'd rather the books get lost forever than be looked at by an AI company.
szszrk•30m ago
My wife loves bulk book hauls. Those places that we frequent, work like a literal permanent discount warehouse - books in high shelves, on pallets, everywhere. Often hundreds of issues of the same one.

But there is so many books there that no one want's to read. Hundreds of the same book lying there for months or years.

Same for public book-sharing "libraries" (small shelves that look like bird house, usually in parks etc). People really like them and there are many in my city, but most books there are products of a gone era and a gone mindset. No one want's that even for free.

We were taught respect for books, but not everything is worth preserving.

ghaff•10m ago
I didn't go this year but my local town library has a book sale where books go for something like $10/bag. I donate some books to them throughout the year. There are still a lot of books available on the last or second to last day of the sale. I'm sure a huge number get pulped.
tsylba•49m ago
Ah yes, the litteral destruction of culture and physical media for a centralised subscription service. I love the liberal world of techno enclosures of our new overlords, viva el free market economy.
Analemma_•47m ago
What do you think happened to all these used books before the AI companies showed up?
alightsoul•32m ago
They sat in boxes
ghc•30m ago
Not quite:

> But what happens when sales numbers don't meet projections? The book is discounted. Then, at the publisher's discretion, the bookstore will receive a directive to rip the covers off the books, recycle the remainder of the book to be "pulped" or turned into other forms of paper, such as notebook paper and toilet paper. The bookstore is expected to mail the book covers to the publisher as evidence that the book has been destroyed.

https://www.offthebeatenshelf.com/blog/pulp-fiction-is-real

sebmellen•21m ago
TIL that's where the term "pulp fiction" comes from!
alightsoul•8m ago
patall•47m ago
Can someone explain why the old books could really be relevant. I get the pre-nuclear steel analogy, but why is this relevant given how much more modern texts exist. A few years ago, millions of yahoo groups were erased but now a few thousand books are what is needed to run a successful AI company? I mean, it can barely be about the information in those books (that would be very often outdated), but just for a little more text (with ever less marginal gain), what is the benefit?
nemomarx•38m ago
Writing style maybe?
layer8•31m ago
Diversity. Modern books with modern content and in modern styles are overrepresented, old ones underrepresented.
aisenik•31m ago
Cognition is encoded in language, they weren't brain-damaged yet. There's better (real, not token-exchange) thinking, which LLMs can copy and reproduce in novel arrangements.
aprilthird2021•31m ago
The information density is a lot less for millions of yahoo groups. They are far more likely to cover the same topics and not have new information in them.

Books are more likely to be about a specific topic or story or time or setting and be more information dense

hyperhello•
YVoyiatzis•23m ago
Pretty much as it happened with vinyl records twenty years ago. I remember seeing photos of this guy somewhere in Brazil standing atop heaps of vinyl records which he had amassed with HDLR intention. Now books. Some of us hold on forever.
mistercheph•13m ago
You are a moron: vinyl records came and went in about 25 years, books have been the engine of human progress for at least the last 3 millenia, the books being burned by these misanthropic lunatics are not available in any other medium, this is not about fascination with some particular mediumn of transmission, it's about the contents
ChickeNES•8m ago
"the engine of human progress"? I think you mean capitalism.

> the books being burned by these misanthropic lunatics are not available in any other medium

prove it, name one title

> this is not about fascination with some particular mediumn of transmission

it very much is. this fetishism of books should really stop, especially when ebooks are more useful, durable, etc.

sly010•27m ago
Well, they could turn the bad faith story into a good faith story by making them available for everyone to download perhaps. (AI companies "saving" old books!) But that would require giving a s*t which they don't and that is the real problem imho.
Aurornis•25m ago
They legally cannot do this.
theroadnotbacon•21m ago
That certainly hasn’t stopped them before… IP theft is kind of their whole thing, isn’t it?
Aurornis•18m ago
Distributing copyrighted works (prior to expiration of their copyright) verbatim is illegal.

Training an LLM on copyrighted works is not illegal.

This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.

sly010•5m ago
That would require effort (to sort, acquire copyright, etc) which they wouldn't put in. Because they don't care.

People obviously feel bad about companies doing this. People reading these stories don't care what's legal, they care what's ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.

icantevenhold•24m ago
Why do they shred them at all instead of donating or selling them again?
quickthrowman•24m ago
A judge ruled it was OK to scan books and save the scanned copy if you shred the physical book afterwards.
shagie•10m ago
The judge ruled that format shifting was fair use. The fair use argument was enhanced because the original was destroyed. Hypothetically, one could do only the format shifting and keep the original - but the original could not be sold or donated or otherwise given away because then the format shifted copy would not be fair use.

I can burn DVD copies of my old VHS tapes. I cannot then give away the old VHS tapes or sell them at a garage sale. If I keep them, they're cluttering the shelf... so the VHS tape gets thrown away afterwards.

andrew_lettuce•21m ago
My understanding is they cut off the binding for scanning. They'd have to resell it donate by the page
laybak•16m ago
I have a similar thought too each time I walk past piles of discount books.

I'm in the camp that perhaps it's healthy to not grasp onto every bit of information. that some artifacts dying a natural death is maybe just the way things are

palmotea•7m ago
> Now they go to AI labs for scanning. Would we rather they were read, or at least owned, by someone? Yes. Is scanning better than turning them into toilet paper? Yes, even if only marginally.

Yeah, now instead old books being turned into toilet paper, we'll get turned into toilet paper.

And Sam Altman will become richer than God, and isn't that what really matters?

But don't worry! You'll still have access to ChatGPT until your savings run out.

ChickeNES•2m ago
Do you have an actual point about scanning the books, or are you just using this as a soapbox to rant about AI and Sam Altman?
That's an American take. In Latin America they sit in boxes and you can see it, because the pages will have been discolored after so many years of sitting somewhere waiting to be sold. They're only recycled if they're given away at book exchange events and no one wants them.
ghc•32m ago
"Once great literature—now great litter."
quickthrowman•26m ago
Feel free to buy books by the ton and preserve them yourself. The simple fact they’re being sold by weight implies they’re not rare or unique.
mistercheph•9m ago
The simple fact that the AI labs are spending billions of dollars to acquire and scan them implies they are rare and unique.
31m ago
They are the memories of the productive part of society. You can leaf through them and get the feel of what it was like. You don’t need most of your memories, personality, or core ideals to be productive to the State.
Legend2440•30m ago
It sounds like it is about the information in the books. The titles they're looking for are all nonfiction. Not everything is on the internet, and just because it's a few years old doesn't mean it's outdated.

Speaking from experience, the information density of published books is a lot higher than most internet text. It's very high quality training data.

The goal here is to have all human knowledge in a single file, which is pretty neat IMO.

ChickeNES•19m ago
> the information density of published books is a lot higher than most internet text

I'm not convinced. I think you are under-weighing the massive volumes of stuff like self-help books, romance novels, etc.

criemen•17m ago
They're targeting non-fiction books, so romance novels would be out.
scottyah•3m ago
Is that a policy change after o4 got a little out of hand?
newsy-combi•19m ago
The internet basically never delivered on the promise of replacing textbooks or even education as a whole. Wikipedia sucks on many topics, has insane internal politics, and is a tertiary source by design (redigesting blogs and books), whereas textbooks are generally secondary.
zardo•14m ago
Also the ability to set the training input limit in the past could be useful.
timcobb•26m ago
I'm guessing the more integration tables they consume, the better they become at integration.

I also guess that they're targeting languages that aren't tier one for them yet. Like, Japanese is probably a relatively small corpus for them.

Aurornis•20m ago
> but now a few thousand books are what is needed to run a successful AI company?

They're scanning millions of books.

It's the diversity of text that helps. One of the lessons we've learned is that more training data leads to better models. Even old books have different mixes of word sequences that will improve the model. The returns are diminishing, but when you have the pipeline set up to ingest it you might as well keep adding to the dataset.

theroadnotbacon•18m ago
I also wonder if it’s used for text generation in image models! Awful lot of typefaces, sizes, orientations, and words in those books.
mistercheph•16m ago
> why the old books could really be relevant > it can barely be about the information in those books (that would be very often outdated),

LOL, not being rude: have you ever read a book outside of what they forced you to read in school? Most old books are not O'reilly's manuals for Visual Studio 2014, they don't go out of date.

They are interesting to human beings for the same reason they are interesting to the labs. If it was just about quantity of text then the labs could generate text with the prev. gen model and use that alone to scale to the next model, there is something of immeasurable value contained in books (hint: it starts with an i and rhymes with bin formation).

ijk•7m ago
Because Yahoo in particular was very good at destroying goldmines shortly before they became ultra valuable.

There's a number of places that destroyed vast amounts of data in the wind down of ZIRP that probably regret it now.