frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Scientific Literature Is Poisonous to LLMs

https://www.reinvent.science/p/the-scientific-literature-is-poisonous
25•surprisetalk•52m ago

Comments

ghostly_s•39m ago
> Imagining a two-by-two matrix with the axes honest vs. dishonest and right vs. wrong, the scientific literature is splashed haphazardly across all four boxes.

Why should we give an extraordinarily claim like this any credence from the authors of a newly-minted substack who can't be assed to write more than a blurb on the topic, and whose listed credentials amount to a cagey statement that isn't even clear on whether they hold degrees?

cwmoore•30m ago
How does your appeal to authority impact their claim?
pandinus•38m ago
This just in: fallible Humans (un)knowingly produce unreliable data. LLM training on unreliable data impacts accuracy. Some sources of Human-produced data are measurably more reliable than others.

"Trust the science!"

sebastianconcpt•36m ago
Then is poisonous to us too.
nemomarx•31m ago
Isn't everything like this? Ask an AI about news and it has to deal with contradictory and poorly labeled accounts of an event. Ask it about mechanics in a game and it has to deal with every version of it, maybe different editions or remakes, etc. What dataset exactly is pure and nicely labeled for correctness? Is it large enough for training?

isn't this why we went to synthetic corpuses anyway

iLoveOncall•31m ago
Not as much as LLMs are poisonous to the scientific literature (or, really, any literature).
nekusar•30m ago
This article is barely even legible, and chains unrelated papers in some big scare.

Regardless of slop, it's definitely no/low quality and a bunch of breathless garbage.

Flagging and warning others.

light_hue_1•25m ago
Ironic that "Reinvent science" would publish such a trash article. Did they even read the original paper? This is not at all what it says.

https://arxiv.org/pdf/2305.13169

The "poisonous scientific literature" removal is on page 14. Look at the table. Removing academic pieces changes performance by on average 0.44%! You could sneeze and change the performance by that much. You could rerun with a different seed and change the performance by more than that. You could retrain on a different GPU that orders floating point operations differently and change performance by that much. etc. This is meaningless.

Also, the authors explain exactly why this happens! It's on that page even. The academic data hurts a little on datasets which aren't academic. It hurts on common sense reasoning like SocialIQA (Q: "Jordan wanted to tell Tracy a secret, so Jordan leaned towards Tracy. Why did Jordan do this?" A: "Make sure no one else could hear"). Shocking that you can't learn this from academic publications.

This is just substandard blog slop that give science a bad name.

The scary part is: "Dan Recht and Ben Reinhardt trained as scientists and now coach scientists. They do other work but don’t link to it here." As someone who has advised plenty of PhD students I can't find the words to express my disdain at that line and these jokers.

tstactplsignore•20m ago
This is really honestly an alarmingly wrong substack post.

1. The claim of the authors is that the scientific literature is on average so dishonest and/or wrong that access to it (pre or post training, I assume?) actually harms LLM accuracy. Hopefully we can all agree this is a claim so extraordinary that would require some very uniquely powerful evidence. After all, we know that LLMs also train on very large corpuses of text with varying degrees of both accuracy and dishonesty - not the least, most of the internet! So the claim of the authors must be that the scientific literature is so bad that it is actually uniquely bad for LLMs, like worse than the general internet.

2. They provide one citation of evidence supporting their claim, which is this paper [0]. If you actually read this paper (which is mostly not about this actual question, but related questions about prefiltering), it provides no evidence for their claims. For example, in Figure 5, removing PubMed, aka the entire biological scientific literature, has by far the strongest negative impact on evaluation of Biomedical questions. It even has a strong negative impact on evaluation of questions in the "Common sense" category! To quote:

"Performance degrades when we remove domains with close alignment between the pre- training and downstream data sources: removing PubMed hurts the BioMed QA evaluations"

3. This post is maybe what you would get if you prompted an LLM: "please provide a citation for the claim that the scientific literature harms LLMs". It's really quite worrying that we're looking at AI generated propaganda designed to convince the reader that the scientific literature is worse than useless.

4. As a scientist, the primary use I want from LLMs outside of code generation is to be a fast and comprehensive search engine. I want to see the papers behind their claims and evaluate them. Literally the only time they are useful in scientific research is when they provide citations for the claims, so I can read the papers and evaluate.

5. Somehow we have gotten to the point where many people, especially people in tech, believe that the scientific literature is mostly junk or mostly useless. This contrasts so starkly with the current rapid pace of genuine, meaningful, society-impacting scientific progress in pretty much every major field. How we have gotten to the point where this myth is so pervasive, I do not know. Is it really spurred just by a few recurring news stories about reproducibility? By just interpersonal bitterness and feelings of anti-'elite' sentiment? (how on Earth your local state university climate scientist is a member of the 'elite' but software engineers making $500k a year or X influencers with millions of followers are not will always simply be beyond me).

[0]. https://aclanthology.org/2024.naacl-long.179/

How to design NATS subject hierarchies

https://www.synadia.com/blog/designing-nats-subject-hierarchies
1•jonzu•33s ago•0 comments

Const_cast: A Necessary Evil

https://www.elbeno.com/blog/?p=1858
1•jandeboevrie•57s ago•0 comments

Two new sermons by St Augustine discovered

https://www.uni-wuerzburg.de/en/news-and-events/einblick/single/news/two-new-sermons-augustine-di...
1•leopoldj•3m ago•0 comments

Matching files to apps: how UTIs work

https://eclecticlight.co/2026/07/29/matching-files-to-apps-how-utis-work/
1•Brajeshwar•3m ago•0 comments

Advice for tech writers on a job hunt in the AI age

https://passo.uno/job-hunt-tech-writers-ai/
1•theletterf•5m ago•0 comments

Show HN: College application essay inspiration and brainstorm

https://essaycompass.net/
1•boveyking•5m ago•0 comments

I Want to Leave the Internet

https://chupacabra.bearblog.dev/i-want-to-leave-the-internet/
2•Looky1173•7m ago•0 comments

Ask HN: Experience with MSFT Security Copilot and the new response capabilities?

1•tty46•7m ago•0 comments

Archaeopteryx: A Ruby MIDI Generator [video]

https://www.infoq.com/presentations/archaeopteryx-bowkett/
1•e-topy•7m ago•0 comments

Amazonian civilization had estimated 3M people in 3% of forest area

https://www.science.org/content/article/odd-shapes-hidden-dense-amazon-rainforest-reveal-sprawlin...
1•marojejian•8m ago•1 comments

My Mom Teaches a Course on ADHD

https://theborderofnormal.substack.com/p/my-mother-teaches-a-course-on-adhd
1•Netherland4TW•8m ago•0 comments

Why We're Dropping Basecamp (2023)

https://blogs.library.duke.edu/blog/2023/11/30/why-were-dropping-basecamp/
2•ksec•8m ago•0 comments

Microsoft Faces UK Probe over Copilot-Linked Price Hikes

https://www.bloomberg.com/news/articles/2026-07-29/microsoft-faces-uk-probe-over-copilot-linked-p...
1•sbulaev•8m ago•0 comments

Moneion Build Production-Ready Apps with AI

https://moneion.com
1•Developer_2•9m ago•0 comments

Rust: Stabilize and Model Polonius Alpha

https://github.com/rust-lang/rust-project-goals/issues/118
1•ksec•9m ago•0 comments

Oryxflow – cheaper and more reliable AI data analysis in Python and Claude Code

https://github.com/oryxintel/oryxflow
1•citynorman•9m ago•0 comments

GitHub Actions Native Egress Firewall – Early Access

https://github.com/github-early-access/actions-native-egress-firewall
1•cebert•10m ago•0 comments

MinIO pitches persistent memory for agents with work to finish

https://www.theregister.com/storage/2026/07/29/minio-pitches-persistent-memory-for-agents-with-wo...
2•bjflanne•10m ago•0 comments

Why does every mammal get 1B heartbeats in their life? [video]

https://www.youtube.com/watch?v=tL9Lw250spc
2•Brajeshwar•10m ago•0 comments

Gemini 2.5 models are being deprecated and discontinued

https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/model-versions
1•thebestmoshe•10m ago•0 comments

Typing latency benchmarks across modern code editors

https://koieditor.com/benchmarks/
1•hackermanai•10m ago•0 comments

Covid.gov – Lab Leak: The True Origins of Covid-19

https://www.whitehouse.gov/lab-leak-true-origins-of-covid-19/
3•moralestapia•11m ago•1 comments

Show HN: Rivora – An open-source memory layer for engineering tools

https://github.com/rivora-dev/rivora
1•sgr0691•11m ago•0 comments

Who Cares About the Model?

https://ampcode.com/news/who-cares-about-the-model
1•tosh•13m ago•0 comments

PwC published reports on AI marred by AI hallucinations

https://www.ft.com/content/7e149ac8-2ce2-4266-8940-192f9821b33c
1•1vuio0pswjnm7•13m ago•0 comments

Arbitrary file read and remote code execution in Active Storage

https://discuss.rubyonrails.org/t/cve-2026-66066-possible-arbitrary-file-read-and-remote-code-exe...
2•baggy_trough•17m ago•0 comments

What if competitor research pointed at the empty seat instead of crowded ones?

https://aswespeak.bsct.so/
1•aminkhorrami•18m ago•0 comments

Show HN: RepoInPeace – A marketplace to buy and sell abandoned startup IP

https://repoinpeace.app/
1•chelski•18m ago•0 comments

Three Things You're Getting Wrong About Religion Data

https://www.graphsaboutreligion.com/p/three-things-youre-getting-wrong
1•toomuchtodo•19m ago•2 comments

Kimi K3 and GLM 5.2 can create undetectable malware for $2

https://www.incalmo.ai/blog/2026/06/26/glm-malware/
7•aoli-al•19m ago•0 comments