frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

The Scientific Literature Is Poisonous to LLMs

https://www.reinvent.science/p/the-scientific-literature-is-poisonous
25•surprisetalk•1h ago

Comments

ghostly_s•47m ago
> Imagining a two-by-two matrix with the axes honest vs. dishonest and right vs. wrong, the scientific literature is splashed haphazardly across all four boxes.

Why should we give an extraordinarily claim like this any credence from the authors of a newly-minted substack who can't be assed to write more than a blurb on the topic, and whose listed credentials amount to a cagey statement that isn't even clear on whether they hold degrees?

cwmoore•38m ago
How does your appeal to authority impact their claim?
pandinus•46m ago
This just in: fallible Humans (un)knowingly produce unreliable data. LLM training on unreliable data impacts accuracy. Some sources of Human-produced data are measurably more reliable than others.

"Trust the science!"

sebastianconcpt•44m ago
Then is poisonous to us too.
nemomarx•39m ago
Isn't everything like this? Ask an AI about news and it has to deal with contradictory and poorly labeled accounts of an event. Ask it about mechanics in a game and it has to deal with every version of it, maybe different editions or remakes, etc. What dataset exactly is pure and nicely labeled for correctness? Is it large enough for training?

isn't this why we went to synthetic corpuses anyway

iLoveOncall•39m ago
Not as much as LLMs are poisonous to the scientific literature (or, really, any literature).
nekusar•38m ago
This article is barely even legible, and chains unrelated papers in some big scare.

Regardless of slop, it's definitely no/low quality and a bunch of breathless garbage.

Flagging and warning others.

light_hue_1•33m ago
Ironic that "Reinvent science" would publish such a trash article. Did they even read the original paper? This is not at all what it says.

https://arxiv.org/pdf/2305.13169

The "poisonous scientific literature" removal is on page 14. Look at the table. Removing academic pieces changes performance by on average 0.44%! You could sneeze and change the performance by that much. You could rerun with a different seed and change the performance by more than that. You could retrain on a different GPU that orders floating point operations differently and change performance by that much. etc. This is meaningless.

Also, the authors explain exactly why this happens! It's on that page even. The academic data hurts a little on datasets which aren't academic. It hurts on common sense reasoning like SocialIQA (Q: "Jordan wanted to tell Tracy a secret, so Jordan leaned towards Tracy. Why did Jordan do this?" A: "Make sure no one else could hear"). Shocking that you can't learn this from academic publications.

This is just substandard blog slop that give science a bad name.

The scary part is: "Dan Recht and Ben Reinhardt trained as scientists and now coach scientists. They do other work but don’t link to it here." As someone who has advised plenty of PhD students I can't find the words to express my disdain at that line and these jokers.

tstactplsignore•28m ago
This is really honestly an alarmingly wrong substack post.

1. The claim of the authors is that the scientific literature is on average so dishonest and/or wrong that access to it (pre or post training, I assume?) actually harms LLM accuracy. Hopefully we can all agree this is a claim so extraordinary that would require some very uniquely powerful evidence. After all, we know that LLMs also train on very large corpuses of text with varying degrees of both accuracy and dishonesty - not the least, most of the internet! So the claim of the authors must be that the scientific literature is so bad that it is actually uniquely bad for LLMs, like worse than the general internet.

2. They provide one citation of evidence supporting their claim, which is this paper [0]. If you actually read this paper (which is mostly not about this actual question, but related questions about prefiltering), it provides no evidence for their claims. For example, in Figure 5, removing PubMed, aka the entire biological scientific literature, has by far the strongest negative impact on evaluation of Biomedical questions. It even has a strong negative impact on evaluation of questions in the "Common sense" category! To quote:

"Performance degrades when we remove domains with close alignment between the pre- training and downstream data sources: removing PubMed hurts the BioMed QA evaluations"

3. This post is maybe what you would get if you prompted an LLM: "please provide a citation for the claim that the scientific literature harms LLMs". It's really quite worrying that we're looking at AI generated propaganda designed to convince the reader that the scientific literature is worse than useless.

4. As a scientist, the primary use I want from LLMs outside of code generation is to be a fast and comprehensive search engine. I want to see the papers behind their claims and evaluate them. Literally the only time they are useful in scientific research is when they provide citations for the claims, so I can read the papers and evaluate.

5. Somehow we have gotten to the point where many people, especially people in tech, believe that the scientific literature is mostly junk or mostly useless. This contrasts so starkly with the current rapid pace of genuine, meaningful, society-impacting scientific progress in pretty much every major field. How we have gotten to the point where this myth is so pervasive, I do not know. Is it really spurred just by a few recurring news stories about reproducibility? By just interpersonal bitterness and feelings of anti-'elite' sentiment? (how on Earth your local state university climate scientist is a member of the 'elite' but software engineers making $500k a year or X influencers with millions of followers are not will always simply be beyond me).

[0]. https://aclanthology.org/2024.naacl-long.179/

Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac

https://github.com/drumih/turbo-fieldfare
135•gitpusher42•1h ago•28 comments

KOReader

https://koreader.rocks/
464•Cider9986•5h ago•153 comments

Handbook.md shows that long policy documents do not reliably govern agents

https://arxiv.org/abs/2607.25398
169•spIrr•3h ago•101 comments

Hugging Face: Anatomy of a frontier-lab agent intrusion

https://huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space/index.html
41•dn2k•1h ago•5 comments

Launch HN: Tokenless (YC S26) – Automatic model switching to save money

https://usetokenless.com/
9•rohaga•28m ago•4 comments

Hamburg's Stadtpark: A Park Built to Be Used

https://alsterrunde.com/hamburgs-stadtpark-a-park-built-to-be-used/
17•mertbio•2d ago•1 comments

Shipping Godot VR and Porting to PSVR2: A Partial Post Mortem

https://www.claire-blackshaw.com/blog/2026/07/shipping-godot-vr-and-porting-to-psvr2-a-partial-po...
56•ibobev•3h ago•0 comments

Hunter-gatherers introduced fish to a mountain lake 7000 years ago

https://www.newscientist.com/article/2580119-hunter-gatherers-introduced-fish-to-a-mountain-lake-...
75•stevenwoo•2d ago•48 comments

Cesium DevCon 2026 talks are up, including a keynote from SQLite's creator

https://cesium.com/events/cesium-developer-conference/2026/
14•jasteinerman•1h ago•1 comments

Darktable

https://www.darktable.org/
127•siatko•3h ago•66 comments

Document-borne AI worms can self-propagate through Copilot for Word

https://enklypesalt.com/posts/context-collapse-part3-ai-worming-through-word/
202•Canopy9560•4h ago•173 comments

More Tailscale tricks for your jailbroken Kindle

https://tailscale.com/blog/jailbroken-kindle-proxy-tun-modes
344•Error6571•11h ago•100 comments

Self-hosting Kimi K3: 20% more hardware cost, 20% better task resolution

https://aistack.imec-int.com/blog/gpu-self-hosting
18•flifenstein•1h ago•2 comments

GPT-5.6 vs. Claude Fable 5 for Physical AI, which performs best?

https://juliahub.com/blog/frontier-models-physical-ai-evaluation
22•mbauman•1h ago•4 comments

Amiga Graphics Archive

https://amiga.lychesis.net/index.html
110•Bluestein•6h ago•17 comments

Disrupting supply chain attacks on NPM and GitHub Actions

https://github.blog/security/supply-chain-security/disrupting-supply-chain-attacks-on-npm-and-git...
57•nyku•4h ago•18 comments

User Interfaces of the Demo Scene

https://www.datagubbe.se/scenegui/
342•zdw•11h ago•57 comments

Superlogical – Mitchell Hashimoto

https://mitchellh.com/writing/superlogical
8•tambourine_man•38m ago•0 comments

Show HN: Qwen Scribe – local transcription and dictation for Apple Silicon

https://github.com/VladUZH/qwen-scribe
13•sidclaw•1h ago•3 comments

AI in Linux

https://drewdevault.com/blog/AI-in-Linux/
22•surprisetalk•2h ago•11 comments

SQLite in Production: Optimizing WAL Mode, Concurrency, and VFS Layers

https://micrologics.org/blog/sqlite-in-production-optimizing-wal-mode-concurrency-and-vfs-layers-...
192•ankitg12•9h ago•58 comments

SpecForge – A Platform for Authoring Formal Specifications

https://docs.imiron.io/v/0.5.10/en/tour.html
61•agnishom•5h ago•7 comments

A.I. Companies Are Recruiting Electricians and Carpenters by the Thousands

https://www.nytimes.com/2026/07/29/business/economy/data-center-electricians-training.html
54•thm•1h ago•29 comments

Show HN: Write, simulate and synthesize VHDL/Verilog in the browser

https://risingedge.pro
14•wozniakpawel•6d ago•4 comments

Ask HN: My domain registrar (Hover) rug-pulled me for $3000

13•shrinks99•37m ago•8 comments

A Texture Lookup Approach to Bézier Curve Evaluation on the GPU (JCGT)

https://jcgt.org/published/0015/02/01/
29•ibobev•3h ago•4 comments

Lisp moving Forth moving Lisp

https://letoverlambda.com/textmode.cl/guest/chap8.html
92•fallat•2d ago•22 comments

Google shuts down Nobel Prize winning AlphaFold

https://www.engadget.com/2225849/google-shuts-down-alphafold/
30•NordStreamYacht•1h ago•1 comments

San Francisco: Don't Fall for Industry Defense of Surveillance Pricing

https://www.eff.org/deeplinks/2026/07/san-francisco-dont-fall-industry-defense-surveillance-pricing
89•hn_acker•3h ago•35 comments

Graph Engineering Needs a Compiler

https://fluxtion-playground.dev/blog/2026-07-29-graph-engineering-needs-a-compiler
8•v12technology•1h ago•1 comments