frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Another researcher says OpenAI trained on conversations, then claimed breakthrou

https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/post/3mv4mt4ikss2d
84•ColinWright•54m ago

Comments

techblueberry•53m ago
But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?
Robotbeat•39m ago
Neither? Competitive academic researchers are susceptible to exaggeration and self-aggrandizing, and CEOs are that and also mostly psychopaths. I tend to think there isn’t systematic spying on researchers looking for breakthroughs. A lot of people are looking for the same things using similar approaches.
techblueberry•32m ago
I mean the accusation is that they were using private ChatGPT conversations. Given the extent to of the gold rush and the long history of Silicon Valley stealing ideas, and arguing it’s not immoral, It almost seems like your making the exceptional claim that this is the one time where Silicon Valley didn’t use information that was at their disposal.

Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.

bigstrat2003•10m ago
If Sam Altman tells you the sky is blue, you should double check. I certainly hope nobody believes him when he claims controversial things from which he stands to benefit.
ColinWright•43m ago
Here's the original Mastodon post:

https://mathstodon.xyz/@andreasthom/117240535270608201

hn1rig3rak•43m ago
the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.
spindump8930•30m ago
The canary string was more about inadvertent scraping or analysis in other papers. Not direct training on user data. And the use of BB has eroded quite a bit, with BB-Hard or other variants being typically used.
jrflo•34m ago
I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:

> Improve the model for everyone

> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.

It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

omnicognate•32m ago
Not unticking a box in settings doesn't constitute consent in my opinion. I'd never put anything I value into ChatGPT anyway, though.
rfgplk•23m ago
Under EU rules it doesn't constitute consent.
nmfisher•25m ago
There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem".

I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).

spindump8930•22m ago
"Improve the model for everyone" can be implemented in so many ambiguous ways.

https://news.ycombinator.com/item?id=49643513

rfgplk•24m ago
Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
jeremyjh•12m ago
Public domain doesn’t mean anyone can assert copyright. It specifically means no one can.
cyanydeez•8m ago
also, none of it means anything without the lawyers to back it up. Just like you can be a pedophile in the highest office of democracy and escape persecution.
spindump8930•23m ago
Reminder that there are degrees of "trained on conversations". From John Schulman:

> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper

> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this

> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"

source: https://x.com/johnschulman2/status/2097440545853637108

rfgplk•22m ago
This would cease to be a problem if OpenAI remained true to their founding motto and... actually open sourced their training/inference pipeline.
Ydarbleoj•9m ago
This is a reminder based on believing what these companies say.

I’ve lived long enough to know what they say and what they do are often quite different; and it is not our job to trust but to verify.

mrbluecoat•23m ago
"Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]"

Welcome to the party, with the rest of humanity.

shevy-java•18m ago
Considering how Apple today announced that its products will contain a spying-on-conversation anti-feature by default, it seems reasonable to assume that all this spying is primarily done to train their AI model; and secondarily also to gather information about The People.

AI is becoming more evil by the day.

azinman2•4m ago
What are you talking about?
gentlerain•18m ago
So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training?

How do people become that trusting?

The phrasing itself is guilt tripping

quentindanjou•10m ago
We are asking people to become experts in all domains rather than providing a safe context through regulations and laws. I don't like thinking the issue is people, I am a person myself, and I often do mistakes on things I don't want to be an expert at but I do believe I should be in a safe context and not have to worry about every single thing.

Or at least: tell me I should be careful/worry about those particular things.

cyanydeez•9m ago
The grift economy requires all marks to be responsible for the fraud perpetrated by others.
speak_plainly•8m ago
Coincidentally, a tweet from OpenAI's Tibo yesterday:

https://x.com/thsottiaux/status/2097746417012166816

jarofgreen•10m ago
Also discussed here: https://news.ycombinator.com/item?id=49638353

Original posts

https://mathstodon.xyz/@andreasthom/117240535270608201

https://mathstodon.xyz/@andreasthom/117240536885387540

https://mathstodon.xyz/@andreasthom/117240537520615623

This story is someone on bluesky screenshotting someone on twitter without a link. The person on twitter screenshotted the source on the fediverse without a link. This is getting daft.

qg127•9m ago
There are so many naive academics. They still believe an "opt-out" button.

Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.

Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.

alansaber•8m ago
I think the heart of this issue is: people assume they have anonymity in numbers, but we have the tools to make it easy to scoop your data if it's interesting to the company.
postalcoder•5m ago
[delayed]
square_usual•4m ago
I think this is stupid, for three reasons:

1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.

2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.

3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)

ColinWright•19m ago
I refer you to this:

https://news.ycombinator.com/item?id=49643556

Quoting:

> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."

asimpleusecase•10m ago
Old Facebook trick - likely resetting that box each time the app is updated.
fithisux•11m ago
Ok, you shut it down, or that is what they make you believe. You give the instruction to shut down, you can't know if it has been applied.

Hitachi launches CO2 heat pump water heaters with solar-friendly tariff controls

https://www.pv-magazine.com/2026/09/07/hitachi-launches-co2-heat-pump-water-heaters-with-solar-fr...
102•thelastgallon•23h ago•77 comments

Show HN: The same nine streaming subscriptions cost $702/year more than in 2021

https://honestlyranked.com/guides/streaming-price-increases/
263•honestlyranked•3h ago•227 comments

What algorithm did Windows XP use to choose your initial user picture?

https://devblogs.microsoft.com/oldnewthing/20260909-00/?p=112683
188•soheilpro•4h ago•90 comments

List of references on Sony websites to players "owning" their digital games

https://consumerrights.wiki/w/Sony_PlayStation_digital_game_ownership_lawsuit
72•haunter•1h ago•9 comments

iPhone Duo

https://www.apple.com/iphone-duo/
1286•thecosmicfrog•19h ago•2242 comments

DeepSeek v4.1 Flash

https://twitter.com/deepseek_ai/status/2097930608790167907
625•Liwink•7h ago•335 comments

Tell HN: OpenAI keeps re-enabling the 'allow training' setting

24•jacquesm•21m ago•4 comments

Another researcher says OpenAI trained on conversations, then claimed breakthrou

https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/post/3mv4mt4ikss2d
86•ColinWright•54m ago•36 comments

Show HN: What if the speed of light was 5 km/h?

https://rivendell.dmitrybrant.com/relativity/
467•dmitrybrant•12h ago•197 comments

Stockfish 19

https://stockfishchess.org/blog/2026/stockfish-19/
139•atiedebee•2d ago•94 comments

Liesegang Rings

https://chillphysicsenjoyer.substack.com/p/liesegang-rings
13•surprisetalk•1d ago•1 comments

Shopify acquires Tailwind

https://tailwindcss.com/blog/tailwind-is-joining-shopify
1090•EdwinHoksberg•1d ago•419 comments

What do Visa and Mastercard do? An intro to card networks

https://tautology.town/2026/06/01/card-networks.html
595•evakhoury•1d ago•361 comments

Object storage is all you need

https://www.tigrisdata.com/blog/object-storage-all-need/
42•jpsaccount•1d ago•25 comments

Thanks to Siri Recaps, your Apple Watch is always listening

https://www.techradar.com/health-fitness/smartwatches/thanks-to-siri-recaps-your-apple-watch-is-a...
98•geox•3h ago•62 comments

Larger Pacific Striped Octopus

https://en.wikipedia.org/wiki/Larger_Pacific_striped_octopus
99•olalonde•1d ago•48 comments

To write non-fiction, draw the trunk, then the rest of the tree

https://devz.cl/posts/how-to-write/
9•DanielVZ•1d ago•1 comments

Who People Talk to When They're Struggling

https://www.graphsaboutreligion.com/p/who-do-you-talk-to-when-youre-struggling
16•toomuchtodo•1h ago•8 comments

Growing proof that autonomous cars save lives

https://spectrum.ieee.org/are-self-driving-cars-safe
392•bookofjoe•20h ago•680 comments

No Man's Sky Cosmos

https://www.nomanssky.com/cosmos-update/
421•Limb•22h ago•421 comments

Samsung Debuts zHBM Prototype, Stacking Memory Directly on AI Accelerators

https://www.thelec.net/news/articleView.html?idxno=12835
42•peter_d_sherman•3d ago•7 comments

GPT-6 Astra, looped transformers, and hidden reasoning

https://magazine.sebastianraschka.com/p/gpt-6-astra-looped-transformers-and
470•ModelForge•23h ago•153 comments

Show HN: Art – draw one stroke, let symmetry complete it

https://mrdee.in/mandala/
32•cyb0rg0•4d ago•17 comments

AirPods 5

https://www.apple.com/newsroom/2026/09/apple-introduces-airpods-5-with-best-in-class-open-ear-act...
486•awad•20h ago•413 comments

iPhone 18 Pro and iPhone 18 Pro Max

https://www.apple.com/newsroom/2026/09/apple-debuts-iphone-18-pro-and-iphone-18-pro-max/
390•meetpateltech•20h ago•435 comments

Automattic's board forces CEO Matt Mullenweg into leave of absence

https://techcrunch.com/2026/09/09/automattics-board-forces-ceo-matt-mullenweg-into-leave-of-absence/
377•LeoPanthera•14h ago•248 comments

Aardman (Wallace and Gromit) Is Selling Its Original Movie Puppets

https://gizmodo.com/aardman-is-selling-its-original-movie-puppets-this-month-2000807968
71•Gaishan•3d ago•18 comments

Desert Ant Labs: local, fast models that run on device

https://desertant.com/blog/introducing-desert-ant-labs/
472•willwhitedc•1d ago•97 comments

Show HN: Compute polynomials twice as fast

https://thomasahle.com/fast-polynomials/
110•thomasahle•1d ago•34 comments

Qwen 3.8 follows GPT-5.5 Pro reasoning prefills

https://gist.github.com/wsxiaoys/e0286dc6bb624ff5fdf49e7f4c528ba3
228•wsxiaoys•20h ago•89 comments