frontpage.
newsnewestaskshowjobs

Open Source @Github

fp.

Open in hackernews

Another researcher says OpenAI trained on conversations, then claimed breakthrou

https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/post/3mv4mt4ikss2d
64•ColinWright•40m ago

Comments

techblueberry•40m ago
But who are you going to believe? Multiple independent academic researchers or the CEO who was fired two years ago for gross dishonesty?
Robotbeat•25m ago
Neither? Competitive academic researchers are susceptible to exaggeration and self-aggrandizing, and CEOs are that and also mostly psychopaths. I tend to think there isn’t systematic spying on researchers looking for breakthroughs. A lot of people are looking for the same things using similar approaches.
techblueberry•19m ago
I mean the accusation is that they were using private ChatGPT conversations. Given the extent to of the gold rush and the long history of Silicon Valley stealing ideas, and arguing it’s not immoral, It almost seems like your making the exceptional claim that this is the one time where Silicon Valley didn’t use information that was at their disposal.

Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.

ColinWright•30m ago
Here's the original Mastodon post:

https://mathstodon.xyz/@andreasthom/117240535270608201

hn1rig3rak•29m ago
the fix is boring and known: BIG-bench shipped a canary GUID for exactly this, and you publish your decontam n-gram threshold (gpt-3 used 13-grams). no threshold disclosed, no claim.
spindump8930•16m ago
The canary string was more about inadvertent scraping or analysis in other papers. Not direct training on user data. And the use of BB has eroded quite a bit, with BB-Hard or other variants being typically used.
jrflo•21m ago
I pay for the Pro ChatGPT plan, and if you go to settings > data controls this is the first setting:

> Improve the model for everyone

> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.

It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.

omnicognate•19m ago
Not unticking a box in settings doesn't constitute consent in my opinion. I'd never put anything I value into ChatGPT anyway, though.
rfgplk•10m ago
Under EU rules it doesn't constitute consent.
nmfisher•11m ago
There's a difference between "this is allowed under their ToS" and "it is academically unethical to fail to credit the people whose specific conversations were fed into a model that was used to solve a problem".

I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).

spindump8930•9m ago
"Improve the model for everyone" can be implemented in so many ambiguous ways.

https://news.ycombinator.com/item?id=49643513

rfgplk•11m ago
Under current understanding of the law, anything produced purely by LLMs (with no substantive human input, which is what OpenAI claimed in their post) is firmly in the public domain. So OpenAI can "claim" anything they want, it doesn't make it reality. In fact if I were the original authors I would just take their 400k lines of lean proof and relicense it under their own names/terms.
spindump8930•10m ago
Reminder that there are degrees of "trained on conversations". From John Schulman:

> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper

> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this

> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"

source: https://x.com/johnschulman2/status/2097440545853637108

rfgplk•9m ago
This would cease to be a problem if OpenAI remained true to their founding motto and... actually open sourced their training/inference pipeline.
mrbluecoat•10m ago
"Another researcher[/artist/writer/musician/programmer/doctor/director/etc] says OpenAI trained on conversations[/imagery/books/songs/code/classifications/videos/etc], then claimed breakthrou[gh/original art/bestselling books/chart-topping songs/unique applications/medical advice/free special effects/etc]"

Welcome to the party, with the rest of humanity.

shevy-java•5m ago
Considering how Apple today announced that its products will contain a spying-on-conversation anti-feature by default, it seems reasonable to assume that all this spying is primarily done to train their AI model; and secondarily also to gather information about The People.

AI is becoming more evil by the day.

gentlerain•5m ago
So people genuinely believe that toggling that "Improve the model for everyone" button makes their data safe from being used for training?

How do people become that trusting?

The phrasing itself is guilt tripping

ColinWright•6m ago
I refer you to this:

https://news.ycombinator.com/item?id=49643556

Quoting:

> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."

Using a 1930 Teletype as a Linux Terminal (2020) [video]

https://www.youtube.com/watch?v=2XLZ4Z8LpEE
1•Phileosopher•52s ago•0 comments

Paul Riechers – The shape of beliefs and abstraction – UCLA [video]

https://www.youtube.com/watch?v=WUnCb0TyHVA
1•binyu•52s ago•0 comments

Bending Spoons to buy Miro in $1.36B cash deal

https://www.reuters.com/legal/transactional/bending-spoons-buy-miro-136-billion-cash-deal-2026-09...
1•matthieu_bl•1m ago•0 comments

Show HN: Orchestrator, spawn and manage Claude Code instances remotely

https://github.com/markusbug/Orchestrator
1•mhaaseth•1m ago•0 comments

Ryze Wave Hacked wide open

https://github.com/davidbuzz/RyzeWaveWatch
1•davidbuzz•3m ago•0 comments

Agon, Alea, Mimicry and Ilinx

1•boxesnlines•4m ago•0 comments

I build Garmin watch faces with AI agents

https://myday24.com/blog/how-i-build-garmin-watch-faces-with-ai-agents/
1•tomasslavicek•5m ago•0 comments

Ask HN: Did you ever worship your career or identity?

1•general_reveal•5m ago•1 comments

Population ethics is a big deal, which is why I made this population ethics quiz

https://mdickens.me/2026/08/31/population_ethics_quiz/
1•surprisetalk•6m ago•0 comments

Your Agent Speaks MCP. Give It a Computer.

https://fly.io/blog/sprites-mcp/
1•torutofu•6m ago•0 comments

Kagi Translate Is Back

https://blog.kagi.com/translate-is-back
2•twapi•6m ago•0 comments

Opinion: We have started losing control of AI. It's time to shut it down

https://www.theguardian.com/commentisfree/2026/sep/10/ai-control-sci-fi
2•drayfield•7m ago•0 comments

Tell HN: OpenAI keeps re-enabling the 'allow training' setting

4•jacquesm•8m ago•0 comments

Rust Is Tier-1 Language at Microsoft

https://rustfoundation.org/media/guest-post-rust-is-tier-1-language-at-microsoft/
3•mmastrac•8m ago•0 comments

Go 1.27 Release Notes

https://go.dev/doc/go1.27
2•nkjoep•8m ago•0 comments

LRU is harder to beat than the KV-cache papers suggest

https://github.com/gauravapiscean/agentic-kv-cache
2•gauravapiscean•8m ago•0 comments

Hiring Staff Machine Learning Engineer

https://careers.elcompanies.com/careers/job/1168273194779
1•jfeola•8m ago•0 comments

The Same Crypto Exchange Matching Engine Measured 130,000/s and 106/s

https://www.ovasylenko.com/blog/matching-engine-performance-benchmarking
1•_alphageek•11m ago•1 comments

Show HN: JavaScript grid and pivot library built for coding agents

2•yonl•12m ago•0 comments

Grafista: Graphic design platform fighting AI slop

https://grafista.io/
1•GiorgosGennaris•12m ago•0 comments

Before NTP there were Time and Daytime

https://www.jeffgeerling.com/blog/2026/rfc-867-868-time/
1•corvad•14m ago•0 comments

Aubusson Weaves Tolkien

https://www.cite-tapisserie.fr/en/le-musee/les-aventures-tissees/aubusson-tisse-tolkien
2•nfeutry•14m ago•0 comments

Some (mostly historical) issues with the Unix load average

https://utcc.utoronto.ca/~cks/space/blog/unix/LoadAverageHistoricalIssues
1•signa11•16m ago•0 comments

I ported a VB6 game to Flutter and generated all 14 art themes

https://apps.apple.com/cz/app/kirian/id6774868017
1•lioil•16m ago•0 comments

Show HN: Implementing Embedding Gemma in PyTorch

https://www.youtube.com/watch?v=rxOUmSJTQBk
2•prasoon21•17m ago•0 comments

OpenAI Gives US Agencies 50% Off Models, Ending $1 per Year Deal

https://www.bloomberg.com/news/articles/2026-09-10/openai-gives-us-agencies-50-off-models-ending-...
1•helsinkiandrew•19m ago•1 comments

Privacy on Bitcoin: what works and what doesn't

https://www.learnbitcoin.com/rabbit-hole/bitcoin-privacy
1•granya•21m ago•0 comments

iTerm 2 Companion App

https://iterm2.com/companion-app.html
1•fredley•23m ago•0 comments

Comparing performance of different offline first solutions

https://github.com/bkniffler/offline-sync-bench
2•quambo•24m ago•0 comments

AI Doomlord Jacob Coxon's Media Tour Has Begun

https://gizmodo.com/ai-doomlord-jacob-coxons-media-tour-has-begun-2000809720
11•jjj123•26m ago•1 comments