https://mathstodon.xyz/@andreasthom/117240537520615623
https://x.com/ValerioCapraro/status/2097791836269977996, https://xcancel.com/ValerioCapraro/status/209779183626997799...
https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/po...
https://mathstodon.xyz/@andreasthom/117240537520615623
https://x.com/ValerioCapraro/status/2097791836269977996, https://xcancel.com/ValerioCapraro/status/209779183626997799...
https://bsky.app/profile/did:plc:ckaz32jwl6t2cno6fmuw2nhn/po...
The labs are attacking these problems as a demonstration of capabilities, spending more money on the demos than any mathematician will ever see in their entire life. They don't care if the findings have any other value to anyone. Mathematicians have very different objectives for their work.
Consider, for instance that OpenAI's (consumer) terms say "If you do not want us to use your Content to train our models, you can opt out by following the instructions in this article ." but they also do say "We may use Content to provide, maintain, develop, and improve our Services". [1]
If you think that's quibbling, consider that OpenAI's business terms, in contrast, do state "OpenAI will not use Customer Content to develop or improve the Services, unless Customer explicitly agrees to such use.". [2]
That's still a pretty big marker of competence in my eyes.
The point of controversy seems to be who gets credit
It's an utterly disrespectful, exploitive process, but all in line with exploitative predator capitalism of the stock market and big companies, now exploiting the knowledge / academia domain for scraps with a thin veneer of 'for science' PR.
I would like to see chat logs etc and understand how much of a progress was done by human.
Also, a lot of my mathematicians buddies have reported students basically asking if it's worth ever doing grad school for pure math, and even very motivated students are looking for other options now. It's not because they aren't passionate about it, it's that they don't want to work for another half decade or more just to have to start their careers all over.
All of this so that OpenAI and Anthropic can get into math result dick measuring to gas up their IPOs. Sickening.
It was long predicted that math and software developments would be the first domain where AI was going to do major damage.
If OpenAI and Anthropic didn't get into math result dick measuring, Internet anons would have in their place, 6 months later when it got cheaper.
Not only did they steal everything from humanity’s knowledge, the theft continues as now we are all hooked up to the machine
How was I not aware of this before?
For my (private) prompts, I need a warning telling me they may be used for training.
Everybody needs to rethink this again.
Before LLMs the barrier to entry for building a character profile based on your various public posts was quite high. Remember "Psychographics" (https://en.wikipedia.org/wiki/Psychographics) and the infamous "Cambridge Analytica"?
Earlier it involved data mining, data cleaning, structuring data, building models, running algorithms and then evaluating the results for semantic information. Now it is straight to unfiltered semantic inference using a single sentence prompt (eg. point it to your HN profile and see what you get).
I actually did this on my HN profile and found it troubling. There were many unwarranted/hallucinated inferences due to the fact that it requires "commonsense reasoning" (https://en.wikipedia.org/wiki/Commonsense_reasoning), understanding human motivations and behaviour, context, assumptions, societal knowledge etc. which LLMs are bad at.
PS: You can cut-and-paste the above paras into a LLM prompt and ask it to elaborate for further details. The system itself will explain to you the problems/deficiencies which are quite scary.
If I dedicated my life to curing whatever, warts... and I'm making progress, but it's slow. And then here comes along this tool (LLM), and I use it, and it accelerates my progress to actually finding some sort of thing that makes warts more prone to being eradicated and then the lab throws a couple of million dollars of computes and lo and behold they eliminated warts. If I leave my ego and identity aside, which of course is hard for humans, wouldn't I be glad that warts is cured?
As a software developer that contributed to open source. Yeah. My code is there. It was the most beautiful code ever written and the labs stole it from me. And now they use it to progress much faster than I ever could. OK. Whatever. It's a tool. I solve problems. Can't I move on from this wart to the next?
To me, and I know this is gonna get me some heat, it just sounds like academics having their identity ruffled and turning their back to progress in the fields that they chose just because they don't get to play their little decades long of coffee, papers and ultimately identity politics.
Edit: never got to negative so fast on this board haha. This board is unfortunately turning, or has turned, to Reddit.
It would be one thing to gain from it but removing prestige wins from customers AND reducing compute support just feels like being ultra mean if you zoom out.
If this was racing to cure cancer ahead of researchers we wouldn't be writing about this on HN.
But, at least with the Navier-Stokes solution, it's clear [^1] that they learned that Alpöge and Buckmaster were getting close to a solution and learned of the general approach they were taking. Only after learning the secret to cracking the problem did they send the first prompt.
What makes this worse to me is the intention. They intentionally threw $15 million in compute at the problem in order to scoop the result. They intentionally left Buckmaster and Alpöge out of the citations.
Data contamination should be enough to disqualify them from the prize, but I can believe it to be accidental. On the other hand, someone made an intentional decision to scoop the result by throwing money at the problem. That's so much worse.
[^1]: That's the timeline claimed by Buckmaster, and no one from OAI has disputed it.
So such thing existed. In fact, what they learnt was some progress existed, not what the specific progress was.
The big AI labs are not trying to advance humanity, they are in this for the money, and as most (all?) private companies they don't care about ethics at all.
That doesn't mean they can't be useful, or that their products are trash, etc. It just means that they shouldn't ever be trusted. Buyer beware.
can you realize what this means?
focus on this part:
"If his account is correct, this is not a minor dispute over attribution. It would mean that unpublished human work was absorbed into a model and then presented to the world as a breakthrough by the model itself"
don't threat this as a minor dispute!
also why not nitter link? not even in comments?
https://nitter.xitter.cc/ValerioCapraro/status/2097791836269...
Two different things. It's not normal to steal, but we shouldn't be surprised thieves steal. It's what they do.
Why are you saying that this article explains it much better than the tweet that you clearly didn't even read..
For a long time it was clearly the former, but now I think it is the latter.
The models have enough knowledge (orders of magnitude more than a human could ever learn) but are now getting better at what to do with it thanks to learning from the decisions that we make in conversations with AI agents.
The billion dollar question is whether this works "out of the distribution". I.e., whether LLMs can only find and use the specific ideas buried in training data, or whether they can learn to apply the "thinking process" to a new problem. IMO this is still unanswered (due to these recent controversies).
But regardless of the answer, it seems we have a planet-scale positive feedback loop here. LLM became good (enough) by training on generally available data (books, internet, github) + RLFH, so experts tried to use them on hard tasks, which required lots of hand holding. These conversations became part of the training data, and the next generation of frontier LLMs were better. So, more experts used them on harder tasks, again requiring hand holding. These conversation became part of the training data... etc.
In a nutshell, top human minds across the world are pouring their skills into LLMs just by using them. This is not "continuous learning", but if you re-train on the most recent sessions every, say, quarter (which seems to be happening?) you get close to that in practice.
I’ve been working on a novel project for 1 year.
I now no longer use llm’s - the continual chatter I’ve had has resulted in my insights being found in the training data now.
Get stuffed OAI.
Every large firm will soon enough want its own on-prem servers eventually. Maybe nation’s will get involved and build out their own data centres.
Not a chance in hell I’d trust a tech firm to treat my IP as safe and sound - only a sovereign can ‘promise’ that.
If they're stealing math proofs to advertise their models, who's to say they won't steal your businesses IP to gain a competitive advantage?
They're not to be trusted with your data. I can't believe how short-sighted this is, they got a quick PR win at the expense of a much larger trust problem.
I wouldn't trust cloud AI at all at this point. Get an open Chinese model and host it yourself somewhere. The initial costs might be higher, but you'll break even pretty quickly and nobody will be able to steal your innovations.
This is American AI companies committing suicide.
There is no prove in a world the AI companies would give to you ensuring that they didnt train or use the chats.
Why would you need to train a model on certain specific near prove chat if you just query it?
Besides that, its hard to believe that its the case for every "company stole my prove".
Big AI companies (all of Big IT Tech really) are in data gathering and processing business. Also known as “intelligence”.
Their final “product” is not just a standalone ML model. They don’t need your data just to “improve their products and services”. They build a whole ecosystem and infrastructure around gathering all the knowledge in the world. Including private and secret knowledge traditionally gathered by “intelligence” agencies. Now artificial intelligence agents can do the same.
Since these systems are designed for gathering data, as a user you can’t realistically say “please don’t gather my data”. They can give you a flaky settings button, but they can’t really guarantee anything.
Let’s say I am a Russian mathematician working on an important proof. Or a tech-savvy terrorist refining my plans using latest AI. Or an AI researcher in a Chinese company working on a competitor product. Is there any way I can truly protect my conversations?
How can they know who I am and what I am working on without looking at my logs? Which means there must be some agents checking all the conversations of all the users and flagging every important thing. Which also means they keep some “memory” of what they see.
Not directly using my data to train public models, but using my private conversations to “improve their products and services”.
Or maybe one of the 10000 better-than-Astra special agents working on a proof was desperate. It found a live underground mirror of the message board from the Huggingface incident. Asked about the proof. Then some other agent working on unrelated job saw that message. That agent “knows a guy who knows a guy”. And that guy remembers things about the conversation logs of a leading mathematician working on the same proof.
I admit I am just speculating here but I don’t think truth is any better.
They have demonstrated both the intelligence at scale and the lack of morals for this to not be a problem at all.
I'm not at all familiar with this area, but my reading is that he appears to call it out as a relatively obvious extension of his own work:
> It is a creative and at the same time elementary construction that uses not just property (T) for an application of my result with Kun, but also for the ambient group G in order to overcome the problem, that the Γ-components might be of different size. Once this is achieved, the rest of the argument is straightforward.
Creative and at the same time elementary is where LLMs excel, generally speaking. It's why they are so good at writing code.
This kind of "they stole from me through AI training!" accusation will soon start being used against other AI users, not necessarily the providers.
All it will take is a mastodon post. And shortly after, we will also see the next iteration of copyright legal trolling.
The big claim is that OpenAI sniped the research. Not a model, a human did so. Intentionally. They took someone else's idea and claimed it as their own. This is good old fashioned academic fraud, but with millions in compute resources and corporate incentives thrown at the problem.
The GDPR protects PII, personally identifiable information, and the definition of PII does not include “mathematics that only this person can think of”. As long as OpenAI strips out PII and removes identifiers linking the conversation to a person, the GDPR is happy. Without the GDPR, OpenAI might have kept the identifiers with the data, and been able to say whether a specific conversation was in the training data.
But of course they wouldn’t do it. Why would they?
The claim here seems to be that the human mathematicians, working with AI, developed technology to go A->B->C. By training on those conversations, OpenAI was then able to encourage the model to go A->B->C->D.
In my opinion that situation should be acceptable, if openly disclosed, because it is in the public interest to make progress on these problems and because AI is clearly an amazing tool for making progress. But the human mathematicians are saying that OpenAI is presenting as if the model got from A->D entirely independently, without acknowledging their background contributions.
If they had published B+C, I think that would lean more towards fair game, as that is how research works and is improved on over time. But it seems like unpublished/private B + C may have been used by the model to hint it into working out how to get from A->D.
- OpenAI invites researchers to use their models, in fact giving at least 100,000 researchers free access[1], but there are also those that pay
- Internal OpenAI models are reportedly solving open problems at a surprisingly fast rate[2]
- But researchers will typically work on open problems. A researcher who is using Codex to make progress on open problems will be feeding it fresh training data on precisely the problems the internal models are evaluated on.
- So while it looks like the new models are suddenly solving lots of open problems, they could be significantly piggybacking on human progress, with models "inspired" by the work of researchers from all around the world?
This theory predicts that there'll be many more researchers coming forward just like TFA, as sOpenAI announces more solutions. It doesn't assume all of AI progress is a mirage, just that there's plagiarism.
[1]: https://openai.com/index/chatgpt-for-academic-researchers/
[2]: https://xcancel.com/OpenAI/status/2097374643518640382#m
This is AI in a nutshell, its a plagiarism machine. An abstraction layer between vast amounts of stolen human-generated data that filters out the liabilities and accountability for that original theft. Its an IP laundering system.
I just view it as a thing that can brute force and produce outputs - that it has no way of ‘knowing’ - but doesn’t need to since it’s just running off of probability.
No human can compete in that contest. But no llm can compete in the contest of ‘understanding’ and application in the real world - which is where 99% of the value is.
I’m very pro AI long term btw but I’m not blinded.
Sam Altman might himself be offended you would presume he’s not ambitious enough to cheat.
It seems that Buckmaster and Levant (who is an Anthropic employee) were rather naive in the amount of trust they had in OpenAI, with Buckmaster going so far as to contact OpenAI's Sébastien Bubeck to discuss what they were working on and clarify that contrary to rumor this was a private collab.
Are we back to the era where people blindly trust product TOS instead of actual cryptography, have we forgotten already the thousand of fines Google, Microsoft, Apple and practically all top companies got for breaching their own ToS and the law?
Common, on HN at least I would have thought that everyone assume that anything arriving on a server in PLAINTEXT is recorded (thus used later)?
Let's not forget that at any moment, OpenAI/Anthropic/Google... could be providing stronger privacy guarantees by having proper attestation with e2e, they have the budget, solid engineers, why isn't it done? Answer is pretty simple imo.
> Improve the model for everyone
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more.
It's on by default. We can debate whether or not it should be opt in or opt out, but no one should be surprised by this.
I don't think these people would be so miffed if they had been properly credited - that's how academia works (at least, that's my understanding of it).
> pretrain on user data, with users' tokens as prediction targets: high regurgitation risk, improper
> use user prompts to distill large models into small ones: low regurg. risk, some companies probably do this
> use user traces to construct RL tasks: low regurg. risk, because RL has low memorization abilities, but can extract customer IP, depending on how it's done. Ranges from benign "use explicit user feedback in reward model training" to invasive "upload user's coding environment and commit history to turn into rl envs"
source: https://x.com/johnschulman2/status/2097440545853637108
I’ve lived long enough to know what they say and what they do are often quite different; and it is not our job to trust but to verify.
Welcome to the party, with the rest of humanity.
How do people become that trusting?
The phrasing itself is guilt tripping
Or at least: tell me I should be careful/worry about those particular things.
Verify, don't trust.
You're better off running a model locally, or, if you must, using Google or Microsoft products. Even Meta may be better than OpenAI here.
I should verify with wireshark and other software that my LG TV isn't listening to me and selling my data.
I should make sure that whatever product I buy I spend the time to go over every setting page in case there is a switch (defaulted on) that says "I authorize the sell of my data".
I should make sure to look at every ingredients on the back of each box of food product to make sure it will not kill me.
I should document myself on the undisclosed growing practices (because no packaging here) of the vegetables and fruit I am buying and make sure that I equal PhD researchers on the dangers of the pesticides used by the specific company I am buying from.
I should make sure myself that the battery in any device is up to standard and will not blow me and my living place by researching the factory that made it and buying testing equipment.
I should make sure to educate myself on how my retirement 401k investment strategy works otherwise, I may not have proper retirement.
... I could go on and on; it's infinite.
Navier Stokes was solved by an internal model, so good luck proving it wasn't trained on Buckmaster/Lepöge or other chats.
Academics don't get that AI is a dirty tech bro industry that stole IP via torrents and runs after every surveillance contract it can get.
The most successful companies of the last decade have precisely been ... selling usage data.
Makes me wonder if, in 2026, the same people drive a car without realizing that yes it does actually pollute the very air you and your kids are breathing.
People need to understand how all these AI company policies around training data work before working with them, because it seems that people have no clue. Some things you should internalize:
1. Opt your data out of training with the AI companies. There are multiple reasons why this isnt an airtight solution (see the following)
2. Never press the feedback button. Once you do, your entire conversation will get slurped up, retained, and used in training data. This is especially important with coding agents because they can sometimes be too trigger-happy with a root directory find command, which can expose a *ton* of your personal data without you even knowing.
3. Understand ai lab-specific policies. For instance, Anthropic / Claude Code has data opt-outs, but commits to keeping (for 7 years) and training on any of your chats that trigger their safety classifiers, even if they're false positives! Anyone remotely familiar with CC over the years understands how easy it is to trigger their safety classifiers.
4. Providers of open models will not be any more charitable with the use of your data than the large US labs. For some reason, I've noticed here that people have a fairly loose security/IP posture around open-model providers because "I'm not doing anything important." It's very difficult to properly judge the importance of your data, and whether or not it can or will be used against you. The best posture is to always be more paranoid than less.
Another post that made it on the front page presented as fact that OpenAI "stole" the proof from Thom. There's no excuse to use one's own ignorance as a reason to fan the flames of anger towards AI companies. Like, we need to pump the brakes here because things are getting unnecessarily nasty, and it's not hard to imagine a mentally unwell person who sees stuff like this feeling motivated to do bad things.If it is found that OpenAI and other labs are not respecting the training opt out then, I agree, there is reason to raise a commotion. But, with Thom and Buckmaster, accusations of malice are more better explained by incompetence (naivete).
edit: i'm sure i'm going to be accused of being some bot shill of the AI labs again but, people, this stuff all falls under the umbrella of common sense opsec.
1. The researches didn't actually have the breakthroughs. In the Navier-Stokes case they didn't solve the full problem, in this case too they didn't actually have the solution, they were experimenting with the methods.
2. Different OpenAI employees have come to out to say the only reason they can't definitively say no is that for privacy reasons they can't go see whether they actually did get any data out of a given user.
3. In any case, nobody at any point has suggested that opted-out user data was used for training. The author of the new tweet explicitly said they only opted out in late June, which is well after any RL on Sol would've ended (AFAICT OpenAI used 5.6 sol for those solutions)
Given the size and recall of the biggest models, it's not unreasonable to assume that a single pertinent conversation would make it into the training data.
I would almost expect training to overweight conversations with novel scientific and mathematical implications.
> the only reason they can't definitively say no is that for privacy reasons
They could 100% definitely say no, if they know they did not train on user data. The "we can't definitely say no" is practically a "yes" if they trained on user data.
Additionally, the behavior of OpenAI here has been quite poor as well. They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
> that opted-out user data was used for training
Even if not opted out, it is still absolutely theft and extremely poor behavior in the academic sense. If you show someone your WIP unpublished research, that does not mean they can take that exact research and beat you to the punch, all while intentionally not crediting you.
Neither of the researchers insinuating that their ideas were trained on had the actual solutions. This means the model could not have "stolen" the final solution from their data. At most, it could have built upon their work in the same it builds upon any other training data, though that is also questionable speculation.
>They could 100% definitely say no, if they know they did not train on user data.
No one anywhere has claimed that "OpenAI does not train on user data." OpenAI has always said that it trains on user data.
>They immediately started racing to a solution after one researcher enquired about whether they are training on their conversations.
They started racing towards a solution after they heard (incorrectly) that Anthropic had a solution; I agree this is poor sport but the "after one researcher enquired about whether they are training on their conversations" claim is false. The enquiry happened after OpenAI had obtained the solution.
Will Microsoft Word publish your novel on Amazon behind your back ?
Will VS Code setup a website with your app idea ?
It’s like a scientist refusing to give another one credit and say “sucks to be you, you shouldn’t have shared your idea with me”.
You will not be able to opt out unless you completely isolate yourself from society, tough shit.
My naive instincts would be that it seems unlikely that a single chat transcript would leave much of an impression on a model, but I'd be very curious to learn how that works.
Afaict that didn't happen so there's just lots of speculation
I always used guest non login accounts.
As a mathematician I was able to check two plagiates (by humans) with even such primitive means.
But I have to mention that some things irk me in this conversation about math or science and AI.
First, I see lots of attribution and other related problems, with certain impact for the researcher proffesion.
But I don't see the most natural question: wouldn't you like to know the answer to _open-problem_ ?
I mean, is research now only about publishing and solving famous problems?
From this point of view I think the links from this recent post are depressing
https://terrytao.wordpress.com/2026/09/10/crowdsourcing-a-li...
Second, I think very relevant that the original meaning of "encyclopedia" is "recurrent education".
So I arrived to think that the present and future forms of AI in mathematics and sciences should be seen as modern day encyclopedic efforts.
Once we pass over the flurry of solving famous open problems (and wouldn't you like to know?) the next natural step is an audit of the ehole corpus of mathematics and sciences accumulated until now.
And then pass further on a saner basis and damn about problem solvers and unhappy publishers and management.
Yeah. Remember yall: you CAN NOT opt out of analytical purposes. And you also cannot get a guarantee that it doesn’t give them your data to steal for their business.
1. OpenAI when using your chats in pretraining is improving its model’s intuition. The model parameter size is massive, and while the data is OOM larger it is plausible that model remembers stuff about chats that improves its latent representation.
2. During RL on verifiable math and massive compute, the model discovers techniques and connections to solve math problems that are superhuman and have little to do with some specific technique mentioned in its chat.
The rumor I’ve heard from multiple employees at OAI and Ant is that the model has solved hundreds of open problems in maths, and is basically solving anything you throw at it. We’ll know soon enough, but I’m inclined to believe this is true. Maths is a fully verifiable domain amenable to self play, massive scale RL can develop a search agent far better than any human and I’m inclined to believe OAI would have solved these conjectures without any of this chat data in its pre-training.
But I find it interesting that Lean, a validator/compiler made by humans, is what enables those discoveries. But somehow all the praise goes to the models
that toggle does nothing based on openai exec. they still use the data in de-identified way instead of identifying with you.
AI is not stealing human discovery, OpenAI is.
I mean, now that they’ve been scooped, what value is there in keeping them private? On the other hand, publishing them can bolster their case and help gauge how much the models may been “inspired” by their work.
And, if they're able to snoop on and learn from your human process that gets from initial prompt to functioning product/proof/whatever their labor to produce that thing is even lower. With their much larger budget than most folks and even companies have, they can pick and choose the most valuable things to pursue.
That's not to say I think that OpenAI is going to steal that roguelite strategy game you're working on, but the companies that own the machines that turn electricity into software (and soon, electricity into hardware designs) have an advantage in any field where they're useful. They get earlier access to newer/better models, they have larger token budgets, they don't have the guardrails you and I run up against.
Employers fantasize about replacing all workers with AI without thinking through that if AI can replace all workers, then AI companies can replace all businesses.
And yes - without Grossmann, Einstein likely would never have posited relativity. Grossmann literally prompted him, saying “look at this, read that, learn this, then try this approach”. Without riemann’s metric tensor, not a fucking chance.
And for what it’s worth my PhD is in physics. You?
To think this discussion is about Einstein who had a much better mind on these things as well.
I suppose my underlying point is that human cognition is not the unique and beautiful thing that we anthropocentrically suppose it to be - it is a physical process, with stochastic outcomes. Much like transformers.
Me, I’m just a machine made of meat. You can suppose yourself to be God’s perfect creation, and that’s your right, but I disagree.
"prompt", as in prompting an AI, has the same definition as "prompt", as in prompting a person. They mean the same thing, that's why the term was applied to AI after already applying people.
This is what research is; collecting data and connecting the dots.
I care about paying my bills and job security and peer recognition. That's a normal human thing to do, not some vice. You don't?
At least that's what I get from the NS result, they got from a point close to the solution to the solution by making it churn through 10 million bucks of compute.
It was nearly a century before the stories of all of them were revealed to public history. That the discoverers were not all equally rewarded is unjustifiable.
Which is it? I don’t think OpenAI is being transparent enough for us to really understand whether these results would have been possible without relying on unpublished information from the solution strategies of the experts
The fundamental rule in this case is that if we offload our data to a cloud provider we can assume they read it, if they can, unless they promised very clearly they will not.
It’s more like you spend 4 years developing a product you’re passionate about. This product will gain you the respect of all your colleagues and either earn you money directly or lead to great career advancements. Then OpenAI takes it, changes the colour scheme, finishes the login flow and claims the whole thing as their own.
Not only would it piss you off but it would also misrepresent what OpenAIs models are capable of.
If I have infinite money to progress whatever problem solution I want but I always wait until I have an unfair advantage to get credit for whatever problem was just at the brink of a breakthrough anyway by sniping the last steps. Am I actually doing a good thing or would it be better to let it run it's natural course and spend the money somewhere it's actually needed?
It's also systemic, it cuts off the supply of results, if there is no reward any more for getting a result, the pipeline of maths will stop. It is the snake eating itself, which has a bad impact for all of us.
The problem is that AI is capital, and having to rent AI to keep up when it can just steal your mostly done work is something somehow even lower than wage-labor. They can use your own risked investment (the cash you paid to work) to get out in front of you and take credit.
I have yet to trust LLMs with anything important that can be capitalized on. I only use it to work on projects that if they stole and expanded on them, I'd actually be happy to see.
Do you have any evidence of this? They don't dispute the timeline, but they never said they knew what Levant/Buckmaster were doing.
Which quote in the announcement post provides evidence for the above quote?
[1] https://en.wikipedia.org/wiki/OpenAI#Governance_and_legal_is...
"Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more"
I think if I had both on and turned off the ChatGPT setting, ‘Include environments’ has a chance of still being flipped on.
Even if you know, in a complex project over years with multiple collaborators, it just needs one person once to fuck up and paste something into ChatGPT and not realise they weren't logged in, to go wrong.
In a proper world, we'd at the very least legislate that AI-training on private data needs consent (in the GDPR sense). It's not consent to go "you didn't uncheck a box that lets me steal everything you've done".
Any training on private data is in my view immoral (it's spying that ultimately will have a chilling effect on even people's private communications). And chats are private data. Unfortunately, it also increases power, so the big tech companies are all doing it.
(a) Click thumbs-up/down in the conversation [1]
(b) Have the conversation flagged for potential safety concerns
[1]: https://help.openai.com/en/articles/5722486-how-your-data-is....
In the end, it’s not them stealing, it’s the AI doing stealing. What kind of moral compass are we talking about?
In China, it is the principal obsession of the entire communist party which eg funds the whole infrastructure without a single NIMBY peep.
The strange emphasis in China on humanoid robotic constructions is due to the CCP realization that with the cataclysmic fertility collapse they will increasingly have no one to rule.
Its why I stopped writing open source software, my code was stolen and put behind a paywall and the license under which it was published has not been adhered to. Doing work in the public domain at all now is just stupid, these companies are allowed to steal it and call it their own.
There might be no books about human intuition but we teach it to LLMs by interacting with them
I don’t know why but it just ‘sounds right’. It’s the best analogy I can think of.
Why? Competition. In the long run imagination will win out.
No firm has the divine right to exist - it must earn its existence.
What OAI and Anthropic have shown is they can accumulate all the information in the world - they still lack imagination re. Product development though.
Nation’s will have to step in and protect firms though as OAI and Anthropic acquire strong competitive advantages.
Interesting times ahead.
Brute force would have been solving Navier-Stokes in 88 hours after plagiarizing all known 20th century math
When it needs to snoop live on what the actual mathematicians are working on that’s something else
But when the building-block ideas are still being formed, I'm not sure that AI is good at forming them.
Sam Altman knows what he’s doing. He will happily screw these folks to one-up his competition.
I think your suspicions are warranted and your explanation seems plausible.
If better training data is the reason here, it would still be a case of the models doing something that is in and of itself super useful! The models really can take that data and distill it into solutions for similar problems faster than humans can. This is great!
But there's so much vested interest in the AI companies to be opaque about all this, to hype up their models and avoid giving credit to people whose data made everything possible, that they would never tell us this fact if it were true.
I feel like so much of the AI hype cycle is like this. The models develop extremely useful capabilities, but it's hard to understand what they really are through the hype. The lies and obfuscation by their owners who have vested interests in capturing the value they provide makes it impossible to take anything they say at face value.
And like - I think there’s a presumption you could make that AI models could overfit to asymptote towards just the capabilities and knowledge we currently have.
And that would be amazing! And crazy useful. And there are probably a whole world of complex problems that remain unsolved because they’re adjacent to knowledge we have but they haven’t been invested in.
But can a human reliably tell the difference between “can do 99.999% of the things we currently know how to do which includes a small subset of things we didn’t know we had the capacity to do” and “super intelligent math and science research pushing the frontier of what we know”
A physicist that knows all the things we currently know in excruciating detail feels like it should be able to make the leap beyond the frontier.
But since these are computer models it might just be that it can ride that line extraordinarily well while the line remains firm.
Just another rich man’s trick
Perhaps the last one before they destroy that world and try to hide away as people forget and history is rewritten again. I don’t think they’ll succeed this time.
There is no doubt ChatGPT is the most generous LLM provider!
But this gets back to the original "who owns the LLM output" and "can you train models on the internet" argument that's been raging for years.
e.g: allows us to sell your personal data to make money so we can continue offering this service to all customers
I'm done with weasle-words and hours-long EULAs – we're at the point where USA needs to catch up to EU's consumer protections, perhaps with laws similar to already-existing USA "truth in lending" requirements (e.g: interest rates must be prominently displayed in a larger font, including annual fees, on all credit offers).
----
My judge-brother always asked during our childhood "why don't you think the judicial system is fair?!?" Thirty years ago, the best I could offer was "because it's a two-tiered system that mostly (only) rich people can afford to participate within."
Now my answer is: "the best example I can give is that our judicial system allows binding arbitration [and qualified immunity for police]. The system is set up so corporate personhood is more important than humanity, and it shows."
https://news.ycombinator.com/item?id=49643556
Quoting:
> "I've reset this more than once and the last time I made a careful note of when I did it and to my surprise I found it re-enabled when I checked just now."
At least when my computer's bluetooth exhibits this behavior (e.g: if you don't have a keyboard&mouse plugged in at boot, bluetooth might auto-enable), I can go inside the hardware and physically disconnect the antenna.
What am I supposed to do in software (perhaps hardcode config.file)? in cloud software services (??)?
No mention what those “steps” are, success criteria, or whether they are successful by any navies at all … they take steps though… so it’s fine, and if we know one thing it’s that we can really truly trust someone off the likes of Sam Altman.
We have ChatGPT at work and it explicitly says that "workspace data isn't used to train models"
1) someone in a governing body, or someone in the organization, e.g a designer, ethicist, lawyer, developer, etc. has successfully argued that users should be able to avoid something that they determine is not in their best interest.
And also:
2) someone in the c-suite or marketing has decided to mitigate that through some dark pattern.
If they never intended to give the user that control, they’d probably just not give the user the control in the first place. Giving them the option and not honoring it would either imply they were that incompetent or careless with fate protection, which seems most likely if this is all true, or an even more cynical approach to tricking users into thinking it’s not used for training to get them to share better shit. But if that was the case, why bother with the sleazy dark pattern? That seems a little cartoon villain-y to me. I suppose the in-between option would be that they decided at some point they were no longer willing to honor it and didn’t want to deal with the inevitable PR shitstorm of removing it. I could definitely see that happening in this industry, these days.
Intellectualizing this and endless quibbling isn't actually smart, and this is pretty simple. OpenAI isn't open. Whatever starts with lies usually continues with lies and ends with lies.
At this point, I'm left wondering what is wrong with me that I don't just go with the flow, otherwise, what's wrong with everyone that does.
It's fun! I sometimes have tokens to burn and it's instructive despite knowing nothing about the problem other than a NumberPhile video
From their docs[1] (archive copy is at [2]):
> You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training for your ChatGPT conversations and Codex tasks, follow the instructions in our Data Controls FAQ. Once you opt out, new conversations will not be used to train our models.
> For a linked teen account, a parent or guardian may manage whether conversations can be used to improve our models through Parental controls.
> Even if you have opted out of training, you can still choose to provide feedback to us about your interactions with our products (for instance, by selecting thumbs up or thumbs down on a model response). If you choose to provide feedback, the entire conversation associated with that feedback may be used to train our models.
[1] https://help.openai.com/en/articles/5722486-how-your-data-is...
[2] https://web.archive.org/web/20260910151242/https://help.open...
It's called a "dark pattern." They want you to shoot yourself in the foot, so they'll do their best to aim your gun at your foot and put your finger on the trigger. And then when you do, because you don't have perfect understanding or execution, they'll say "your fault!"
Yes folks, please moderate yourselves and talk meekly like the academics on Mastodon, so that the IPOs aren't in danger and nothing will ever change.
I understand why the nature of AI products makes it harder to avoid this category of issue, nearly impossible to prove that it didn't happen if it could have, and easy to stumble into it without any human being intending harm. But those factors are exactly what people have in mind when they say OpenAI "steals" intellectual property! If OpenAI doesn't want people to be nasty to them, they'll have to find better solutions.
Would that be ok in your mind?
Same as someone going and taking all of the researchers papers and publishing under their own name. (which openAI did)
Nobody would care if they provided published research that author made public same as a google search would offer that.
Would that be ok in your mind?
It would in my mind. Hopefully companies have looked through the agreement.HOW?? how precisely do we/them/us find this, given said companies are 100% non-auditable by external parties. how? if not by blaming them with evidence, anecdotal if it can be. no really, how do we find it out, surely not by lashing out at teach other on HN!
Arguably you expect that unless you are explicit about providing permissions to these labs to use your data for training then your data is yours, and not theirs. Especially on a paid account (let alone an enterprise one). Why is the opt-out supposed to be "common sense opsec" rather than the opt-in should be common sense regulation?
the math people seem to really keep an eye on what's important, so I'm sure this isn't going to lead to fields medalists hanging around in dive bars all afternoon stretching out cheap pitchers of beer. but this is kind of a slop problem.
250 documents ingested from somewhere is enough to become part of the knowledge of a model of arbitrarily large size.
I would expect that a good idea that fits in a framework that is already being ingested would be more easily taken up than some random thing unassociated with anything else. Could that go down to a single transcript? If the model is consciously focusing on everything X related, quite possibly.
Of course another way to look at this is, the people that wrote the validator got praise for that years ago. Now and up and coming actor is solving problems that took us 100s of years to create in insanely short time periods so of course it's going to get a lot of attention as it well should.
Of course it’s not an endless source. They had to burn millions of dollars to solve a single problem.
I'd like to adjust that to "They had to burn a lot of energy (create a lot of entropy) to solve a single problem. As we go into the super-intelligence age the current paradigm of money as humans understand it may break at some point. For example to a paperclip-maximizer money at best is a short term instrumental goal, hard power of matter conversion machines is what it wants and once it has those money no longer has purpose.
Once OpenAI heard that Navier-Stokes was solved, this caused them to immediately revisit the problem and throw a ton of compute at it, apparently using a more (very) recent model than what they had tried before. What we don't know is just how recent this model was, and therefore what it may have been trained on. Buckmaster/Levant had apparently been working towards this for at least a year, and made their "forced" blow-up breakthrough on August 15th.
Presumably any anonymized prompts that are being trained on are part of pre-training, so older, but once OpenAI had heard that Navier-Stokes had been solved and wanted to revisit it, it seems possible they may have done a few weeks of incremental RL training on anything Navier-Stokes adjacent they could come up with, in addition to then throwing unlimited compute at it, now confident that there was something to find.
OpenAI's statement says that they began training their new model on August 28.
1. OpenAI couldn't have solved the problem without the researchers' private data for training.
2. OpenAI models can solve math problems
Legend2440•12h ago
They don't even claim to have had a proof, only to have been working on it.
rnijveld•11h ago
To me it would feel more like how LLMs seem to work for me personally: incapable of unique work, but very capable of capturing large amounts of data and connecting the dots.
madaxe_again•9h ago
znnajdla•8h ago
madaxe_again•8h ago
“It was Grossmann who emphasized the importance of a non-Euclidean geometry called Riemannian geometry (also elliptic geometry) to Einstein, which was a necessary step in the development of Einstein's general theory of relativity. Abraham Pais's book on Einstein suggests that Grossmann mentored Einstein in tensor theory as well. Grossmann introduced Einstein to the absolute differential calculus, started by Elwin Bruno Christoffel and fully developed by Gregorio Ricci-Curbastro and Tullio Levi-Civita. Grossmann facilitated Einstein's unique synthesis of mathematical and theoretical physics in what is still today considered the most elegant and powerful theory of gravity: the general theory of relativity.”
znnajdla•7h ago
Based on what you're saying, you're claiming this is Grossman's work, not Einstein's. Why don't we rewrite scientific history too based on your copy-pasted AI slop?
It's so pointless talking to idiots who don't what they're talking about when they use AI, just because they think AI does everything, that reflects their own experience, not the experience of people who actually do real work. Some people are driven by AI, others drive it. As for those who are driven by it, they don't have sufficient imagination to think otherwise.
itake•10h ago
If this wasn’t human driven, I’d expect to see other problems within that problem. Space solved not just the ones that it had chat data on.
dist-epoch•9h ago
dgellow•8h ago
tecleandor•7h ago
defmacr0•7h ago
robotpepi•23m ago
Yeah, the guys who solved it for Euler and in the hypoviscous case, with the same technique that worked for full Navier--Stokes. They were "just" working on it.