Edit: Apt domain.
If local LLMs get "good" enough, people will soon paying for subscriptions to ChatGPT and Claude, which hurts their revenue.
Think you missed a word there.
Given the recent deepseekv4.1 advances - how good of a 3B model can we make to run on an iphone natively? is it good enough to match common muse/dot use cases for consumers? the phone is already always on.. no need for a cloud server.
It is vanishingly rare I ask an older model to do any task. Newer bigger and smarter models will just do the task better.
Therefore, I believe we are nowhere near 'good enough'.
I never drive my steam engine to work these days. It isn't good enough.
Since I have no idea what "spawn-camped" means I gave up reading the rest.
In a PvP (player vs player) game, if you kill a player the moment they spawn into the game arena, that's called "spawn-camping".
It's not that niche, if you've been online a little bit you'd know this expression.
I've been online since circa 1995 (earlier if you count BBSs), and I can't say I did. It's possible to infer its meaning but assuming everyone is on the same circles as one is, is silly.
No way. You'd need to be pretty well versed in gamer lingo. Even more specifically, combative, likely FPS gamer lingo.
Does anyone know if there are any distillation datasets available? I'd love to see these distributed on BitTorrent. I think it's critical that AI be democratized and not isolated in the hands of a few private companies.
It’s sort of like gun nuts arguing that more guns is the answer. I mean, ok, maybe you’re a responsible gun owner or AI user but relying on personal responsibility doesn’t fix systemic problems. There are bad people out there.
They do claim that it violates their ToS, which we can assume is simply correct, since they get to put whatever they want in their ToS.
Given all that, I don't know what the fuss is. Are they supposed to not use the advances that were openly published by Chinese labs? The entire industry is built on a discovery made at Google, which was published openly. Should Chinese labs therefore not use transformers? Should US labs not try to prevent distillation of their models?
Ah brings back Halo 2 memories
Feels like Anthropic crying do as I say not as I do.
it's not complex. there's hundreds of billions of investor dollars counting on vendor lock in and walled gardens
If I'm a business and I need something done today, and bc Anthropic has the best model, there's a 99.9 chance it will be completed successfully for $1000. And using Deepseek there's a 70% chance it will, for $10 - you or me will go for the $10. Big businesses don't. Bc 1000 per task is nothing to them.
I don’t think it is intentional but this is actually quite bad for the western labs.
The entire booster narrative has been “look at how their revenue is growing! $10bn to $100bn ARR in under a year! This’ll be a multi-trillion IPO!” and the extrapolated future growth from $100bn to $500bn and $500bn to $1tn justified future investment… but that revenue was just because inference was expensive.
The revenue growth story is all that matters pre-IPO. If revenue falls from $100bn to $50bn that’s very very bad optics for OpenAI and Anthropic even if they are now profitable, it completely destroys the growth narrative.
"Thus the expert in battle moves the enemy, and is not moved by him."
They figured out a clever method for avoiding excessive training costs via distillation. That forces the hand of frontier labs to move faster, produce better models, etc. (to avoid embarrassment and 'falling behind'—all the while shouldering most of the cost), which they can just keep distilling—or applying other techniques against—much to the dismay of said frontier labs.
Checkmate.
I think this explains why they are open sourcing broadly. It's not to be nice. It's a strategic play by the Chinese government to help ensure there are many players in this race and not too much power accumulates to American labs (even if American labs benefit in the process)
I realize "going as fast as we can" is not the most popular position atm. But I'm far more interested in what good we can do than 10% apocalypse scenarios. I volunteer with a charity for childhood brain cancer and I do not want to see another 4 year old die. I'm willing to risk anything to stop this.
How co-designed are these optimizations with the model itself? I'd imagine you can't just stick post-training adapters onto existing architectures for these things, or am I wrong?
I really want to explore the inference space, but it seems like many of the inference optimizations are coming from model-hardware codesign. I don't seem to recall many generic "inference engine" optimizations since prefill/decode disagg a year ago.
This matters for me since I want to break in but the bar seems to be understanding the actual theory of the training process now too given the codesign happening, and I'm not the richest guy on the block lol
Because contrarily to the author's assumption, all labs, Western or not, have sufficient skills to discover the optimizations anyway, and publishing or not is not actually that important?
I'm glad people are saying this out loud, because that is what they want. Not for the good of the world, but for the good of their pockets.
Not that OpenAI, Anthropic or SpaceX aren't doing the same.
Do you do business in China?
I'm curious what you mean by this? Because in my experience, you can only do business in China by doing "China" things.
I'd be interested in picking your brain as to how you get around those issues?
I'm talking like, getting 100x, VC sized returns. Of course you can sell widgets in China, it's a major world economy.
The premise in this article is: Western companies do a ton of expensive work building new models, meanwhile the Chinese companies just wait for a Western release and then they immediately grab and distill it and announce it as their own model. That’s the spawn-camp.
"spawn-camping" is the process of taking out your enemies at the point they spawn (or appear) in a game without giving them a chance to regroup. In this case I think the writer is saying that the news implies that western models are getting distilled on release. Not the perfect analogy but it gives some color.
Do you have many mini tantrums like this per day? Probably makes you very difficult to work with Mr I was a CTO.
People did use global agreements and regulation to fix the ozone hole, though, so I think there’s a chance.
You can do this with cars, tools, computers, ... whatever you want. So, no, I think your point is wrong.
Now, what I want to regulate are accordions.
Other countries have governments that have earned that level of trust. I'd like mine get there eventually, but I'm not going to give it additional high risk responsibilities when it can't even handle the ones it already has.
Replace “gun” with anything and you will see how your comment falls apart.
What’s next? A registry for food purchases? Your beer gut is starting to show.
Worse for whom?
The only effective defense against predatory corporate and government AI is personal protective AI.
Anything else is unilateral disarmament. It's the only way individuals can survive in the worse case scenario.
> gun nuts
Guns are different. They can't protect you against the government, contrary to gun nut claims.
You ask about distillation but I wonder, is there any training datasets (~ TB-order) available that startup folks in SV use or is it so that everyone has to create their own scraping pipeline ?
Or did they not pull back when their models allegedly became highly capable, with the whole mythos debacle ?
https://www.anthropic.com/research/glm-5-3-and-the-spread-of...
It's VERY clear that the US companies are trying to push for regulation to kill open models and open weights. I see this as much more hostile and authoritarian response than what we're seeing come out of China right now.
So is China going to always publish in the open? No clue. But right now they're modeling much better behavior.
The point is get what you can from both to develop open models, data and tools.
An argument can be made that Anthropic is also only doing the last 5% of the work (because the content they are training on was the other 95%) but that's a bit more philosophical.
Anthropic’s anger here seems mostly rooted in their annoyance that this exposes they don’t really have core IP that’s not just easily replicated. And that’s clearly a problem for a deeply unprofitable company trying to convince people they’re worth $2 trillion.
The original authors of all the text, creators of the media and developers of the software did far more work than Anthropic.
The word itself is the pivot, not anything else.
I don't know anyone with even a passing understanding of how LLM training works that thinks that is the appropriate analogy.
I never take them seriously, I just assume they are coming from countries that don't understand how capitalism works or are operating out of bad faith. The underlying reality of the market is always changing and needs are always changing. Some AI companies will fail, that is a given. Remember alta-vista? Yahoo? Did search go away? How about Microsoft phones? Nokia? Motorola?
OpenAI and Anthropic are not in the inference business. That is a commodity. They need to sell products and solutions.
amelius•36m ago
Any ideas?
ambicapter•33m ago
curuinor•32m ago
Because of the basic huge recession going on in China, you can't actually make money in China doing China things. So they gotta gird up their export stuff and try to export. That entails strong relations with American companies, American PR, English stuff, etc.
If you want an essay about this from a VC, read this one
https://earnedintuition.substack.com/p/involution-without-ex...
dabedee•26m ago
curuinor•21m ago
ajkjk•19m ago
iamnothere•17m ago
Enshittification and related problems can be a result of market forces just as much as they can be a result of monopoly/duopoly or a small cartel. Excess competition sometimes results in all firms scraping the barrel to squeeze out pennies, especially with technology (such as large online marketplaces) making pricing more transparent.
Marx actually predicted that ever-intensifying competition would destroy markets through overproduction, although he did not use the term involution.
DrewADesign•12m ago
foul•32m ago
pj_mukh•30m ago
twoodfin•21m ago
Once upon a time, everyone had a secret sauce in network or data encoding or query optimization, but in the last ~10 years computational physics and economics have basically decided the “correct” architecture and everyone (including OSS) has converged.
carbonguy•26m ago
mpalmer•24m ago
TrackerFF•23m ago
Basically, western labs are in it for the money / commercial monopoly. Chinese labs are in it for the tech? As long as they can keep distilling models, and get access to research other ways, they benefit. And if they can push western labs forward, they'll benefit from that themselves.
jollyllama•23m ago
corford•23m ago
chrismarlow9•19m ago
I can't even fathom the trend these days of "we don't review the code" from security team perspective.
Just my guess though.
Windchaser•18m ago
Unpopular, maybe, but what about the normal reasons? The researchers are looking to make a name for themselves, and/or they genuinely care about AI advancement.
feverzsj•16m ago
The weird ideology here is to dominate the market at ANY COST, even it benefits the opponents.
teekert•16m ago
Why did we (the west) ever start open sourcing anything? Maybe we just like sharing? Maybe humanity only grows on pre-competitive layers like Linux and clean water. Maybe, the chinese government is closer to their people, and does not let large companies influence them and just doesn't like closed private hyperscalers with a lot of power?
(Some points assume the government has a role in the openness, which I think is likely)
Catloafdev•15m ago
thefourthchime•14m ago
We don't know either way, so I find the whole thing silly to speculate on.
seydor•14m ago
audunw•10m ago
Put another way: if they were not cheaper and open, they would simply not be competitive. They would already be dead.
I don’t think this ends well for the Chinese labs. This is going pretty much like I thought. Western labs is just copying their improvements (I don’t think publishing the techniques matter here.. they’d just hire to gain the knowledge or figure it out themselves), and they have access to more GPUs and have better branding, so in the end where can the Chinese labs compete? Even lower cost? Open weights? I’m not sure open is a sustainable way to compete either. Eventually there will be some fully open source AI models that cuts out that avenue of competition as well.
HeavenFox•6m ago