This looks very similar to the claim that distilling a model from Anthropic is the same thing as Anthropic distilling the model from information on the internet.
Which is very flawed, since distillation requires the thing to exist in the thing it’s distilled from. And no LLM model existed in the information Anthropic used to train the model. Instead the model was built using information and utilizing new technology including hardware, software, transformer architecture, etc.
In any case, I doubt Kimi was trained without "stealing" the same data. Assembling all of your training data from Claude responses seems infeasible. It's much more likely that Kimi's base model was trained similarly to any other base model, with terabytes of data from all imaginable sources. Then the model was fine-tuned with "high-quality" data, followed by reinforcement learning. Throwing in lots of chat transcripts from other chatbots into the "high-quality" dataset would be expected, and is done to some degree by everyone, but maybe a lot more for Kimi. And likely they did a lot of reinforcement learning against the Claude API
The model would exist without Claude, it just wouldn't be nearly as coherent or smart
If training on copyrighted data without authors consent is ok then distilling is ok as well.
I find both claims unsound, though. Knowledge or model behavior itself is not copyrightable, so all these claims just boil down to the "I am not happy with that" argument. You cannot claim someone is stealing something you don't own in the first place.
Edit: No that doesn't seem to be what's happening here. I think it's some kind of word frequency analysis.
But I don’t grok training enough to know that’s silly.
If my new prior is you can…that’s a pretty thin moat that’s essentially indefensible.
Fable 5 -> 0.42
Opus 4.8 -> 0.45
Sonnet 5 -> 0.45
Opus 4.7 -> 0.46
Grok 4.3 -> 0.52
There's an obvious jump at Grok 4.3, and it would not surprise me that the similarity there is because Grok used Anthropic models for training too (you can get the similarity list for Grok and it does look like the top most similar models for Grok are either Anthropic models or some Chinese models).
The damning evidence that K3 used Anthropic models for training is that K3 is more similar to those models than it is to K2.6. If you look at the Anthropic, OpenAI or Google models, they are most similar with their own other models. Not so with K3, where K2.6 is less similar than 15 other models.
Now, why is Fable 5 the most similar to K3 and not Opus 4.8. I think it's quite likely that K3 did some fine tuning at the end, when Fable 5 became available. They probably had all the infrastructure in place, and Mythos had been announced for months, so they were probably waiting for the second the newest Anthropic model was released to start using it for synthetic data generation.
I don’t feel sorry for the model companies
Keep in mind, Anthropic started limiting access and introduced anti-distillation measures around 4.5-4.6 (?). So the majority of distillation should have happened on earlier models.
Maybe a better explanation is that they have access to the same training datasets? Which if private can again raise questions about theft, but on a very different level.
K3-to-Fable is blue at 0.42. Is 0.42 meaningful, or did we set 0.4 as the lower bound because it makes 0.42 look significant?
Sol-to-Fable is 0.69. It's dark yellow, making this look VERY different from 0.42. But is it? What do these numbers mean in absolute terms?
great_psy•35m ago
There is an optimal answer to any question. Something that maximizes utility and minimizes tokens.