There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.
Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre-cursor acquisition.
And lastly, kimi architecture is vastly different than that of fable as it uses mechanisms developed by... kimi themselves. US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.
Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
edit: (moved this to bottom) The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens.
How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies?
I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies
So here robbers are blaming robbers?
These claims are just pointless, everytime
"What was I supposed to do? Call him for cheating better than me in front of the others?!"
Said in response to being out-cheated at a high-stakes poker game.
Except in this case, it sounds like that's exactly the path they have chosen.
https://getyarn.io/yarn-clip/7612c4ce-1077-479f-a7bf-617dbc6...
> "Well, Steve [Jobs]… I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."
Source: https://www.goodreads.com/quotes/824084-well-steve-jobs-i-th...
The economic viability of Anthropic and OpenAI rely on their being able to charge more for model access than their R&D and inference costs. If the market price for SOTA model access drops below that level, then these businesses will have to decide whether to continue to lose money or to reduce spending on R&D.
Moonshot's papers [1] claim that their training load was primarily from synthetic data and model self-teaching rather than RLHF and therefore keep their costs low. If Moonshot genuinely does not rely on human-led training, they will surpass US closed-source model providers. The United States government considers US supremacy in "AI" as a national security consideration.
This announcement is noteworthy because it implies that Moonshot's success is in fact due to distillation. It's in the interest of US frontier labs to place barriers to this if they find themselves in the position of subsidizing rival labs' research.
1. Kimi K2, https://arxiv.org/html/2507.20534v1
> @MehdiKarech
> I don't remember letting Anthropic or Open Ai scrapping my GitHub, my research gate and all my online writings L O L
https://xcancel.com/MehdiKarech/status/2080000779859939678#m
none of the frontier labs provide probability distributions over the tokens which is the actual method of distillation you use to train a smaller model based on a larger one. they don't even provide all the tokens.
therefore this so-called distillation the frontier labs whine about is just a set of clever methods to work the existing LLM into the training process for a new model. methods like having the existing model grade the output of the new model and work those grades into the RL method. give the new models structured tasks and use the existing model as a source of truth for those tasks and a myriad of other hacks.
efficiency scales with the gap between the models and generally allows an efficient bootstrap process. the implication that distillation wouldn't allow further advancement is false however, you can then start doing the same thing the frontier labs have been doing: dumping cash on humans to provide the signals or burning tokens on exploratory paths and grading the results.
what openai and anthropic don't like is that fact that all the cash they burned can be used to benefit everyone and not just them. and that no matter how much more cash they burn to build up the gap it will closed at a small fraction of the price.
Other US labs cannot directly distill from OpenAI/Anthropic as it’s a violation of the terms of service. It holds other US labs back. Leading them to build second tier models And in the end OpenAI/Anthropic may be unable to prevent distillation.
Why fight it when there’s clear money to make here?
https://typebulb.com/u/lab/you-re-relatively-right/full
According to these results GLM 5.2 is very similar to Google Gemini and Kimi K3 is very similar to Fable 5.
The American frontier labs are not similar to each other.
No, seriously: first of all, that's not the AI labs' data, it's ours. And if the AI labs think they can rake in tons of money using our data, then I'm actually glad if someone comes along and at least offers us a good product at reasonable prices.
Mendel doesn't get a cut every time somebody uses the principles of heritability he discovered, and Einstein's family aren't getting royalties if you compute relative speeds. I think the frontier labs should expect to be treated more like scientists than artists in this regard.
A more interesting part of this discussion is that consistently these Chinese models are held up as a great achievement, and that they're "catching up" when in reality they're just using the work of Anthropic and OpenAI to try and keep up with them. This isn't even to say it's not a valid tactic, but it definitely colors these announcements and proclamations about foreign companies catching up to American ones.
If I get a 1600 on the SAT and you copied my answers and got a 1540, your achievement isn't that significant.
Sounds like the opposite of the conversation Anthropic would want to have.
I love using Claude but Fable's unusable wrt useful work like cryptography, biology, &c.
Kneecapping my productivity when I pay $100/month is annoying af.
If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model, would this be considered illegal?
I ask because this happens all the time, synthetic datasets have basically become a key aspect of training a model at this point. I even generated a synthetic set from DeepSeek v4 to aid in fine-tuning a classifier just a few weeks ago.
So I just wonder on what grounds any of this makes sense, I wouldn't be surprised if some of these American labs were using open models on their own self hosted infrastructure to generate training data, but by nature of them being open nobody has to know.
I'll make a prediction: I don't think we will ever see any of the evidence of this "distillation" before they end up implementing some type of ban.
Because there is no way in hell I'm going to make an effort creating quality content for existing platforms. The website should be entirely my own without moderation subject only to my local legal system.
Can just insert this comment as a prompt and vibe code everything in a few days⸮
https://www.reddit.com/r/ClaudeCode/comments/1tqaist/opus_48...
(don't take this too seriously)
What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training, only distilling the complicated traces they need, originating from real user prompts and traces. These inference providers are explicitly blocked in the claude cli if you reverse engineer it.
The real picture is that these Chinese labs have figured out how to get exactly what they need, at a high quality, directly from distinct and unique real user prompts.
It’s only “covert” because Anthropic doesn’t like it, while simultaneously being perfectly fine to do.
Ba-dum-tss
You're using available information (copyrighted works, or the output of another model) to train a model to encode the information in a new form. Why is the former not theft, but the latter is theft?
Distillation itself, however, is still clearly valuable - else competitors wouldn't pay so much to their rival on distillation campaigns or try to circumvent anti-distillation defenses.
As for the morality of it, if you paid for the tokens they're yours. It is already understood that you own the output. Seems to me like a variation of ordinary business arbitrage. Providers might object to certain use-cases or intention and try to craft terms around that, but that's hard to enforce at scale.
The US gov and AI providers when funny chinese people steal their data to train their models: >:(
clowns
edit: TIL you can't use emojis on HN
It also leads me to think about things like the original release of Fable 5, people were complaining that it was safeguarded too much - if you lock the models down too much they cease to be useful. So it’s going to be increasingly difficult to protect a model from competition while ALSO keeping it useful.
Doesn't bode well for the valuations of these labs.
By all means use whatever works for you, I’m not even going to try to make an argument on ethics (and honestly I’m not even sure where I stand, given the behavior of American AI companies).
But I just cringe every time I see people acting like any of this is done in good faith.
Open source coming out of China is a state-sponsored criminal enterprise, built only for the benefit of the Chinese regime, one of the worst to exist in human history.
Side note, didn't they stop releasing real thinking tokens for Fable? Or is it still part of some subs or API usage?
That said, I doubt the "they distilled Fable" is the reason why K3 is as good as it is, considering the timelines involved, and that Anthropic hides thinking traces, and their overly aggressive "safety" filters.
This constant FUD spread by Anthropic is so tiring.
American AI corporations are pushing up the prices for computing, making it unaffordable for the common man. Additionally, they have built their entire business on stealing(yes, stealing) work from us.
So fuck em
Like, goddamn ya’ll are hypocrites.
Fable level performance, for much lower price.
But really, this is the USA getting ready to bring AI companies completely under the control of the Trump administration for ‘national security’
They definitely used closed private saas products to train their own models, to prove that just drop random small screenshots of any popular product behind a login screen and see how well it's able to identify all of them. ex: https://x.com/michalwols/status/2079968211865330165
or other similar "AI" startups https://x.com/envconfig/status/2079613455296827402
As of a couple months ago, when using Claude to write adult content through the API, sometimes it will silently inject a system prompt giving the model a bunch of guidelines on exactly what kind of adult content it's allowed to write, steering it away from anything "questionable" ("Claude will not write etc etc").
Moonshot distilled Claude so hard recently, they actually ended up distilling this prompt injection, too. Using K3 to write adult content results in it randomly hallucinating the injected Claude prompt during thinking, and it will quote parts of that prompt, complete with the name "Claude".
Not that I think distillation is a bad thing, just thought this was funny.
And now a regime best known for lying to their own people is the one trying to convince me?
Go, China!
The Chinese are not gonna deterred, but the posturing by the Americans is so blatantly hypocritical that everybody is cheering for their demise. See, for example, one of Francis Fukuyama's latests videos on youtube.
- US AI companies
We have information that Fable was distilled from humans.
If it works it works. Isn't that the argument?
AI outputs are not copyrightable, so distillation is fair use.
It may be a TOS violation, but that's a private matter. Cancel the accounts used for distillation and be done.
Second: Post is rich with allegations but light with evidence. Can very well be bullshit.
Waiting for the whataboutism....junk away...
> we're entering the most geopolitically volatile moment since the trinity test lit up the alamogordo desert and the only US policy prescription is a big button labeled sinophobia
https://bsky.app/profile/thebadcode.com/post/3mr3skoyass2k , and,
> every vendor cranking the big dial labeled "sinophobia" and looking back at the us government for approval
The government itself doing the propaganda here, skipping the vendors. Sinophobia intensifies. War drums of "be afraid be afraid be afraid" beat louder.
It's so bad, it's so stupid. Kimi lands one showing pretty clearly this was absolutely the determining concern happening at vast scale, that they can just a lot of this themselves, and this noise pollution from the most hopelessly lost aggro administration ever still gets blared out the trumpets of war & discord. What a joke. Give me a break, give it a rest.
War here is less winnable than the Iran war they started. They're going to make America itself so much worse, these people so hungry to put down free and good models. This pathetic attempt is not going to work, you are just going to once again hold the US citizens hostage & make their lives worse, for sick political games.
See also, don't trust anyone in Trump's government who says "we have information".
"they distilled us" is fast becoming standard US FUD.
The same as people telling me with a serious face that the Chinese models are distilled just because it says "I am Claude".
I am not the only one, look at this post on interconnects about Kimi K3 for example:[1]
It should be clear looking at this model that if adversarial distillation from the closed frontier models in the U.S. contributed, it is at most to a relatively small degree. AI observers who followed the distillation panic and came away with the wrong conclusion that Chinese AI labs are only producing good models due to IP theft are in for an awakening – that Chinese companies are extremely good at building models in the same way the leading American companies are.
[1] https://www.interconnects.ai/p/kimi-k3-the-open-weights-esca...The problem is… what are you going to do about it?
This is obviously an idiotic and dangerous Cold War and has no happy ending.
Nobody cares. This is neither a controversy nor news, and that would be the case even if Anthropic hadn’t just settled a 1.5 billion dollar lawsuit where they trained Claude on thousands of books without permission lol.
To be clear I’m not taking a jab at OP - I’m saying the labs crying about distillation have neither a legal nor a moral leg to stand on. There’s nothing wrong with distillation.
The current US administration is known to be collection of BS artists and liars.
And before the Chinese astroturfing starts (it already started, that’s clear from the comments and voting): the point is not even that the US companies have the right to intelectual property over their models (they should, but ok, that’s not even the point). The point is that China is incapable of innovation and any innovation into AI we can expect, will always come from the US.