Some experts say there is no evidence that Moonshot AI “distilled” Anthropic’s artificial intelligence models to develop Kimi K3.
On July 22, Michael Kratsios, top technology advisor to US President Donald Trump, evaluated Moonshot AI, the Chinese company behind Kimi K3, using the “distillation” method from Anthropic’s Fable 5 product.
However, according to SCMP, Many global experts do not support this view. Professor Michael Chau, an expert in innovation and information management at Hong Kong University Business School (HKU Business School), said that proving that a large language model (LLM) is developed from another model is very difficult. There is currently not enough evidence to prove whether Kimi K3 took training data from Fable 5 or not.
According to Stephen Hsu, professor of computational mathematics, science, and engineering at Michigan State University (USA), assessing Kimi K3’s progress thanks to the refinement of the American AI model is “too sketchy”. He believes that the breakthrough on Kimi K3 lies in fundamentally restructuring computing capabilities through a series of improvements, including mechanisms such as Kimi Delta Attention and Attention Residuals with architecture “designed to improve the way information flows across chain length and model depth”.
“I don’t believe a model can achieve such strong performance based solely on distilling data from Fable,” Braden Hancock, a researcher at the Laude Institute and co-founder of artificial intelligence company Snorkel AI, told TechCrunch. “In terms of time, that is unlikely. Fable was just announced on July 1. Collecting and ‘distilling’ a large enough amount of data, then training and releasing the model in just two weeks is difficult to do.”
“The ‘distillation’ method is becoming less and less effective, especially as the Chinese model is accelerating reinforcement learning capabilities,” Nathan Lambert, an AI researcher at the non-profit Allen Institute for AI, also said on the podcast Interconnects broadcast on July 23. “If that happens, other models can easily catch up with the leading competitor by using their data to ‘distill’. But we haven’t, or won’t see this.”
According to him, China has very good AI engineers, but “distillation” cannot happen overnight. Even applying the most advanced technology, a company will need a large infrastructure. “Using a lab’s API to do this is extremely expensive and can become a time bottleneck,” he emphasized.
Moonshot AI’s Kimi K3 logo. Image: Bao Lam
Others criticized the White House’s attempt to classify “distillation” as intellectual property theft. According to Kevin Xu, founder of technology investment fund Interconnected Capital (USA), the output or chain of thought of an AI model will usually not be protected by copyright nor considered a trade secret.
“Therefore, accusing the ‘distillation’ process of intellectual property theft is not a serious legal argument,” he wrote on X. “It is just political arguments and lobbying disguised as legal arguments.”
Theo The Information On July 22, the US Department of Industry and Security is investigating whether Chinese artificial intelligence companies, including Moonshot, are accessing advanced US AI chips. A day earlier, Treasury Secretary Scott Bessent said the US would consider the possibility of Chinese AI models stealing features from US products.
Paul Triolo, head of technology policy at risk consulting firm DGA-Albright Stonebridge Group (USA), said that Bessent’s threat “seems a bit reckless”, because the reason for using sanctions in this case is quite vague.
Others assess that the US preventing domestic companies from accessing open source AI from abroad could slow down the development of artificial intelligence. According to Politicoa group of 179 US businesses sent a letter to Mr. Kratsios and US Secretary of Commerce Howard Lutnick with the content: “Denying US startups access to existing models abroad will stifle competition to strengthen the position of existing businesses and act as a tax on intelligence, increasing costs and narrowing options. Meanwhile, foreign competitors still retain access to the entire global market.”
According to Xiaoyin Qu, founder of Tycoon AI (USA), punishing Kimi or similar open source models could cause US consumers and businesses to suffer, including increasing operating costs and slowing down the application of AI in the US.
Kimi K3 is attracting great attention globally. Announced on July 17, this is the world’s largest open weight model (publicizing the parameters used in the training process, allowing others to download and modify) with 2.8 trillion parameters, and is expected to be fully released on July 27. According to Business Insider, Kimi K3 can now process hundreds of pages of text with a single command line, suitable for the task of analyzing long documents and large code bases. Its programming capabilities and agents are competitive with leading models from US companies such as OpenAI and Anthropic, but at a much lower cost.
Immediately upon launch, Kimi K3 received positive feedback. Box CEO Aaron Levie called the release of Kimi K3 a “huge win” for companies building on AI platforms, saying it was “absolutely amazing” to see the performance of the open model. Nvidia CEO Jensen Huang said the US should not ban China’s open source AI model. “The models are excellent,” he told Axios July 22. “We should use excellent open source models.”
Moonshot AI was co-founded by Yang Zhilin, an AI researcher who studied for a doctorate at Carnegie Mellon University (USA), in 2023. According to Reutersa startup that raised more than 2 billion USD in May, bringing its total capital to more than 5.5 billion USD, is currently allowed to continue to mobilize an additional 2 billion USD and reach a valuation of 30 billion USD. The Chinese company was once accused by Anthropic of using “distillation” to train its AI model, along with DeepSeek and MiniMax, through creating 24,000 fake accounts.
In the world of AI, the concept of “distillation” refers to the “transfer of knowledge” from one model to another in a teacher-student fashion. “Distillation is a technique designed to transfer the knowledge of a large pre-trained model (teacher) into a smaller model (student), allowing the student model to achieve performance equivalent to the teacher model,” scientists Vishal Yadav and Nikhil Pandey once told Forbes. “The technique helps users take advantage of the quality of the LLM, while reducing the cost of inference.”