Artificial intelligence researchers are questioning allegations that Chinese AI company Moonshot developed its powerful Kimi K3 model by copying capabilities from Anthropic’s Fable system.
The controversy began after White House science adviser Michael Kratsios accused Moonshot of conducting large-scale industrial distillation using Anthropic’s technology. He also alleged that the company relied on advanced computing chips that were not authorized for export to China.
Moonshot, which created Kimi K3, has not publicly explained the model’s complete training process or responded in detail to the allegations.
Kimi K3 has attracted significant attention because of its strong performance and open-weight design, which allows developers to download, study and customize parts of the model. Its capabilities have also raised questions about how Chinese laboratories continue to produce competitive AI systems despite restrictions on access to advanced American chips.
Several AI experts argue that the timeline makes it unlikely that Moonshot built Kimi K3 primarily by copying Fable.
Fable reportedly became publicly available on July 1, while Kimi K3 was released roughly two weeks later. Researchers say this would have left Moonshot with very little time to generate training data from Fable, process that information, train a large model and complete the testing required for a public launch.
Distillation generally involves sending large numbers of queries to an advanced model and using its responses to improve another system. The technique can help a smaller model imitate certain behaviors, writing styles or reasoning patterns.
However, experts say modern frontier models require more than a collection of copied answers. Their strongest capabilities increasingly depend on reinforcement learning, complex evaluation systems and substantial computing infrastructure.
This means that even if Moonshot used outputs from other models during development, those outputs alone would probably not explain Kimi K3’s overall performance.
Supervised fine-tuning is one of the most common methods used in model distillation. A developer collects examples of prompts and responses, then trains another model to produce similar answers.
This process can cause a model to adopt the language patterns, tone or habits of the system that generated the training data. In some cases, a model may even incorrectly identify itself as the original system.
Researchers caution that such behavior does not prove the entire model was copied. It may only indicate that some training examples originated from another AI assistant.
As models become more advanced, supervised fine-tuning appears to provide fewer competitive advantages on its own. Reproducing deeper reasoning and agentic capabilities may require reinforcement learning, where models repeatedly attempt tasks, receive evaluations and adjust their behavior.
Running these programs at scale can involve millions of interactions and enormous computing costs. Using a commercial model through an application programming interface for every interaction would likely be expensive, slow and difficult to conceal.
The dispute also highlights uncertainty around what should be considered improper model copying.
AI companies frequently use synthetic data, which is information generated by other AI systems, to train or improve their models. Distillation is also widely used to transfer capabilities from a larger system to a smaller and more efficient one.
The boundary between acceptable synthetic-data generation and unauthorized capability extraction is not always clear.
Anthropic previously accused Moonshot, DeepSeek and MiniMax of systematically querying its models to extract capabilities. The company said it identified unusual activity involving millions of interactions linked to users associated with the Chinese laboratories.
Those earlier allegations suggest that Anthropic’s models may have contributed data to some development programs. They do not necessarily establish that Fable was responsible for Kimi K3’s strongest capabilities.
Researchers also warn that focusing exclusively on copying allegations may underestimate the technical expertise of Chinese AI teams. Moonshot employs experienced researchers and engineers who are capable of developing original training methods, datasets and model architectures.
The allegations concerning computing hardware may be more difficult to separate from the debate over model training.
Kratsios claimed that Moonshot gained access to Nvidia Grace Blackwell 300 chips and servers equipped with the advanced processors. Such hardware is subject to strict export restrictions intended to prevent China from obtaining the most powerful American AI technology.
Experts have warned that black markets and overseas data centers can create routes around these controls. A company may not need to import restricted chips directly if it can rent computing capacity from a facility in another country.
This has increased calls for stronger know-your-customer requirements at data centers. Under such rules, operators conducting large AI training runs could be required to verify the identity of their customers and report how advanced hardware is being used.
Supporters believe these measures would make it more difficult for restricted organizations to access high-end computing resources through intermediaries.
The debate over Kimi K3 comes as Chinese AI developers continue to narrow the performance gap with leading American laboratories.
Export controls may increase costs and slow development, but researchers say they are unlikely to stop progress completely. Chinese teams can improve model efficiency, use older chips more effectively, develop specialized infrastructure or access computing resources outside mainland China.
For policymakers, the Kimi K3 dispute presents two different challenges. The first is determining whether American AI systems are being systematically exploited to train foreign competitors. The second is enforcing hardware restrictions across a global network of chip suppliers, data centers and cloud-computing providers.
Experts remain skeptical that distillation from Fable alone could explain Kimi K3’s rapid arrival and advanced performance. The model may have benefited from earlier systems, synthetic data or outside computing resources, but researchers say its success also reflects the growing sophistication of China’s domestic AI industry.
Share your thoughts about this article.
Be the first to post a comment!