Independent Testing Finds a 51% Hallucination Rate in Kimi K3, a Figure Absent From Moonshot's Own Benchmarks
Artificial Analysis independently measured a 51% hallucination rate in Kimi K3, the 2.8 trillion parameter open model from China's Moonshot AI, up from 39% in predecessor Kimi K2.6. The figure does not appear anywhere in Moonshot's published benchmark charts, and it lands just days before the model's full weights are due for release on July 27, 2026.
What Artificial Analysis Measured
Independent benchmarking firm Artificial Analysis measured a 51% hallucination rate in Kimi K3, the 2.8 trillion parameter open weight model released by China's Moonshot AI on July 17, 2026, at WAIC, the World Artificial Intelligence Conference, in Shanghai. The figure captures how often the model delivers a wrong answer with a confident tone, and it climbed from the 39% recorded on the prior model in the same family, Kimi K2.6.
A Paradox: More Accurate and More Confidently Wrong
According to Artificial Analysis, K3 improved overall factual accuracy, from 33% to 46%, on the same test battery where the hallucination rate rose. In other words, the model now gets more questions right, but it also answers the ones it still gets wrong with more confidence, instead of flagging uncertainty or declining to answer. That is the pattern Artificial Analysis's methodology captures: the hallucination rate is calculated as incorrect answers divided by all non correct responses, not as a share of every answer given.
How K3 Compares to Other Frontier Models
Put in perspective, Claude Fable 5 itself, today one of the most advanced models available, posts a 54.9% hallucination rate under the same methodology, slightly above K3's rate. That context shows elevated hallucination is not unique to Kimi K3, but rather a challenge that persists even in models considered state of the art, which does not exempt K3 from scrutiny but does put the number in context.
Missing From Moonshot's Own Charts
The most striking point raised by coverage from Memeburn and TechTimes is that this 51% hallucination rate does not appear anywhere in the benchmark charts Moonshot officially released for the K3 launch. The company highlighted coding performance and other technical tasks, areas where K3 scores strongly, but did not surface this specific factual reliability metric, which only came to light through Artificial Analysis's independent evaluation.
Full Weights Arrive in Days, Under a Modified MIT License
Moonshot promised to release Kimi K3's full weights by July 27, 2026, under a modified MIT license, which would make K3 the largest open weight model in the world available for download and fine tuning by any developer. That means agencies and companies will soon be able to run K3 on their own infrastructure, making the hallucination rate an even more relevant data point for anyone evaluating the model for production use.
Why This Matters for Brazilian Agencies and SMBs
K3's API pricing, 3 dollars per million input tokens and 15 dollars per million output tokens, is competitive and has been drawing attention from Brazilian agencies looking for cheaper alternatives to Western AI models. But a 51% hallucination rate is a practical warning: for customer service tasks or content generation where factual accuracy matters, such as answering questions about prices, deadlines, return policies or product information, the cheapest model is not always the safest one without added verification layers. Before switching models purely on price, it is worth running a pilot test with the business's own real questions and measuring the error rate in practice, rather than trusting the vendor's published benchmark alone.