ByteDance, the Chinese technology company behind TikTok, is reportedly training an artificial intelligence model with up to 10 trillion parameters. The Financial Times reports the move could make it the largest Chinese AI model ever built.
Three insiders told the FT that the model is currently in pretraining, a phase that typically lasts three to six months. If completed as described, the model would be three times the size of Moonshot's Kimi K3, which currently holds the title of largest Chinese model.
A model on par with Anthropic's top system
Industry estimates place the ByteDance model in the same ballpark as Anthropic's Mythos 5, which is believed to have around eight trillion parameters. Anthropic has not disclosed its own parameter numbers.
Parameters determine how much a model can store, though performance also depends on data quality and training methods. That distinction matters as companies race to scale up their systems.
One source says ByteDance has avoided distillation, a training method that relies on outputs from other companies' models, for over a year. That approach suggests the company is building its models from original data rather than borrowing from rivals.
Founder pushes for world-leading capabilities
ByteDance founder Zhang Yiming told the 2,000-person Seed team internally to aim for world-leading model capabilities over the long term. The Seed team is ByteDance's internal AI research group and is responsible for developing the model.
The instruction signals a strategic focus on AI leadership that extends beyond short-term releases. It also places ByteDance in direct competition with both Chinese and American AI developers.
xAI trains similarly sized models
ByteDance is not alone in pursuing massive parameter counts. Elon Musk, CEO of xAI, disclosed that his company is training Grok variants with six and ten trillion parameters on its Colossus 2 cluster.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
That puts xAI in the same scale range as ByteDance, though the two companies are pursuing different paths. Musk's disclosure came as part of broader updates on xAI's infrastructure and training efforts.
A shifting competitive landscape
The reported model would put ByteDance in direct contention with Anthropic's top system, which is widely considered one of the most capable AI models available. It would also leapfrog Moonshot's Kimi K3, which had been the largest Chinese model.
The pretraining phase is early, and the model's final capabilities remain unknown. But the scale alone signals that ByteDance intends to compete at the highest level of AI development.
The article was published on Aug 7, 2026, and the model was still in pretraining at the time of the report. The timeline for completion remains unclear, though the typical pretraining window of three to six months offers a rough frame.
What parameter counts mean
Parameter counts are a common shorthand for model size, but they do not tell the whole story. A model's performance depends heavily on the quality of its training data and the methods used to train it.
ByteDance's reported avoidance of distillation for over a year suggests a focus on original training data. That choice could shape how the model performs once it moves beyond pretraining.
The company has not commented publicly on the report. The FT's sources remain anonymous, and the details have not been independently verified.

