ByteDance Reportedly Training AI Model With Up to 10 Trillion Parameters
ByteDance is aiming for the largest model scale among multiple Chinese research institutes training models comparable to American-made ones, using an independent approach rather than distilling existing models.
Reporting from 1 source: GIGAZINE.
The Financial Times reports that ByteDance is in the early stages of training an AI model with up to 10 trillion parameters. The training phase typically takes three to six months, followed by fine-tuning and release. The report says the model would surpass Anthropic's Claude Mythos 5, which industry estimates put at 8 trillion parameters.
ByteDance's development team is said to be behind other companies because they are using an independent approach rather than distilling and retraining other models, according to the Financial Times report. The number of parameters indicates the scale of an AI model, and a larger number generally means more knowledge, though it does not directly translate to higher processing capability.
The report notes that a 2.4 trillion parameter model has just appeared in China recently, with Alibaba announcing Qwen3.8-Max. The Financial Times cited Claude Mythos 5, which has been widely discussed for its high ability to find vulnerabilities, and pointed out that industry estimates put it at 8 trillion parameters.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.