On August 8th, according to the Financial Times on August 7thByteDanceWe're training up to 10 trillionLarge ModelIt is still in the early stages of pre-training. According to three informed sources, this size will be more than three times the size of the largest model published in China, Kimi K3, the dark side of the moon, and close to the volume of Mythos 5, the most advanced model of Anthropic, estimated by industry at around 8 trillion parameters。

It was reported that the project was led by a byte by the head of the Seed Foundation, reflecting the idea that CEOs were promoting original research rather than imitating competitors ' models. The pre-training phase usually takes three to six months, after which, if progress is successful, the fine-tuning and eventual release of the exact parameters will take a later stage to determine。
As previously reported in LateLatePost, byte beats are discussing the training of a model of parameters in excess of 5 trillion, more than Ali Qwen 3.8-Max (2.4 trillion parameters) and the dark side of the moon, Kimi K3 (2.8 trillion parameters), which is the largest training programme of known parameters in the country。