Message bytes beats begin pre-training 10 trillion-billion-billion-parameter large model, size or near the Anthropic flagship

On August 8th, according to the Financial Times on August 7thByteDanceWe're training up to 10 trillionLarge ModelIt is still in the early stages of pre-training. According to three informed sources, this size will be more than three times the size of the largest model published in China, Kimi K3, the dark side of the moon, and close to the volume of Mythos 5, the most advanced model of Anthropic, estimated by industry at around 8 trillion parameters。

Message bytes beats begin pre-training 10 trillion-billion-billion-parameter large model, size or near the Anthropic flagship

It was reported that the project was led by a byte by the head of the Seed Foundation, reflecting the idea that CEOs were promoting original research rather than imitating competitors ' models. The pre-training phase usually takes three to six months, after which, if progress is successful, the fine-tuning and eventual release of the exact parameters will take a later stage to determine。

As previously reported in LateLatePost, byte beats are discussing the training of a model of parameters in excess of 5 trillion, more than Ali Qwen 3.8-Max (2.4 trillion parameters) and the dark side of the moon, Kimi K3 (2.8 trillion parameters), which is the largest training programme of known parameters in the country。

statement:The content of the source of public various media platforms, if the inclusion of the content violates your rights and interests, please contact the mailbox, this site will be the first time to deal with.
Information

Ari's whole new generation video generation model Wan3.0 Public measure: single generation can produce 30 seconds of video, known as everything

2026-8-7 9:39:50

Information

OpenAI calls Astra or "significant" network security risk line, suspends some internal development and tightens controls

2026-8-8 15:52:48

Search