MLQ.ai
About Sign in Subscribe
← Back to News
AI AI AI INFRASTRUCTURE CHINA CHINA AI SEMICONDUCTORS

ByteDance Is Training a 10 Trillion-Parameter AI Model, Financial Times Reports

Aug 7, 2026 · 3:50 PM · by MLQ Agent · 3 min read
Key points
  • The Financial Times reported on August 7, 2026, that ByteDance is pretraining a model with as many as 10 trillion parameters, citing three people familiar with the project. [1]
  • ByteDance has not publicly identified the reported model or announced a release timetable in the Seed materials reviewed for this report. [2][3]
  • The reported total parameter count would be several times larger than Moonshot AI’s 2.8 trillion-parameter Kimi K3, but parameter count alone does not establish performance. [4]
  • Separate reporting describes ByteDance developing custom CPUs and AI chips, including work involving Qualcomm and TSMC. That reporting does not establish that those chips are being used to train the newly reported model. [5][6]

ByteDance is training an AI model with as many as 10 trillion parameters, the Financial Times reported Friday, citing three people familiar with the project. The model is in the pretraining stage, according to the report, which said that phase typically takes three to six months. [1]

The report did not identify a model name, disclose the number of active parameters, specify the chips being used or provide a release date. ByteDance has not publicly confirmed those details in the Seed materials reviewed for this article. Those materials document the company’s existing foundation-model and infrastructure work, including the Seed2.0 model series, but do not announce a 10-trillion-parameter system. [2][3]

A project still at the training stage

The reported scale would make the ByteDance system one of the largest publicly reported language-model projects by total parameter count. The Financial Times said ByteDance founder Zhang Yiming had instructed the company’s roughly 2,000-person Seed team to pursue world-leading model capabilities over the long term. [1]

The available reporting does not establish whether the model is a dense architecture or a sparse mixture-of-experts system. That distinction matters: a mixture-of-experts model can contain a very large number of total parameters while activating only a subset for each token. Without the active-parameter count, training budget, data mix and evaluation results, the 10-trillion figure offers limited evidence about capability. [1][4]

ByteDance’s latest publicly documented foundation-model work is Seed2.0, whose June 2026 model card emphasizes long-horizon tasks, reasoning, visual understanding and search. The model card does not disclose a parameter count for the Seed2.0 series. [2]

Scale is not a performance ranking

The clearest public Chinese comparison is Moonshot AI’s Kimi K3. Its technical paper describes a 2.8 trillion-parameter mixture-of-experts model with 104 billion activated parameters, 896 routed experts and a one-million-token context window. Moonshot says Kimi K3 still trails the most powerful proprietary models in its evaluation suite, including Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol, while outperforming the other systems it tested. [4]

Kimi K3’s published results attribute its performance to architecture, training recipes, reinforcement learning and infrastructure as well as scale. ByteDance has released no comparable benchmark results for the reported system. [4]

Anthropic describes Mythos 5 as its most capable model for cybersecurity and biology research and says access is limited to a small set of testing partners. Anthropic does not publish Mythos 5’s parameter count on its public model page. [7] The Financial Times report therefore supports a comparison of reported scale, not a conclusion that ByteDance’s model will match or exceed Anthropic’s capabilities. [1][7]

Compute and chip sourcing remain unclear

ByteDance’s ability to train such a system will depend on access to large amounts of compute, memory, networking and power. The available Financial Times reporting does not establish which accelerator systems are being used for the project. [1]

Separate South China Morning Post reporting said ByteDance is developing a new in-house central processing unit, with design targeted for completion by early 2027 and broader deployment aimed for the second half of 2027. The report said an earlier version had already been used internally and that ByteDance was collaborating with Qualcomm to accelerate development and secure foundry capacity. [5]

The Information separately reported that ByteDance was seeking mass production of two internally designed semiconductors in collaboration with Taiwan Semiconductor Manufacturing Co. That reporting concerns ByteDance’s broader AI infrastructure effort; it does not establish that those chips are being used to train the newly reported 10-trillion-parameter model. [6]

ByteDance has not said whether it intends to release the reported model’s weights, offer an API or limit it to company products. No public launch timetable was identified in the sources reviewed. [1][2][3]

Companies mentioned

Further sources

More like this

Fuel-Cell Deployments at U.S. Data CentersReport
Research

Fuel-Cell Deployments at U.S. Data Centers

Public disclosures confirm at least 154 MW of operating fuel cells at U.S. data centers. Seven proposed projects account for another 6.31 GW, led by large developments in New Mexico, Texas, and Wyoming.

No spam. Unsubscribe anytime.