TL;DR
- DeepSeek Plan: Chinese AI lab DeepSeek reportedly plans at least 160,000 Huawei Ascend 950DT chips to serve models at an Inner Mongolia data center.
- Serving Scale: The proposal would expand domestic capacity for serving finished models, but the wider gigawatt-scale site requires more than chips alone.
- Supply Constraint: Bloomberg’s sources estimate that limited Huawei output and high-bandwidth memory supply could stretch fulfillment beyond a year.
Chinese AI lab DeepSeek is reportedly planning an installation of at least 160,000 of Huawei’s Ascend 950DT accelerators at a data center it is building in Inner Mongolia. The chips would run inference, producing answers from finished AI models, and their scale would test whether China’s domestic hardware supply can support a major user-facing service.
As of September 9, 2026, neither DeepSeek nor Huawei had publicly confirmed the plan. Bloomberg says DeepSeek is building the site and places the 160,000 accelerators within a wider gigawatt-scale infrastructure goal whose remaining hardware mix is undisclosed.
Why Inference Still Requires Major Compute
Training builds a model’s behavior by adjusting its internal parameters across large datasets. Inference applies that trained model to a new prompt or other input. This includes processing the prompt and then decoding an answer token by token. A service handling many simultaneous requests can therefore need a large fleet even though it is no longer creating the base model.
Huawei’s own roadmap makes the workload choice less surprising than a training-only description would suggest. The company says the Ascend 950DT is designed for both inference decoding and model training. It lists 144 GB of HiZQ 2.0 high-bandwidth memory, 4 TB per second of memory bandwidth and 2 TB per second of interconnect bandwidth.
Memory bandwidth helps feed model data to the processors quickly, while the interconnect carries data among accelerators working on the same service. Both matter when one request, or a high volume of requests, must be distributed across many chips. Huawei specifically assigns the 950DT to the decoding stage that generates model output.
Supply Connects the Plan to Real Scale
Bloomberg’s sources estimate that Huawei could produce 950DT accelerators only in the low hundreds of thousands during 2026 and that limited supplies of the required high-bandwidth memory add another constraint. They said fulfilling DeepSeek’s intended quantity could take more than a year.
Huawei’s roadmap schedules the 950DT for availability in the fourth quarter of 2026. The company has not publicly linked that schedule to DeepSeek, disclosed an allocation for the project or confirmed that customer deliveries have started. Even a completed chip would still need memory, boards, networking, cooling, power equipment and system software before the facility could turn it into a service.
What 160,000 Chips and Gigawatt Scale Describe
DeepSeek’s final server design, interconnect layout, software stack, utilization and remaining chip mix have not been disclosed yet. A large accelerator fleet for DeepSeek or other AI labs will still depend on software, memory, packaging and interconnects to deliver useful work. Bloomberg says the 950DT tranche would be one part of a larger buildout.
Inner Mongolia already has a defined place in China’s national computing geography. It is one of eight approved national computing hubs in a program designed to move more data-center work from high-demand eastern regions toward resource-rich western areas. Energy supply, land and a cooler climate can make those locations attractive for facilities that must remove large amounts of heat.
If supplied and completed, the installation would expand the amount of Huawei-based infrastructure available for DeepSeek inference. Huawei hardware would then handle the recurring work performed each time users or applications call a DeepSeek model. For users, the project would add capacity to serve DeepSeek models, possibly at reduced cost.
Inference Diversification Is Not Training Independence
So far, DeepSeek had relied on Nvidia hardware for training and did not plan to train models on this 950DT tranche. In April 2026, DeepSeek V4 was said as training on Nvidia hardware while supporting inference on Huawei Ascend systems.
The proposed data center would be a major domestic inference deployment if the chips are procured, delivered and integrated into a working facility. The reported number of at least 160,000 chips shows how much Huawei hardware DeepSeek wants to dedicate to model serving. Whether that ambition becomes available compute depends on the supply and system work that neither company has yet publicly confirmed.

