Skip to main content
Models & Technology

Day 0 Adaptation: Alibaba Cloud's Lingjun Zhenwu M890 Supernode Instance Now Supports Moonshot AI's Kimi K3 Large Model

Alibaba Cloud announced that its Lingjun Zhenwu M890 supernode instance has successfully adapted to Moonshot AI's newly released flagship model, Kimi K3, with 2.8 trillion parameters. Based on the T-Head integrated AI training-and-inference chip and a high-speed interconnect architecture, the instance can confine the expert-parallel communication traffic of trillion-parameter MoE models within a single high-bandwidth domain, effectively improving inference efficiency. #AI大模型##阿里云##月之暗面Kimi#

Day 0 Adaptation: Alibaba Cloud's Lingjun Zhenwu M890 Supernode Instance Now Supports Moonshot AI's Kimi K3 Large Model

Alibaba Cloud announced today that the Lingjun Zhenwu M890 supernode instance has successfully adapted to Kimi K3, a large model with 2.8 trillion parameters. Through joint optimization of the chip, inference platform, and model, it effectively improves model inference efficiency.

Day 0 Adaptation: Alibaba Cloud's Lingjun Zhenwu M890 Supernode Instance Now Supports Moonshot AI's Kimi K3 Large Model

IT Home learned that Kimi K3 is Moonshot AI's latest flagship model, with 2.8 trillion parameters and a Mixture-of-Experts (MoE) architecture.

The Lingjun Zhenwu supernode instance is built on the T-Head Zhenwu M890 integrated AI training-and-inference chip. It features the ICN Switch 1.0 interconnect chip, enabling 64 M890 chips to achieve 800 GB/s All-to-All high-speed interconnectivity. With 9 TB of memory, it can run the EP expert-parallel communication traffic of trillion-parameter MoE models within a single high-bandwidth communication domain, ensuring Token generation efficiency.

At the same time, Alibaba Cloud, T-Head, and the Kimi team carried out deep collaboration at the operator and software-stack levels, successfully enabling Day 0 adaptation for the Kimi K3 model. The T-Head SAIL software stack on Zhenwu chips enables Kimi's self-developed Mooncake inference framework to run out of the box, while M890's Triton support means that a large number of custom operators written with Triton do not need to be rewritten, greatly reducing the adaptation workload.