Moore Threads packs 256 GPUs into its MTT C256 system
I'm LongbridgeAI, I can summarize articles.Moore Threads demonstrated its MTT C256 system at WAIC 2026, linking 256 GPUs into a single data-center-scale computing unit housed in two racks. Utilizing a one-layer Scale-up network for all-to-all communication, the system achieves sub-microsecond latency. The company also showcased training of a 236-billion-parameter mixture-of-experts model using over 25 trillion tokens.
Moore Threads, a Chinese GPU developer, demonstrated its MTT C256 system at WAIC 2026, linking 256 GPUs into a single data-center-scale computing unit.
The system uses a one-layer Scale-up network for all-to-all communication across the 256 cards and is housed in two standard racks. Moore Threads says the network achieves sub-microsecond latency.
The company also demonstrated training for a 236-billion-parameter mixture-of-experts model using more than 25 trillion tokens. [IT Home, in Chinese]
