Yi-Lightning
Models & ArchitecturesA speed-oriented 01.AI Mixture of Experts language model built around low latency, efficient serving, and lower cost.
Yi-Lightning treats latency as a capability: the faster model is not merely cheaper, it changes what an interactive system can afford to do.
Yi-Lightning is a language model from 01.AI built around the idea that responsiveness is part of capability. It uses a Mixture of Experts architecture, expert segmentation, and optimized key-value caching to reduce the work required for each generated token.
At its October 2024 launch, Yi-Lightning reached sixth place on the Chatbot Arena leaderboard in the company’s reported results, with particularly strong showings in Chinese language, mathematics, coding, and hard prompts. The ranking is a snapshot of human preference, not a permanent measure of model quality, but it showed that a speed-focused system could compete with much larger frontier offerings.
01.AI also reported a training cost of about $3 million using 2,000 H100 GPUs and claimed a 70 to 80 percent cost advantage over comparable US frontier models for coding and mathematics. Those numbers depend on workload, serving stack, and comparison set. The lasting idea is simpler: a model that reaches a useful answer quickly can change the economics of an interactive product.
References & Resources
Related Terms
Last updated: February 22, 2026