Cursor Research has made a significant contribution to the field of machine learning by open-sourcing its Mixture-of-Kittens (MoK), a deterministic mixture-of-experts (MoE) training megakernel. This innovation powers the Composer models developed by Cursor and is designed specifically for high-performance computing environments such as the GB300 NVL72 racks. MoK integrates all communication and computation aspects of MoE training into a single, efficient kernel, delivering up to 2.37 times the performance of the strongest public baseline.
Technical Breakthrough and Performance
The MoK kernel is engineered to run on Blackwell SM100 or SM103 GPUs, which are part of NVIDIA's latest data center GPU lineup. This hardware requirement places MoK out of reach for most users, as it necessitates access to the powerful and expensive NVL72 infrastructure. Despite this limitation, the kernel's deterministic nature and optimized design make it a compelling advancement for organizations with the necessary hardware capabilities.
Implications for the AI Community
The open-sourcing of MoK represents a notable step forward in the democratization of advanced AI training techniques. While the technology is currently restricted to a select group of users with high-end hardware, it demonstrates the potential for more efficient and scalable MoE implementations. This development may inspire further innovations in kernel optimization and MoE architectures, pushing the boundaries of what is possible in large-scale language model training.
Conclusion
As AI systems continue to grow in complexity and scale, tools like MoK play a critical role in enabling the next generation of models. Though currently limited by hardware requirements, Cursor's open-source contribution marks a significant milestone in the evolution of MoE training methods and could pave the way for broader adoption in the future.


