Tag
7 articles
Alibaba's Qwen team introduces Qwen3.8-Flash-Next, a cost-efficient model that uses only 6% of its parameters per token, outperforming competitors like Claude Opus 4.6 and DeepSeek-V4-Flash.
Cursor Research has open-sourced Mixture-of-Kittens (MoK), a deterministic MoE training megakernel designed for high-performance computing environments like GB300 NVL72 racks, delivering up to 2.37x faster performance than public baselines.
Learn how to set up and run inference with AMD's Instella-MoE-16B-A3B, a 16B parameter Mixture-of-Experts language model that activates only 2.8B parameters per token.
Learn about Laguna S 2.1, an efficient AI coding model that performs well despite its relatively small size, using advanced techniques like Mixture-of-Experts and open-weight design.
This article explains the advanced technical concepts behind Meituan's LongCat-2.0, a 1.6 trillion-parameter Mixture-of-Experts model with native 1-million-token context and LongCat Sparse Attention.
This explainer article dives into NVIDIA's Nemotron-Cascade 2, an advanced Mixture-of-Experts (MoE) model that demonstrates how strategic parameter allocation can enhance reasoning capabilities while maintaining computational efficiency.
Learn how Yuan 3.0 Ultra, a new AI model, uses Mixture-of-Experts to work more efficiently than traditional models, while still delivering top performance.