Tag
16 articles
This article explains FreeToken, an edge-native serving engine that enables large MoE models to run on a single workstation GPU by intelligently managing cache misses through PCIe bandwidth and CPU execution.
Learn how NVIDIA's new AI models Nemotron 3.5 Lightning and NeMo Switchyard work together to make AI systems more efficient and powerful by using only the right AI experts for each job.
Cursor Research has open-sourced Mixture-of-Kittens (MoK), a deterministic MoE training megakernel designed for high-performance computing environments like GB300 NVL72 racks, delivering up to 2.37x faster performance than public baselines.
Learn about Alibaba's new AI model Qwen3.8-Max, a powerful 2.4 trillion parameter system that can understand text, images, and video.
This article explains the advanced concepts behind Inkling-Small, a 276 billion parameter multimodal MoE model that achieves efficiency through sparse activation and open weights.
Learn how to set up and run inference with AMD's Instella-MoE-16B-A3B, a 16B parameter Mixture-of-Experts language model that activates only 2.8B parameters per token.
Moonshot AI has open-sourced MoonEP, a new Expert Parallelism library designed to improve the efficiency of Mixture-of-Experts (MoE) training at scale. The tool is released under an MIT license and aims to reduce communication overhead in distributed AI workloads.
This article explains the advanced concepts behind Alibaba's Qwen3.8-Max, a 2.4 trillion-parameter multimodal model, including multimodal capabilities, mixture-of-experts architecture, and parameter scaling effects.
This article explains Mixture of Experts (MoE) AI models, how they work like teams of specialists, and why they're important for efficient AI performance.
Soofi Consortium releases Soofi S 30B-A3B, an open hybrid Mamba-Transformer MoE model for German and English. The model leverages 3.2 billion active parameters out of 31.6 billion for efficient multilingual processing.
NVIDIA introduces Nemotron-Labs-3-Puzzle-75B-A9B, a compressed hybrid MoE LLM delivering 2.03x server throughput, leveraging hardware-aware compression and knowledge distillation.
JetBrains has released Mellum2, a 12-billion parameter MoE model trained on 10.6 trillion tokens, designed to accelerate specialized AI tasks in multi-model pipelines.