Tag
12 articles
OpenAI's new Jalapeño inference chip delivers industry-leading speed and efficiency in AI inference, with higher throughput and lower latency for modern AI models.
OpenAI's new Jalapeño chip promises faster AI responses and improved efficiency, positioning it ahead of competitors in the rapidly evolving AI hardware market.
Apple's new M5 Ultra chip delivers 4x faster AI performance than M3 Ultra, featuring up to 512GB of unified memory and support for up to eight external displays.
OpenAI has released GPT-5.6 in its Kiro platform, enhancing developer tools with improved price-performance for software planning, building, reviewing, and testing.
Sakana AI's Fugu Ultra v1.1 claims up to 7.9 points improvement over v1.0 and introduces a Claude Code-compatible endpoint, though independent verification is pending.
Anthropic's Claude Opus 5 achieves near-Fable 5 performance at half the token price, excelling in coding and problem-solving benchmarks.
Gigatoken, a new Rust-based BPE tokenizer, achieves text encoding speeds of up to 24.53 GB/s, outperforming HuggingFace tokenizers by 989x and tiktoken by 681x.
This article explains how Meta's Muse Spark 1.1 outperforms GLM-5.2 in coding and cost-efficiency, focusing on advancements in hallucination reduction and model reliability.
This explainer explores the technical advances behind Fable 5's AI work automation performance, examining the sophisticated mechanisms enabling autonomous task execution and their implications for human-AI collaboration.
Anthropic's Claude Fable 5 delivers only a 5.7% performance boost over Opus 4.8, but at twice the token price. Safety features contribute to the increased cost.
Learn to implement and evaluate a hybrid MoE-diffusion model that demonstrates the performance benefits of converting autoregressive LLMs into diffusion models for improved inference speed.
This explainer explores the technical advancements in GPT-5.4, OpenAI's latest AI language model that outperforms humans by 83% in professional tasks while reducing errors by 33%. We examine the underlying architecture improvements and implications for AI reliability.