Tag
25 articles
Liquid AI introduces LFM2.5-DSpark draft models that accelerate decoding by up to 3.18x without altering model outputs, using speculative decoding techniques.
OpenAI launches Ultrafast mode for GPT-5.6 Sol, delivering up to 750 output tokens per second using Cerebras hardware. The move introduces a three-tier pricing model centered on inference speed.
OpenAI introduces Ultrafast mode, a new API service tier that runs GPT-5.6 Sol up to 14× faster using Cerebras hardware, delivering up to 750 output tokens per second.
AMD acquires Canadian startup Taalas, which specializes in embedding AI model weights directly into silicon for ultra-fast inference. Google is reportedly pursuing a similar strategy for its Gemini models.
Google has released LiteRT.js, a JavaScript binding of its LiteRT library that enables running .tflite models in browsers via WebGPU, offering up to 60x performance gains over CPU-only execution.
OpenAI has unveiled its first custom AI processor, Jalapeño, developed in partnership with Broadcom. The chip is designed for AI inference tasks and marks a move toward vertical integration in the AI industry.
Learn how to create and run a simple AI inference example, understanding the core concepts behind AI model deployment that companies like Baseten are building upon.
Xiaomi's MiMo team, with TileRT, has achieved over 1000 tokens per second on a 1-trillion-parameter model using a single 8-GPU commodity node, marking a significant leap in LLM inference performance.
Perplexity AI introduces a hybrid local-server inference orchestrator that automatically routes AI tasks between on-device and cloud models, enhancing both performance and privacy.
NVIDIA introduces Dynamo Snapshot, a CRIU-based system that accelerates AI inference on Kubernetes by enabling fast startup and restoration of vLLM workers.
Perplexity AI has introduced an intelligent system that dynamically splits AI workloads between local PCs and cloud servers, optimizing performance and cost.
AI chip startup Groq is raising $650 million in internal funding as it pivots from hardware to focus more on AI inference, the process of refining how AI models respond to user prompts.