Tag
14 articles
Cartesia's new Sonic-3.6 streaming TTS model leads two major artificial analysis speech leaderboards, using state space models for superior performance and speed.
Learn to create and use AI voice models with Fish Audio's technology stack, including setting up environments, training custom voices, and building API integrations.
Alibaba's Qwen Audio 3.0 TTS Plus has topped the Artificial Analysis Speech Arena leaderboard, showcasing advanced multilingual capabilities and expressive controls, though it lags in speed compared to competitors.
Learn to use MisoTTS, an 8B emotive text-to-speech model with open weights, to generate emotionally expressive speech by conditioning on both text and audio context.
Learn how to create a text-to-speech application using Python and the TTS library, from setting up your environment to generating and customizing speech output.
Supertone has released Supertonic v3, an on-device text-to-speech model with 31-language support, improved reading stability, and expressive voice tags.
Learn to build a basic speech-to-speech conversational AI system that processes voice input, generates intelligent responses, and speaks back to users.
This article explains how the Deepgram Python SDK enables developers to integrate advanced voice AI capabilities like transcription, text-to-speech, and asynchronous audio processing into Python applications.
Google introduces Gemini 3.1 Flash TTS, a new text-to-speech model that enhances speech quality, expressive control, and multilingual generation. This release marks a shift toward more controllable and natural AI voice outputs.
Learn how to use Google's new Gemini 3.1 Flash Text-to-Speech model to convert text into natural-sounding speech in over 70 languages with precise control over style, pace, and tone.
Google introduces Gemini 3.1 Flash TTS, a new text-to-speech technology that delivers more natural and expressive AI-generated voices. The advancement represents a significant step forward in making AI interactions more human-like and emotionally nuanced.
Learn what Microsoft VibeVoice is, how it uses AI to understand and generate human speech, and why it's important for the future of voice technology.