Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
Back to Home
ai

Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas

August 18, 202630 views2 min read

Cartesia's new Sonic-3.6 streaming TTS model leads two major artificial analysis speech leaderboards, using state space models for superior performance and speed.

Cartesia has made a significant leap in the text-to-speech (TTS) landscape with the launch of Sonic-3.6, a cutting-edge streaming TTS model that is now leading two major artificial analysis speech leaderboards. Unlike many of its competitors, Sonic-3.6 is built using state space models rather than the widely adopted transformer architecture, marking a notable shift in approach within the industry.

Performance and Technical Breakthroughs

The model has achieved a top ranking on both the Provider Voice and Controlled Voice leaderboards, scoring 1,283 Elo on the former and 1,123 on the latter. The Controlled Voice leaderboard is particularly rigorous, as it standardizes performance by cloning each model onto the same eight reference voices to isolate the core synthesis engine’s quality. Sonic-3.6’s impressive results highlight its ability to produce high-fidelity speech with minimal latency, reportedly achieving sub-90ms time-to-first-audio—a critical metric for real-time applications.

Implications for the TTS Industry

This advancement positions Cartesia at the forefront of a new wave of TTS innovation. By leveraging state space models, which are known for their efficiency and scalability in sequence modeling, Sonic-3.6 could offer advantages in processing speed and resource consumption compared to traditional transformer-based systems. The model is currently available in beta through Cartesia’s API, signaling the company’s intent to gather feedback and refine the system before a full rollout.

As streaming TTS becomes increasingly critical for applications ranging from virtual assistants to immersive media experiences, tools like Sonic-3.6 are setting new standards for quality and responsiveness. With this latest release, Cartesia not only demonstrates technical prowess but also reinforces its role in shaping the future of synthetic speech.

Source: MarkTechPost

Related Articles