Tag
12 articles
This article explains Google's new Gemini 3.5 Transcribe speech-to-text model, detailing its dual-endpoint architecture, technical mechanisms, and implications for developers building voice agents and transcription systems.
Google's new Gemini 3.5 Transcribe supports 85 languages and features real-time auto-correction of verbal stumbles with a 4.0% word error rate.
Particle's new Radar platform transcribes and analyzes over 130,000 podcasts, making them searchable and accessible to AI agents through API and MCP integration.
Learn how AI-powered meeting transcription works and how Google Meet's new feature automatically creates meeting notes for you.
Meetily offers free, open-source meeting transcription and summarization without requiring a subscription, challenging the premium model of competing tools.
Learn how GPT Transcribe works and why speech recognition technology matters in everyday life.
A Zoom incident has highlighted the privacy risks of automatic meeting transcription, where private conversations may be recorded without explicit consent. The issue raises important questions about user privacy and data handling in virtual environments.
This article explains speech-to-text technology and how Microsoft's new MAI-Transcribe-1.5 model improves speed, accuracy, and language support for converting spoken words into text.
A recent test of Wispr Flow and other AI transcription tools reveals that while free services suffice for basic needs, paid solutions offer significant advantages for professionals requiring accuracy and advanced features.
Cohere launches an open-source voice model for transcription with just 2 billion parameters, designed for consumer-grade GPUs and supporting 14 languages.
Learn to build an AI-powered meeting transcription and summarization system using Python and OpenAI's API. This tutorial teaches you how to process audio files, generate transcriptions, and create structured meeting notes with action items.
Learn to implement and compare speech-to-text capabilities using Google Cloud and ElevenLabs APIs, including audio processing, transcription functions, and service evaluation.