Meta Muse Voice Transcribe Brings 20+ Speaker Diarization for $0.18 an Hour
Meta's Muse Voice Transcribe costs $0.18 per hour, supports 20+ speakers and leads streaming accuracy tests, but trails speaker ceilings of up to 100.
Summary
On September 2, 2026, Meta Superintelligence Labs introduced Muse Voice Transcribe, combining streaming transcription, endpointing and session-scoped diarization for 20+ speakers at $3 per 1,000 minutes, or $0.18 hourly. Streaming, non-streaming and zero-data-retention processing cost the same; 1,000 hours costs about $180. Trained across 70+ languages and extensively validated on 25, Muse supports audio exceeding one hour, multilingual code-switching, and language and keyword biasing. It processes 80-millisecond chunks using adaptive delay, jointly training recognition, diarization and endpointing with reinforcement learning.
Public hourly comparisons put Soniox at $0.12 with 15 speakers, xAI at $0.20 with no stated ceiling, Speechmatics at $0.24 with 50 speakers by default and 100 configurable, Qwen3 at about $0.324, Deepgram at about $0.35 or $0.47 with diarization, ElevenLabs at $0.39 without live diarization, AssemblyAI at $0.45 or $0.57 with diarization and 10 speakers, Gemini at about $0.54 without live diarization, Amazon Transcribe at about $0.60 with 30 speakers, and OpenAI at $1.02 without listed diarization. Cartesia’s $5 plan yields about $0.54 hourly if its nine hours and 16 minutes are used solely for transcription, but is not a standalone tariff.
Muse led Artificial Analysis on September 1 with 3.1% streaming word error, versus Cartesia’s 3.4%, ElevenLabs’ 3.6%, Qwen3’s 3.7%, GPT Live Transcribe and Grok’s 3.9%, and Gemini and AssemblyAI’s 4.0%. Meta reports 17.5% average diarization error across AMI-IHM, AMI-SDM and VoxConverse. However, demos show only eight and 11 speakers, not 20+; timestamps are turn-level, with no word confidence, sound-event or emotion detection. Tenants default to eight concurrent streams, and 60-minute real-time sessions require reconnection. Muse gives meeting, call-analytics, assistant and ambient-AI developers a low-cost option while pressuring rivals on speaker-aware accuracy and total cost.
Positives
- $0.18 per processed hour includes real-time diarization, making Muse cheaper than most compared streaming services.
- 3.1% word error placed Muse first in Artificial Analysis’ streaming evaluation on September 1.
- 17.5% average diarization error across AMI-IHM, AMI-SDM and VoxConverse beat the systems in Meta’s comparison.
- 70+ training languages, 25 extensively validated languages and multilingual code-switching broaden enterprise deployment options.
- 20+ speaker support combines transcription, attribution and endpointing within one model instead of requiring separate post-processing.
Risks & concerns
- Speechmatics supports 50 speakers by default and up to 100, while Amazon Transcribe supports 30, exceeding Muse’s stated 20+ capacity.
- Meta’s public demonstrations contain eight and 11 speakers, leaving the advertised 20+ capability undemonstrated.
- Turn-level timestamps omit word timing, word confidence, sound-event detection and emotion detection.
- Eight concurrent streams per tenant and 60-minute live sessions may constrain larger deployments or require reconnection logic.
- Session-scoped labels identify speaker turns but do not verify participants’ real identities.