Wednesday, September 23, 2026
Tech Beat
Sep 23, 2026, 1:17 PMArtificial Intelligence

NVIDIA Nemotron 3 Diarization Tops VoiceArena for Real-Time Speaker Tracking

NVIDIA Nemotron 3 Diarization tops VoiceArena with 14.72% DER, supports eight speakers, and live streams with latency as low as 0.32 seconds.

Listen to this briefingAudio briefing

Summary

NVIDIA released Nemotron 3 Diarization on September 23, 2026, an open-weight, 100M-parameter model that assigns timestamps to as many as eight anonymous speakers across live or recorded audio, including overlapping speech. It processes 16 kHz mono audio in chunks, uses arrival-ordered channels plus speaker and recent-context caches, and offers recommended input-buffer latencies of 30.4, 1.04, 0.64 and 0.32 seconds. Training covered public and licensed data, including David AI material used for English and multilingual mixtures spanning 21 languages, cutting compound DER from 11.19% to 10.42%.

Positives

  • 14.72% DER placed Nemotron 3 Diarization first among 12 systems and 17 configurations in VoiceArena’s initial benchmark.
  • Eight speaker channels double the four-speaker capacity of NVIDIA Streaming Sortformer while supporting simultaneous speech.
  • 41.0% mean relative DER improvement was recorded across eight evaluation conditions at 1.04-second input-buffer latency.
  • 15,113 times RTFx at 30.4 seconds beat the previous model’s 2,619 times on an RTX PRO 5000.
  • Argmax Pro SDK 3 adds real-time, on-device support and a pre-diarized transcription API for complex conversations.

Risks & concerns

  • Eight channels are anonymous labels, so the model cannot identify real people without separate metadata or speaker verification.
  • 14.72% DER still represents missed speech, false alarms and speaker confusion under VoiceArena’s scoring protocol.
  • VoiceArena’s initial ranking may change after its Version 1 evaluation and paired statistical analysis are completed.
  • More than eight speakers, long recordings, noise, reverberation, far-field capture and domain shift can reduce accuracy.
  • 0.32-second latency excludes computation, networking, ASR and application processing, so complete systems will respond more slowly.
Primary sourceHugging Face - Bloghttps://huggingface.co/blog/nvidia/nemotron-diarization
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceSep 24

Ringg Claims 65% Call Resolution and 90% Lower AI Costs With GPT-5.6

Artificial IntelligenceSep 23

OpenAI Academy Marks Two Years and Expands AI Skills Access

Artificial IntelligenceSep 23

Google Gemini 3.8 TTS Launches Custom Voices in 100-Plus Languages