$NVDA Open-Sources Nemotron 3 Diarization to Enhance Real-Time Voice Agents

$NVDA Open-Sources Nemotron 3 Diarization to Enhance Real-Time Voice Agents
Nvidia has released Nemotron 3 Diarization, a 100M-parameter open-weight model designed to track active speakers, timestamps, and overlapping speech in real time. The technology addresses a key limitation in standard speech-to-text systems, which transcribe words but often struggle to identify which individual made a specific request or commitment.
The model tracks up to eight simultaneous speakers—double the capacity of Nvidia's previous streaming model—while maintaining consistent speaker labels throughout a conversation. Operating across both live and recorded audio with recommended streaming latency as low as 0.32 seconds, Nemotron 3 ranked first among 12 systems on VoiceArena’s initial benchmark with a 14.72% diarization error rate.
While the model labels participants generically as speaker 1 or speaker 2, applications can link these identifiers to individual user profiles. This capability enables multi-person AI assistants to separate comments and preserve individual context across meetings, customer service calls, and robotics applications.