Nvidia Releases Nemotron 3 Diarization AI Model
Nvidia has released Nemotron 3 Diarization, an AI model that identifies which speaker is talking at any given moment in a conversation. The model, with about 100 million parameters, can tell apart up to eight speakers and detect when multiple people talk at the same time.
While it's not perfect, Nemotron 3 outperforms its predecessor, Streaming Sortformer, by 41% in the VoiceArena Diarization Benchmark v1. The model also leads with a DER of 14.7%, significantly lower than other systems on the benchmark.
The model can be used in conjunction with speech recognition systems like Parakeet to produce transcripts with speaker labels, although these are currently anonymous.