Nemotron 3 Diarization
Nemotron 3 Diarization is NVIDIA's open-weight speaker diarization model that identifies who spoke when in multi-speaker audio, tracking up to 8 simultaneous speakers with streaming and offline inference.
NVIDIASpeechOpen weightsLaunched (date only)
Why it isn't ranked
Speaker diarization model: no ranked speech category yet.
Launches outside the ranked categories are tracked, not scored: they never appear on the charts or in the rankings.
Banner
On the “New in past 24 hours” banner for 24 hours after launch, until .
Official page
huggingface.co ↗Sources
“A 100M-parameter speaker diarization model designed to determine "who spoke when" in real-world audio.”