Skip to content
Model Pareto

Nemotron 3 Diarization

Nemotron 3 Diarization is NVIDIA's open-weight speaker diarization model that identifies who spoke when in multi-speaker audio, tracking up to 8 simultaneous speakers with streaming and offline inference.
NVIDIASpeechOpen weightsLaunched (date only)

Why it isn't ranked

Speaker diarization model: no ranked speech category yet.

Launches outside the ranked categories are tracked, not scored: they never appear on the charts or in the rankings.

Banner

On the “New in past 24 hours” banner for 24 hours after launch, until .

Official page

huggingface.co ↗

Sources

  • “A 100M-parameter speaker diarization model designed to determine "who spoke when" in real-world audio.”
    OfficialHugging Face (nvidia)Sep 23, 2026huggingface.co ↗

All other launches →