Nemotron 3 Diarization tested: 100M-param real-time speaker diarization, better than expected locally

Japanese developer @ouchi tests NVIDIA Nemotron 3 Diarization (open-sourced Sep 23): real-time streaming accuracy is quite good. The ~100M-parameter model is built on the Streaming Sortformer architecture (31-layer Transformer + Arrival-Order Speaker Cache) and does diarization in a single forward pass: up to 8 speakers, 10 ms resolution, streaming latency down to about 320 ms, OpenMDW 1.1 license, runs on 4GB GPUs. The author observes the model accumulates per-speaker voice characteristics as it listens, so offline analysis of long files is slightly more accurate than real time; the HF demo caps at 2-minute inputs which limited quality, but running locally exceeded expectations. Tops VoiceArena Diarization-Bench (DER 14.72%) and scores 9.8% on AISHELL-4 (predecessor 27.2%). A key missing piece for voice agents knowing who is talking - directly relevant to robot voice interaction.





