NVIDIA Releases Nemotron 3 Diarization As An Open Model
NVIDIA is competing in real-time transcription with an open license and a speaker-ID benchmark claim, which puts an eight-speaker model in reach of anyone building meeting, captioning, or archive tools without a hosted API.
Reporting from 1 source: GIGAZINE.
NVIDIA released Nemotron 3 Diarization, a speech recognition model that identifies speakers while transcribing in real time. It handles up to eight speakers and labels them in order of appearance as Speaker 1, Speaker 2, and so on, without registering voices in advance. Training used public datasets and data licensed from David AI. The model is open, under the OpenMDW-1.1 license, and available on Hugging Face.
The model splits speaker identification from transcription instead of running both as one pass, and it does not need each voice registered beforehand. It labels appearing speakers in sequence, so a recording with unknown participants can be divided by voice as it is transcribed.
NVIDIA says it trained the model on public datasets plus datasets licensed from David AI. The release includes a Hugging Face Space demo and the weights under the OpenMDW-1.1 license, with a YouTube demonstration showing speaker changes recognized during transcription.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.