Google's 4B model cuts speaker-labelling errors significantly
Google has released a new 4B-parameter model called DiarizationLM-Gemma-4-E4B-v1 under its Hugging Face organization on October 4. The model improves speaker-labelling accuracy on all four standard diarisation test sets, with statistically significant reductions in errors at p < 0.0001. Google emphasizes that this is not an officially supported product but rather a set of code and weights.
The model focuses on correcting the output of existing speech recognition and diarisation systems rather than replacing them. It reduces the word diarisation error rate (WDER) significantly on telephone conversations but shows smaller gains in meeting transcripts. For example, WDER dropped from 5.32% to 2.99% on Fisher English and from 7.74% to 4.92% on Callhome American English. In meeting corpora like ICSI and AMI, the improvements were more modest, at -0.60 and -0.79 points respectively.
Despite having half the parameters of its predecessor, the 4B model outperformed the 8B version on telephone speech tasks. Google attributes these gains to its training method, Locality-Preserving Oracle Supervision, which refines turn boundaries and speaker identities. The training involved 10,000 steps using a LoRA adapter on Google Cloud TPU v5p chips. The model is licensed under Apache 2.0, with no downloads logged yet.
While the model shows promise, Google has not published latency figures or on-device performance metrics. For comparison, Nvidia released a 100M-parameter diarisation model on September 28 that handles labelling directly, rather than correcting existing outputs. The underlying framework, DiarizationLM, was introduced at Interspeech 2024.