IBM's Granite Speech 5.0 Models Transcribe 3.5 Hours of Speech in One Second
IBM has released two compact English speech recognition models that can process more than 3.5 hours of speech in one second, significantly surpassing previous open models. The models, known as Granite Speech 5.0 TurboCTC and a non-commercial counterpart, have just 470 million parameters each and achieve an aggregate throughput of over 12,600 RTFx on a single NVIDIA H200 GPU.
The pair outperforms earlier Granite Speech releases by dropping the language model entirely, instead using a 16-layer Conformer encoder trained with connectionist temporal classification. This change enables the models to emit fewer, longer tokens at a quarter of the frame rate, resulting in significantly faster processing times.
While IBM claims competitive accuracy for these models, their performance is vendor-reported and awaits confirmation from the OpenASR Leaderboard. The strongest independent signal comes from far-field audio evaluations, where the non-commercial model ranks fifth in accuracy and the Apache model ninth on the FFASR Leaderboard.