IBM Unveils Compact Granite Speech Models for High-Speed English Transcription
IBM has released two new compact Granite Speech 5.0 models for enterprises that need to transcribe large volumes of English audio quickly, either on their own infrastructure or edge devices.
The release marks a shift in IBM's approach to speech-to-text, focusing on downloadable open-weight models and self-hosted deployment, rather than offering a broad hosted service spanning multiple languages and voice functions.
The two models have 470 million parameters and can process more than 3.5 hours of speech per second during batched inference on a single NVIDIA H200 GPU, though this was achieved under ideal circumstances.
IBM's Granite Speech 5.0 is an English transcription release only, not a multilingual or translation model, which may limit its suitability for certain applications such as multilingual contact centers or global media operations.