Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 43m ago ·

Hayamimi Runs Real-Time Multilingual Speech Recognition on CPU Alone

A CPU-only pipeline that hits 3.8% character error rate on Japanese broadcast audio removes the GPU and cloud dependency that most real-time speech recognition still assumes, which puts live subtitles within reach on ordinary hardware.

Reporting from 1 source: GIGAZINE.

Hayamimi Runs Real-Time Multilingual Speech Recognition on CPU Alone

Hayamimi is a real-time multilingual speech recognition system that runs on CPU without GPUs or cloud APIs. It starts showing subtitles while speech is in progress and finalizes Japanese output about 100ms after a speaker stops. It works under 2GB of memory, supports live subtitles, speaker labeling, and translated subtitles, and uses INT8-quantized ONNX models on sherpa-onnx.

Hayamimi routes each utterance to a dedicated model after determining its language, an approach that departs from the single-model setup typical of CPU-based speech recognition. The models are INT8-quantized ONNX files running on sherpa-onnx, so PyTorch and CUDA are not required. On a 6-core desktop CPU it reaches speeds 10 to 50 times real time. A two-pass correction step re-decodes the preceding utterance after two seconds of silence, which improves Japanese character error rate from 15.5% to 12.0%. It requires Python 3.10 or higher and ffmpeg on the PATH.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources