Third Door Develops Japanese Speech Recognition Engine That Runs on CPU Alone
The engine's CPU-only operation and accuracy gains could lower the barrier for speech recognition in applications where GPUs or cloud APIs are impractical, such as phone response automation and on-premise transcription.
Reporting from 1 source: ASCII.jp.
LLC Third Door has developed a Japanese speech recognition engine that runs on a standard PC's CPU without a GPU. In a same-condition comparison using all 5,000 utterances of the JSUT basic5000 dataset, it achieved a character error rate of 1.86%, outperforming OpenAI's Whisper large-v3 (3.15%) and Google Speech-to-Text (3.29%). The engine operates at about 13 times real-time speed on an Apple M4 notebook CPU and is available for testing in a browser.
LLC Third Door has developed a Japanese speech recognition engine that runs on the CPU of a standard PC alone, without requiring a GPU. In a same-condition comparison using all 5,000 utterances of the public evaluation data JSUT basic5000, the engine recorded a character error rate of 1.86%, outperforming OpenAI's Whisper large-v3 (3.15%) and Google Speech-to-Text (3.29%).
The engine operates at approximately 13 times real-time speed on the CPU of a notebook PC (Apple M4), and it is available on a demo page that can be tried from a browser. The company notes that many high-accuracy speech recognition systems assume GPUs or cloud APIs, which has been a barrier in applications where it is difficult to prepare a GPU for each line or send voice data externally.
Applying the same design approach to English, the engine recorded a word error rate of 1.91% on the LibriSpeech test-clean dataset, surpassing Whisper large-v3 (2.47%) under the same scoring conditions.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.