Anime, manga, and games, with a take · A Yukimedia publication

← all stories other 1 sources · 1h ago ·

Meta Releases Muse Voice Transcribe, Its First Real-Time Speech Recognition Model

Meta's first real-time speech recognition model tops Artificial Analysis rankings and is priced below rival transcription models, giving the Muse Spark family a direct entry into the streaming voice segment.

Reporting from 1 source: GIGAZINE.

Meta Releases Muse Voice Transcribe, Its First Real-Time Speech Recognition Model

Meta Superintelligence Labs released Muse Voice Transcribe, its first real-time speech recognition model. It handles streaming recognition, identifies 20 or more speakers, and detects the end of speech. It ranks first on Artificial Analysis with a final transcription error rate of 3.1 percent. Pricing runs $0.18 per hour or $3 per 1000 minutes. Trained on more than 70 languages, it is available from September 2 via Meta Model API, Meta AI for Mac, and Muse Code.

Muse Voice Transcribe is an autoregressive multimodal model in the Muse Spark family, Meta's line of reasoning models. Its speaker identification score on Artificial Analysis also shows a lower misrecognition rate than other models, and it natively supports code-switching, where bilingual speakers alternate languages within or across sentences.

The model handles audio longer than one hour and conversations with more than 20 participants without any post-recording speaker labeling or level adjustment. Twenty-five languages are extensively validated, including Japanese, English, Mandarin Chinese, Korean, and Hindi.

Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.

Sources