01 / STREAMING ASR
Words arrive while the audio is still moving.
Muse processes audio in 80 ms chunks and balances recognition accuracy with adaptive delay, so live transcripts feel immediate without throwing context away.
The real-time speech perception model built to transcribe as people talk, keep track of who said what, and recognize when a turn is complete.
Independent guide · verified Sep 2026
01 / Core capabilities
Muse VoiceTranscribe brings recognition, speaker context, and turn boundaries into one real-time model instead of stitching together separate batch tools.
01 / STREAMING ASR
Muse processes audio in 80 ms chunks and balances recognition accuracy with adaptive delay, so live transcripts feel immediate without throwing context away.
02 / DIARIZATION
Live attribution works for 20+ speakers inside the recognition model, including overlapping and messy real-world conversations.
03 / ENDPOINTING
Speech onset and end-of-speech detection are emitted as part of the stream, making voice interfaces easier to hand off and respond to.
02 / Evidence from Meta



Benchmark values and rankings are reproduced from Meta’s September 1, 2026 announcement. They are reported claims, not an independent evaluation by this site.
03 / Model details
The model is trained on 70+ languages, validates 25 for the initial release, and supports seamless code-switching, context biasing, long audio, and more than 20 speakers.
70+
languages trained
25
validated languages
20+
speakers tracked
1h+
long audio
04 / Where it fits
Live notes with speaker-aware turns
Fast handoff after a user finishes speaking
Long-form transcripts with attribution
System-wide voice input on Mac
05 / Access
Muse Voice Transcribe is listed in Meta Model API as `muse-voice-transcribe-1.0` at $3 per 1,000 minutes. Availability and account requirements can change, so use Meta’s developer page as the source of truth.
06 / FAQ
Muse Voice Transcribe is Meta Superintelligence Labs’ first real-time audio perception model. It combines streaming automatic speech recognition, speaker diarization, and endpointing in one model.
Yes. Meta lists muse-voice-transcribe-1.0 in the Meta Model API. It is also used by Meta AI for Mac and Muse Code. Check the official developer page for current access requirements.
The published price is $3 per 1,000 minutes, or $0.18 per hour, for muse-voice-transcribe-1.0.
The model is trained on 70+ languages, with 25 languages validated for the initial release. It supports 20+ speakers and audio longer than one hour.
No. This is an independent information site. The official sources are linked throughout the page and should be treated as the authority for documentation, availability, and pricing.
Stay close to the release