Skip to main content
Meta Superintelligence Labs/Real-time audio perception

Muse
VoiceTranscribe

The real-time speech perception model built to transcribe as people talk, keep track of who said what, and recognize when a turn is complete.

Model snapshot01.0 / API
streaming_asrlive
speaker_attribution20+
endpoint_detectionnative
API price$0.18 / hour

Independent guide · verified Sep 2026

01 / Core capabilities

One stream. Three signals. A better ear for software.

Muse VoiceTranscribe brings recognition, speaker context, and turn boundaries into one real-time model instead of stitching together separate batch tools.

01 / STREAMING ASR

Words arrive while the audio is still moving.

Muse processes audio in 80 ms chunks and balances recognition accuracy with adaptive delay, so live transcripts feel immediate without throwing context away.

02 / DIARIZATION

Every turn keeps its speaker context.

Live attribution works for 20+ speakers inside the recognition model, including overlapping and messy real-world conversations.

03 / ENDPOINTING

The model knows when a thought is done.

Speech onset and end-of-speech detection are emitted as part of the stream, making voice interfaces easier to hand off and respond to.

02 / Evidence from Meta

Built for the speed-accuracy tradeoff.

See benchmark methodology
Meta benchmark showing Muse Voice Transcribe streaming word error rate
Streaming final-transcription WER comparison published by Meta Research.
Meta benchmark showing Muse Voice Transcribe diarization error rate
Diarization error rate across public multi-speaker benchmarks.
Meta benchmark showing Muse Voice Transcribe adaptive delay tradeoff
Adaptive delay moves the model along the speed-accuracy frontier.

Benchmark values and rankings are reproduced from Meta’s September 1, 2026 announcement. They are reported claims, not an independent evaluation by this site.

03 / Model details

Designed for real conversations, not clean test audio.

The model is trained on 70+ languages, validates 25 for the initial release, and supports seamless code-switching, context biasing, long audio, and more than 20 speakers.

70+

languages trained

25

validated languages

20+

speakers tracked

1h+

long audio

04 / Where it fits

01

Meetings

Live notes with speaker-aware turns

02

Voice agents

Fast handoff after a user finishes speaking

03

Podcasts

Long-form transcripts with attribution

04

Dictation

System-wide voice input on Mac

Meta says system-wide dictation is available on Mac by holding Fn in any application.

05 / Access

Ready to build with the signal.

Muse Voice Transcribe is listed in Meta Model API as `muse-voice-transcribe-1.0` at $3 per 1,000 minutes. Availability and account requirements can change, so use Meta’s developer page as the source of truth.

06 / FAQ

The short version.

What is Muse VoiceTranscribe?

Muse Voice Transcribe is Meta Superintelligence Labs’ first real-time audio perception model. It combines streaming automatic speech recognition, speaker diarization, and endpointing in one model.

Is Muse VoiceTranscribe available now?

Yes. Meta lists muse-voice-transcribe-1.0 in the Meta Model API. It is also used by Meta AI for Mac and Muse Code. Check the official developer page for current access requirements.

How much does the Meta Model API cost?

The published price is $3 per 1,000 minutes, or $0.18 per hour, for muse-voice-transcribe-1.0.

How many languages and speakers does it support?

The model is trained on 70+ languages, with 25 languages validated for the initial release. It supports 20+ speakers and audio longer than one hour.

Is musevoicetranscribe.net an official Meta website?

No. This is an independent information site. The official sources are linked throughout the page and should be treated as the authority for documentation, availability, and pricing.

Stay close to the release

The model is here. The ecosystem is next.