Meta has thrown down a gauntlet in the real-time speech-to-text market with the launch of Muse Voice Transcribe, a unified audio perception model that packages streaming transcription, speaker diarization for more than 20 participants, and endpoint detection into a single API priced at $0.18 per hour of processed audio. The offering, developed by Meta Superintelligence Labs, processes speech as it happens rather than waiting for a complete recording, and arrives at a moment when enterprise developers are increasingly demanding transcription systems that can track who said what without breaking the bank. At that price point and with those capabilities, Muse Voice Transcribe does not merely enter the market; it reshapes the competitive calculus for meeting intelligence, live agent infrastructure, and real-time voice analytics.
Muse Voice Transcribe Brings Streaming Diarization, Code-Switching, and Adaptive Delay Into One Model
Muse Voice Transcribe is designed to solve a problem that has long plagued speech-to-text systems: the tradeoff between latency and accuracy in multi-speaker environments. Traditional speech recognition answers only the question of what was said. Diarization adds the critical question of who said it. Meta’s new model addresses both simultaneously within an autoregressive multimodal architecture that processes audio in 80-millisecond chunks, or 12.5 chunks per second. Each chunk is transformed into a soft token, and at every step the model decides whether to consume more audio or emit text.
Meta calls this mechanism “adaptive delay.” Rather than applying a fixed latency budget to every word, Muse waits longer when speech is ambiguous and commits earlier when it has sufficient context. Reinforcement learning combines word-error-rate and delay rewards to train this behavior, giving the model a built-in ability to balance speed against precision depending on the acoustic situation.
The model supports long audio exceeding one hour, seamless multilingual code-switching, language and keyword biasing, and speaker diarization that operates without a separate post-processing pipeline. Muse was trained across more than 70 languages, with 25 extensively validated for the initial release. For enterprise developers building meeting systems, call analytics, live assistants, or ambient AI, that combination of features in a single model could matter more than any single headline number.
Speaker Attribution and Endpointing Are Built Into the Token Sequence
Meta has integrated speaker attribution and endpoint detection directly into the token-generation process rather than running speaker clustering as an unrelated downstream step. A special <|start_of_turn|>codecodecode token marks a potential new speaker turn, tokens such as <|speaker_A|>codecodecode identify the speaker, and separate onset and endpoint tokens define speech boundaries. This unified approach means that ASR, diarization, and endpointing are trained together, which Meta argues produces more reliable speaker attribution in real-time streaming conditions.
The Model API for Muse Voice Transcribe exposes diarization as a first-class operating mode alongside push-to-talk and endpointing. Speaker labels such as A and B are scoped to a session rather than verified identities, and the API provides turn-level rather than word-level timestamps. These design decisions reflect a focus on practical meeting and conversation scenarios where absolute speaker identity matters less than consistent attribution across a single session.
How Many Speakers Can Muse Voice Transcribe Handle? The 20-Plus Claim in Context
The 20-plus speaker figure that Meta advertises for Muse Voice Transcribe is substantial by market standards, but a review of current vendor documentation shows that it does not establish a new global record. Several competing systems publish higher maximums for real-time speaker diarization.
Speechmatics currently makes the strongest explicit capacity claim, with its real-time STT documentation indicating speaker diarization is available live and its FAQ specifying support for 50 speakers by default, with the limit configurable up to 100. Amazon Transcribe differentiates a maximum of 30 unique speakers and provides explicit instructions for speaker partitioning in streaming transcription. Soniox documents a maximum of 15 speakers per session for both real-time and asynchronous processing. AssemblyAI lets developers set max_speakers between one and 10 for its streaming diarization system, with both companies noting that live speaker attribution is inherently more difficult than offline processing because streaming systems must make decisions with less future audio context. XAI’s current Speech-to-Text API supports speaker diarization in streaming mode, but its documentation does not publish a maximum speaker count, making a direct comparison with Muse impossible.
It would therefore be inaccurate to describe Muse’s 20-plus capability as a world record. The highest explicitly documented real-time number identified in this survey is Speechmatics’ configurable 100-speaker ceiling. Meta also does not demonstrate 20-plus simultaneous participants in its launch material. The principal live demonstration uses eight speakers, while the long-form recording contains 11 labeled participants. The 20-plus number is a stated model capability rather than a participant count verified in public demos.
Speaker capacity and diarization error rate should not be conflated. A platform capable of representing 100 people is not automatically better at correctly attributing speech than one supporting 20, and Meta’s benchmark does not test every competitor operating at its advertised maximum speaker count.
At $0.18 Per Hour, Meta Puts Pricing Pressure on the Entire Market
Meta’s pricing for Muse Voice Transcribe makes the competitive picture significantly more interesting. According to the Muse Voice Transcribe developer page, the service costs $3 per 1,000 minutes, or $0.18 per hour. Streaming and non-streaming transcription cost the same, and zero-data-retention processing is priced at parity with standard processing. Billing applies to audio actually processed and is rounded down to whole seconds.
Standardizing publicly posted rates to one hour of streaming audio reveals where Muse stands relative to the field. The comparison is necessarily imperfect because pricing models vary widely across vendors. Qwen’s price varies by deployment geography, with its international real-time rate of $0.00009 per second working out to approximately $0.324 per hour. Google’s Gemini figure is an estimated blended token cost rather than a flat hourly tariff. AWS prices vary by region and usage tier. ElevenLabs lists $0.39 per hour on its API pricing page but advertises $0.28 per hour or lower on annual Business plans.
Deepgram’s pricing illustrates why feature-level comparisons matter. Its current Nova-3 Multilingual streaming rate is about $0.35 per hour, but speaker diarization costs an additional $0.002 per minute, bringing the comparable total to roughly $0.47 per hour. AssemblyAI similarly lists $0.45 per hour for Universal-3.5 Pro Realtime and another $0.12 per hour for streaming diarization. Cartesia is harder to normalize because Ink-2 is packaged through monthly credit plans rather than a simple metered PAYG hourly rate, but its $5 Pro plan includes roughly nine hours and 16 minutes of Ink-2 transcription, which works out to about $0.54 per transcription hour if every credit is consumed exclusively on speech-to-text.
Even with those caveats, Muse’s positioning is clear. It is not the absolute cheapest streaming transcription service — Soniox currently publishes a lower equivalent rate — but $0.18 per hour with diarization included puts Meta toward the low end of the market, especially against providers that charge separately for speaker attribution. At 1,000 hours of processed audio, Meta’s public rate implies roughly $180 in transcription charges, compared to $470 or more from providers that unbundle diarization as an add-on.
What Is the True Cost of Real-Time Speech-to-Text With Speaker Diarization?
For developers evaluating Muse Voice Transcribe against the competition, the total cost of ownership depends heavily on whether diarization is included in the base price or sold as an add-on. Meta’s $0.18 per hour includes streaming transcription and speaker diarization for up to 20 speakers in a single price. By contrast, Deepgram charges approximately $0.35 per hour for base transcription plus another $0.12 per hour for diarization, bringing the effective rate to $0.47 per hour. AssemblyAI charges $0.45 per hour for real-time transcription plus $0.12 per hour for streaming diarization, for a total of $0.57 per hour. At scale, the difference is substantial: 10,000 hours of audio processed with Meta would cost $1,800, while the same volume with Deepgram or AssemblyAI would cost $4,700 or $5,700 respectively, assuming all diarization features are enabled. Muse Voice Transcribe thus offers a significant cost advantage for workloads that require speaker attribution as part of the core transcription pipeline.
Muse Leads Meta’s Supplied Streaming Accuracy Benchmarks
Price matters less if accuracy suffers, and Meta’s benchmark material argues that Muse performs competitively at the top of the field. On the Artificial Analysis AA-WER Streaming Index supplied with the launch, Muse records a 3.1% final-transcription word error rate. That figure places it ahead of Cartesia Ink-2 at 3.4%, ElevenLabs Scribe v2 Realtime at 3.6%, Qwen3 ASR Flash Realtime at 3.7%, GPT Live Transcribe and Grok Speech to Text Streaming at 3.9%, and Gemini 3.5 Transcribe Live and AssemblyAI U3.5 Realtime Pro at 4.0%.
Meta points out that Muse took the number one spot on third-party independent AI benchmarking firm Artificial Analysis’ streaming speech-to-text evaluation as of September 1. The benchmark charts published in Meta’s launch post show Muse leading on final transcription word error rate across the tested competitors.
The diarization result may be even more relevant to the product’s positioning. Meta reports an average 17.5% diarization error rate across the AMI-IHM, AMI-SDM, and VoxConverse datasets, lower than the competing systems shown in its chart. This suggests that Muse’s unified architecture, which trains ASR, diarization, and endpointing together, may confer a real advantage in correctly attributing speech compared to systems that run speaker clustering as a separate post-processing step.
Technical Architecture: How Adaptive Delay and Unified Training Differentiate Muse
The technical decisions behind Muse Voice Transcribe reflect a deliberate architectural bet. Meta describes audio arriving in 80-millisecond chunks, each transformed into a soft token. At every step, the model decides whether to consume more audio or emit text. This adaptive delay mechanism contrasts with fixed-latency approaches that apply the same time budget to every utterance regardless of acoustic clarity.
Reinforcement learning combines word-error-rate and delay rewards to train the model’s timing behavior. When speech is clear and unambiguous, the model can commit to a transcription early, producing low-latency output. When speech is noisy, accented, or context-dependent, the model can wait for additional audio before deciding. This approach mirrors how human listeners handle ambiguous speech: they pause, gather more information, and then commit to an interpretation.
The unification of ASR, diarization, and endpointing within a single token sequence represents another architectural differentiator. Traditional speech-to-text pipelines run speaker clustering as a separate process after transcription is complete, which introduces latency and can produce artifacts when speaker turns overlap or change rapidly. By incorporating speaker tokens and turn-boundary tokens directly into the generation process, Muse can attribute speech to speakers at the same time it transcribes the words, reducing both latency and error propagation.
Meta’s approach does come with tradeoffs. The API currently provides turn-level timestamps rather than word-level timestamps, and it does not expose word-level confidence scores, sound-event detection, or emotion detection. The documentation specifies eight concurrent streams per tenant by default and real-time sessions of up to 60 minutes before an application must reconnect. These constraints may matter for certain use cases, particularly those requiring fine-grained timing analysis or long-running continuous streaming.
Competitive Implications: Who Benefits From Meta’s Entry and Who Feels the Pressure
Meta’s entry into the real-time speech-to-text market with Muse Voice Transcribe creates pressure on multiple fronts. The $0.18 per hour pricing with diarization included challenges vendors that have built pricing models around separating these capabilities. Deepgram, AssemblyAI, and ElevenLabs, among others, now face the question of whether to adjust their pricing structures or to differentiate on features that Meta does not yet offer, such as word-level confidence scores, sound-event detection, or higher maximum speaker counts.
For enterprise developers, the arrival of a low-cost, high-accuracy option from a major infrastructure player reduces the risk of committing to a single vendor. Muse Voice Transcribe provides a credible alternative to established providers, and its inclusion in Meta’s broader developer ecosystem means that organizations already using other Meta AI services may find integration straightforward. The model’s support for 70-plus languages and 25 validated at launch also widens the addressable market for real-time transcription in languages that smaller vendors may not cover as thoroughly.
The competitive picture is not entirely one-sided. Speechmatics holds the documented speaker-count advantage with a configurable 100-speaker ceiling, and Amazon Transcribe offers deeper integration with the AWS ecosystem that many enterprise customers already use. xAI’s Grok Speech to Text API and Google’s Gemini Transcribe Live both represent significant competitors with their own infrastructure advantages. Muse’s 20-plus speaker ceiling, while high, may not satisfy use cases that require tracking very large numbers of participants, such as certain conference room scenarios or legislative proceedings.
Still, for the vast majority of practical applications — team meetings, customer service calls, medical dictation, classroom transcription, voice agents — 20 speakers is more than sufficient. The combination of that capacity with competitive accuracy and aggressive pricing creates a compelling package that many enterprise teams will want to evaluate.
Enterprise Use Cases Where Muse Voice Transcribe Makes Immediate Sense
Several categories of enterprise application stand to benefit directly from Muse’s feature set and pricing. Meeting intelligence is perhaps the most obvious: systems that automatically transcribe and summarize team meetings require reliable speaker attribution to generate useful records. A meeting assistant that correctly transcribes every sentence but attributes an approval, commitment, or objection to the wrong participant creates an unreliable corporate record. Muse’s integrated diarization reduces that risk.
Customer-service analytics represents another strong use case. Call center recordings and live call analysis benefit from knowing which party said what, and the need to process audio in real time while calls are still in progress makes Muse’s streaming architecture relevant. The model’s multilingual code-switching capability is particularly valuable for customer service environments where agents and customers may switch between languages within a single conversation.
Live assistant and voice-agent infrastructure is a third category. AI agents that operate in rooms where several people can speak need to understand not just what was said but who said it, and they need that information with low latency to respond appropriately. Muse’s endpoint detection and speaker attribution within a single model reduce the complexity of building such systems.
Compliance workflows, particularly in financial services and healthcare, require accurate speaker-attributed transcripts for regulatory record-keeping. Muse’s zero-data-retention option, priced at parity with standard processing, addresses privacy requirements while keeping costs manageable.
What Muse Voice Transcribe Does Not Yet Offer
A complete evaluation of Muse Voice Transcribe requires acknowledging its current limitations alongside its strengths. The API provides turn-level timestamps but does not expose word-level timestamps, which may be a constraint for applications that require precise timing alignment between audio and text. Word-level confidence scores are similarly absent, making it harder for downstream systems to gauge the reliability of individual transcription outputs.
Sound-event detection and emotion detection are not part of the current Muse feature set. Vendors such as AssemblyAI and Deepgram offer these capabilities as add-ons or as part of their extended API feature sets. For use cases that require identifying laughter, applause, silence, or emotional tone alongside transcription, Muse may need to be supplemented with additional processing.
The eight-concurrent-streams-per-tenant default and the 60-minute session limit before reconnection are operational constraints that may affect deployment patterns for high-volume or long-running applications. Developers planning to process many simultaneous streams or continuous audio feeds will need to account for these limits in their architecture.
Speaker labels are scoped to a session rather than being verified identities. This means that Speaker A in one session is not guaranteed to be the same person as Speaker A in another session. For applications that require persistent speaker identification across sessions, additional identity-matching logic must be built on top of the API output.
These limitations are not unusual for a first-generation API offering, and Meta may address them in future releases. But enterprise developers evaluating Muse should consider whether the current feature set meets their specific requirements or whether the gaps push them toward more established providers with more mature feature sets.
At $0.18 per hour, with 20-plus-speaker diarization inside the same real-time model that currently leads Meta’s supplied streaming accuracy benchmarks, Muse Voice Transcribe gives enterprise teams a serious new option for meeting intelligence, live transcription, and voice-agent infrastructure. The model does not break every record in the category, but it may not need to. For the majority of real-world multi-speaker scenarios, a service that preserves speaker attribution, accurate text, and usable turn boundaries while a complicated conversation is still unfolding — at a price that dramatically undercuts the competition — represents a threshold moment in the commoditization of real-time speech intelligence. Competitors will now have to justify premiums on speaker-aware accuracy and total operating cost, not merely raw speech recognition performance.