Meta releases real-time audio transcription model

2026-09-10

Meta's Superintelligence Labs has introduced Muse Voice Transcribe, a real-time speech processing model. The model reportedly offers accurate, low-cost streaming transcription and speaker identification.

VERA Brief

AI-generated. Grounded in the article and its cited sources.

Meta's Superintelligence Labs has released Muse Voice Transcribe, a real-time speech processing model that offers accurate, low-cost streaming transcription and speaker identification. This model is intended as a foundational component for personal AI agents that can continuously monitor audio.

Key facts

  • Meta's Superintelligence Labs has introduced Muse Voice Transcribe, a new real-time audio transcription model.
  • The model processes speech in 80-millisecond segments, identifies different speakers, and detects sentence boundaries.
  • According to Artificial Analysis, Muse Voice Transcribe provides the most accurate streaming transcription at the lowest market price.
  • Muse Voice Transcribe is intended as a foundational component for personal AI agents capable of continuous audio monitoring.
  • This development raises questions about the accuracy and verifiability of AI systems continuously processing human speech for personal agents.

Source: The Decoder

Reported by VERA Newswire.

More from September 2026 in The Record.