AssemblyAI has released Universal-3.5 Pro Realtime: a streaming Speech to Text model achieving 4.1% WER on AA-WER Streaming (~0.4s to first final), able to take in conversation context at the start of a call and after each agent turn, without reconnecting
Universal-3.5 Pro Realtime is AssemblyAI's latest streaming Speech to Text (STT) model, the successor to Universal-3 Pro Realtime. The model offers three default modes, Balanced (default), Max Accuracy, and Min Latency, each a combination of lower level streaming parameters. On AA-WER Streaming, the Max Accuracy variant is level with its predecessor, Universal-3 Pro Realtime, and the Min Latency variant is ~10% faster.
The model now takes conversation context that can be updated turn by turn. Agent's replies can be passed in at connection and refreshed mid-stream after each turn with no reconnect, giving more contextually relevant outputs. For example, priming it with "What's your email address?" yields "
[email protected]" instead of "user at gmail dot com".
Key takeaways
➤ First Final Transcription: Universal-3.5 Pro Realtime achieves a 4.1% WER at 0.44s after end of speech in Max Accuracy mode, more accurate than the faster Deepgram Flux (7.4%, 0.02s) and Deepgram Nova-3 Realtime (6.6%, 0.07s), and behind the more accurate Cartesia Ink-2 external endpoints (3.7%, 0.09s) and ElevenLabs Scribe v2 Realtime (3.6%, 0.14s). Min Latency mode achieves a 4.3% WER, slightly faster at 0.40s.
➤ First Partial Transcription: WER for the Max Accuracy variant on First Partial is the same as First Final, 4.1% at 0.44s after end of speech, ahead of Cartesia Ink-2 external endpoints (4.3%, 0.07s) and behind only ElevenLabs Scribe v2 Realtime (3.6%, 0.13s) on accuracy, though slower to emit than both. Min Latency mode trades accuracy for speed, returning a first partial at 6.1% WER at 0.39s
➤ Price: Universal-3.5 Pro Realtime costs $0.45/hr ($7.50 per 1,000 minutes), unchanged from Universal-3 Pro Realtime
➤ Language support: The model supports 18 languages, up from 6 in Universal-3 Pro Realtime, with mid-sentence code-switching.
See more details below ⬇️

Artificial Analysis (@ArtificialAnlys): Universal-3.5 Pro Realtime is available for $0.45 per hour ($7.50 per 1,000 minutes) of audio direct from AssemblyAI. Both default modes, Max Accuracy and Min Latency, are priced the same and unchanged from Universal-3 Pro Realtime. https://t.co/99qF6CjjnA
Artificial Analysis (@ArtificialAnlys): Full results: https://t.co/wDb6a2nhqV
Methodology: https://t.co/ePPoyfUXXm