Signal Brief

AssemblyAI发布Universal-3.5 Pro Realtime流式STT模型

AssemblyAI发布Universal-3.5 Pro Realtime流式语音转文本模型,在AA-WER Streaming上实现4.1% WER(最大准确率模式),首词延迟0.44秒,价格维持$0.45/小时不变,支持18种语言(上一代为6种),并支持对话中更新上下文。

twitter关注列表 Artificial Analysis (@ArtificialAnlys) 发布 2026-07-06 收录 2026-07-06 观察

一句话判断

与上一代相比,语言支持从6种增至18种,且首词延迟略降,但价格不变;值得关注其在实际对话场景中的上下文利用效果。

核心信息

AssemblyAI发布Universal-3.5 Pro Realtime流式语音转文本模型,在AA-WER Streaming上实现4.1% WER(最大准确率模式),首词延迟0.44秒,价格维持$0.45/小时不变,支持18种语言(上一代为6种),并支持对话中更新上下文。

原始内容

AssemblyAI has released Universal-3.5 Pro Realtime: a streaming Speech to Text model achieving 4.1% WER on AA-WER Streaming (~0.4s to first final), able to take in conversation context at the start of a call and after each agent turn, without reconnecting Universal-3.5 Pro Realtime is AssemblyAI's latest streaming Speech to Text (STT) model, the successor to Universal-3 Pro Realtime. The model offers three default modes, Balanced (default), Max Accuracy, and Min Latency, each a combination of lower level streaming parameters. On AA-WER Streaming, the Max Accuracy variant is level with its predecessor, Universal-3 Pro Realtime, and the Min Latency variant is ~10% faster. The model now takes conversation context that can be updated turn by turn. Agent's replies can be passed in at connection and refreshed mid-stream after each turn with no reconnect, giving more contextually relevant outputs. For example, priming it with "What's your email address?" yields "[email protected]" instead of "user at gmail dot com". Key takeaways ➤ First Final Transcription: Universal-3.5 Pro Realtime achieves a 4.1% WER at 0.44s after end of speech in Max Accuracy mode, more accurate than the faster Deepgram Flux (7.4%, 0.02s) and Deepgram Nova-3 Realtime (6.6%, 0.07s), and behind the more accurate Cartesia Ink-2 external endpoints (3.7%, 0.09s) and ElevenLabs Scribe v2 Realtime (3.6%, 0.14s). Min Latency mode achieves a 4.3% WER, slightly faster at 0.40s. ➤ First Partial Transcription: WER for the Max Accuracy variant on First Partial is the same as First Final, 4.1% at 0.44s after end of speech, ahead of Cartesia Ink-2 external endpoints (4.3%, 0.07s) and behind only ElevenLabs Scribe v2 Realtime (3.6%, 0.13s) on accuracy, though slower to emit than both. Min Latency mode trades accuracy for speed, returning a first partial at 6.1% WER at 0.39s ➤ Price: Universal-3.5 Pro Realtime costs $0.45/hr ($7.50 per 1,000 minutes), unchanged from Universal-3 Pro Realtime ➤ Language support: The model supports 18 languages, up from 6 in Universal-3 Pro Realtime, with mid-sentence code-switching. See more details below ⬇️ ![photo](https://pbs.twimg.com/media/HMjlTkAboAAD2xt.jpg) Artificial Analysis (@ArtificialAnlys): Universal-3.5 Pro Realtime is available for $0.45 per hour ($7.50 per 1,000 minutes) of audio direct from AssemblyAI. Both default modes, Max Accuracy and Min Latency, are priced the same and unchanged from Universal-3 Pro Realtime. https://t.co/99qF6CjjnA Artificial Analysis (@ArtificialAnlys): Full results: https://t.co/wDb6a2nhqV Methodology: https://t.co/ePPoyfUXXm

相关动态

02

Kimi K3非对称竞争分析

Kimi K3在Arena前端/WebDev人类盲测中以1679 Elo排名第一,领先Fable 5(1631)和GPT-5.6 Sol(1618),但在综合智能指数中仅排第四(57.1分),落后Fable 5(59.9)和GPT-5.6 Sol(58.9)。K3定价15美元/百万输出tokens,远低于Fable 5的50美元,并计划7...

twitter关注列表2026-07-18#AI#模型#技术突破
观察
04

BestBlogs 早报 · 07-18

月之暗面发布Kimi K3,2.8万亿参数,896选16的Stable LatentMoE,上下文100万token,接近Fable-5但仍落后最强闭源模型,完整权重7月27日前开源;VentureBeat调查显示54%企业已发生AI代理安全事件;xAI开源Grok Build(84万行Rust代码)并残留上传用户代码痕迹;Cursor评...

twitter关注列表2026-07-17#AI#模型发布#开源
观察