AssemblyAI

Universal 3.5 Pro Demo: Smarter Speech-to-Text with Contextual Awareness

Jul 1, 2026 10 min
speech-to-textassemblyaicontextual promptingvoice agentllm
Watch on YouTube Follow AssemblyAI on Rundown — free

Summary

AI summaries can be incomplete or wrong. Verify anything important against the original video.

This video demonstrates AssemblyAI's Universal 3.5 Pro model, focusing on new prompting and conversation context features that significantly improve speech-to-text accuracy. It highlights how providing contextual information, either through detailed prompts or by passing conversation history, helps the model understand and transcribe speech more effectively, especially in challenging audio conditions.

The video introduces AssemblyAI's Universal 3.5 Pro model, emphasizing its enhanced contextual awareness and prompting capabilities for improved speech-to-text accuracy. The presenter walks through two main features: contextual prompting and conversation context. Contextual prompting works by providing the model with information about the audio's domain, scenario, or detailed descriptions, which helps it handle uncommon names and terms. Key term prompting is also discussed, where specifying important terms boosts their recognition accuracy. A demonstration in the playground shows how adding a key term like 'Klebanoff' ensures accurate transcription. The video then explains the 'conversation context' feature, which passes previous transcriptions and agent replies as context to the model. This allows the model to understand the ongoing conversation and anticipate user intent, providing more relevant responses and improving accuracy, particularly in voice agent use cases. Statistical improvements on voice agent datasets are briefly shown. Finally, a detailed demo in the 'Context Carryover Playground' illustrates how providing agent context dynamically, even with poor audio, leads to better transcriptions and more intelligent agent interactions.

Concepts introduced

Locked

Key Points

Locked

Worth watching if: You are a developer or product manager looking to improve the accuracy and intelligence of speech-to-text applications, especially for voice agents. This video provides practical examples and demonstrations of how to leverage contextual prompting and conversation context features.

Sign in to unlock the full extract

Every claim, key point, and timestamp for this AssemblyAI video — plus a daily email of every channel you follow.

Sign in with Google

No credit card. Free tier forever.

Watch on YouTube