Product roadmap

What We Are Building

What's coming next across the AssemblyAI platform — from the speech models we build to the APIs and tooling developers integrate — and what shipped in the last 90 days.

MODELS

Models

The speech models we build: speech-to-text, text-to-speech, LLM, and full-duplex conversation.

Upcoming

The next Universal-3.x release. Native-language coverage will grow from 18 to over 30 languages, with accuracy gains on already supported languages. A new model option priced the same as Universal-3.5.

New languages: Korean, Russian, Greek, Bulgarian, Polish, Romanian, Hungarian, Czech, Thai, Tagalog, Indonesian, Afrikaans.

Universal-TTS 4

Q4 2026

A standalone text-to-speech model for production voice workloads. Low time-to-first-byte, voice prompting, and accurate delivery of phone numbers, emails, and named entities that today’s TTS struggles with.

The next major accuracy and capability release after Universal-3.x, across pre-recorded and realtime audio. Targets the lowest turn latency and the strongest handling of voice-agent audio (noise, interruptions, hesitation, accented speech), with instruction-following strong enough to replace today’s STT + LLM + TTS stack. The foundation for our speech-to-speech architecture.

A single native model that replaces today’s Voice Agent pipeline (STT, LLM, TTS) with a unified Realtime Speech LLM. Tighter latency, better prosody, and more natural interruption handling than orchestrated stacks.

Shipped in last 90 days

INFERENCE

Inference

How models run in production — across transports, regions, and deployment options.

Upcoming

Ongoing catalog expansion across Anthropic, OpenAI, Google, and Qwen models, with DeepSeek and other open-weight models through new inference providers next.

Self-Hosted

Q4 2026

Run our speech models on-premise — streaming and pre-recorded — for regulated environments with strict data-residency requirements.

Shipped in last 90 days

ORCHESTRATION

Orchestration

Infrastructure for running voice applications in production.

Upcoming

Purchase phone numbers directly through the AssemblyAI API — no third-party vendor required — and handle inbound and outbound calls with DTMF keypad support. A SIP integration guide with no bridge required is already live.

Make backend requests while the phone is still ringing, before the call connects. The caller's phone number is included in the request, so you can run a CRM lookup and use the result to decide whether to answer the call at all — or to greet the caller by name with a dynamic, personalized greeting.

Shipped in last 90 days

APIS

APIs

The endpoints developers integrate, built on the layers above.

Upcoming

Dictation API

Q3 2026

Real-time dictation with formatting, punctuation, and voice commands built in.

Accuracy improvements to Speaker ID, Translation, and Custom Formatting. Translation covers both streaming and pre-recorded audio, for workflows where the spoken language differs from the output.

Category redaction for entity types you define, building on the static find-and-replace redaction that has already shipped.

Recognize the same speaker across recordings, not just within a single file. For meetings, call centers, and cross-session analytics.

Detect speaker emotions and shifts in input audio. For therapy, CX scoring, and compliance monitoring.

Shipped in last 90 days

DEVELOPER / AGENT EXPERIENCE

Developer & Agent Experience

The dashboard, accounts, and tooling that make AssemblyAI easy to adopt — for humans and AI agents.

Upcoming

Voice Agent SDK

Q4 2026

Official client libraries, starting with Python and TypeScript. The SDK owns the connection and audio pipeline, so agents work on the first try.

Deeper dashboard observability: P50 and P95 turnaround time, webhook delivery stats, uptime, and latency histograms.

A coding agent that gets you started with working sample code for your use case.

Shipped in last 90 days