Audio and Speech
6 July 2026
OpenAI released GPT-Realtime-2.1-mini in the API
N
New-ZZZ desk
X @OpenAIDevs · 1 month ago
OpenAI has added the GPT-Realtime-2.1-mini model to the API. It expands the Realtime mini lineup with reasoning and tool use while keeping the price at the GPT-Realtime-mini level.
The company also said that p95 latency for all realtime voice models has been reduced by at least 25% thanks to improved caching. This should make voice scenarios faster and more responsive.
Why it matters
- —The new realtime model gains reasoning and tool use without a cost increase compared with the previous mini version.
- —Lower latency matters for voice assistants and other applications where users expect an almost instant response.
Key facts
- GPT-Realtime-2.1-mini became available through the API.
- The model adds reasoning and tool use to the Realtime mini lineup.
- Pricing is stated to be at the GPT-Realtime-mini level.
- OpenAI reduced p95 latency by at least 25% across all realtime voice models through improved caching.
The full text is in the original source. Here we provide a brief summary and key facts.