GPT-Audio-1.5 vs GPT-Realtime-2.1
Side-by-side comparison of API pricing, specs, benchmarks, and capabilities
GPT-Audio-1.5 is 51.4% cheaper on blended cost ($4.38 vs $9.00/Mtok)
| Specification |
by OpenAI
|
by OpenAI
|
|---|---|---|
| Overview | ||
| Status | Current flagship | Current flagship |
| Released | Feb 23, 2026 | Jul 6, 2026 |
| Pricing per million tokens | ||
| Input | $2.50/Mtok | $4.00/Mtok |
| Output | $10.00/Mtok | $24.00/Mtok |
| Blended avg | $4.38/Mtok | $9.00/Mtok |
| Cached input | — | $0.400/Mtok |
| Price basis | OpenAI prices this model per 1M tokens on two separate modalities: text at $2.50 input / $10.00 output, and audio at $32.00 / $64.00. The tracked figures are the TEXT rates, matching how the realtime rows are tracked. OpenAI publishes no cached-input rate for this model on either modality, so the cached-input column is empty rather than estimated. | OpenAI prices this model per 1M tokens on three separate modalities: text at $4.00 input / $0.40 cached input / $24.00 output, audio at $32.00 / $0.40 / $64.00, and image input at $5.00 / $0.50. The tracked figures are the TEXT rates, which is how GPT-Realtime-2 is tracked too, so the two rows compare like for like. Every text and audio rate is identical to GPT-Realtime-2; only the model generation differs. |
| Specifications | ||
| Context window | 128K tokens | 128K tokens |
| Parameters | Proprietary | Proprietary |
| Speed (TPS) | — | — |
| Modalities | ||
| Input |
audiotext
|
audiotextimage
|
| Providers | ||
| Available from |
OpenAI — $2.50/$10.00/Mtok
|
OpenAI — $4.00/$24.00/Mtok
|
Cost at scale
1M tokens · 50/50 input/output| Volume | GPT-Audio-1.5 | GPT-Realtime-2.1 | Savings |
|---|---|---|---|
| 1M tokens | $4.38 | $9 | $4.63 (51.4%) |
| 10M tokens | $43.75 | $90 | $46.25 (51.4%) |
| 100M tokens | $437.5 | $900 | $462.5 (51.4%) |
| 1000M tokens | $4375 | $9000 | $4625 (51.4%) |
When to pick which
distilled from the pricing and spec data above
Pick GPT-Audio-1.5 if…
- cost dominates: $4.38/Mtok blended vs $9.00 — 51.4% less on the same 50/50 token mix
Pick GPT-Realtime-2.1 if…
- your workload re-reads context (agents, RAG, long chats): cached input costs $0.400/Mtok — 90% off its list input price, a discount GPT-Audio-1.5 doesn't offer
Try GPT-Audio-1.5 on OpenAI
The cheaper option here — GPT-Audio-1.5 costs $4.38/Mtok blended on OpenAI.
Summary
GPT-Audio-1.5 by OpenAI costs $2.50/Mtok input and $10.00/Mtok output, with a 128K-token context window. It supports audio, text input.
GPT-Realtime-2.1 by OpenAI costs $4.00/Mtok input and $24.00/Mtok output, with a 128K-token context window. It supports audio, text, image input.
On a blended cost basis, GPT-Audio-1.5 is 51.4% cheaper than GPT-Realtime-2.1.
Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.
More comparisons
pairs sharing a model with this page
GPT-Realtime-2.1 mini vs GPT-Realtime-2.1
88% price gap
GPT-Realtime-2.1 vs GPT-Realtime-2
GLM-4.7-Flash vs Ministral 3 3B
100% price gap
Solar Pro 4 vs Gemini 3.5 Flash-Lite
94% price gap
Solar Pro 4 vs DeepSeek V4 Flash
92% price gap
Claude Fable 5 vs DeepSeek V4 Pro
90% price gap
GPT-5.6 Terra vs GPT-5.6 Luna
90% price gap
Llama 3.3 70B vs Claude Sonnet 4.6
89% price gap