OpenAI API Pricing & Model Reference
Explore upstream provider rates and model pricing for OpenAI models routed through Valstorm AI Platform, including batch execution and context caching rates.
Official Provider Documentation: OpenAI API Pricing | OpenAI Models Overview
Overview & Gateway Architecture
This reference document outlines the upstream provider base pricing for OpenAI models.
Like open model gateways and routing platforms (such as OpenRouter or OpenCode), the Valstorm AI Platform provides unified access to these upstream models alongside advanced enterprise capabilities:
- Intelligent Routing & Failover: Dynamic routing across models and providers with zero-downtime fallback.
- Enterprise Tool Execution: Sandboxed compute environments, web search grounding, vector retrieval (VFS), and database connectors.
- Context Caching & Optimization: Transparent prompt caching and token minimization pipelines.
- Multi-Tenant Security & Observability: Granular cost allocation, audit logs, and policy-driven rate limiting.
Note: Base provider rates shown below represent upstream pass-through pricing in USD per 1 Million (1M) tokens unless specified otherwise. Platform usage includes upstream consumption plus applicable platform value-add markup.
1. Flagship & Reasoning Models
All prices are in USD per 1M tokens (1,000,000 tokens).
Standard Execution Tier
| Model ID | Short Context Input | Short Context Cached Input | Short Context Cache Writes | Short Context Output | Long Context Input | Long Context Cached Input | Long Context Cache Writes | Long Context Output |
|---|---|---|---|---|---|---|---|---|
gpt-5.6-sol | $5.00 | $0.50 | $6.25 | $30.00 | $10.00 | $1.00 | $12.50 | $45.00 |
gpt-5.6-terra | $2.00 | $0.20 | $2.50 | $12.00 | $4.00 | $0.40 | $5.00 | $18.00 |
gpt-5.6-luna | $0.20 | $0.02 | $0.25 | $1.20 | $0.40 | $0.04 | $0.50 | $1.80 |
gpt-5.5 | $5.00 | $0.50 | — | $30.00 | $10.00 | $1.00 | — | $45.00 |
gpt-5.5-pro | $30.00 | — | — | $180.00 | $60.00 | — | — | $270.00 |
gpt-5.4 | $2.50 | $0.25 | — | $15.00 | $5.00 | $0.50 | — | $22.50 |
gpt-5.4-mini | $0.75 | $0.075 | — | $4.50 | — | — | — | — |
gpt-5.4-nano | $0.20 | $0.02 | — | $1.25 | — | — | — | — |
gpt-5.4-pro | $30.00 | — | — | $180.00 | $60.00 | — | — | $270.00 |
2. Processing Modes & Execution Tiers
OpenAI provides multiple execution modes optimized for cost, throughput, and latency.
Batch & Flex Modes (50% Cost Reduction)
For non-blocking asynchronous jobs (Batch API) or interruptible compute workloads (Flex Tier), base rates are discounted by 50%:
| Model ID | Batch/Flex Input | Batch/Flex Cached Input | Batch/Flex Cache Writes | Batch/Flex Output | Long Context Input | Long Context Cached Input | Long Context Output |
|---|---|---|---|---|---|---|---|
gpt-5.6-sol | $2.50 | $0.25 | $3.125 | $15.00 | $5.00 | $0.50 | $22.50 |
gpt-5.6-terra | $1.00 | $0.10 | $1.25 | $6.00 | $2.00 | $0.20 | $9.00 |
gpt-5.6-luna | $0.10 | $0.01 | $0.125 | $0.60 | $0.20 | $0.02 | $0.90 |
gpt-5.5 | $2.50 | $0.25 | — | $15.00 | $5.00 | $0.50 | $22.50 |
gpt-5.5-pro | $15.00 | — | — | $90.00 | — | — | — |
gpt-5.4 | $1.25 | $0.13 | — | $7.50 | $2.50 | $0.25 | $11.25 |
gpt-5.4-mini | $0.375 | $0.0375 | — | $2.25 | — | — | — |
gpt-5.4-nano | $0.10 | $0.01 | — | $0.625 | — | — | — |
gpt-5.4-pro | $15.00 | — | — | $90.00 | $30.00 | — | $135.00 |
Fast Mode / Priority Processing
Fast Mode (formerly Priority Processing, configurable via service_tier: "fast" or service_tier: "priority") provides low-latency dedicated throughput:
| Model ID | Fast Mode Input | Fast Mode Cached Input | Fast Mode Cache Writes | Fast Mode Output |
|---|---|---|---|---|
gpt-5.6-sol | $10.00 | $1.00 | $12.50 | $60.00 |
gpt-5.6-terra | $4.00 | $0.40 | $5.00 | $24.00 |
gpt-5.6-luna | $0.40 | $0.04 | $0.50 | $2.40 |
gpt-5.5 | $12.50 | $1.25 | — | $75.00 |
gpt-5.4 | $5.00 | $0.50 | — | $30.00 |
gpt-5.4-mini | $1.50 | $0.15 | — | $9.00 |
Data Residency & Regional Uplift
- Regional Processing: Dedicated regional data residency endpoints incur a 10% premium over standard global base rates for supported models.
3. Multimodal & Audio Models
Realtime & Voice APIs
Prices per 1M tokens unless specified as per-minute rates.
| Model ID | Modality | Input Rate | Cached Input Rate | Output / Time Rate |
|---|---|---|---|---|
gpt-realtime-2.1 | Audio | $32.00 / 1M | $0.40 / 1M | $64.00 / 1M |
| Text | $4.00 / 1M | $0.40 / 1M | $24.00 / 1M | |
| Image | $5.00 / 1M | $0.50 / 1M | — | |
gpt-realtime-2.1-mini | Audio | $10.00 / 1M | $0.30 / 1M | $20.00 / 1M |
| Text | $0.60 / 1M | $0.06 / 1M | $2.40 / 1M | |
| Image | $0.80 / 1M | $0.08 / 1M | — | |
gpt-realtime-translate | Audio Stream | — | — | $0.034 / minute |
gpt-live-transcribe | Live Audio | — | — | $0.017 / minute |
gpt-realtime-whisper | Whisper Stream | — | — | $0.017 / minute |
Transcription Models
| Model ID | Mode | Input (1M tokens) | Output (1M tokens) | Effective Rate |
|---|---|---|---|---|
gpt-transcribe | Async Transcription | — | — | $0.0045 / minute |
gpt-live-transcribe | Live Streaming Transcription | — | — | $0.0170 / minute |
gpt-4o-transcribe | Multimodal Transcription | $2.50 | $10.00 | ~$0.0060 / minute |
gpt-4o-mini-transcribe | Lightweight Transcription | $1.25 | $5.00 | ~$0.0030 / minute |
4. Visual Generation Models
Image Generation (gpt-image)
Prices per 1M tokens.
| Model ID | Tier | Image Input | Cached Input | Image Output | Text Input | Text Output |
|---|---|---|---|---|---|---|
gpt-image-2 | Standard | $8.00 | $2.00 | $30.00 | $5.00 | — |
| Batch | $4.00 | $1.00 | $15.00 | $2.50 | — | |
gpt-image-1.5 | Standard | $8.00 | $2.00 | $32.00 | $5.00 | $10.00 |
| Batch | $4.00 | $1.00 | $16.00 | $2.50 | $5.00 | |
gpt-image-1-mini | Standard | $2.50 | $0.25 | $8.00 | $2.00 | — |
| Batch | $1.25 | $0.13 | $4.00 | $1.00 | — |
Video Generation (sora-2)
Prices per second of generated video.
| Model ID | Resolution | Aspect / Orientation | Standard Price / Sec | Batch Price / Sec |
|---|---|---|---|---|
sora-2 | 720p | 720x1280 (Portrait) / 1280x720 (Landscape) | $0.10 / sec | $0.05 / sec |
sora-2-pro | 720p | 720x1280 / 1280x720 | $0.30 / sec | $0.15 / sec |
| 1024p | 1024x1792 / 1792x1024 | $0.50 / sec | $0.25 / sec | |
| 1080p | 1080x1920 / 1920x1080 | $0.70 / sec | $0.35 / sec |
5. Specialized & Coding Models
| Model ID | Category | Standard Input | Standard Cached Input | Standard Output | Fast Mode Input | Fast Mode Output |
|---|---|---|---|---|---|---|
chat-latest | Conversational | $5.00 | $0.50 | $30.00 | — | — |
gpt-5.3-codex | Code Generation & Debugging | $1.75 | $0.175 | $14.00 | $3.50 | $28.00 |
gpt-5.4-cyber | Security & Threat Analysis | Custom | Custom | Custom | Custom | Custom |
6. Built-in Tools & Runtime Infrastructure
| Tool / Service | Description | Pricing Rate |
|---|---|---|
| Web Search | Standard live web search (all models) | $10.00 / 1,000 calls + Search content tokens at model rates |
| Image Web Search | Visual search grounding | $10.00 / 1,000 calls + Search content tokens at model rates |
| Web Search (Reasoning) | Reasoning models (gpt-5, o-series) | $10.00 / 1,000 calls + Search content tokens at model rates |
| Web Search (Non-Reasoning) | Non-reasoning models preview | $25.00 / 1,000 calls (Search content tokens included) |
| Containers (Hosted Shell / Code Interpreter) | Sandboxed execution runtime (20-min session): • 1 GB RAM: $0.03 / session • 4 GB RAM: $0.12 / session • 16 GB RAM: $0.48 / session • 64 GB RAM: $1.92 / session | Billed by the minute (5-minute minimum per session) |
| File Search (Vector Store) | Document indexing & retrieval storage | $0.10 / GB per day (1 GB free) |
| Tool Calling | Native tool orchestration requests | $2.50 / 1,000 calls |
| Agent Kit | File and image asset storage | $0.10 / GB-day (1 GB free per account per month) |
7. Model Fine-Tuning
| Model ID | Training Cost | Inference Input (1M tokens) | Inference Cached Input | Inference Output (1M tokens) |
|---|---|---|---|---|
o4-mini-2025-04-16 | $100.00 / hour | $4.00 | $1.00 | $16.00 |
o4-mini-2025-04-16 (Data Sharing) | $100.00 / hour | $2.00 | $0.50 | $8.00 |
o4-mini-2025-04-16 (Batch) | $100.00 / hour | $2.00 | $0.50 | $8.00 |
Reference Information & Inquiries
- Official Upstream Pricing: OpenAI Developer Pricing Documentation
- Status & Deprecations: OpenAI Model Deprecations Guide
- Platform Inquiries: Contact your Valstorm account team for custom volume commitments, dedicated enterprise routing, or hybrid private cloud deployments.