Pricing and Modelsv1.0.03 min read

OpenAI API Pricing & Model Reference

Official upstream pricing, reasoning tokens, realtime voice APIs, batch execution, and context caching rates for OpenAI models on Valstorm.

Official Provider Documentation: OpenAI API Pricing | OpenAI Models Overview


Overview & Gateway Architecture

This reference document outlines the upstream provider base pricing for OpenAI models routed through the Valstorm AI Platform.

Operating as an enterprise AI gateway, Valstorm provides unified access to these models alongside enterprise infrastructure:

  • Intelligent Routing & Failover: Dynamic routing across models and providers with zero-downtime fallback.
  • Enterprise Tool Execution: Sandboxed compute environments, web search grounding, vector retrieval (VFS), and database connectors.
  • Context Caching & Optimization: Transparent prompt caching and token minimization pipelines.
  • Multi-Tenant Security & Observability: Granular cost allocation, audit logs, and policy-driven rate limiting.

Note: Base provider rates shown below represent upstream pass-through pricing in USD per 1 Million (1M) tokens unless specified otherwise.


1. Flagship & Reasoning Models

All prices are in USD per 1M tokens (1,000,000 tokens).

Model IDShort Context InputShort Context Cached InputShort Context OutputLong Context InputLong Context Output
gpt-5.6-sol$5.00$0.50$30.00$10.00$45.00
gpt-5.6-terra$2.00$0.20$12.00$4.00$18.00
gpt-5.6-luna$0.20$0.02$1.20$0.40$1.80
gpt-5.5$5.00$0.50$30.00$10.00$45.00
gpt-5.4$2.50$0.25$15.00$5.00$22.50
gpt-5.4-mini$0.75$0.075$4.50
gpt-5.4-nano$0.20$0.02$1.25

2. Processing Modes & Execution Tiers

Batch & Flex Modes (50% Cost Reduction)

For non-blocking asynchronous jobs (Batch API) or interruptible compute workloads (Flex Tier), base rates are discounted by 50%:

Model IDBatch/Flex InputBatch/Flex Cached InputBatch/Flex Output
gpt-5.6-sol$2.50$0.25$15.00
gpt-5.6-terra$1.00$0.10$6.00
gpt-5.6-luna$0.10$0.01$0.60
gpt-5.5$2.50$0.25$15.00
gpt-5.4$1.25$0.13$7.50
gpt-5.4-mini$0.375$0.0375$2.25

3. Realtime Voice & Audio Models

Model IDModalityInput Rate (1M tokens)Cached Input RateOutput Rate (1M tokens)
gpt-realtime-2.1Audio$32.00$0.40$64.00
Text$4.00$0.40$24.00
gpt-realtime-2.1-miniAudio$10.00$0.30$20.00
Text$0.60$0.06$2.40
gpt-transcribeAsync Audio$0.0045 / minute
gpt-live-transcribeStreaming Audio$0.0170 / minute