Anthropic Claude API Pricing & Model Reference
Comprehensive reference for Anthropic Claude API pricing, token caching rates, model availability, and enterprise features on Valstorm.
Official Provider Documentation: Anthropic Pricing | Anthropic Claude Models Overview | Anthropic Rate Limits & Billing
Overview & Gateway Architecture
This reference document details the upstream provider base pricing for the Anthropic Claude model family.
Operating as an enterprise AI gateway and developer enablement platform (similar to OpenRouter or OpenCode), the Valstorm AI Platform provides unified access to Claude models alongside comprehensive enterprise features:
- Intelligent Routing & Prompt Caching: Automatic cache breakpoint management (5m/1h writes) and context optimization.
- Agentic Runtimes & Tooling: Stateful execution environments, computer use orchestration, automated web search, and sandboxed code interpreters.
- Enterprise Governance: Granular multi-tenant metering, policy-based spend caps, audit logging, and single sign-on.
- Resilient Fallback Pipelines: Multi-region and multi-provider failover without client code changes.
Note: Base provider rates shown below represent upstream pass-through pricing in USD per 1 Million (1M) tokens / MTok unless specified otherwise. Platform usage includes upstream consumption plus applicable platform value-add markup.
1. Core Model Pricing
All rates are listed in USD per Million Tokens (MTok).
| Model Name | Model ID | Base Input | 5-Minute Cache Write (1.25x) | 1-Hour Cache Write (2.0x) | Cache Hit / Refresh (0.1x) | Base Output |
|---|---|---|---|---|---|---|
| Claude Fable 5 | claude-5-fable | $10.00 | $12.50 | $20.00 | $1.00 | $50.00 |
| Claude Mythos 5 (Limited Availability) | claude-5-mythos | $10.00 | $12.50 | $20.00 | $1.00 | $50.00 |
| Claude Opus 5 | claude-5-opus | $5.00 | $6.25 | $10.00 | $0.50 | $25.00 |
| Claude Opus 4.8 | claude-4-8-opus | $5.00 | $6.25 | $10.00 | $0.50 | $25.00 |
| Claude Opus 4.7 | claude-4-7-opus | $5.00 | $6.25 | $10.00 | $0.50 | $25.00 |
| Claude Opus 4.6 | claude-4-6-opus | $5.00 | $6.25 | $10.00 | $0.50 | $25.00 |
| Claude Opus 4.5 | claude-4-5-opus | $5.00 | $6.25 | $10.00 | $0.50 | $25.00 |
| Claude Sonnet 5 | claude-5-sonnet | $3.00 | $3.75 | $6.00 | $0.30 | $15.00 |
| Claude Sonnet 4.6 | claude-4-6-sonnet | $3.00 | $3.75 | $6.00 | $0.30 | $15.00 |
| Claude Sonnet 4.5 | claude-4-5-sonnet | $3.00 | $3.75 | $6.00 | $0.30 | $15.00 |
| Claude Haiku 4.5 | claude-4-5-haiku | $1.00 | $1.25 | $2.00 | $0.10 | $5.00 |
| Claude Opus 4.1 (Deprecated) | claude-4-1-opus | $15.00 | $18.75 | $30.00 | $1.50 | $75.00 |
| Claude Haiku 3.5 (Retired) | claude-3-5-haiku | $0.80 | $1.00 | $1.60 | $0.08 | $4.00 |
Note: Models from Claude 4.7 onwards utilize a high-efficiency tokenizer that produces approximately 30% more tokens for identical raw text payloads.
2. Prompt Caching Architecture
Prompt caching drastically reduces latency and cost for long-context prompts, recurring system prompts, and multi-turn agent conversations.
| Cache Operation | Multiplier Relative to Base Input | Cache Time-to-Live (TTL) | Cost Efficiency Break-Even |
|---|---|---|---|
| 5-Minute Cache Write | 1.25x base input rate | 5 minutes | Pays off after 1 cache hit (90% savings per read) |
| 1-Hour Cache Write | 2.00x base input rate | 1 hour | Pays off after 2 cache hits |
| Cache Read (Hit) | 0.10x base input rate (90% discount) | Matches cache lifetime | Continuous 90% discount on recurring context |
3. Execution Tiers & Modifiers
Batch Processing API (50% Discount)
Non-latency-critical background jobs processed asynchronously receive a 50% discount across base input, base output, and cache write tokens.
Data Residency & Regional Pinning
For Claude 4.6 and newer models:
inference_geo: "global"(Default): Standard base pricing.inference_geo: "us": Applies a 1.1x multiplier (+10%) across all token operations (inputs, outputs, and cache writes).
Cloud Platform Marketplaces
- Claude Platform on AWS: Metered in Claude Consumption Units (CCU) at $0.01 per CCU (100 CCU = $1.00 USD).
- Amazon Bedrock & Google Cloud Vertex AI: Regional / multi-region pinned endpoints include a 10% premium over standard global routing.
4. Built-in Tools & Agentic Features
| Tool / Capability | Pricing Model | Usage Details & Metering |
|---|---|---|
| Web Search Tool | $10.00 / 1,000 searches | Search result tokens retrieved into the conversation context are billed as standard input tokens. |
| Web Fetch Tool | Free (no tool fee) | Standard input token rates apply to retrieved web page content. |
| Code Execution Tool | Free with Search/Fetch • Standalone: 1,550 free hrs/mo, then $0.05 / hr | Minimum 5 minutes execution time per container session. Free when bundled with Web Search / Web Fetch. |
| Computer Use (Beta) | Standard Token Rates | Adds ~466–499 system prompt tokens + 735 tool definition tokens. Screenshots billed at standard vision token rates. |
5. Claude Managed Agents Runtime
Claude Managed Agents combine model inference with managed stateful compute environments:
| Billing Dimension | SKU Rate | Metering Rules |
|---|---|---|
| Token Consumption | Standard Model Token Rates | Input, output, reasoning, and prompt cache multipliers apply identically to base model rates. |
| Session Runtime | $0.08 per session-hour | Metered to the millisecond strictly while status is running. No runtime charges during idle (waiting for user input) or terminated states. Replaces container-hour charges. |
Worked Example: 1-Hour Interactive Session
An interactive coding session using Claude Opus 5 consuming 50,000 input tokens (with 40,000 cached reads) and 15,000 output tokens:
- 10,000 Uncached Input Tokens: $0.050
- 40,000 Cache Read Tokens: $0.020
- 15,000 Output Tokens: $0.375
- 1.0 Hour Session Runtime: $0.080
- Total Base Provider Cost: $0.525
Reference Information & Inquiries
- Official Upstream Pricing: Anthropic Claude Pricing Documentation
- Model Overview & Capabilities: Anthropic Documentation: About Claude Models
- Platform Inquiries: Contact your Valstorm account team for custom routing policies, provisioned capacity, or high-volume agent arrangements.