Pricing6 min read

Anthropic Claude API Pricing & Model Reference

Comprehensive reference for Anthropic Claude API pricing, token caching rates, model availability, and enterprise features on Valstorm.

Official Provider Documentation: Anthropic Pricing | Anthropic Claude Models Overview | Anthropic Rate Limits & Billing


Overview & Gateway Architecture

This reference document details the upstream provider base pricing for the Anthropic Claude model family.

Operating as an enterprise AI gateway and developer enablement platform (similar to OpenRouter or OpenCode), the Valstorm AI Platform provides unified access to Claude models alongside comprehensive enterprise features:

  • Intelligent Routing & Prompt Caching: Automatic cache breakpoint management (5m/1h writes) and context optimization.
  • Agentic Runtimes & Tooling: Stateful execution environments, computer use orchestration, automated web search, and sandboxed code interpreters.
  • Enterprise Governance: Granular multi-tenant metering, policy-based spend caps, audit logging, and single sign-on.
  • Resilient Fallback Pipelines: Multi-region and multi-provider failover without client code changes.

Note: Base provider rates shown below represent upstream pass-through pricing in USD per 1 Million (1M) tokens / MTok unless specified otherwise. Platform usage includes upstream consumption plus applicable platform value-add markup.


1. Core Model Pricing

All rates are listed in USD per Million Tokens (MTok).

Model NameModel IDBase Input5-Minute Cache Write (1.25x)1-Hour Cache Write (2.0x)Cache Hit / Refresh (0.1x)Base Output
Claude Fable 5claude-5-fable$10.00$12.50$20.00$1.00$50.00
Claude Mythos 5 (Limited Availability)claude-5-mythos$10.00$12.50$20.00$1.00$50.00
Claude Opus 5claude-5-opus$5.00$6.25$10.00$0.50$25.00
Claude Opus 4.8claude-4-8-opus$5.00$6.25$10.00$0.50$25.00
Claude Opus 4.7claude-4-7-opus$5.00$6.25$10.00$0.50$25.00
Claude Opus 4.6claude-4-6-opus$5.00$6.25$10.00$0.50$25.00
Claude Opus 4.5claude-4-5-opus$5.00$6.25$10.00$0.50$25.00
Claude Sonnet 5claude-5-sonnet$3.00$3.75$6.00$0.30$15.00
Claude Sonnet 4.6claude-4-6-sonnet$3.00$3.75$6.00$0.30$15.00
Claude Sonnet 4.5claude-4-5-sonnet$3.00$3.75$6.00$0.30$15.00
Claude Haiku 4.5claude-4-5-haiku$1.00$1.25$2.00$0.10$5.00
Claude Opus 4.1 (Deprecated)claude-4-1-opus$15.00$18.75$30.00$1.50$75.00
Claude Haiku 3.5 (Retired)claude-3-5-haiku$0.80$1.00$1.60$0.08$4.00

Note: Models from Claude 4.7 onwards utilize a high-efficiency tokenizer that produces approximately 30% more tokens for identical raw text payloads.


2. Prompt Caching Architecture

Prompt caching drastically reduces latency and cost for long-context prompts, recurring system prompts, and multi-turn agent conversations.

Cache OperationMultiplier Relative to Base InputCache Time-to-Live (TTL)Cost Efficiency Break-Even
5-Minute Cache Write1.25x base input rate5 minutesPays off after 1 cache hit (90% savings per read)
1-Hour Cache Write2.00x base input rate1 hourPays off after 2 cache hits
Cache Read (Hit)0.10x base input rate (90% discount)Matches cache lifetimeContinuous 90% discount on recurring context

3. Execution Tiers & Modifiers

Batch Processing API (50% Discount)

Non-latency-critical background jobs processed asynchronously receive a 50% discount across base input, base output, and cache write tokens.

Data Residency & Regional Pinning

For Claude 4.6 and newer models:

  • inference_geo: "global" (Default): Standard base pricing.
  • inference_geo: "us": Applies a 1.1x multiplier (+10%) across all token operations (inputs, outputs, and cache writes).

Cloud Platform Marketplaces

  • Claude Platform on AWS: Metered in Claude Consumption Units (CCU) at $0.01 per CCU (100 CCU = $1.00 USD).
  • Amazon Bedrock & Google Cloud Vertex AI: Regional / multi-region pinned endpoints include a 10% premium over standard global routing.

4. Built-in Tools & Agentic Features

Tool / CapabilityPricing ModelUsage Details & Metering
Web Search Tool$10.00 / 1,000 searchesSearch result tokens retrieved into the conversation context are billed as standard input tokens.
Web Fetch ToolFree (no tool fee)Standard input token rates apply to retrieved web page content.
Code Execution ToolFree with Search/Fetch
• Standalone: 1,550 free hrs/mo, then $0.05 / hr
Minimum 5 minutes execution time per container session. Free when bundled with Web Search / Web Fetch.
Computer Use (Beta)Standard Token RatesAdds ~466–499 system prompt tokens + 735 tool definition tokens. Screenshots billed at standard vision token rates.

5. Claude Managed Agents Runtime

Claude Managed Agents combine model inference with managed stateful compute environments:

Billing DimensionSKU RateMetering Rules
Token ConsumptionStandard Model Token RatesInput, output, reasoning, and prompt cache multipliers apply identically to base model rates.
Session Runtime$0.08 per session-hourMetered to the millisecond strictly while status is running. No runtime charges during idle (waiting for user input) or terminated states. Replaces container-hour charges.

Worked Example: 1-Hour Interactive Session

An interactive coding session using Claude Opus 5 consuming 50,000 input tokens (with 40,000 cached reads) and 15,000 output tokens:

  • 10,000 Uncached Input Tokens: $0.050
  • 40,000 Cache Read Tokens: $0.020
  • 15,000 Output Tokens: $0.375
  • 1.0 Hour Session Runtime: $0.080
  • Total Base Provider Cost: $0.525

Reference Information & Inquiries