Contacts
1207 Delaware Avenue, Suite 1228 Wilmington, DE 19806
Let's discuss your project
Business Address: 1207 Delaware Avenue, Suite 1228 Wilmington, DE 19806

OpenAI API Pricing 2026: Every Model, Every Tier, and What a Real Workload Costs

OpenAI API pricing 2026 per million tokens for GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna OpenAI API cost by processing tier and workload, Axis Intelligence Research Effective Workload Rate

OpenAI API Pricing 2026

By Axis Intelligence Research

Co-author: Sarah Mitchell | Last updated: October 1, 2026 | License: CC BY 4.0

GPT-6 Astra, OpenAI’s most capable API model, lists at $10 per million input tokens and $50 per million output tokens at Standard processing, according to OpenAI’s developer pricing page (verified October 1, 2026). GPT-6.1 Sol lists at $2 / $10 and GPT-6 Luna at $0.10 / $0.50.


Quick Answer: How Much Does the OpenAI API Cost?

OpenAI API pricing in October 2026 runs from $0.10 to $10 per million input tokens and $0.50 to $50 per million output tokens across the current flagship family (GPT-6 Luna to GPT-6 Astra), at Standard processing for prompts under 272K tokens. Axis Intelligence Research’s Effective Workload Rate (EWR) puts a typical support-chatbot workload at $0.13, $2.47 and $12.53 per million tokens on Luna, 6.1 Sol and Astra respectively.

Key Findings

  1. GPT-6 Astra costs 100× more than GPT-6 Luna per million tokens of real chatbot traffic, according to Axis Intelligence Research’s EWR calculation from OpenAI’s October 1, 2026 rate card ($12.53 vs $0.13).
  2. Processing tier swings the bill 12×, on the same model. GPT-6 Astra output costs $25 per million tokens on Batch and $300 on Ultrafast, per OpenAI’s developer pricing page.
  3. GPT-6.1 Sol halved the cached-input rate of GPT-6 Sol within seven days — $0.20 at the September 22 launch, $0.10 at the September 29 launch, per OpenAI’s changelog.
  4. OpenAI’s two public price pages disagree. On October 1, 2026, openai.com/api/pricing still listed GPT-5.6 Sol at $5 / $30, while the developer docs listed $4 / $20 following the August 21 promotional cut.
  5. Caching is no longer free to set up. Since GPT-5.6, explicit cache writes bill at 1.25× the uncached input rate, per OpenAI’s prompt-caching guide — the first time OpenAI has charged for writing to the cache.

OpenAI API Pricing per 1M Tokens: The Current Flagship Models

Three models make up OpenAI’s current flagship line. The naming is astronomical — Astra (frontier), Sol (workhorse), Luna (volume) — and the price gaps between them are wide enough that picking the wrong one is the single largest cost decision most teams make.

GPT-6 Astra, GPT-6.1 Sol and GPT-6 Luna pricing (Standard tier)

ModelInputCached inputCache writeOutputContext bandVerifiedSource
GPT-6 Astra$10.00$1.00$12.50$50.00Short (<272K)2026-10-01OpenAI
GPT-6 Astra$20.00$2.00$25.00$75.00Long (>272K)2026-10-01OpenAI
GPT-6.1 Sol$2.00$0.10$2.50$10.00Short2026-10-01OpenAI
GPT-6.1 Sol$4.00$0.20$5.00$15.00Long2026-10-01OpenAI
GPT-6 Luna$0.10$0.01$0.125$0.50Short2026-10-01OpenAI
GPT-6 Luna$0.20$0.02$0.25$0.75Long2026-10-01OpenAI

Read the long-context rows carefully. Crossing the 272K-token line doubles input and cache rates but raises output by only 1.5×. For document-heavy work — contract review, codebase ingestion — that asymmetry means the long-context penalty lands almost entirely on the prompt, so trimming context is worth more than trimming answers.

One honest caveat on the threshold itself: OpenAI’s changelog describes short-context pricing as applying to prompts “up to 272K input tokens,” while the openai.com pricing page says “under 270K.” We use 272K, the figure in the developer documentation, and flag the gap.

GPT-6 Astra price tiers: Batch, Flex, Fast and Ultrafast

ModelTierInputCached inputOutput (short ctx)vs StandardSource
GPT-6 AstraBatch$5.00$0.50$25.000.5×OpenAI
GPT-6 AstraFlex$5.00$0.50$25.000.5×OpenAI
GPT-6 AstraStandard$10.00$1.00$50.001×OpenAI
GPT-6 AstraFast$20.00$2.00$100.002×OpenAI
GPT-6 AstraUltrafast$60.00$6.00$300.006×OpenAI
GPT-6.1 SolBatch / Flex$1.00$0.05$5.000.5×OpenAI
GPT-6.1 SolFast$4.00$0.20$20.002×OpenAI
GPT-6 LunaBatch / Flex$0.05$0.005$0.250.5×OpenAI
GPT-6 LunaFast$0.20$0.02$1.002×OpenAI

All figures from OpenAI’s developer pricing page, verified October 1, 2026. Ultrafast is listed for GPT-6 Astra only.

The tier ladder is now a clean set of multipliers — 0.5×, 1×, 2×, 6× — which makes it easier to reason about than the old Priority Processing scheme it replaced. The OpenAI Fast mode guide says Fast delivers up to 2.5× faster speeds than Standard. So on paper you pay 2× for up to 2.5× speed. Whether that trade holds for your traffic depends on your token mix and on whether latency is actually your bottleneck — a voice agent and a nightly classification job sit at opposite ends of that question.

Ultrafast is the outlier. At $300 per million output tokens, a GPT-6 Astra response of 2,000 tokens costs $0.60 before any input. That’s a product priced for latency-critical paths where a human is waiting on a long agentic run, not for bulk traffic.

The Effective Workload Rate (EWR): What OpenAI’s Rate Card Costs in Practice

List prices answer the wrong question. Nobody buys “input tokens” — teams buy a workload, and every workload mixes fresh input, cache reads, cache writes and output in its own proportions. Since GPT-5.6 introduced explicit cache writes at 1.25× the input rate, the gap between the sticker price and the real price has a new component that no rate card shows.

So Axis Intelligence Research built one number that does.

Effective Workload Rate (EWR): the blended cost, in US dollars, of 1 million tokens of a defined workload, weighting each OpenAI price component by its share of that workload’s traffic.

Formula: EWR = (s_input × P_input) + (s_cached × P_cached) + (s_write × P_write) + (s_output × P_output), where s is each token type’s share of total tokens (summing to 1) and P is OpenAI’s list price per million tokens.

The three workload profiles

ProfileUncached inputCached inputCache writeOutputTypical shape
Support chatbot40%40%5%15%Shared system prompt, short replies
RAG / document Q&A75%10%5%10%New retrieved chunks every call
Coding agent10%70%5%15%Long, repeated context across loop steps

These shares are Axis Intelligence Research’s modeling assumptions, not measured OpenAI data. They are disclosed so anyone can substitute their own shares into the formula; the dataset below lets you rerun every cell.

EWR by model and workload (Standard tier, short context)

ModelSupport chatbotRAG / document Q&ACoding agentSource
GPT-6 Astra$12.53$13.23$9.83Axis Intelligence Research (EWR)
GPT-6.1 Sol$2.47$2.64$1.90Axis Intelligence Research (EWR)
GPT-6 Luna$0.13$0.13$0.10Axis Intelligence Research (EWR)

Worked example, GPT-6 Astra support chatbot: (0.40 × $10) + (0.40 × $1) + (0.05 × $12.50) + (0.15 × $50) = $4.00 + $0.40 + $0.625 + $7.50 = $12.525 per million tokens.

Three things fall out of this table.

The coding agent is the cheapest workload per token on every model, despite being the most expensive in absolute spend. Agents re-read the same context on each loop step, and a 70% cache-hit share pulls Astra’s blended rate below its own $10 input sticker. That is the economic reason long agentic runs are viable at all at frontier prices.

RAG is the most expensive per token. Retrieval feeds the model new text on every call, so cache reads barely help. If your bill looks high relative to your traffic and you run retrieval, this is usually why.

And Astra’s chatbot EWR sits 25% above its own input price. Teams that budget by multiplying traffic by the input rate — a common shortcut in procurement spreadsheets — will undershoot on any output-heavy product.

What does 1 billion tokens a month cost on the OpenAI API?

ModelEWR (support chatbot)Monthly cost at 1B tokens3-year cost at 1B tokens/monthSource
GPT-6 Astra$12.525$12,525$450,900Axis Intelligence Research estimate
GPT-6.1 Sol$2.465$2,465$88,740Axis Intelligence Research estimate
GPT-6 Luna$0.12525$125.25$4,509Axis Intelligence Research estimate

Method: EWR × 1,000 (1 billion tokens is 1,000 million) × months. These are Axis Intelligence Research estimates holding October 1, 2026 list prices constant for 36 months — which they won’t be. OpenAI cut the price of three existing API models between July 30 and August 21, 2026 alone (see the change log below). Treat the three-year column as a budgeting ceiling at today’s rates, not a forecast.

The spread is what matters for procurement: the same billion tokens costs $125 or $12,525 a month depending on model choice. Routing — sending easy requests to Luna and escalating only hard ones to Sol or Astra — is where most of the savings in an OpenAI deployment live, well before any negotiation.

EWR by processing tier (support chatbot profile)

ModelBatch / FlexStandardFastUltrafastSource
GPT-6 Astra$6.26$12.53$25.05$75.15Axis Intelligence Research (EWR)
GPT-6.1 Sol$1.23$2.47$4.93—Axis Intelligence Research (EWR)
GPT-6 Luna$0.06$0.13$0.25—Axis Intelligence Research (EWR)

A notable crossover: GPT-6 Astra on Batch ($6.26) costs more than GPT-6.1 Sol on Fast ($4.93) for the same workload. Paying for speed on a cheaper model can undercut paying for intelligence on a slower tier — the right call depends on whether the task needs Astra at all.

OpenAI API Price Change Log: Every 2026 Change, Dated

OpenAI moves prices often, and rarely with a press release. This log is assembled from OpenAI’s API changelog and is the record we update on every refresh.

DateModel / productChangeSource
2026-09-29GPT-6.1 SolLaunched at $2 input / $0.10 cached / $2.50 cache write / $10 outputOpenAI changelog
2026-09-29GPT-6 AstraUltrafast mode added in the Responses APIOpenAI changelog
2026-09-22GPT-6 Sol, GPT-6 LunaLaunched: Sol $2 / $0.20 cached / $10; Luna $0.10 / $0.01 / $0.50OpenAI changelog
2026-09-10GPT-Live 1Generally available at $0.05 per minute, billed per secondOpenAI changelog
2026-09-08GPT Image 2.5 Sunburst / FlareReleased at GPT Image 2 token ratesOpenAI changelog
2026-09-03GPT-6 AstraReleasedOpenAI changelog
2026-08-21GPT-5.6 SolCut to $4 input / $20 output (−20% / −33%), promotional through at least Nov 21, 2026OpenAI changelog
2026-08-13Ultrafast modeAnnounced for GPT-5.6 Sol, “up to 14x faster” than Standard, limited previewOpenAI changelog
2026-08-05GPT-5.6 familyFast mode extended to prompts over 272K tokensOpenAI changelog
2026-07-30GPT-5.6 Luna / TerraLuna −80%, Terra −20%; Fast mode replaces Priority ProcessingOpenAI changelog
2026-07-09GPT-5.6 Sol / Terra / LunaModel family released; explicit prompt caching controls introducedOpenAI changelog
2026-06-02ContainersBilling moved to per-minute with a 5-minute minimumOpenAI changelog
2026-05-29Prompt caching24h retention became the default for non-ZDR organizationsOpenAI changelog

The pattern is worth naming. OpenAI’s 2026 price moves have been cuts on existing models and new models launching at the old price points. GPT-6 Sol and GPT-6.1 Sol launched a week apart at identical input and output rates; only the cached-input rate moved, from $0.20 to $0.10. For an agent that spends 70% of its tokens on cache reads, that single line item is the whole story — and the headline rate didn’t change.

Upcoming dates that will move your bill

  • October 5, 2026 — billing begins for GPT-Rosalind (gpt-rosalind-research) at $5 / $0.50 / $25, per OpenAI’s pricing page.
  • November 21, 2026 — earliest end date for GPT-5.6 Sol’s promotional $4 / $20 rate, per OpenAI’s changelog.
  • February 26, 2027 — shutdown of whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize, per OpenAI’s changelog.

Why Do OpenAI’s Two Pricing Pages Show Different Prices?

On October 1, 2026, openai.com/api/pricing headlined the GPT-5.6 family — Sol at $5 / $30, Terra at $2 / $12, Luna at $0.20 / $1.20 — and did not list GPT-6 Astra, GPT-6.1 Sol or GPT-6 Luna. The developer pricing page, at developers.openai.com/api/docs/pricing, led with the GPT-6 family and listed GPT-5.6 Sol at $4 / $20 in its Cyber models table.

The developer page matches OpenAI’s August 21 changelog entry for GPT-5.6 Sol. The marketing page does not. We could not determine from either page whether the GPT-5.6 Terra and Luna figures shown on openai.com reflect the July 30 cuts, so we report them as displayed rather than reconciling them.

The practical rule: budget from developers.openai.com/api/docs/pricing, and verify the per-model page or your Usage dashboard before committing. Your invoice reflects what the billing system charges, not what a marketing page last rendered.

GPT-5.6 Sol, StandardInputCached inputOutputSource
openai.com/api/pricing$5.00$0.50$30.00openai.com
developers.openai.com pricing$4.00$0.40$20.00OpenAI developer docs

How Much Do OpenAI’s Voice, Image and Transcription APIs Cost?

Realtime voice models

ModelModalityInputCached inputOutputSource
GPT-Realtime-2.1Audio$32.00$0.40$64.00OpenAI
GPT-Realtime-2.1Text$4.00$0.40$24.00OpenAI
GPT-Realtime-2.1 miniAudio$10.00$0.30$20.00OpenAI
GPT-Realtime-2.1 miniText$0.60$0.06$2.40OpenAI

Per-minute audio pricing

ModelUsePrice per minuteSource
GPT-Live 1Full-duplex voice session (backend model billed separately)$0.05OpenAI
GPT-Realtime-TranslateLive speech translation$0.034OpenAI
GPT-Live-TranscribeStreaming transcription$0.017OpenAI
GPT-Realtime-WhisperStreaming transcription$0.017OpenAI
GPT-4o-transcribeFile transcription (deprecated Feb 2027)$0.006OpenAI
GPT-TranscribeFile transcription$0.0045OpenAI
GPT-4o-mini-transcribeFile transcription (deprecated Feb 2027)$0.003OpenAI

GPT-Live’s $0.05 per minute is the voice layer only. The reasoning model behind it bills at its own token rate, so a voice agent’s real cost is the session fee plus a full EWR on top.

Image generation

GPT Image 2.5 Sunburst and GPT Image 2.5 Flare both list at $8 per million image-input tokens, $2 cached and $30 per million image-output tokens, with text input at $5, per OpenAI’s developer pricing page. On Batch, GPT Image 2 drops to $4 / $1 / $15.

What Do OpenAI’s Built-In Tools Cost?

ToolPriceNotesSource
Web search$10.00 per 1,000 callsPlus search content tokens at model ratesOpenAI
Web search preview (non-reasoning models)$25.00 per 1,000 callsSearch content tokens freeOpenAI
File search — tool call$2.50 per 1,000 callsResponses API onlyOpenAI
File search — storage$0.10 per GB per dayFirst 1 GB freeOpenAI
Containers (Hosted Shell, Code Interpreter)$0.03 (1 GB) to $1.92 (64 GB) per 20-minute sessionBilled by the minute, 5-minute minimumOpenAI
Regional processing / data residency+10%Models released on or after March 5, 2026; FedRAMP endpoints also +10%OpenAI

Web search is the line item agent builders underestimate. An agent that searches five times per task spends $0.05 on calls before a single answer token — a figure that dwarfs the token cost of the whole task on GPT-6 Luna.

What Moves Your OpenAI API Bill (Beyond the Rate Card)

Output share. Output costs 5× input on every current flagship model. The fastest way to cut spend is shorter, structured answers, not cheaper prompts.

Cache hit rate — and now cache writes. Since GPT-5.6, writes on explicit cache breakpoints cost 1.25× input, and reads cost 0.1× (0.05× on GPT-6.1 Sol), per OpenAI’s prompt-caching guide. A breakpoint that gets written but rarely read is a net loss. Prefixes must be at least 1,024 tokens to cache on GPT-5.6 and later.

The 272K line. Crossing it doubles input pricing. Compaction or retrieval that keeps prompts under the threshold pays for itself on long-context workloads.

Processing tier. Batch and Flex halve the bill for anything that can wait; Fast doubles it; Ultrafast multiplies it by six.

Hidden reasoning tokens. Reasoning models bill their internal reasoning as output tokens. Higher reasoning_effort settings raise output share — and output is the expensive side. Our AI inference cost analysis covers why per-token price cuts often fail to lower total bills.

Data residency. Regional endpoints add 10% for models released after March 5, 2026 — which is every current flagship.

Which OpenAI Model Should You Pay For?

Analysis by Sarah Mitchell, AI analyst, Axis Intelligence Research.

The rate card makes GPT-6 Astra look like the default “best” choice and Luna like a toy. The EWR table says something more useful: Astra is 100× Luna’s price on chatbot traffic, so Astra has to be meaningfully better on your task — not on a launch-day benchmark — to earn its slot. The benchmark said state of the art at every launch this year; the interesting number is how often the cheaper model passes your own eval on the fifth try.

A defensible default for most production teams in October 2026:

  • GPT-6 Luna for classification, extraction, routing and high-volume chat where answers are short. At roughly $125 per billion chatbot tokens, it is close to a rounding error.
  • GPT-6.1 Sol for coding, agents and professional writing. Its $0.10 cached-input rate makes it the strongest value for cache-heavy agent loops in the current lineup.
  • GPT-6 Astra for the hard tail — multi-step tasks where a Sol failure costs more than the 5× price gap.
  • Batch for anything with no user waiting. It halves every model’s cost.

Run your own eval before switching. A price table tells you what a token costs; only your test set tells you how many tokens a correct answer costs.

OpenAI API Pricing FAQ

Is GPT-6.1 Sol cheaper than GPT-6 Sol?

On cached input, yes. GPT-6.1 Sol launched September 29, 2026 at $0.10 per million cached input tokens, half GPT-6 Sol’s $0.20 from September 22, per OpenAI’s changelog. Input ($2) and output ($10) are identical. For cache-heavy agents, 6.1 Sol is the cheaper model.

Does ChatGPT Plus or Business include API credits?

No. OpenAI states on its API pricing page that API usage is billed separately from ChatGPT Plus, Business, Enterprise and Edu subscriptions.

What is the cheapest way to call GPT-6 Astra?

Batch or Flex processing, both at $5 input / $25 output per million tokens — half the Standard rate — per OpenAI’s developer pricing page. Batch runs asynchronously within 24 hours; Flex trades lower cost for slower responses and occasional resource unavailability.

Is Fast mode worth twice the price?

OpenAI prices Fast mode at 2× Standard and claims up to 2.5× faster speeds. It pays off when latency directly affects revenue or user retention — voice, live coding assistance — and rarely for background jobs.

Does prompt caching still save money if cache writes cost extra?

Yes, if the cached prefix is read more than once. On GPT-6 Astra a write costs $12.50 per million tokens and each read $1, versus $10 uncached. One write plus one read ($13.50) already beats two uncached passes ($20).

Do I pay for Playground usage?

Yes. OpenAI bills Playground calls at the same per-token rates as API calls, according to its pricing page FAQ.

Do OpenAI models cost the same on Azure or Amazon Bedrock?

Not necessarily. OpenAI’s pricing page states that its models on Amazon Bedrock and Microsoft Azure are billed through those services, under their own rate cards.

How do I cap my OpenAI API spend?

Set a monthly hard spend limit at the organization or project level. Since July 22, 2026, requests return a 429 error once tracked spend reaches the cap, per OpenAI’s changelog. Spend alerts can warn you before traffic stops.

Methodology

Price collection. Every list price on this page was read from OpenAI’s developer pricing page and openai.com/api/pricing on October 1, 2026. Price-change dates and percentages come from OpenAI’s API changelog. Cache-write and cache-read multipliers come from OpenAI’s prompt-caching guide; Fast mode speed claims from the Fast mode guide. No third-party price tracker was used as a source.

Conflict rule. Where the two OpenAI pages disagree, the developer documentation is treated as authoritative because it matches the dated changelog; both values are published.

Effective Workload Rate (EWR). EWR = Σ (token-type share × list price). Workload shares (support chatbot 40/40/5/15; RAG 75/10/5/10; coding agent 10/70/5/15 across uncached input / cached input / cache write / output) are Axis Intelligence Research modeling assumptions. Monthly and three-year figures multiply EWR by 1,000 per billion tokens and hold October 1, 2026 prices constant. EWR covers token charges only — tool calls, containers, storage and data-residency uplifts are excluded and should be added separately.

About This Dataset

The full dataset behind this page — 223 rows covering list prices by model, tier, context band and token type, per-minute audio prices, tool prices, dated price changes and every EWR and cost figure — is available as openai-api-pricing.csv under a CC BY 4.0 license. Every row carries its source URL, retrieval date and a flag for Axis-calculated values.

Cite this page:

  • APA: Axis Intelligence Research. (2026, October 1). OpenAI API pricing 2026: Every model, every tier, and what a real workload costs. Axis Intelligence. https://axis-intelligence.com/openai-api-pricing/
  • MLA: Axis Intelligence Research. “OpenAI API Pricing 2026: Every Model, Every Tier, and What a Real Workload Costs.” Axis Intelligence, 1 Oct. 2026, axis-intelligence.com/openai-api-pricing/.
  • Chicago: Axis Intelligence Research. “OpenAI API Pricing 2026: Every Model, Every Tier, and What a Real Workload Costs.” Axis Intelligence, October 1, 2026. https://axis-intelligence.com/openai-api-pricing/.

Related Axis research: AI inference cost statistics · AI model release tracker · DeepSeek statistics · Google Cloud statistics · Agentic commerce statistics

Recent Posts

Starlink Cost in 2026: Every Plan, Hardware Price and the True Monthly Cost Over 3 Years

Starlink Cost in 2026 By Axis Intelligence Research Co-author: Alex Rivera | Last updated: October 1, 2026 | License: CC

How Much Does It Cost to Install Solar Panels in 2026? Price per Watt by State, From 450,000 Real Installs

Install Solar Panels Cost 2026 By Axis Intelligence Research Co-author: Aidan Jad | Last updated: October 1 | License: C

SoundCloud Statistics 2026: Users, Revenue, Creators — and Which Numbers Actually Hold Up

SoundCloud Statistics 2026 By Axis Intelligence Research Co-author: Alex Rivera | Last updated: October 1, 2026 | Licens

Axis Intelligence Research

Stay ahead on tech & data

Get notified when we publish or update datasets, trackers, research, and reports across technology, business, AI, cybersecurity, finance, infrastructure, energy, and more.

Research updates only. No spam. Unsubscribe anytime.