Contacts
1207 Delaware Avenue, Suite 1228 Wilmington, DE 19806
Let's discuss your project
Business Address: 1207 Delaware Avenue, Suite 1228 Wilmington, DE 19806

Multimodal AI Statistics 2026: What Voice, Image and Video Really Cost Per Token

Multimodal AI statistics 2026 chart showing audio tokens cost 3.33 times text tokens Modality Premium Ratio comparing API prices for text, image, audio and video tokens

Multimodal AI Statistics 2026

By Axis Intelligence Research

Co-author: Sarah Mitchell | Last updated: September 25, 2026 | License: CC BY 4.0

Across 17 same-model price pairs published by OpenAI and Google, Axis Intelligence Research finds that a non-text token costs a median 2.67× the price of a text token. Audio input is the steepest routine surcharge at a median 3.33×, and the spread runs from parity (1.0×) to 16.67× depending on which model a team picks.


Quick Answer

Multimodal AI is now priced, not just promised. As of September 24, 2026, Axis Intelligence Research’s Modality Premium Ratio (MPR™) — the price of a voice, image or video token divided by the price of a text token on the same model — sits at a median 2.67×. Audio input costs a median 3.33× text; image input only 1.29×. Meanwhile Gartner forecasts that 80% of enterprise software will be multimodal by 2030, up from less than 10% in 2024.

Key Findings

  1. Axis Intelligence Research finds a median Modality Premium Ratio (MPR™) of 2.67× across 17 same-model text-versus-non-text API price pairs, as of September 24, 2026.
  2. Axis Intelligence Research finds audio input carries a median 3.33× premium over text input, ranging from 1.0× (Gemini 3.5 Flash-Lite) to 16.67× (gpt-realtime-2.1-mini).
  3. Axis Intelligence Research finds image input is the cheapest non-text modality, at a median 1.29× text across four model pairs from OpenAI and Google.
  4. Generated image tokens are the most expensive output in the dataset: Google prices Nano Banana 2 image output at 20× its text output, per its September 2026 pricing page.
  5. Gartner forecast in September 2024 that 40% of generative AI solutions will be multimodal by 2027, up from 1% in 2023.

What Is Multimodal AI, and Why Measure It by Price?

Multimodal AI is a model that takes in or produces more than one data type — text, images, audio, video — inside a single system. Gartner’s July 2025 definition covers inputs and outputs such as images, video, speech, text and numerical data within one generative model.

Most multimodal “statistics” pages recycle market-size projections from research resellers whose methods nobody can check. We took a different route. Vendors publish per-token prices by modality, and those price sheets are the most honest adoption signal in the industry: a provider only discounts audio when it has the inference capacity and the demand to fill it. So this page measures multimodal AI where the money changes hands — the API rate card — and pairs that with the few adoption figures that come from filings and earnings calls rather than surveys.

Sarah Mitchell: The keynote says “natively multimodal.” The rate card says audio is a different product at a different price. When those two stop disagreeing — when a voice token costs what a text token costs — the model really is one model to the people paying for it. Flash-Lite is already there. The realtime tier is nowhere close.

How Much Do Multimodal AI Models Cost Per Token in 2026?

Every price below was read directly from the vendor’s live pricing page on September 24, 2026: Google’s Gemini Developer API pricing and OpenAI’s API pricing. Standard tier, paid, per 1 million tokens, USD.

Realtime voice model pricing: text vs. audio vs. image

ModelText inputAudio inputImage inputText outputAudio outputSource
gpt-realtime-2.1$4.00$32.00$5.00$24.00$64.00OpenAI, Sep 2026
gpt-realtime-2.1-mini$0.60$10.00$0.80$2.40$20.00OpenAI, Sep 2026
gemini-3.8-live$0.75$3.00$1.00 (image/video)$4.50$12.00Google, Sep 2026

General-purpose Gemini models: does audio still cost extra?

ModelText / image / video inputAudio inputAudio premiumSource
gemini-2.5-flash$0.30$1.003.33×Google, Sep 2026
gemini-3-flash-preview$0.50$1.002.0×Google, Sep 2026
gemini-3.1-flash-lite$0.25$0.502.0×Google, Sep 2026
gemini-3.5-flash-lite$0.30 (all modalities)$0.301.0×Google, Sep 2026

The newest Flash-Lite generation lists a single $0.30 input rate covering text, image, video and audio. Two older Flash models still charge double for audio, and Gemini 2.5 Flash charges 3.33×. Within one vendor’s own catalog, audio parity arrived with the newest cheap model first.

Sarah Mitchell: Parity showing up at the bottom of the lineup is the tell. The lab isn’t cutting audio prices as a gesture; the small model’s inference path handles audio tokens cheaply enough that separate pricing stopped being worth the billing complexity. Expect the Flash tier to follow once its next version ships.

The Modality Premium Ratio (MPR™): Axis Intelligence Research’s Multimodal Price Index

Modality Premium Ratio (MPR™) = price per 1M non-text tokens ÷ price per 1M text tokens, same model, same direction (input vs. input, output vs. output), standard paid tier.

An MPR of 1.0× means the vendor treats the modality exactly like text. Above 1.0×, the modality carries a surcharge. The index uses only pairs where the vendor publishes both prices on the same model — no cross-vendor mixing, no estimated rates.

MPR™ by model and modality (as of September 24, 2026)

ModelModalityDirectionCalculationMPR™Source
gpt-realtime-2.1-miniAudioInput$10.00 ÷ $0.6016.67×OpenAI
gemini-3.1-flash-image (Nano Banana 2)ImageOutput$60.00 ÷ $3.0020.0×Google
gemini-3-pro-image (Nano Banana Pro)ImageOutput$120.00 ÷ $12.0010.0×Google
gpt-realtime-2.1-miniAudioOutput$20.00 ÷ $2.408.33×OpenAI
gpt-realtime-2.1AudioInput$32.00 ÷ $4.008.0×OpenAI
gemini-3.8-liveAudioInput$3.00 ÷ $0.754.0×Google
gemini-2.5-flashAudioInput$1.00 ÷ $0.303.33×Google
gpt-realtime-2.1AudioOutput$64.00 ÷ $24.002.67×OpenAI
gemini-3.8-liveAudioOutput$12.00 ÷ $4.502.67×Google
gemini-3-flash-previewAudioInput$1.00 ÷ $0.502.0×Google
gemini-3.1-flash-liteAudioInput$0.50 ÷ $0.252.0×Google
gemini-omni-1.1-flashVideoOutput$17.50 ÷ $9.001.94×Google
gpt-realtime-2.1-miniImageInput$0.80 ÷ $0.601.33×OpenAI
gemini-3.8-liveImage/videoInput$1.00 ÷ $0.751.33×Google
gpt-realtime-2.1ImageInput$5.00 ÷ $4.001.25×OpenAI
gemini-2.5-flashImage/videoInput$0.30 ÷ $0.301.0×Google
gemini-3.5-flash-liteAudioInput$0.30 ÷ $0.301.0×Google

What the MPR™ baseline reading says

  • Median MPR™, all 17 pairs: 2.67×. A typical multimodal token costs roughly two and two-thirds text tokens.
  • Median audio-input MPR™: 3.33× across seven pairs.
  • Median image-input MPR™: 1.29× across four pairs.
  • Widest single gap: gpt-realtime-2.1-mini audio input at 16.67×. The mini model’s text rate is cheap enough ($0.60) that the $10.00 audio rate becomes the dominant cost line.

This is the index’s baseline reading. The per-pair inputs, formula and all 17 values are in the free CSV, so anyone can recompute it.

Sarah Mitchell: Vision is effectively solved as a pricing problem — a screenshot costs about what the words describing it would. Voice is where the inference economics still bite, and the gap is widest on the “mini” realtime tier, which is exactly where startups building phone agents go first to save money. The cheap model is cheap for text. It is not cheap for the thing a voice agent spends its day doing.

How Much Does a 10-Minute AI Voice Call Cost?

Google publishes per-minute equivalents for Gemini 3.8 Live: $0.005 per minute of audio input and $0.018 per minute of audio output. OpenAI bills its GPT-Live 1 voice sessions at $0.05 per minute, with backend model and tool usage charged separately.

Axis Intelligence Research estimate (assuming the caller and the agent each talk for half of a 10-minute call):

  • Gemini 3.8 Live: (5 min × $0.005) + (5 min × $0.018) = $0.115 in audio tokens.
  • GPT-Live 1: 10 min × $0.05 = $0.50 in session fees, before backend model charges.

These are not like-for-like products — OpenAI’s session fee covers orchestration and still bills the model underneath — so we publish them side by side rather than as a ranking. What the math does show: a production voice agent handling thousands of calls a day is a line item measured in cents per call, which is why contact-center software is where multimodal revenue is landing first.

Speech-to-text and translation price per minute

ModelUse casePrice per minuteSource
gpt-transcribeFile transcription$0.0045OpenAI, Sep 2026
gpt-realtime-translateLive translation$0.034OpenAI, Sep 2026
gemini-3.8-liveAudio input$0.005Google, Sep 2026
gemini-3.8-liveAudio output$0.018Google, Sep 2026

How Much Does It Cost to Send an Image to an AI Model?

Image input is the one modality where pricing has largely converged with text. Anthropic’s documentation shows Claude counts an image in 28×28-pixel patches billed at the model’s ordinary input rate — a 1000×1000 image uses 1,296 visual tokens. Anthropic’s vision guide puts that at about $1.30 per thousand images on Claude Haiku 4.5 and $6.48 per thousand on Claude Opus 5.

ModelWhat is pricedCostSource
Claude Haiku 4.51,000 images at 1000×1000 px~$1.30Anthropic, Sep 2026
Claude Opus 51,000 images at 1000×1000 px~$6.48Anthropic, Sep 2026
gemini-3-pro-image1 input image$0.0011Google, Sep 2026

The practical read for teams running document or screenshot pipelines: the bill is driven by resolution and model tier, not by the fact that the input is an image. Downsampling before upload changes cost more than switching modality does.

How Much Does AI Image and Video Generation Cost Per Output?

Output is where multimodal gets expensive. Google prices generated image tokens at $60 per million on Nano Banana 2 versus $3 for its text output — the highest MPR™ in our dataset at 20×.

OutputPriceSource
1K image, Nano Banana 2 (gemini-3.1-flash-image)$0.067 per imageGoogle, Sep 2026
1K–2K image, Nano Banana Pro (gemini-3-pro-image)$0.134 per imageGoogle, Sep 2026
720p video, Gemini Omni Flash~$0.10 per secondGoogle, Sep 2026
gpt-image-2 image output$15.00 per 1M tokensOpenAI, Sep 2026

For deeper market data on these two generative categories, see our dedicated AI image generation statistics and AI video generation statistics pages; this page stays on the cross-modality pricing picture.

Sarah Mitchell: Notice that video output (1.94×) carries a smaller premium than image output (10–20×). That’s not video being cheap — Omni Flash’s text output is already priced at $9.00. The ratio is only as meaningful as the text rate it’s divided by, which is why we publish the raw prices next to every MPR™.

How Many People Use Multimodal AI in 2026?

Consumer-scale multimodal usage is best documented in Alphabet’s earnings commentary, the one place a major lab reports numbers under investor-disclosure obligations.

MetricValueAs ofSource
Gemini app monthly active users950 millionQ2 2026Alphabet Q2 2026 call
Ask YouTube (Gemini questions about videos) users140 million+June 2026Alphabet Q2 2026 call
Increase in daily users creating videos in Gemini app since Omni launch40%Q2 2026Alphabet Q2 2026 call
Gemini model API throughput~22 billion tokens per minute (up from 16 billion)Q2 2026Alphabet Q2 2026 call
Songs generated by Lyria 3 in the Gemini app150 million+Q1 2026Alphabet Q1 2026 call
Speech-to-text language coverage70 languagesQ1 2026Alphabet Q1 2026 call

Ask YouTube is the cleanest example of multimodal AI hiding inside an everyday product: a user asks a text or voice question, and the model answers from video content. At 140 million engaged users in a single month, it is larger than most standalone AI apps.

Multimodal AI hardware: AI glasses unit sales

Multimodal AI has a hardware form factor, and it sells. EssilorLuxottica reported that AI-glasses units — Ray-Ban Meta and Oakley Meta — were above 7 million in full-year 2025, per its Q4/FY 2025 results release. A camera and microphones on the face, feeding a model that answers out loud, is multimodal AI by definition.

What Share of Enterprise Software Will Be Multimodal?

Gartner has published two forecasts that bracket the enterprise trajectory:

ForecastBaselineTargetSource
Generative AI solutions that are multimodal1% (2023)40% by 2027Gartner, Sept 9, 2024
Enterprise software and applications that are multimodalLess than 10% (2024)80% by 2030Gartner, July 2, 2025

Axis Intelligence Research calculation: Gartner’s 2027 forecast implies a 40-fold rise from the 2023 baseline — an average of 9.75 percentage points per year over four years.

Sarah Mitchell: Forecasts like these are easy to repeat and hard to check, so read them with the price data beside them. If 80% of enterprise apps really go multimodal by 2030, audio has to get much cheaper than 3.33× text — nobody wires voice into every CRM screen at eight times the token rate. The MPR™ is the variable that tells you whether the Gartner line is on schedule.

Which Multimodal AI Model Is Cheapest for Each Use Case?

Using only the September 24, 2026 rate cards:

  • Cheapest audio input on a general model: gemini-3.5-flash-lite at $0.30 per 1M tokens, at text parity.
  • Cheapest per-minute transcription listed: OpenAI gpt-transcribe at $0.0045 per minute.
  • Lowest audio surcharge on a realtime voice model: gemini-3.8-live at 4.0× on input and 2.67× on output, versus 8.0× and 2.67× on gpt-realtime-2.1.
  • Cheapest per-image input listed: gemini-3-pro-image at $0.0011 per image.

Rankings here follow the published prices only; model quality, latency and tool support are separate decisions.

Methodology

Data collection. Axis Intelligence Research fetched and read every source on September 24, 2026. Prices come directly from vendor pricing pages (Google Gemini Developer API; OpenAI API) and Anthropic’s vision documentation. Adoption figures come from Alphabet earnings-call remarks (Q1 and Q2 2026), EssilorLuxottica’s FY 2025 results release, and two Gartner press releases. No figure on this page comes from a market-research reseller or from model memory.

Modality Premium Ratio (MPR™) formula. MPR™ = (price per 1M tokens of the non-text modality) ÷ (price per 1M tokens of text), computed on the same model, same direction (input/input or output/output), standard paid tier, short-context pricing. A pair enters the index only when the vendor publishes both prices for that model. Medians are computed across all qualifying pairs (n = 17), and within modality groups (audio input n = 7; image input n = 4). Arithmetic was verified in two independent passes: floating point and exact rational (Python fractions.Fraction).

Scope. The MPR™ covers list prices, not negotiated enterprise rates, batch discounts or caching. Token counts per second of audio or per image differ by vendor, so the MPR™ compares price per token, not price per minute of content; per-minute and per-image costs are published separately above. Anthropic’s Claude is not included in the MPR™ pairs because it bills image tokens at the standard input rate (an implicit 1.0×) and does not accept audio input natively. Gartner figures are forecasts and are labeled as such.

Voice call estimate. The 10-minute call figures assume a 50/50 speaking split and include audio tokens only (Gemini) or the session fee only (GPT-Live 1). They are Axis Intelligence Research estimates, not vendor quotes.

About This Dataset

What it contains: 84 rows covering 2026 multimodal AI API prices by modality, all 17 MPR™ pair calculations with formulas, per-minute and per-image costs, and verified adoption figures from Alphabet, EssilorLuxottica and Gartner. File: multimodal-ai-statistics.csv — every row carries source organization, document, URL, retrieval date, primary-source flag and, for calculated rows, the formula. License: CC BY 4.0. Use, share and adapt with attribution: “Axis Intelligence Research, Multimodal AI Statistics 2026.” Also available on: Hugging Face · Kaggle · GitHub (Axis Intelligence Research).

How to Cite This Page

APA: Axis Intelligence Research, & Mitchell, S. (2026, September 25). Multimodal AI statistics 2026: What voice, image and video really cost per token. Axis Intelligence. https://axis-intelligence.com/multimodal-ai-statistics/

MLA: Axis Intelligence Research, and Sarah Mitchell. “Multimodal AI Statistics 2026: What Voice, Image and Video Really Cost Per Token.” Axis Intelligence, 25 Sept. 2026, axis-intelligence.com/multimodal-ai-statistics/.

Chicago: Axis Intelligence Research, and Sarah Mitchell. “Multimodal AI Statistics 2026: What Voice, Image and Video Really Cost Per Token.” Axis Intelligence, September 25, 2026. https://axis-intelligence.com/multimodal-ai-statistics/.

Multimodal AI Questions Builders and Buyers Ask

Is an audio token more expensive than a text token?

Usually, yes. Axis Intelligence Research’s MPR™ puts the median audio-input surcharge at 3.33× text across seven model pairs (September 2026). The exception is Google’s gemini-3.5-flash-lite, which prices audio at text parity ($0.30 per 1M tokens).

Why do realtime voice models cost more than transcription?

Realtime models keep a live audio stream open and generate spoken replies with low latency. OpenAI lists gpt-realtime-2.1 audio input at $32 per 1M tokens, while its gpt-transcribe file transcription runs $0.0045 per minute — two different jobs priced as two different products.

Does sending screenshots to an LLM cost more than sending text?

Only slightly. Image input shows a median MPR™ of 1.29× in our dataset, and Anthropic bills Claude image tokens at the normal input rate. A 1000×1000 image uses 1,296 tokens on Claude — resolution drives cost more than modality does.

How much does one AI-generated image cost through the API in 2026?

Google lists $0.067 per 1K image on Nano Banana 2 and $0.134 per 1K–2K image on Nano Banana Pro, as of September 24, 2026. Image output tokens carry the highest MPR™ in our dataset: 20× text on Nano Banana 2.

What does a minute of AI-generated video cost?

Google’s Gemini Omni Flash bills 720p video output at approximately $0.10 per second, per its pricing page. At that rate, a 60-second clip is about $6 in output tokens before input and retries.

Will multimodal become the default for enterprise software?

Gartner forecasts 80% of enterprise software and applications will be multimodal by 2030, up from less than 10% in 2024. Treat it as a forecast; the falling audio premium on newer models is the measurable signal to watch.

Which is the most-used consumer multimodal AI feature right now?

Among publicly disclosed figures, Alphabet reported more than 140 million users engaged with Ask YouTube — Gemini answering questions about video — in June 2026. The Gemini app itself reached 950 million monthly active users in Q2 2026.

What is the Modality Premium Ratio (MPR™)?

The MPR™ is Axis Intelligence Research’s price index for multimodal AI: the price of a non-text token divided by a text token on the same model and direction. A reading of 1.0× means parity; the September 2026 baseline median is 2.67×.


Recent Posts

Smart Grid Statistics 2026: 128 Million Smart Meters, Only 14 in 100 Activated

Smart Grid Statistics 2026 By Axis Intelligence Research Co-author: Aidan Jad | Last updated: September 25, 2026 | Licen

Live Streaming Statistics 2026: Hours Watched, Platform Share & the Viewer Density Gap

Live Streaming Statistics 2026 By Axis Intelligence Research Co-author: Noah Mc | Last updated: September 25, 2026 | Lic

AI Data Center E-Waste Statistics 2026: The 87% Nobody Counted

AI Data Center E-Waste Statistics 2026 By Axis Intelligence Research Co-author: Aidan Jad | Last updated: September 25,

Axis Intelligence Research

Stay ahead on tech & data

Get notified when we publish or update datasets, trackers, research, and reports across technology, business, AI, cybersecurity, finance, infrastructure, energy, and more.

Research updates only. No spam. Unsubscribe anytime.