Contacts
1207 Delaware Avenue, Suite 1228 Wilmington, DE 19806
Let's discuss your project
Business Address: 1207 Delaware Avenue, Suite 1228 Wilmington, DE 19806

AI Voice Generator Statistics 2026: What One Hour of Synthetic Speech Costs, Who Is Earning From It, and the Rules Now in Force

AI voice generator statistics 2026 chart showing cost per audio hour by vendor Voice-Hour Cost Index comparing ElevenLabs, Google Chirp 3 HD and Amazon Polly text to speech pricing

AI Voice Generator Statistics 2026

By Axis Intelligence Research

Co-author: Sarah Mitchell | Last updated: September 24, 2026 | License: CC BY 4.0

According to Axis Intelligence Research, the median cost of one finished hour of current-generation AI voice audio is $1.30 across five expressive tiers from Amazon, Google and ElevenLabs, as of September 24, 2026. The cheapest current tier produces an hour for $0.90; the priciest costs $4.32, a 4.8x spread for the same hour of speech.


Quick Answer

An hour of AI-generated speech now costs between $0.17 and $6.92 at list price, depending on the voice tier, with the Axis Voice-Hour Cost Index (VHCI) at $1.30 per audio hour. Voicing a full-length novel of about 600,000 characters costs $18 on Amazon Polly Generative or Google Chirp 3 HD and $60 on ElevenLabs v3. ElevenLabs, the category’s largest independent vendor, passed $500 million in annual recurring revenue in the first four months of 2026.

Key Findings

  1. According to Axis Intelligence Research, the Voice-Hour Cost Index (VHCI) stood at $1.30 per finished audio hour on September 24, 2026, the median of five current-generation expressive voice tiers.
  2. ElevenLabs surpassed $500 million in annual recurring revenue within the first four months of 2026 and was valued at $11 billion in its February 2026 Series D, per the company.
  3. According to Axis Intelligence Research, vendors disagree by 38.8% on how much speech one million characters produces: AWS implies 43,228 characters per audio hour, ElevenLabs implies 60,000.
  4. ElevenLabs creators on its voice marketplace had earned more than $22 million by May 2026, doubling from $11 million in November 2025, across 10,400+ earning creators, per ElevenLabs.
  5. Consumer Reports found in March 2025 that four of six voice cloning products tested relied only on a self-attestation checkbox to confirm consent from the person being cloned.

How Much Does an AI Voice Generator Cost Per Hour of Audio?

Every major vendor bills text-to-speech by the character or by the audio token. Nobody buys characters, though. Producers, developers and localization teams buy finished minutes. So Axis Intelligence Research converted every published rate card into the unit buyers actually think in: dollars per finished hour of speech.

The conversion factor comes from Amazon itself. The Amazon Polly pricing page states that one million characters produces roughly 23 hours and 8 minutes of speech. That works out to 43,228 characters per audio hour, and Axis applies it to every character-billed tier below. Google’s newest Gemini TTS models bill by audio token instead, at 25 tokens per second of output according to the Google Cloud Text-to-Speech pricing page, which lets us compute their hourly cost directly with no conversion at all.

AI voice generator cost per audio hour, by vendor and tier (September 2026)

VendorVoice tierList priceCost per finished audio hourSource
Google CloudStandard$4 per 1M characters$0.17Google Cloud pricing page
Amazon PollyStandard$4 per 1M characters$0.17AWS pricing page
Google CloudNeural2$16 per 1M characters$0.69Google Cloud pricing page
Amazon PollyNeural$16 per 1M characters$0.69AWS pricing page
Google CloudGemini 2.5 Flash TTS$10 per 1M audio tokens$0.90Google Cloud pricing page
Google CloudChirp 3 HD$30 per 1M characters$1.30Google Cloud pricing page
Amazon PollyGenerative$30 per 1M characters$1.30AWS pricing page
Google CloudGemini 2.5 Pro TTS$20 per 1M audio tokens$1.80Google Cloud pricing page
ElevenLabsFlash/Turbo, v3 Conversational$0.05 per 1K characters$2.16ElevenLabs pricing page
Google CloudInstant custom voice$60 per 1M characters$2.59Google Cloud pricing page
Amazon PollyLong-Form$100 per 1M characters$4.32AWS pricing page
ElevenLabsv3, v2 Multilingual$0.10 per 1K characters$4.32ElevenLabs pricing page
Google CloudStudio$160 per 1M characters$6.92Google Cloud pricing page

Hourly costs are Axis Intelligence Research calculations: list price per million characters x 43,228 / 1,000,000. Gemini rows: price per million audio tokens x 25 tokens per second x 3,600 seconds, excluding the small text-input token charge. Retrieved September 24, 2026.

The table reads like a price ladder with a gap in the middle. Legacy voices sit at 17 cents an hour, a figure that has not moved in years because nobody is competing for that business anymore. The action is in the $0.90 to $2.16 band, where Google, Amazon and ElevenLabs’s low-latency models now overlap. According to Axis Intelligence Research, the premium ElevenLabs v3 tier costs 25 times more per character than the Standard voices at Google and Amazon.

Sarah Mitchell, analyst comment: The price gap is no longer a quality gap you can hear in a demo. It’s a gap in controllability: emotional direction, multi-speaker dialogue, long-form stability. Those are the things that break on the fifth generation, not the first. A $0.90 hour that needs three regenerations costs more than a $2.16 hour that lands on the first pass, and none of these rate cards tell you the regeneration rate.

What Is the Voice-Hour Cost Index (VHCI)?

The Voice-Hour Cost Index (VHCI) is an Axis Intelligence Research metric that tracks what one finished hour of current-generation, expressive AI speech costs at public list price. It is the median cost per audio hour across a fixed basket of five tiers, each being the newest natural-voice offering in its vendor’s catalog.

VHCI components and reading (September 24, 2026)

Basket componentVendorCost per audio hourBasis
Gemini 2.5 Flash TTSGoogle Cloud$0.90Audio-token billing
Chirp 3 HDGoogle Cloud$1.30Character billing, AWS duration basis
Polly GenerativeAmazon$1.30Character billing, AWS duration basis
Flash/Turbo and v3 ConversationalElevenLabs$2.16Character billing, AWS duration basis
v3ElevenLabs$4.32Character billing, AWS duration basis
VHCI (median)$1.30
Premium-to-floor ratio4.8x$4.32 / $0.90

Source: Axis Intelligence Research calculation from vendor pricing pages retrieved September 24, 2026.

Why a median rather than an average? One vendor’s premium tier would drag a mean upward and make the typical hour look pricier than what most buyers pay. The median holds steady when a single vendor reprices, and moves only when the middle of the market moves. Legacy tiers (Standard, WaveNet, Neural2, Neural) sit outside the basket because they no longer represent what a buyer shopping for natural speech in 2026 would choose. Studio and Long-Form tiers are excluded as specialty products.

According to Axis Intelligence Research, the VHCI will be recomputed whenever any basket vendor changes a list price, and each reading will carry its date. The first reading, published here, is the baseline.

Why Do AI Voice Vendors Disagree on How Much Audio a Character Buys?

This is the finding most buyers miss, and it changes budgets by more than a third.

Amazon says one million characters equals about 23 hours and 8 minutes of speech. ElevenLabs’s own API pricing page labels its $0.10 per 1,000 characters as roughly $0.10 per minute, which implies 1,000 characters per minute, or 60,000 characters per hour. According to Axis Intelligence Research, that is a 38.8% gap in assumed speech density between two vendors describing the same language.

Characters per audio hour: vendor-implied conversion factors

Vendor basisCharacters per audio hourHours per 1M charactersSource
Amazon Polly pricing example43,22823.13AWS pricing page
ElevenLabs minute equivalence60,00016.67ElevenLabs pricing page

The practical consequence: on ElevenLabs’s own basis, an hour of v3 audio costs $6.00, not the $4.32 Axis computes on the AWS basis. Neither figure is wrong. Speaking rate depends on the voice, the text and the model’s pacing, and a fast conversational voice burns through characters faster than a measured narrator. We publish both readings rather than averaging them, because a blended number would describe no real voice. Buyers should time a representative sample of their own script and apply that rate, not either vendor’s.

How Much Does It Cost to Turn a Book Into an AI Audiobook?

Amazon’s pricing examples include a full novel: Adventures of Huckleberry Finn, about 600,000 characters and about 13 hours and 50 minutes of speech. That makes a useful yardstick, because a single title is how authors and publishers budget.

Cost to voice a 600,000-character novel, by tier

VendorVoice tierCost for the full novelSource
Amazon PollyStandard$2.40AWS pricing example
Amazon PollyNeural$9.60AWS pricing example
Google CloudGemini 2.5 Flash TTS$12.49Axis calculation
Amazon PollyGenerative$18.00AWS pricing example
Google CloudChirp 3 HD$18.00Axis calculation
Google CloudGemini 2.5 Pro TTS$24.98Axis calculation
ElevenLabsFlash/Turbo, v3 Conversational$30.00Axis calculation
ElevenLabsv3, v2 Multilingual$60.00Axis calculation
Amazon PollyLong-Form$60.00AWS pricing example
Google CloudStudio$96.00Axis calculation

Axis calculations: price per million characters x 0.6, or hourly cost x 13.88 hours for Gemini. Retrieved September 24, 2026.

ElevenLabs’s own marketing sits inside this range. The company wrote in May 2026 that independent authors produce full-length audiobooks with marketplace voices for under $100 in credits, which matches the $60 Axis computes for its top tier. The raw generation cost of an AI audiobook has fallen to the price of a restaurant dinner. What still costs money is everything around generation: pronunciation fixes, retakes on misread dialogue, mastering, and the distribution rules below.

For a tool-by-tool quality comparison rather than a cost view, see our hands-on best AI voice generators tested in 2026.

How Big Is ElevenLabs? Revenue, Valuation and Headcount

ElevenLabs is the only pure-play AI voice company that discloses revenue milestones, which makes it the category’s de facto public benchmark. Its disclosures also contain the clearest example on this page of why every number needs a date and a document.

ElevenLabs business metrics (2025 to 2026)

MetricValueAs ofSource
ARR at end of 2025 (first report)Over $330MDec 31, 2025ElevenLabs Series D announcement, Feb 4, 2026
ARR at end of 2025 (later report)$350MDec 31, 2025ElevenLabs $500M ARR post, May 5, 2026
ARR milestoneOver $500MFirst four months of 2026ElevenLabs, May 5, 2026
Series D valuation$11BFeb 4, 2026ElevenLabs, Feb 4, 2026
Total funding at Series D announcement$781M across five roundsFeb 4, 2026ElevenLabs, Feb 4, 2026
Employees530, across 50+ countriesMay 5, 2026ElevenLabs, May 5, 2026
Employee tender offer$100MMay 5, 2026ElevenLabs, May 5, 2026
Valuation-to-ARR multiple22xAxis calculation$11B / $500M
ARR per employee$0.94MAxis calculation$500M / 530

The same company reported its year-end 2025 figure two ways. In February it said it closed 2025 with over $330 million in ARR; in May it said $350 million. “Over $330 million” and “$350 million” are not strictly contradictory, but they are not the same number either, and we do not pick one. Growth from year-end to the $500 million mark is 42.9% on the $350 million base and 51.5% on the $330 million base, according to Axis Intelligence Research, in no more than four months.

Sarah Mitchell, analyst comment: Watch where the growth is attributed. The May post credits enterprises deploying voice agents in support, sales, hiring and marketing, not creators making voiceovers. The text-to-speech engine that built the company is now the input to a conversational product billed per minute of call. That is a different business with different competitors, and the model labs shipping native voice are the ones to watch.

Voice agents now have their own list price. ElevenLabs charges $0.08 per minute for its Speech Engine agent pipeline, which Axis converts to $4.80 per hour of live conversation.

How Much Do Voice Actors Earn From AI Voice Marketplaces?

The licensing model is shifting from a one-time session fee to per-use royalties, and ElevenLabs is the only platform publishing payout totals. According to its May 2026 marketplace update, creators had earned $11 million by November 2025 and more than $22 million six months later, with 10,400+ creators earning and voices spanning 32 languages.

ElevenLabs Voice Library creator payouts

MetricValueAs ofSource
Cumulative creator payouts$11MNovember 2025ElevenLabs
Cumulative creator payoutsOver $22MMay 22, 2026ElevenLabs
Creators earning10,400+May 22, 2026ElevenLabs
Languages in the Voice Library32May 22, 2026ElevenLabs
Mean cumulative payout per creator$2,115May 22, 2026Axis calculation: $22M / 10,400

That $2,115 average is a mean across a heavily skewed distribution. ElevenLabs’s own examples include creators earning a full-time living alongside thousands earning little. Treat it as the size of the pool divided by headcount, not as a typical creator’s income.

We considered expressing payouts as a share of ElevenLabs revenue and declined. Cumulative payouts since launch and an annualized run rate measure different things over different windows, and dividing one by the other would produce a ratio that means nothing. The calculation is logged as a retracted row in the dataset.

How Fast Are AI Voice Prices Falling?

ElevenLabs cut self-serve API prices on May 7, 2026. Per its pricing announcement, the Flash model on the Creator plan dropped from $0.11 to $0.05 per 1,000 characters, a 55% reduction. Speech-to-text fell 45% and agent pricing fell 20%, from $0.10 to $0.08 per minute.

ElevenLabs self-serve price cuts, May 7, 2026

ProductBeforeAfterCutSource
Text to Speech (Flash, Creator plan)$0.11 per 1K characters$0.05 per 1K characters55%ElevenLabs
Speech to Text (Scribe v2, Starter plan)$0.40$0.2245%ElevenLabs
Agents (Starter plan)$0.10 per minute$0.08 per minute20%ElevenLabs

The market leader cutting its core product by more than half, three months after a financing round at a record valuation, tells you where the pressure comes from. Google’s Gemini 2.5 Flash TTS already produces an hour for $0.90. A cut of this size looks less like a promotion and more like a repositioning of text-to-speech as the entry point, with margin moving to agents and enterprise contracts.

How Many Audiobooks Use AI Narration?

No platform publishes a clean count of AI-narrated titles, and figures circulating in secondary coverage vary widely. What the platforms do publish is the plumbing.

AI narration on the major audiobook platforms

PlatformMetricValueAs ofSource
SpotifyAudiobook catalog700,000+ titlesMay 21, 2026Spotify
SpotifyAudiobook markets22May 21, 2026Spotify
SpotifyAudiobook listening hours growth60% year over yearMay 21, 2026Spotify
SpotifyAudiobooks+ subscribers1 millionMay 21, 2026Spotify
SpotifyElevenLabs narration languages accepted via Findaway Voices29Feb 20, 2025 [older data]Spotify
AudibleAI-generated voices offered to publishers100+May 13, 2025 [older data]Audible
AudibleAI narration languages at launch4May 13, 2025 [older data]Audible

At its May 2026 Investor Day, Spotify announced an invite-only beta of audiobook creation tools inside Spotify for Authors, powered by ElevenLabs, launching in English with no exclusivity requirement. The same post reported audiobook listening hours up 60% year over year and Audiobooks+ on track for $100 million in annualized recurring revenue. Audible’s May 2025 announcement offered publishers more than 100 AI voices across English, Spanish, French and Italian.

Sarah Mitchell, analyst comment: The distribution layer is absorbing the production layer. When the store hands authors the generator, the question stops being whether AI narration is allowed and becomes which platform owns the default voice. Growth in listening hours means the demand side can absorb far more titles than human narration capacity could ever record.

What Laws Regulate AI Voice Generators in 2026?

Two rules now shape how synthetic voice can be deployed, one American and one European.

United States. On February 8, 2024, the Federal Communications Commission unanimously ruled that calls made with AI-generated voices are “artificial” under the Telephone Consumer Protection Act. The FCC announcement means a voice-cloned robocall to a consumer needs the same prior express consent as any prerecorded call, and it gave state attorneys general a direct enforcement route. Every outbound AI voice agent product in the US is built on top of that ruling.

European Union. Article 50 of the EU AI Act applies from August 2, 2026. Providers of systems that generate synthetic audio must mark outputs in a machine-readable, detectable format, and deployers must label deepfakes. Systems already on the market before August 2 have until December 2, 2026 for the marking duty, according to the European Commission’s Code of Practice FAQ. The voluntary code supporting compliance was published June 10, 2026, drafted by six independent experts with more than 187 participants, and will be reviewed at least every two years.

Regulatory milestones for AI voice

JurisdictionMeasureDateSource
United StatesFCC ruling: AI voices are “artificial” under the TCPAFeb 8, 2024FCC
European UnionCode of Practice on Transparency of AI-Generated Content publishedJun 10, 2026European Commission
European UnionAI Act Article 50 transparency obligations applyAug 2, 2026European Commission
European UnionMarking deadline for systems already on the marketDec 2, 2026European Commission

For the fraud side of synthetic speech, including complaint and loss data, see our voice cloning scam statistics and AI scam statistics.

Do AI Voice Cloning Tools Verify Consent?

Mostly, they did not when tested. Consumer Reports’ March 2025 assessment examined six voice cloning products: Descript, ElevenLabs, Lovo, PlayHT, Resemble AI and Speechify. Researchers could create a clone from publicly available audio in four of the six, which required only a checkbox or similar self-attestation. Four of the six also let researchers open an account with just a name and email address.

According to Axis Intelligence Research, that is a 66.7% share of tested products with no technical consent check at the time. The EU marking duty and the FCC ruling both address what happens after a clone is made. The consent step before it remains a product decision in most jurisdictions. Detection tools covering the other end of the pipeline are reviewed in our deepfake detection technology guide.

Frequently Asked Questions About AI Voice Generator Statistics

How much does one minute of AI voice cost in 2026?

At list price, one minute of AI speech costs from under a cent on legacy voices to about $0.10 on ElevenLabs v3 by that vendor’s own minute equivalence. According to Axis Intelligence Research, the median current-generation tier works out to about $0.022 per minute ($1.30 per hour).

Is ElevenLabs cheaper than Google or Amazon for text-to-speech?

No, at list price. ElevenLabs Flash costs $50 per million characters and v3 costs $100, versus $30 for Google Chirp 3 HD and Amazon Polly Generative, as of September 2026. Buyers pay the premium for emotional control, voice cloning and a voice marketplace, not raw per-character cost.

How many characters of text make one hour of AI speech?

It depends on whose basis you use. Amazon’s pricing example implies 43,228 characters per hour; ElevenLabs’s minute equivalence implies 60,000. Axis Intelligence Research recommends timing a sample of your own script, since pacing varies by voice and model.

What is ElevenLabs’s annual revenue in 2026?

ElevenLabs reported surpassing $500 million in annual recurring revenue in the first four months of 2026. ARR is a run rate, not audited annual revenue. The company reported year-end 2025 ARR as over $330 million in February 2026 and as $350 million in May 2026.

Can I sell an AI-narrated audiobook on Spotify or Audible?

Spotify accepts disclosed digital-voice narration, including ElevenLabs audiobooks through Findaway Voices, and announced an invite-only ElevenLabs-powered creation beta for June 2026. Audible offers AI narration through its own publisher programs with more than 100 voices. Each platform sets its own disclosure and submission rules.

Are AI-generated voices legal in phone calls?

In the US, AI-generated voices count as “artificial” voices under the TCPA following the FCC’s February 8, 2024 ruling, so non-emergency calls to consumers require prior express consent. In the EU, AI systems interacting with people must disclose that they are AI from August 2, 2026.

Do AI voice generators have to watermark their audio?

In the EU, yes. From August 2, 2026, providers of systems generating synthetic audio must mark outputs in a machine-readable, detectable format, with a transition to December 2, 2026 for systems already on the market. There is no equivalent federal requirement in the US.

How much can a voice actor earn by licensing a voice to an AI marketplace?

ElevenLabs reported more than $22 million in cumulative payouts to 10,400+ earning creators by May 2026, a mean of about $2,115 per creator by Axis Intelligence Research’s calculation. The distribution is highly uneven, with a small number of creators earning full-time incomes.

Methodology

Collection. Every figure on this page comes from a document fetched and read on September 24, 2026: vendor pricing pages (Amazon Web Services, Google Cloud, ElevenLabs), company announcements (ElevenLabs, Spotify, Audible), a Consumer Reports product assessment, the FCC and the European Commission. Each figure appears as a row in the downloadable dataset with its source URL, as-of date and retrieval date. No figure originates from secondary aggregation or market-research estimates.

Formulas. Cost per finished audio hour for character-billed tiers = list price per 1M characters x 43,228 / 1,000,000, where 43,228 characters per hour derives from AWS’s statement that 1M characters is about 23 hours 8 minutes of speech. For Gemini audio-token tiers: price per 1M audio tokens x 25 tokens per second x 3,600 seconds / 1,000,000. VHCI = median of the five basket tiers listed above. Novel cost = price per 1M characters x 0.6, or hourly cost x 13.88 hours for token-billed tiers.

Scope. List prices only; negotiated enterprise discounts, free tiers and taxes are excluded. Gemini costs exclude text-input tokens, which are small relative to audio output. The VHCI measures price, not voice quality. Speech density varies by voice and text, which is why both vendor conversion bases are published. ElevenLabs revenue figures are company-reported run rates. Source discrepancies are published side by side and never averaged.

About This Dataset

The dataset contains 102 rows covering list prices for 16 AI voice tiers across three vendors, Axis-calculated cost per audio hour and per novel, the VHCI components and reading, ElevenLabs revenue, funding, headcount and marketplace payouts, Spotify and Audible AI narration metrics, Consumer Reports consent-test results and US and EU regulatory dates. Every row carries source_org, source_document, source_url, retrieved_date, is_primary, axis_calculated and method_note columns.

  • File: ai-voice-generator-statistics-2026.csv
  • License: CC BY 4.0
  • Updates: relevance-driven, triggered by a list-price change at any VHCI basket vendor, a new ElevenLabs revenue disclosure, or a new platform AI-narration figure. Each VHCI reading is dated.
  • Required attribution: Axis Intelligence Research, AI Voice Generator Statistics 2026, axis-intelligence.com

How to Cite This Page

APA: Axis Intelligence Research. (2026). AI voice generator statistics 2026: What one hour of synthetic speech costs, who is earning from it, and the rules now in force. Axis Intelligence. https://axis-intelligence.com/ai-voice-generator-statistics/

MLA: Axis Intelligence Research. “AI Voice Generator Statistics 2026: What One Hour of Synthetic Speech Costs, Who Is Earning From It, and the Rules Now in Force.” Axis Intelligence, 24 Sept. 2026, axis-intelligence.com/ai-voice-generator-statistics/.

Chicago: Axis Intelligence Research. “AI Voice Generator Statistics 2026: What One Hour of Synthetic Speech Costs, Who Is Earning From It, and the Rules Now in Force.” Axis Intelligence, September 24, 2026. https://axis-intelligence.com/ai-voice-generator-statistics/.


Related Axis research: Best AI voice generators tested in 2026 · Best AI voice cloning tools · Best AI tools compared · Identity theft statistics

Recent Posts

Shein Statistics 2026: Revenue, Customers, Profit and What the IPO Filing Revealed

Shein Statistics 2026 By Axis Intelligence Research Co-author: Mia Scarlett (Business & Capital Markets) | Last upda

API Security Statistics 2026: Exploited Auth Flaws Doubled, Attacks Per Org Up 113%

API Security Statistics 2026 By Axis Intelligence Research Co-author: Marcus Chen (Cybersecurity) | Last updated: Septem

Streaming Statistics 2026: 49% of US TV Time, Who Gets Paid for It, and the Streaming-Linear Crossover Ratio

Streaming Statistics 2026 By Axis Intelligence Research Co-author: Alex Rivera | Last updated: September 24, 2026 | Lice

Axis Intelligence Research

Stay ahead on tech & data

Get notified when we publish or update datasets, trackers, research, and reports across technology, business, AI, cybersecurity, finance, infrastructure, energy, and more.

Research updates only. No spam. Unsubscribe anytime.