AI Voice Generator Statistics 2026
By Axis Intelligence Research
Co-author: Sarah Mitchell | Last updated: September 24, 2026 | License: CC BY 4.0
According to Axis Intelligence Research, the median cost of one finished hour of current-generation AI voice audio is $1.30 across five expressive tiers from Amazon, Google and ElevenLabs, as of September 24, 2026. The cheapest current tier produces an hour for $0.90; the priciest costs $4.32, a 4.8x spread for the same hour of speech.
Quick Answer
An hour of AI-generated speech now costs between $0.17 and $6.92 at list price, depending on the voice tier, with the Axis Voice-Hour Cost Index (VHCI) at $1.30 per audio hour. Voicing a full-length novel of about 600,000 characters costs $18 on Amazon Polly Generative or Google Chirp 3 HD and $60 on ElevenLabs v3. ElevenLabs, the category’s largest independent vendor, passed $500 million in annual recurring revenue in the first four months of 2026.
Key Findings
- According to Axis Intelligence Research, the Voice-Hour Cost Index (VHCI) stood at $1.30 per finished audio hour on September 24, 2026, the median of five current-generation expressive voice tiers.
- ElevenLabs surpassed $500 million in annual recurring revenue within the first four months of 2026 and was valued at $11 billion in its February 2026 Series D, per the company.
- According to Axis Intelligence Research, vendors disagree by 38.8% on how much speech one million characters produces: AWS implies 43,228 characters per audio hour, ElevenLabs implies 60,000.
- ElevenLabs creators on its voice marketplace had earned more than $22 million by May 2026, doubling from $11 million in November 2025, across 10,400+ earning creators, per ElevenLabs.
- Consumer Reports found in March 2025 that four of six voice cloning products tested relied only on a self-attestation checkbox to confirm consent from the person being cloned.
How Much Does an AI Voice Generator Cost Per Hour of Audio?
Every major vendor bills text-to-speech by the character or by the audio token. Nobody buys characters, though. Producers, developers and localization teams buy finished minutes. So Axis Intelligence Research converted every published rate card into the unit buyers actually think in: dollars per finished hour of speech.
The conversion factor comes from Amazon itself. The Amazon Polly pricing page states that one million characters produces roughly 23 hours and 8 minutes of speech. That works out to 43,228 characters per audio hour, and Axis applies it to every character-billed tier below. Google’s newest Gemini TTS models bill by audio token instead, at 25 tokens per second of output according to the Google Cloud Text-to-Speech pricing page, which lets us compute their hourly cost directly with no conversion at all.
AI voice generator cost per audio hour, by vendor and tier (September 2026)
| Vendor | Voice tier | List price | Cost per finished audio hour | Source |
|---|---|---|---|---|
| Google Cloud | Standard | $4 per 1M characters | $0.17 | Google Cloud pricing page |
| Amazon Polly | Standard | $4 per 1M characters | $0.17 | AWS pricing page |
| Google Cloud | Neural2 | $16 per 1M characters | $0.69 | Google Cloud pricing page |
| Amazon Polly | Neural | $16 per 1M characters | $0.69 | AWS pricing page |
| Google Cloud | Gemini 2.5 Flash TTS | $10 per 1M audio tokens | $0.90 | Google Cloud pricing page |
| Google Cloud | Chirp 3 HD | $30 per 1M characters | $1.30 | Google Cloud pricing page |
| Amazon Polly | Generative | $30 per 1M characters | $1.30 | AWS pricing page |
| Google Cloud | Gemini 2.5 Pro TTS | $20 per 1M audio tokens | $1.80 | Google Cloud pricing page |
| ElevenLabs | Flash/Turbo, v3 Conversational | $0.05 per 1K characters | $2.16 | ElevenLabs pricing page |
| Google Cloud | Instant custom voice | $60 per 1M characters | $2.59 | Google Cloud pricing page |
| Amazon Polly | Long-Form | $100 per 1M characters | $4.32 | AWS pricing page |
| ElevenLabs | v3, v2 Multilingual | $0.10 per 1K characters | $4.32 | ElevenLabs pricing page |
| Google Cloud | Studio | $160 per 1M characters | $6.92 | Google Cloud pricing page |
Hourly costs are Axis Intelligence Research calculations: list price per million characters x 43,228 / 1,000,000. Gemini rows: price per million audio tokens x 25 tokens per second x 3,600 seconds, excluding the small text-input token charge. Retrieved September 24, 2026.
The table reads like a price ladder with a gap in the middle. Legacy voices sit at 17 cents an hour, a figure that has not moved in years because nobody is competing for that business anymore. The action is in the $0.90 to $2.16 band, where Google, Amazon and ElevenLabs’s low-latency models now overlap. According to Axis Intelligence Research, the premium ElevenLabs v3 tier costs 25 times more per character than the Standard voices at Google and Amazon.
Sarah Mitchell, analyst comment: The price gap is no longer a quality gap you can hear in a demo. It’s a gap in controllability: emotional direction, multi-speaker dialogue, long-form stability. Those are the things that break on the fifth generation, not the first. A $0.90 hour that needs three regenerations costs more than a $2.16 hour that lands on the first pass, and none of these rate cards tell you the regeneration rate.
What Is the Voice-Hour Cost Index (VHCI)?
The Voice-Hour Cost Index (VHCI) is an Axis Intelligence Research metric that tracks what one finished hour of current-generation, expressive AI speech costs at public list price. It is the median cost per audio hour across a fixed basket of five tiers, each being the newest natural-voice offering in its vendor’s catalog.
VHCI components and reading (September 24, 2026)
| Basket component | Vendor | Cost per audio hour | Basis |
|---|---|---|---|
| Gemini 2.5 Flash TTS | Google Cloud | $0.90 | Audio-token billing |
| Chirp 3 HD | Google Cloud | $1.30 | Character billing, AWS duration basis |
| Polly Generative | Amazon | $1.30 | Character billing, AWS duration basis |
| Flash/Turbo and v3 Conversational | ElevenLabs | $2.16 | Character billing, AWS duration basis |
| v3 | ElevenLabs | $4.32 | Character billing, AWS duration basis |
| VHCI (median) | $1.30 | ||
| Premium-to-floor ratio | 4.8x | $4.32 / $0.90 |
Source: Axis Intelligence Research calculation from vendor pricing pages retrieved September 24, 2026.
Why a median rather than an average? One vendor’s premium tier would drag a mean upward and make the typical hour look pricier than what most buyers pay. The median holds steady when a single vendor reprices, and moves only when the middle of the market moves. Legacy tiers (Standard, WaveNet, Neural2, Neural) sit outside the basket because they no longer represent what a buyer shopping for natural speech in 2026 would choose. Studio and Long-Form tiers are excluded as specialty products.
According to Axis Intelligence Research, the VHCI will be recomputed whenever any basket vendor changes a list price, and each reading will carry its date. The first reading, published here, is the baseline.
Why Do AI Voice Vendors Disagree on How Much Audio a Character Buys?
This is the finding most buyers miss, and it changes budgets by more than a third.
Amazon says one million characters equals about 23 hours and 8 minutes of speech. ElevenLabs’s own API pricing page labels its $0.10 per 1,000 characters as roughly $0.10 per minute, which implies 1,000 characters per minute, or 60,000 characters per hour. According to Axis Intelligence Research, that is a 38.8% gap in assumed speech density between two vendors describing the same language.
Characters per audio hour: vendor-implied conversion factors
| Vendor basis | Characters per audio hour | Hours per 1M characters | Source |
|---|---|---|---|
| Amazon Polly pricing example | 43,228 | 23.13 | AWS pricing page |
| ElevenLabs minute equivalence | 60,000 | 16.67 | ElevenLabs pricing page |
The practical consequence: on ElevenLabs’s own basis, an hour of v3 audio costs $6.00, not the $4.32 Axis computes on the AWS basis. Neither figure is wrong. Speaking rate depends on the voice, the text and the model’s pacing, and a fast conversational voice burns through characters faster than a measured narrator. We publish both readings rather than averaging them, because a blended number would describe no real voice. Buyers should time a representative sample of their own script and apply that rate, not either vendor’s.
How Much Does It Cost to Turn a Book Into an AI Audiobook?
Amazon’s pricing examples include a full novel: Adventures of Huckleberry Finn, about 600,000 characters and about 13 hours and 50 minutes of speech. That makes a useful yardstick, because a single title is how authors and publishers budget.
Cost to voice a 600,000-character novel, by tier
| Vendor | Voice tier | Cost for the full novel | Source |
|---|---|---|---|
| Amazon Polly | Standard | $2.40 | AWS pricing example |
| Amazon Polly | Neural | $9.60 | AWS pricing example |
| Google Cloud | Gemini 2.5 Flash TTS | $12.49 | Axis calculation |
| Amazon Polly | Generative | $18.00 | AWS pricing example |
| Google Cloud | Chirp 3 HD | $18.00 | Axis calculation |
| Google Cloud | Gemini 2.5 Pro TTS | $24.98 | Axis calculation |
| ElevenLabs | Flash/Turbo, v3 Conversational | $30.00 | Axis calculation |
| ElevenLabs | v3, v2 Multilingual | $60.00 | Axis calculation |
| Amazon Polly | Long-Form | $60.00 | AWS pricing example |
| Google Cloud | Studio | $96.00 | Axis calculation |
Axis calculations: price per million characters x 0.6, or hourly cost x 13.88 hours for Gemini. Retrieved September 24, 2026.
ElevenLabs’s own marketing sits inside this range. The company wrote in May 2026 that independent authors produce full-length audiobooks with marketplace voices for under $100 in credits, which matches the $60 Axis computes for its top tier. The raw generation cost of an AI audiobook has fallen to the price of a restaurant dinner. What still costs money is everything around generation: pronunciation fixes, retakes on misread dialogue, mastering, and the distribution rules below.
For a tool-by-tool quality comparison rather than a cost view, see our hands-on best AI voice generators tested in 2026.
How Big Is ElevenLabs? Revenue, Valuation and Headcount
ElevenLabs is the only pure-play AI voice company that discloses revenue milestones, which makes it the category’s de facto public benchmark. Its disclosures also contain the clearest example on this page of why every number needs a date and a document.
ElevenLabs business metrics (2025 to 2026)
| Metric | Value | As of | Source |
|---|---|---|---|
| ARR at end of 2025 (first report) | Over $330M | Dec 31, 2025 | ElevenLabs Series D announcement, Feb 4, 2026 |
| ARR at end of 2025 (later report) | $350M | Dec 31, 2025 | ElevenLabs $500M ARR post, May 5, 2026 |
| ARR milestone | Over $500M | First four months of 2026 | ElevenLabs, May 5, 2026 |
| Series D valuation | $11B | Feb 4, 2026 | ElevenLabs, Feb 4, 2026 |
| Total funding at Series D announcement | $781M across five rounds | Feb 4, 2026 | ElevenLabs, Feb 4, 2026 |
| Employees | 530, across 50+ countries | May 5, 2026 | ElevenLabs, May 5, 2026 |
| Employee tender offer | $100M | May 5, 2026 | ElevenLabs, May 5, 2026 |
| Valuation-to-ARR multiple | 22x | Axis calculation | $11B / $500M |
| ARR per employee | $0.94M | Axis calculation | $500M / 530 |
The same company reported its year-end 2025 figure two ways. In February it said it closed 2025 with over $330 million in ARR; in May it said $350 million. “Over $330 million” and “$350 million” are not strictly contradictory, but they are not the same number either, and we do not pick one. Growth from year-end to the $500 million mark is 42.9% on the $350 million base and 51.5% on the $330 million base, according to Axis Intelligence Research, in no more than four months.
Sarah Mitchell, analyst comment: Watch where the growth is attributed. The May post credits enterprises deploying voice agents in support, sales, hiring and marketing, not creators making voiceovers. The text-to-speech engine that built the company is now the input to a conversational product billed per minute of call. That is a different business with different competitors, and the model labs shipping native voice are the ones to watch.
Voice agents now have their own list price. ElevenLabs charges $0.08 per minute for its Speech Engine agent pipeline, which Axis converts to $4.80 per hour of live conversation.
How Much Do Voice Actors Earn From AI Voice Marketplaces?
The licensing model is shifting from a one-time session fee to per-use royalties, and ElevenLabs is the only platform publishing payout totals. According to its May 2026 marketplace update, creators had earned $11 million by November 2025 and more than $22 million six months later, with 10,400+ creators earning and voices spanning 32 languages.
ElevenLabs Voice Library creator payouts
| Metric | Value | As of | Source |
|---|---|---|---|
| Cumulative creator payouts | $11M | November 2025 | ElevenLabs |
| Cumulative creator payouts | Over $22M | May 22, 2026 | ElevenLabs |
| Creators earning | 10,400+ | May 22, 2026 | ElevenLabs |
| Languages in the Voice Library | 32 | May 22, 2026 | ElevenLabs |
| Mean cumulative payout per creator | $2,115 | May 22, 2026 | Axis calculation: $22M / 10,400 |
That $2,115 average is a mean across a heavily skewed distribution. ElevenLabs’s own examples include creators earning a full-time living alongside thousands earning little. Treat it as the size of the pool divided by headcount, not as a typical creator’s income.
We considered expressing payouts as a share of ElevenLabs revenue and declined. Cumulative payouts since launch and an annualized run rate measure different things over different windows, and dividing one by the other would produce a ratio that means nothing. The calculation is logged as a retracted row in the dataset.
How Fast Are AI Voice Prices Falling?
ElevenLabs cut self-serve API prices on May 7, 2026. Per its pricing announcement, the Flash model on the Creator plan dropped from $0.11 to $0.05 per 1,000 characters, a 55% reduction. Speech-to-text fell 45% and agent pricing fell 20%, from $0.10 to $0.08 per minute.
ElevenLabs self-serve price cuts, May 7, 2026
| Product | Before | After | Cut | Source |
|---|---|---|---|---|
| Text to Speech (Flash, Creator plan) | $0.11 per 1K characters | $0.05 per 1K characters | 55% | ElevenLabs |
| Speech to Text (Scribe v2, Starter plan) | $0.40 | $0.22 | 45% | ElevenLabs |
| Agents (Starter plan) | $0.10 per minute | $0.08 per minute | 20% | ElevenLabs |
The market leader cutting its core product by more than half, three months after a financing round at a record valuation, tells you where the pressure comes from. Google’s Gemini 2.5 Flash TTS already produces an hour for $0.90. A cut of this size looks less like a promotion and more like a repositioning of text-to-speech as the entry point, with margin moving to agents and enterprise contracts.
How Many Audiobooks Use AI Narration?
No platform publishes a clean count of AI-narrated titles, and figures circulating in secondary coverage vary widely. What the platforms do publish is the plumbing.
AI narration on the major audiobook platforms
| Platform | Metric | Value | As of | Source |
|---|---|---|---|---|
| Spotify | Audiobook catalog | 700,000+ titles | May 21, 2026 | Spotify |
| Spotify | Audiobook markets | 22 | May 21, 2026 | Spotify |
| Spotify | Audiobook listening hours growth | 60% year over year | May 21, 2026 | Spotify |
| Spotify | Audiobooks+ subscribers | 1 million | May 21, 2026 | Spotify |
| Spotify | ElevenLabs narration languages accepted via Findaway Voices | 29 | Feb 20, 2025 [older data] | Spotify |
| Audible | AI-generated voices offered to publishers | 100+ | May 13, 2025 [older data] | Audible |
| Audible | AI narration languages at launch | 4 | May 13, 2025 [older data] | Audible |
At its May 2026 Investor Day, Spotify announced an invite-only beta of audiobook creation tools inside Spotify for Authors, powered by ElevenLabs, launching in English with no exclusivity requirement. The same post reported audiobook listening hours up 60% year over year and Audiobooks+ on track for $100 million in annualized recurring revenue. Audible’s May 2025 announcement offered publishers more than 100 AI voices across English, Spanish, French and Italian.
Sarah Mitchell, analyst comment: The distribution layer is absorbing the production layer. When the store hands authors the generator, the question stops being whether AI narration is allowed and becomes which platform owns the default voice. Growth in listening hours means the demand side can absorb far more titles than human narration capacity could ever record.
What Laws Regulate AI Voice Generators in 2026?
Two rules now shape how synthetic voice can be deployed, one American and one European.
United States. On February 8, 2024, the Federal Communications Commission unanimously ruled that calls made with AI-generated voices are “artificial” under the Telephone Consumer Protection Act. The FCC announcement means a voice-cloned robocall to a consumer needs the same prior express consent as any prerecorded call, and it gave state attorneys general a direct enforcement route. Every outbound AI voice agent product in the US is built on top of that ruling.
European Union. Article 50 of the EU AI Act applies from August 2, 2026. Providers of systems that generate synthetic audio must mark outputs in a machine-readable, detectable format, and deployers must label deepfakes. Systems already on the market before August 2 have until December 2, 2026 for the marking duty, according to the European Commission’s Code of Practice FAQ. The voluntary code supporting compliance was published June 10, 2026, drafted by six independent experts with more than 187 participants, and will be reviewed at least every two years.
Regulatory milestones for AI voice
| Jurisdiction | Measure | Date | Source |
|---|---|---|---|
| United States | FCC ruling: AI voices are “artificial” under the TCPA | Feb 8, 2024 | FCC |
| European Union | Code of Practice on Transparency of AI-Generated Content published | Jun 10, 2026 | European Commission |
| European Union | AI Act Article 50 transparency obligations apply | Aug 2, 2026 | European Commission |
| European Union | Marking deadline for systems already on the market | Dec 2, 2026 | European Commission |
For the fraud side of synthetic speech, including complaint and loss data, see our voice cloning scam statistics and AI scam statistics.
Do AI Voice Cloning Tools Verify Consent?
Mostly, they did not when tested. Consumer Reports’ March 2025 assessment examined six voice cloning products: Descript, ElevenLabs, Lovo, PlayHT, Resemble AI and Speechify. Researchers could create a clone from publicly available audio in four of the six, which required only a checkbox or similar self-attestation. Four of the six also let researchers open an account with just a name and email address.
According to Axis Intelligence Research, that is a 66.7% share of tested products with no technical consent check at the time. The EU marking duty and the FCC ruling both address what happens after a clone is made. The consent step before it remains a product decision in most jurisdictions. Detection tools covering the other end of the pipeline are reviewed in our deepfake detection technology guide.
Frequently Asked Questions About AI Voice Generator Statistics
How much does one minute of AI voice cost in 2026?
At list price, one minute of AI speech costs from under a cent on legacy voices to about $0.10 on ElevenLabs v3 by that vendor’s own minute equivalence. According to Axis Intelligence Research, the median current-generation tier works out to about $0.022 per minute ($1.30 per hour).
Is ElevenLabs cheaper than Google or Amazon for text-to-speech?
No, at list price. ElevenLabs Flash costs $50 per million characters and v3 costs $100, versus $30 for Google Chirp 3 HD and Amazon Polly Generative, as of September 2026. Buyers pay the premium for emotional control, voice cloning and a voice marketplace, not raw per-character cost.
How many characters of text make one hour of AI speech?
It depends on whose basis you use. Amazon’s pricing example implies 43,228 characters per hour; ElevenLabs’s minute equivalence implies 60,000. Axis Intelligence Research recommends timing a sample of your own script, since pacing varies by voice and model.
What is ElevenLabs’s annual revenue in 2026?
ElevenLabs reported surpassing $500 million in annual recurring revenue in the first four months of 2026. ARR is a run rate, not audited annual revenue. The company reported year-end 2025 ARR as over $330 million in February 2026 and as $350 million in May 2026.
Can I sell an AI-narrated audiobook on Spotify or Audible?
Spotify accepts disclosed digital-voice narration, including ElevenLabs audiobooks through Findaway Voices, and announced an invite-only ElevenLabs-powered creation beta for June 2026. Audible offers AI narration through its own publisher programs with more than 100 voices. Each platform sets its own disclosure and submission rules.
Are AI-generated voices legal in phone calls?
In the US, AI-generated voices count as “artificial” voices under the TCPA following the FCC’s February 8, 2024 ruling, so non-emergency calls to consumers require prior express consent. In the EU, AI systems interacting with people must disclose that they are AI from August 2, 2026.
Do AI voice generators have to watermark their audio?
In the EU, yes. From August 2, 2026, providers of systems generating synthetic audio must mark outputs in a machine-readable, detectable format, with a transition to December 2, 2026 for systems already on the market. There is no equivalent federal requirement in the US.
How much can a voice actor earn by licensing a voice to an AI marketplace?
ElevenLabs reported more than $22 million in cumulative payouts to 10,400+ earning creators by May 2026, a mean of about $2,115 per creator by Axis Intelligence Research’s calculation. The distribution is highly uneven, with a small number of creators earning full-time incomes.
Methodology
Collection. Every figure on this page comes from a document fetched and read on September 24, 2026: vendor pricing pages (Amazon Web Services, Google Cloud, ElevenLabs), company announcements (ElevenLabs, Spotify, Audible), a Consumer Reports product assessment, the FCC and the European Commission. Each figure appears as a row in the downloadable dataset with its source URL, as-of date and retrieval date. No figure originates from secondary aggregation or market-research estimates.
Formulas. Cost per finished audio hour for character-billed tiers = list price per 1M characters x 43,228 / 1,000,000, where 43,228 characters per hour derives from AWS’s statement that 1M characters is about 23 hours 8 minutes of speech. For Gemini audio-token tiers: price per 1M audio tokens x 25 tokens per second x 3,600 seconds / 1,000,000. VHCI = median of the five basket tiers listed above. Novel cost = price per 1M characters x 0.6, or hourly cost x 13.88 hours for token-billed tiers.
Scope. List prices only; negotiated enterprise discounts, free tiers and taxes are excluded. Gemini costs exclude text-input tokens, which are small relative to audio output. The VHCI measures price, not voice quality. Speech density varies by voice and text, which is why both vendor conversion bases are published. ElevenLabs revenue figures are company-reported run rates. Source discrepancies are published side by side and never averaged.
About This Dataset
The dataset contains 102 rows covering list prices for 16 AI voice tiers across three vendors, Axis-calculated cost per audio hour and per novel, the VHCI components and reading, ElevenLabs revenue, funding, headcount and marketplace payouts, Spotify and Audible AI narration metrics, Consumer Reports consent-test results and US and EU regulatory dates. Every row carries source_org, source_document, source_url, retrieved_date, is_primary, axis_calculated and method_note columns.
- File: ai-voice-generator-statistics-2026.csv
- License: CC BY 4.0
- Updates: relevance-driven, triggered by a list-price change at any VHCI basket vendor, a new ElevenLabs revenue disclosure, or a new platform AI-narration figure. Each VHCI reading is dated.
- Required attribution: Axis Intelligence Research, AI Voice Generator Statistics 2026, axis-intelligence.com
How to Cite This Page
APA: Axis Intelligence Research. (2026). AI voice generator statistics 2026: What one hour of synthetic speech costs, who is earning from it, and the rules now in force. Axis Intelligence. https://axis-intelligence.com/ai-voice-generator-statistics/
MLA: Axis Intelligence Research. “AI Voice Generator Statistics 2026: What One Hour of Synthetic Speech Costs, Who Is Earning From It, and the Rules Now in Force.” Axis Intelligence, 24 Sept. 2026, axis-intelligence.com/ai-voice-generator-statistics/.
Chicago: Axis Intelligence Research. “AI Voice Generator Statistics 2026: What One Hour of Synthetic Speech Costs, Who Is Earning From It, and the Rules Now in Force.” Axis Intelligence, September 24, 2026. https://axis-intelligence.com/ai-voice-generator-statistics/.
Related Axis research: Best AI voice generators tested in 2026 · Best AI voice cloning tools · Best AI tools compared · Identity theft statistics
