Skip to content
← Blogs

What Is ElevenLabs? A 2026 Guide From an Implementation Partner

ElevenLabs is an AI voice platform for speech and voice agents. See Eleven v4 and v4 Turbo, Agents, plans, GDPR, HIPAA and TCPA points, and rollout timelines.

Voice AI11 min read

ElevenLabs is an AI audio company that turns text into natural-sounding speech and lets businesses build voice agents that hold real conversations over the phone, the web and messaging apps. Its products cover text-to-speech, speech-to-text, voice cloning, dubbing, music, sound effects and conversational AI agents.

Most people know ElevenLabs for its voices. For businesses, the bigger story in 2026 is ElevenLabs Agents: AI agents that answer customer calls, resolve routine requests and hand complex cases to a human.

Haass is an ElevenLabs implementation partner. We design, build and deploy ElevenLabs voice agents for businesses in India, the US and Europe. This guide explains what the platform does, which model to use for what, what the plans include, the compliance points US and EU buyers need to check, and where the platform has limits.

Who ElevenLabs Is For

  • Customer support teams that want an AI agent to answer routine calls around the clock and pass complex ones to people
  • Sales and booking teams handling appointment scheduling, lead qualification and reminders
  • Healthcare providers automating appointment and prescription-refill calls, with HIPAA terms in place
  • Media, education and publishing companies producing narration, courses, audiobooks and dubbed video at scale
  • Software companies adding voice to their own products through the API
  • Public sector organizations running information lines and benefits hotlines. ElevenLabs reports that a national hotline in the Czech Republic handles about 5,000 calls a day, with around 85% resolved without a human.

Who it isn't for

  • Organizations that must process all audio on their own hardware, since there's no on-premises option
  • Teams that only need a few fixed phone menu prompts, where a basic text-to-speech service costs less
  • Buyers choosing purely on lowest price per character. Open-source models and budget providers cost less, with trade-offs in quality, language coverage and support.

What ElevenLabs Includes

Text-to-speech

The product ElevenLabs started with. You send text, choose a voice and a model, and get back audio as a file or a live stream. The voice library holds thousands of voices, and you can design new ones or clone your own.

ElevenLabs Agents

Agents combine speech-to-text, a large language model of your choice and ElevenLabs voices into one system that listens and replies in real time. The platform handles turn-taking, interruptions, knowledge bases, tool calls to your own systems, and connections to phone lines and messaging channels. This is the product most businesses come to us for.

Scribe (speech-to-text)

Scribe converts speech to text. There are two versions: Scribe v2 for recorded audio such as call recordings and meetings, and Scribe v2 Realtime for live conversations. Agents use the realtime version automatically.

Voice cloning

Instant voice cloning creates a usable voice from a short sample. Professional voice cloning trains a high-fidelity copy for brand voices, narrators and spokespeople. Both require consent from the person whose voice is cloned.

Dubbing

Dubbing translates video or audio into other languages while keeping the original speaker's voice and tone. Dubbing Studio lets editors adjust the transcript, translation and timing.

Music, sound effects and Studio

ElevenLabs also generates background music and sound effects from text prompts. Studio brings voice, music and effects together in one editor for long-form content such as courses, audiobooks and videos.

ElevenLabs Models: Which One to Use

Choosing the model is the first technical decision on every project, because speed and expressiveness pull in opposite directions.

ModelStrengthBest for
Eleven v4ElevenLabs' flagship and most emotive model, on a new architecture, with 90+ languagesNarration, ads, audiobooks, character voices and multi-speaker dialogue
Eleven v4 TurboThe low-latency version of v4, with a median inference latency of about 100ms, built and tuned alongside ElevenLabs AgentsExpressive real-time voice agents for support, sales and scheduling
Eleven v3Expressive delivery controlled with audio tags, with 70+ languagesExisting v3 projects. New expressive work should move to v4
Multilingual v2Stable, high-quality speech across many languagesLong-form narration and consistent brand voice
FlashVery low latency (around 75ms of model time) and the lowest cost per characterHigh-volume agents where cost per call matters most

What's new in Eleven v4

ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28, 2026, and calls them its fastest and most emotive voice models yet. The main changes:

  • A new architecture. v4 reads tone, pacing, emotion and context from the text, so it performs a script instead of reading it aloud. It's also more consistent than v3 and better at conversations between several speakers.
  • Direction inside the script. Inline tags control emotion, pacing, reactions and sound effects without a second recording session.
  • 90+ languages, up from 70 on v3, and the ability to switch language or accent while keeping the same voice.
  • Better voice cloning. ElevenLabs says Instant Voice Clones can now capture a voice from 10 seconds of audio, and Professional Voice Clones, which v3 didn't support, work with v4's full emotional range.
  • Pronunciation control through the International Phonetic Alphabet, useful for brand names, product names and people's names.
  • Independent ranking. At launch, Eleven v4 ranked first on Artificial Analysis' Speech Arena leaderboard.

Which model we use for voice agents

For new voice agents, we now test Eleven v4 Turbo first. It brings v4's expressiveness to real-time conversation at around 100ms of model latency, which is the combination agents have been missing. Flash stays the choice for very high call volumes where cost per call matters most. The full Eleven v4 model suits pre-produced audio, where a few hundred extra milliseconds don't matter.

One implementation detail: at launch, v4 Turbo streams through ElevenLabs' Text to Dialogue WebSocket rather than the standard text-to-speech WebSocket, so existing custom integrations may need changes before switching models. Agents built on the ElevenLabs Agents platform don't need this work.

ElevenLabs Agents in Production

A demo agent can be running in a day. A production agent that handles thousands of real calls needs more planning. These are the points we work through on every deployment.

Latency

Total response time is the sum of several steps: speech recognition, deciding the caller has finished speaking, the language model's first output, and voice generation. A benchmark published by Deepgram measured ElevenLabs agents at about 530ms in controlled conditions, before Eleven v4 Turbo was released. v4 Turbo's median inference latency of about 100ms should bring that total down, but measure it on your own calls. Real phone calls add network time on top, so test from the regions your callers are in, and track the slowest 5% of responses (P95), not only the average.

Turn-taking

The hardest part of a voice agent is knowing when the caller has stopped talking. ElevenLabs uses a turn-taking model that looks at filler words, rhythm and pauses rather than silence alone. You can set how eager the agent is to reply (Eager, Normal or Patient) and the turn timeout. Callers who pause to look up an account number need a more patient setting than callers ordering a pizza.

Concurrency

Each plan caps how many agent calls can run at the same time. According to ElevenLabs' help center, the limits are 4 on Free, 6 on Starter, 10 on Creator, 20 on Pro and 30 on Scale and Business. Anything above 30 simultaneous calls needs an Enterprise agreement. Size this against your peak hour, not your average day.

Knowledge and tools

An agent is only as accurate as the information it can reach. We load approved knowledge sources, connect read-only tools first (order status, appointment availability, account lookups) and add actions that change data only after testing. Every agent gets clear rules on what it must not answer and when to transfer to a human.

Handover to humans

Escalation design decides whether customers trust the agent. The transfer should pass the conversation summary to the human agent, so the caller never repeats themselves.

One Call From Start to Finish

Here's how a typical ElevenLabs agent handles a call for a US online retailer.

  1. The call connects. A customer calls the support number. Twilio routes the call to the ElevenLabs agent, which answers in under a second.
  2. The agent identifies the caller. It matches the phone number to a customer record and confirms the customer's name and order number.
  3. It checks the order. Using a tool connected to the retailer's order system, the agent finds the parcel is delayed at a regional hub, with delivery expected tomorrow.
  4. The caller pushes back. They need the item today for an event. The agent recognizes it can't promise same-day delivery and offers a transfer to a human.
  5. The handover happens. The call moves to a support rep, who sees the transcript and order details on screen and offers expedited shipping or a partial refund.
  6. The record updates. The conversation summary and outcome are written to the helpdesk, and the call is tagged so the operations team can track delivery complaints.

The agent resolved the lookup in seconds and knew when to stop. That balance is what separates a useful agent from a frustrating one.

ElevenLabs Plans and What Each One Includes

ElevenLabs runs on credits. Each plan includes a monthly credit allowance, roughly one credit per character of text-to-speech, with other products drawing credits at their own rates. Usage above the allowance is billed per model. Check current prices on the ElevenLabs pricing page and the Agents pricing page.

PlanWhat it adds
FreeMonthly credits for testing. No commercial license
StarterCommercial license and instant voice cloning
CreatorProfessional voice cloning and higher monthly credits
ProHigher-quality audio output, including 44.1 kHz PCM through the API, and more agent concurrency
ScaleMultiple seats and higher credit volumes
BusinessLarger credit allowance, more seats and low-latency pricing for high volumes
EnterpriseCustom terms, SLAs, a custom DPA, BAAs for HIPAA, SSO and concurrency above 30 calls

Three points come up in most of our scoping calls:

  1. Commercial use starts on Starter. Audio made on the Free plan can't be used in products or client work.
  2. Concurrency often decides the plan. Teams with call peaks above 30 simultaneous conversations need Enterprise, whatever their credit usage.
  3. Healthcare and regulated work usually means Enterprise. BAAs, custom DPAs and SSO sit there.

ElevenLabs for Indian Businesses

Indian companies are adopting voice agents for support lines, collections reminders, appointment booking and lead qualification. A few points need planning for Indian deployments:

  • Languages. ElevenLabs supports Hindi, Tamil and other Indian languages, and its voice library includes voices with Indian accents. Indian callers often mix English with Hindi or a regional language in the same sentence. Eleven v4 can switch languages and accents while keeping the same voice, but test code-mixed conversations with real callers before launch.
  • Telephony. Agents connect to Indian phone numbers through SIP trunking, which most Indian cloud telephony providers support. Confirm SIP compatibility with your current provider early, because it decides the integration path.
  • WhatsApp. For many Indian businesses, WhatsApp carries more customer conversations than email or web chat. ElevenLabs Agents can handle WhatsApp text, voice notes and supported calls.
  • TRAI rules. Outbound calls are covered by TRAI's commercial communication regulations. Service calls to existing customers and promotional calls follow different rules, so confirm with your telecom provider which category each call type falls into.
  • The DPDP Act, 2023. Recordings and transcripts contain personal data. Tell callers they're speaking with an AI and that calls may be recorded, define a retention period, and confirm where audio and transcripts are processed. ElevenLabs documents EU data residency and Zero Retention Mode, so check how these meet your requirements.

For more detail, read our guides to ElevenLabs for Indian businesses and ElevenLabs vs Sarvam AI vs Gnani AI.

Compliance for US and EU Deployments

This section is general information, not legal advice. Review your deployment with your legal and compliance teams.

What ElevenLabs provides

ElevenLabs states that its platform is SOC 2 Type 2, GDPR, CPRA and HIPAA compliant. It also offers EU data residency endpoints and a Zero Retention Mode, which processes data in memory without storing it. In October 2026, ElevenLabs announced FedRAMP 20x Class A certification for its agents, text-to-speech and speech-to-text products.

European Union

  • GDPR. Call recordings and transcripts contain personal data. You need a lawful basis for processing, a data processing agreement with ElevenLabs, a defined retention period and a way for callers to exercise their rights. EU data residency and Zero Retention Mode help with data minimization.
  • The EU AI Act. The Act's transparency rules require people to be told when they're interacting with an AI system, and AI-generated audio to be identifiable as such. In practice, your agent should say it's an AI at the start of the call, and synthetic audio used in content may need labeling.

United States

  • TCPA and outbound calls. In 2024, the FCC confirmed that AI-generated voices count as "artificial" voices under the Telephone Consumer Protection Act. Outbound calls using an AI voice need prior express consent, and marketing calls need prior express written consent. Inbound calls, where the customer calls you, carry far fewer restrictions.
  • HIPAA. Healthcare providers handling protected health information need a Business Associate Agreement, which ElevenLabs offers on Enterprise. Check that your own call recording and logging don't store health data outside approved systems.
  • CPRA. California residents have rights over their personal information, including recordings and transcripts. Include voice data in your privacy notices and data request processes.
  • Public sector. FedRAMP certification matters for federal agencies, and ElevenLabs now runs a dedicated ElevenLabs for Government program.

Voice cloning consent

Whatever the region, only clone a voice with written consent from the speaker, and keep that consent on record.

Telephony and Channels

ElevenLabs Agents connect to phone systems through a native Twilio integration and SIP trunking, so they can work alongside most cloud contact center platforms and carriers. Agents also run on websites and in mobile apps.

Since December 2025, agents can also work on WhatsApp, handling text, voice notes and supported calls. This matters most for European businesses, where WhatsApp is a common customer channel. Check the number eligibility rules and Meta's template requirements before moving an existing support number.

Strengths and Limits

Where ElevenLabs does well

Voice quality. Eleven v4 ranked first on Artificial Analysis' Speech Arena at launch, and that quality keeps callers engaged instead of asking for a human straight away.

One platform for the whole agent. Speech recognition, the language model connection, voices, knowledge bases, tool calls and telephony sit in one system, which reduces the integration work.

Speech recognition accuracy. Scribe v2 Realtime ranks at the top of the AA-WER v2 accuracy evaluation from Artificial Analysis.

Fast to prototype. A working demo agent takes a day, which makes it easy to test an idea with real callers before committing.

Strong compliance coverage. SOC 2 Type 2, GDPR, HIPAA with BAAs, EU data residency, Zero Retention Mode and FedRAMP cover most enterprise checklists.

Where it has limits

No on-premises deployment. Organizations that need fully air-gapped processing need a different provider.

Concurrency caps. Self-serve plans top out at 30 simultaneous agent calls. High-volume contact centers need Enterprise terms.

Noisy audio needs testing. ElevenLabs hasn't published performance figures for heavy background noise or speakerphone audio, so test with real call recordings from your own customers.

Mixed-language replies. Eleven v4 can switch languages within a voice, but older models work on one language per request. Agents serving bilingual callers still need language detection and testing built into the design.

Pronunciation of specific terms. Product codes, names and addresses sometimes need text normalization or pronunciation rules before they're spoken correctly.

None of these is a reason to avoid ElevenLabs. They're reasons to design the deployment carefully and test it under real conditions.

How Long Does an ElevenLabs Implementation Take?

These are the planning estimates we use when we scope ElevenLabs projects. The actual timeline depends on your integrations, the number of use cases and your compliance review.

ScopeWhat's includedTypical timeline
PilotOne use case, one channel, a knowledge base and a basic integration2 to 4 weeks
Production rolloutSeveral use cases, telephony, a CRM or helpdesk integration and a compliance review4 to 6 weeks
Multi-team or enterpriseSeveral agents, languages or regions, custom integrations and a security review8 to 12 weeks

The things that most often extend a project are legal review of call recording and consent, access to internal systems for tool integrations, and the time it takes to approve knowledge content.

Frequently Asked Questions

What is ElevenLabs used for?

Businesses use ElevenLabs to build AI voice agents for customer support and sales calls, to generate narration for videos, courses and audiobooks, to dub content into other languages, and to transcribe calls and meetings.

Which ElevenLabs model is best for voice agents?

Eleven v4 Turbo for most new agents, because it combines v4's expressiveness with a median inference latency of about 100ms. Flash suits very high-volume agents where cost per call matters most. The full Eleven v4 model suits pre-produced content.

What is Eleven v4?

Eleven v4 is ElevenLabs' newest text-to-speech model, released on September 28, 2026 alongside a low-latency version, Eleven v4 Turbo. It's built on a new architecture, supports 90+ languages, and ranked first on Artificial Analysis' Speech Arena at launch.

How long does it take to deploy an ElevenLabs voice agent?

Our planning estimates are 2 to 4 weeks for a pilot, 4 to 6 weeks for a production rollout and 8 to 12 weeks for multi-team or enterprise projects.

Work With an ElevenLabs Implementation Partner

Haass designs and deploys ElevenLabs voice agents for businesses in India, the US and Europe. We handle model and voice selection, conversation design, knowledge base setup, telephony and CRM integrations, compliance review and testing, then support the agent after launch.

If you're comparing providers first, our guide to the best TTS APIs in 2026 puts ElevenLabs side by side with Google, Amazon, OpenAI and others. For the technical side, read what an AI voice agent is and how it works and our ElevenLabs text-to-speech API integration guide.

See our ElevenLabs implementation services or book a call with our team.