Tavus

Tavus

🔥 Hot
AI Video Personalization
Quick answer

Tavus builds real-time conversational video — a photorealistic digital human that holds a live face-to-face conversation with a user, delivered as an API rather than a content tool. Pricing is per minute of live conversation, and unusually for this category it's published: Free at 20 minutes, Starter at $22 for 60, Builder at $59 for 175, Growth at $397 for 1,300 and Business at $975 for 4,000, with Enterprise custom. Builder and above add uncapped pay-as-you-go beyond the allowance. The billing detail that catches people: a conversation is billed the moment it's created, whether or not anyone joins, though a test mode exists for development. The stack runs on Tavus's own models — Phoenix-4 for rendering, Raven-1 for perception, Sparrow-1 for turn-taking — at sub-600ms latency.

Best for: Developers building products where the conversation itself is the interface
Skip if: You want to make marketing videos — this is infrastructure, not a content tool
Free 20 min · $22 Starter · $59 Builder · $397 Growth · $975 Business
EdGrowsReviewed by EdGrows·Updated Aug 24, 2026
Try Tavus for free

Affiliate link — we may earn a commission

Researched

This analysis is based on documentation, public user reports, and vendor materials — not yet on our own hands-on testing. How we rate

This is not a video tool

The most common mistake with Tavus is putting it in the wrong category. It sits alongside HeyGen and Synthesia on comparison lists, and it's doing something different.

Those tools generate videos. You write a script, an avatar reads it, you get a file. Marketing, training, localization.

Tavus builds a conversation. A user opens your app and talks to a photorealistic digital human that responds in real time, perceives what they're doing on camera, and knows when to speak and when to listen. There's no script and no output file. The conversation is the product.

That distinction determines everything else — the pricing model, the buyer, and whether this is remotely relevant to you.

Pricing, published and per-minute

Unusually for this space, Tavus publishes developer rates.

PlanPriceConversation minutesOverage
Free$020None
Starter$22/mo60None
Builder$59/mo175Uncapped PAYG
Growth$397/mo1,300Uncapped PAYG
Business$975/mo4,000Uncapped PAYG
EnterpriseCustomNegotiatedNegotiated

A minute is time an AI human is actively in a session with a user, connect to disconnect, running the full pipeline.

Note the overage split. Builder and above let you exceed your allowance and pay for it. Free and Starter don't — you stop. For anything user-facing where traffic is unpredictable, that makes Builder the practical floor regardless of what your average usage suggests.

Worth knowing: Tavus's consumer-facing PALs product no longer publishes prices, with that page now redirecting to developer pricing. If someone quotes you a PALs rate from a article, it isn't verifiable anymore.

The billing detail that will catch you

A conversation is billed the moment the creation request is sent — not when a participant joins.

Create a session for a user who never shows up, and you've paid for it. There's a safeguard: a conversation automatically closes after five minutes with no participant present, and those limits are configurable when you create it. But the default behavior means sloppy session management costs real money.

Two things follow. First, wire up test mode during development. Tavus supports setting a test flag so you can exercise the conversation creation flow without incurring charges, and building that in from day one avoids burning your allowance on integration work. Second, be deliberate about when you create sessions in your application flow. Creating one on page load rather than on user action is an expensive architectural choice.

What's actually included

One genuine simplification: the minute price covers the whole pipeline.

Tavus's materials state that a CVI conversation includes the language model, audio, WebRTC transport and all of Tavus's own models, with no additional services or charges required to run end to end.

Compare that to assembling this yourself — speech recognition from one vendor, an LLM from another, text-to-speech from a third, lip sync from a fourth, transport infrastructure of your own, each billing separately and each a potential point of latency. The bundled price is a real argument, and it's the reason a per-minute rate that looks high in isolation can be competitive in practice.

Replica training beyond plan limits does carry separate overage, so check that if you're creating many digital twins rather than a few.

The models are the differentiator

Tavus built four proprietary systems rather than wrapping commodity APIs, and this is the technical case for choosing it.

Phoenix-4, released February 2026, handles visual rendering — the talking face with natural head motion and full-face micro-expressions rather than a static avatar with a moving mouth. It supports programmatic emotion control, letting you specify states that adjust facial geometry.

Raven-1 does perception, reading facial expressions and tone so the agent has awareness of what the user is doing, not just what they said.

Sparrow-1 handles turn-taking — when to speak, when to wait. This is the unglamorous part that most AI video demos get wrong, and getting it right is a large share of why a conversation feels natural.

Hummingbird-0 integrates the above into a coherent real-time interaction.

Reported end-to-end latency is sub-600ms over WebRTC. That number is the whole ballgame. Roughly a second is the threshold where a video conversation stops feeling like waiting and starts feeling like talking, and a technically beautiful avatar that pauses awkwardly is worse than a text box. Latency is the first thing to test in any evaluation.

How we researched this

We haven't built on Tavus, and on a real-time latency-critical API, the things that matter most are exactly the things you learn by building.

This page draws on Tavus's published pricing and documentation, its model release materials for Phoenix-4, independent pricing tracking that verified developer rates against tavus.io in late July 2026, and third-party reviews published between May and July 2026.

The pricing figures here are more solid than most tools we cover, because Tavus publishes them. The performance claims are less so — latency figures, benchmark results and rendering quality are all vendor-reported, and while Tavus publishes benchmarks for its models rather than just asserting quality, we haven't independently verified any of it. Treat the sub-600ms figure as a target to test rather than a guarantee.

One transient detail worth noting for anyone reading older pricing: Builder briefly advertised a first-month discount in mid-July 2026 which was pulled by the end of the month. No tier currently shows an introductory rate.

The concerns worth raising

Company scale. Roughly $58 million raised, backing from Sequoia and Y Combinator, and a team reported around forty people. That's small for something positioning itself as infrastructure. Maintaining four proprietary research models while supporting enterprise deployments is significant surface area for that headcount, and commercial scale trails the larger video AI companies substantially. This is a technology bet, not an established platform.

Regulatory exposure. Raven-1's emotion perception sits in territory the EU AI Act restricts, particularly for emotion detection in workplace contexts. European deployments need legal review of that feature specifically. Separately, synthetic likeness law is developing fast and varies by jurisdiction — this is technology for making a real person appear to say things they didn't say, and the consent chain matters.

Category risk. Real-time conversational video is unproven as a product category. The demos are impressive. Whether users actually prefer talking to a digital human over typing, at scale, across use cases, is not yet established. Building a product on that premise is a bet on user behavior as much as on technology.

Where it fits, and where it doesn't

Right fit: developer teams building applications where the conversation is the interface — interview practice, coaching, guided onboarding, interactive sales demos, healthcare intake. Products where the face adds something rather than decorating something.

Wrong fit: anyone wanting to make videos. If your job is training content, localized marketing or talking-head explainers, HeyGen and Synthesia do that job properly and Tavus does not. Also wrong for teams without developer resources — this is an API, and there's no meaningful no-code path.

Against the alternatives

Against HeyGen: different products despite the surface similarity. HeyGen generates video content at scale with a mature interface and far greater commercial traction. Tavus builds live conversations. If you're comparing them on a feature grid, you've probably misidentified what you need.

Against Synthesia: the enterprise standard for AI video content — training material, corporate communications, localization. Again, generation rather than conversation.

Against D-ID: the closest competitor on real-time conversational avatars, and the direct comparison worth running. Test both on latency and turn-taking with your actual use case.

Against Hedra: character generation with a creative rather than infrastructure orientation.

Against ElevenLabs: voice only, and excellent at it. If the face isn't essential to your product, voice-only conversational AI is dramatically cheaper and simpler, and it's worth honestly asking whether the video layer earns its cost.

Against Sierra AI: both build conversational agents; Sierra sells finished enterprise customer service outcomes while Tavus sells the video conversation layer as infrastructure. Different points in the stack.

Pricing 2026

PlanCostMinutesNotes
Free$020Full CVI pipeline, no overage
Starter$22/mo60No pay-as-you-go
Builder$59/mo175Uncapped PAYG — practical floor for production
Growth$397/mo1,300Uncapped PAYG
Business$975/mo4,000Uncapped PAYG
EnterpriseCustomNegotiatedWhite-label, SLAs, SOC 2, HIPAA

Checked August 2026. Developer CVI rates verified against Tavus's published pricing in late July 2026. A conversation is billed on creation regardless of participant join; test mode avoids charges during development. Included minute price covers the full pipeline — LLM, audio, WebRTC and all Tavus models — with replica training overage separate. Tavus's consumer PALs product no longer publishes a price list. Verify current rates at tavus.io.

Build test mode in first. Billing on conversation creation means integration work will otherwise eat your allowance.

Treat Builder as the production floor. Free and Starter have no overage — traffic spikes stop your product rather than costing you money.

Test latency before anything else. Sub-600ms is the claim and it's the variable that determines whether users tolerate the interface.

Get legal involved early on EU deployments. Emotion perception and synthetic likeness both sit in actively developing regulatory territory.

Our Verdict

Tavus is doing genuinely hard technical work and doing it well. Building four proprietary models to handle rendering, perception, turn-taking and integration — rather than assembling commodity APIs and calling it a platform — is the real differentiator, and the published per-minute pricing that bundles the entire pipeline is refreshingly straightforward for this market.

The pricing is honest and the billing has a trap. Rates are published, the bundled pipeline is a real simplification, and conversations bill on creation whether or not anyone joins. That last point is documented and will still surprise people who didn't read carefully.

Latency is the product. Sub-600ms is what separates a conversation from a slideshow with a face, and it's the only claim worth verifying yourself before committing.

The company is small for its ambition. Roughly forty people maintaining four research models while serving enterprise customers, with commercial scale well behind the video AI leaders. That's a real procurement consideration, not a knock on the technology.

For developers building products where the conversation genuinely is the interface — interview practice, coaching, guided onboarding — Tavus is the strongest technical option and the free tier costs nothing to evaluate. For anyone who wants AI video content, this is the wrong category entirely and HeyGen or Synthesia will serve you better. And for anyone in between, the question worth asking honestly is whether the face adds enough over a voice-only agent to justify the complexity and the per-minute rate.

Note: AIVario earns no commission from Tavus. This page is based on published documentation, vendor model releases and independent pricing verification rather than a build.

Best for: Developers building conversation-as-interface products, interview and coaching applications, interactive onboarding and guided demos, teams that want the full pipeline from one vendor Not ideal for: Marketing and training video production, teams without developer resources, EU workplace deployments using emotion perception, buyers who need an established platform rather than a technology bet Bottom line: Impressive proprietary technology with transparent per-minute pricing, from a small team betting on a category that hasn't proven itself yet. Test the latency, then decide.

  • HeyGen — AI video generation at scale, a different job entirely
  • Synthesia — the enterprise standard for training and corporate video
  • D-ID — the closest competitor on real-time conversational avatars
  • ElevenLabs — voice-only conversational AI, cheaper and simpler
  • Hedra — character generation with a creative orientation
  • Sierra AI — conversational agents sold as enterprise outcomes

Frequently Asked Questions about Tavus

What does Tavus cost?

Developer plans are published and billed on live conversation minutes. Free gives 20 minutes. Starter is $22 a month for 60 minutes, Builder $59 for 175, Growth $397 for 1,300 and Business $975 for 4,000, with Enterprise on a custom quote. Builder, Growth and Business include uncapped pay-as-you-go minutes beyond the included allowance, while Free and Starter have no overage option — you simply stop. Tavus's consumer-facing PALs product no longer publishes a price list, with that page redirecting to developer pricing.

What counts as a billable minute?

Time an AI human is actively in a session with a user, from connect to disconnect, running the full pipeline. The important detail is when billing starts: a conversation is billed the moment the creation request is sent, regardless of whether a participant ever joins. As a safeguard, a conversation automatically closes after five minutes if nobody is present, and those limits are configurable when creating the conversation. For development work you can set test mode to avoid incurring costs entirely, which is worth wiring in early.

Is there anything to pay for on top of the minutes?

Not for the conversation itself. Tavus's pricing materials state that a CVI conversation includes everything needed end to end — the language model, audio, WebRTC transport and all of Tavus's own models — with no additional services required. That's a meaningful simplification against assembling speech recognition, an LLM, text-to-speech and lip sync from separate vendors, each with its own bill. Replica training beyond plan limits does carry overage, so check that against your plan if you're creating many digital twins.

What are Phoenix, Raven and Sparrow?

Tavus's own models, each handling a different part of the problem. Phoenix-4, released in February 2026, does visual rendering — the talking face with natural head motion and micro-expressions. Raven-1 handles perception, reading facial expressions and tone so the agent has some awareness of what the user is doing. Sparrow-1 manages turn-taking, deciding when to speak and when to wait, which is the part most AI video demos get wrong. Hummingbird-0 ties them together. This is a research organization building proprietary layers rather than a vendor wrapping commodity APIs, and that's the technical case for choosing it.

How fast is it really?

Tavus reports sub-600ms end-to-end latency, streaming over WebRTC. That number matters more than it sounds, because roughly a second is the threshold where a video conversation stops feeling like waiting on a chatbot and starts feeling like talking to someone. Latency is the make-or-break variable for this entire product category — a technically impressive avatar that pauses awkwardly is worse than a text interface — and it's the thing to test first in any evaluation.

How do I create a digital replica?

From roughly two minutes of consented video footage, from which Phoenix-4 learns the person's face, micro-expressions and voice. A feature added in 2026 extends replica creation to still photographs. Beyond custom replicas, Tavus provides a library of stock replicas you can use without training anything, which is the faster route for prototyping. Consent and likeness rights are your responsibility, and worth taking seriously — this is technology for making a real person appear to say things they didn't say.

Any regulatory concerns?

Yes, and they're worth raising before deployment rather than after. Raven-1's emotion-perception capability sits in territory the EU AI Act restricts, particularly around emotion detection in workplace contexts, so European deployments need legal review of that specific feature. There's also the general question of synthetic likeness — laws around digital replicas of real people are developing quickly and vary by jurisdiction. Neither is a reason to avoid the platform; both are reasons to involve counsel earlier than you would with ordinary SaaS.

Should I be worried about company size?

It's a fair consideration. Tavus has raised roughly $58 million with backing from Sequoia and Y Combinator, and runs a team reported around forty people. That's small for a company trying to be infrastructure — maintaining four proprietary research models while supporting enterprise deployments is a lot of surface area for that headcount. Commercial scale also trails the better-known video AI companies by a wide margin. Treat Tavus as a strong technology bet rather than an established platform, and weight that accordingly in procurement.

View all →