AI PRODUCT MANAGEMENT

Tavus Griffin for Product Managers: The First Face-to-Face AI Interaction Model

By Institute of AI PM·14 min read·Oct 10, 2026

TL;DR

Tavus launched Griffin on October 1, 2026, calling it the first Human Interaction Model: a full-duplex, video-to-video AI that can see, hear, speak, and react in real time with a 0.43-second latency. In a Tavus-run study, 48% of participants believed they were speaking with a human after a one-minute call. Griffin-Lite is currently a research preview for select testers. For AI PMs, this is an inflection point in what "conversational AI" can mean: not a chatbot, not a voice agent, but a face-to-face video persona. The product opportunity is real. So is the disclosure obligation.

The AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.

What Griffin Actually Is

Most video AI products in 2026 are pipelines: a speech-to-text model feeds an LLM, which feeds a text-to-speech model, which drives a separately trained video avatar. Each handoff adds latency and removes coherence. Griffin eliminates the pipeline. It merges perception (seeing and hearing the user), conversational decision-making, speech, and video generation into a single end-to-end model.

1

Full-duplex video input

Griffin processes the user's video feed in real time, not just their audio. It can react to facial expressions, head nods, and visual context. This is qualitatively different from a voice assistant that happens to have a face.

2

Sub-second audio-to-video latency

Audio-to-video delay averages 0.43 seconds on NVIDIA H100 chips, according to Tavus. That is roughly half the delay of the next-fastest comparable approach. At that latency, conversation feels continuous rather than call-and-response.

3

Turn-taking model

Every sub-second, Griffin reassesses the conversation to decide whether to speak, nod, backchannel ('mm-hmm'), or wait. This models how humans actually regulate conversation, not just how they generate content.

4

Voice cloning

Griffin can clone a voice from approximately 10 seconds of audio. Combined with the video generation capability, this allows a persona to sound and look like a specific person, which amplifies both the use cases and the disclosure requirements.

5

720p generation

Griffin-Lite generates at 720p. This is sufficient for standard video call quality. Higher-resolution generation would require more compute and is likely on the roadmap for a full GA release.

On NVIDIA's VideoFDB benchmark, Griffin-Lite scored 3.83 for video generation quality (human reference: 3.92, next-best system: 2.80). NVIDIA ran the evaluation independently. In a 54-person Turing-style test with one-minute calls, 48% of participants believed they were talking to a human. The previous best was 2.4%, also Tavus's own earlier system. Note: these numbers come from Tavus's announcement and have not yet been replicated by independent researchers.

Product Use Cases: Where Griffin Changes the Calculation

Before Griffin, video-based AI was primarily a novelty or a B2B demo. The latency and realism gap meant users experienced it as a parlor trick, not a product. At 0.43-second latency and 48% human pass rate, that gap is closing. These are the use cases where a video-native AI persona creates genuine product value.

High-stakes customer onboarding

Complex financial products, insurance, and B2B SaaS that currently require a human video call to close. An AI persona that guides users through the process, answers questions in real time, and adjusts based on user reactions can handle peak volume without adding headcount.

When the product requires explanation, not just information

Healthcare intake and triage

Patient intake that captures not just answers but behavioral signals: hesitation, confusion, anxiety. A video-native AI can observe and respond to these in ways that text and voice cannot. Disclosure requirements are strict in healthcare; this use case requires explicit informed consent.

With strict disclosure and clinical validation

Language learning and tutoring

Conversational language practice with a video partner that responds naturally to mistakes, facial expressions, and pace. This is the use case where video-native AI is qualitatively better than voice: the visual channel is how humans learn conversational norms.

High readiness. Low disclosure complexity.

Sales training and role-play

Sales reps practice objection handling with an AI buyer persona that behaves like a skeptical procurement manager. The video channel adds realism that voice-only role-play lacks. Internal training use cases have lower regulatory stakes than customer-facing ones.

High readiness. Internal use avoids disclosure complexity.

Frontline customer support escalation

When text and voice support hit a frustration threshold, a video persona can de-escalate with visual warmth. The 48% human pass rate means users who don't know they're talking to AI may respond differently. Disclosure is not optional here.

Proceed only with explicit disclosure. Regulatory risk is real.

Interactive explainer content

Product walkthroughs, onboarding flows, and documentation that responds to user questions mid-flow. Not a passive video, but a video that answers back. Lower fidelity than a full Griffin persona; consider whether voice-only is sufficient before adding video cost.

Moderate readiness. Evaluate cost-to-value carefully.

The Disclosure Problem You Cannot Ignore

A 48% human pass rate is impressive in an academic context. In a product context, it is a liability disclosure, not a feature. The moment users cannot reliably tell whether they are talking to an AI, you have a legal, ethical, and trust problem that will not stay contained.

Legal and regulatory exposure

EU AI Act Article 52

Requires that AI systems intended to interact with natural persons disclose their AI nature, except for law enforcement applications. Customer-facing Griffin personas require explicit disclosure at the point of interaction. Failure to disclose is a compliance violation, not just bad practice.

FTC guidance on deceptive practices

US FTC has historically treated undisclosed AI impersonation as a deceptive trade practice. While no Griffin-specific ruling exists, the framework is clear: presenting an AI as human without disclosure is deceptive.

State-level bot disclosure laws

California's BOTS Disclosure Act (AB 13) requires bots interacting with consumers in commercial contexts to disclose their AI nature. Similar laws exist or are pending in multiple states. Video bots interacting with consumers for commercial purposes fall within scope.

Voice cloning consent

If you use Griffin's voice cloning to make the AI sound like a real person (e.g., a company executive or a famous spokesperson), you need explicit consent from that person. Unauthorized voice cloning is covered by right-of-publicity law in most US states.

The PM's disclosure design principle

Disclosure is not a legal checkbox you hide in the terms of service. It is a UX element that must be visible at the point of interaction, before the user commits to the conversation. The right design makes the AI nature clear without degrading the product experience. "Hi, I'm Maya, an AI from Acme" is sufficient. A system that waits until after the conversation to inform the user they were talking to an AI is not compliant and will not survive public scrutiny.

Build AI Products That Ship and Scale

The AI PM Masterclass covers AI product design, trust and safety, and how to bring new AI capabilities to market responsibly, taught live by a Salesforce Sr. Director PM.

Technical Constraints to Build Around

Griffin-Lite is a research preview. The technical constraints are real and will shape what you can build with it today versus in 6-12 months.

H100 infrastructure requirement

Griffin's 0.43-second latency benchmark is measured on NVIDIA H100 chips. On different hardware, latency will be higher. For most teams, this means you are dependent on Tavus's hosted infrastructure rather than self-hosting. Build your architecture around API calls, not self-hosted models.

PM action: Design for API-first. Plan on Tavus infrastructure costs, not compute costs.

Research preview access

Griffin-Lite is limited to select trusted testers. There is no public self-service access at launch. This means you need to apply to the waitlist and demonstrate a use case to Tavus before you can evaluate the model with your own data.

PM action: Apply to the waitlist now if you have a serious use case. Include a specific business context.

720p maximum resolution

720p is sufficient for video calls but below what most users expect from video conferencing apps by 2026. For use cases where the video quality needs to feel premium (healthcare consultations, financial advisory), this may be a limiting factor.

PM action: Benchmark user acceptance at 720p before committing to a resolution requirement.

Real-time compute cost

Running a 0.43-second latency video model on H100s is expensive per conversation-minute. The cost model for Griffin will be unlike text or voice AI: expect per-minute pricing, not per-token. Run your economics before scoping a product.

PM action: Get pricing from Tavus before making product commitments. Unit economics may not support all use cases.

What AI PMs Should Do Right Now

Griffin-Lite is in preview and not production-ready for most teams. That is not a reason to ignore it. It is a reason to get ahead of it before the GA release.

1

Audit your product roadmap for video touchpoints

Which customer interactions in your product currently require a human video call that could be AI-native? Map these explicitly. The answer shapes how you should position Griffin when it becomes generally available.

2

Draft your disclosure UX before you need it

Design the disclosure interaction before you start building. What does it look like? Where does it appear? What language do you use? Getting this right takes iteration. Starting early means you are not shipping rushed disclosure copy under time pressure.

3

Talk to your legal team now

Your legal team needs to understand that video-native AI with a 48% human pass rate exists and that you are evaluating it. This is not a normal AI feature. Get their input on your specific jurisdiction, industry, and use case before you write a line of product spec.

4

Watch the independent benchmarks

Tavus's numbers are self-reported. The 48% Turing pass rate comes from their own study. Wait for independent academic or third-party replication before relying on these figures in internal presentations or product decisions.

5

Apply to the waitlist with a specific use case

General interest will not get you early access. Tavus is selecting trusted testers. Articulate a specific, responsible use case: the user population, the disclosure design, and the business context. Researchers and teams with clear safety stories will get in first.

The Bigger Picture: What Video-Native AI Changes

Griffin is not just a better video avatar. It represents a shift in what the human-computer interaction layer looks like for AI products.

Voice AI was a stepping stone

Voice-based AI products like customer service bots and voice assistants succeeded in part because they removed the visual uncanny valley. Griffin collapses that valley. Video-native AI will outperform voice in trust-sensitive contexts once the latency and cost curves improve.

The attention economy gets a new variable

Video is the highest-fidelity communication channel humans use. Products that can hold a user's attention through a face-to-face interaction have qualitatively different engagement curves than chat or voice. Griffin opens a category that did not exist six months ago.

Trust architecture becomes a product differentiator

Companies that figure out how to earn and maintain user trust in face-to-face AI interactions will have a durable advantage. Companies that deploy it without disclosure will create backlash that damages the category for everyone. Trust architecture is now a product skill.

Accessibility applications are underexplored

Sign language interpretation, communication support for people with speech differences, or video-based interfaces for users who cannot easily use text or voice. These are high-value, high-trust applications where the disclosure obligation is a feature, not a constraint.

Stay Ahead of What AI Products Can Do

The AI PM Masterclass covers emerging AI capabilities, product design, and how to bring new interaction models to market responsibly. Taught live by a Salesforce Sr. Director PM.

Before you go: get the AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.