AI Avatars Keep Falling Apart on Camera — Mirage Avatar X Claims It Fixed That

Captions launched Mirage Avatar X, an AI avatar model that claims to fix the consistency and expression problems that have plagued digital twins since the category started.

Share
AI Avatars Keep Falling Apart on Camera — Mirage Avatar X Claims It Fixed That

Captions launched Mirage Avatar X, a new AI avatar model built to fix a problem that's plagued the category since it started: avatars that hold up for the first few seconds, then visibly break as a video runs longer.

The model creates a digital twin — face, voice, and mannerisms — from a single 10-second video clip recorded directly in the Captions app. No extra photos required, though users can add reference images to sharpen quality if they want.

What Avatar X Actually Changes

The core claim is consistency. Captions says Avatar X holds the same fidelity at the 59-second mark as it does at the first second, while showing side-by-side comparisons of a competitor's avatar losing lip-sync accuracy and mouth-shape quality as a video progresses.

The model also targets non-verbal expression, an area Captions says most avatar platforms handle poorly. Laughing, pausing, gasping, and smirking are called out specifically as moments where competing avatars tend to either freeze or keep mouthing words through actions that shouldn't involve speech at all.

Avatar X supports both vertical and horizontal video from the same twin, and generations run up to three minutes long. Voice cloning comes from the same 10-second clip via Captions' Mirage Audio model, and the resulting twin can speak more than 30 languages through AI dubbing and translation.

How the Setup Compares

Captions frames the 10-second requirement as a meaningful advantage over competitors, citing HeyGen's published requirement of 15 seconds and Synthesia's 1-to-5-minute range. Twins are described as consent-first — Captions requires users to confirm they have the right to create a twin of themselves before generating one.

Competitive Context

The AI avatar space has grown crowded fast, with HeyGen and Synthesia as the most established names alongside newer entrants. Runway, Kling, and other AI video platforms have also pushed toward better consistency and identity preservation over the past year, though most of that work has focused on generated characters rather than personal avatars built from a user's own likeness.

Avatar X's positioning — shorter setup time, stronger claimed consistency, non-verbal expression handling — targets the specific pain points that have made avatar tools feel unreliable for longer-form content like tutorials, presentations, or thought-leadership videos.

The Signal in the Noise

Every comparison Captions has published so far is self-reported: Captions' own side-by-side clips against unnamed or lightly-identified "leading competitor" footage. That's standard for a launch campaign, but it means the "no degradation" and expression-handling claims haven't been verified independently yet.

That said, the specific failure modes being addressed — mouths continuing to move during laughter, dead-eyed pauses, mid-video quality drop — are real, commonly cited complaints about avatar tools in general, not invented problems. Whether Avatar X actually solves them at scale, versus in a curated demo, is the thing worth watching for as more people get hands-on access.

Specs & Pricing

  • Minimum footage required: 10 seconds (single continuous clip)
  • Output formats: vertical and horizontal, same twin
  • Max generation length: up to 3 minutes
  • Voice: cloned from the same 10-second clip via Mirage Audio
  • Language support: 30+ languages for dubbing/translation, 100+ for captions
  • Pricing: not listed on the launch page — available through Captions' standard plans at captions.ai/pricing

Resources & Reads