HeyGen has introduced Avatar V, an AI video model designed to create a photorealistic digital twin from a single 15-second webcam or phone recording. The company says the tool can capture a user’s face, motion patterns, and voice profile, then use that identity to generate unlimited studio-style videos without professional production gear.
The launch was announced on April 8, and HeyGen’s post on X drew 472,000 views. The product is aimed at a familiar weakness in earlier AI avatar systems: they often looked convincing for a brief clip, then lost consistency as the video ran longer. HeyGen says Avatar V was built to preserve identity across the full duration of a generated video, whether the output is a short clip or a much longer module.
Identity comes from motion, appearance comes from a photo
According to HeyGen, Avatar V is trained on what it calls a temporally grounded identity embedding created from the 15-second source clip. That recording is used to capture micro-expressions, lip geometry, facial silhouette, gesture patterns, and expression transitions that make a person recognizable across different scenes.
The company separates identity from appearance. The short video defines how a person moves, while a separate base photo defines how that person looks. After that, users can change outfits, settings, and styles through text prompts while keeping the motion signature consistent. HeyGen says wide shots, medium shots, and close-ups remain stable from the same original recording.
175-language output with optional voice cloning
The workflow is straightforward. Users first record a 15-second clip, then optionally create a standalone voice clone, and finally choose a base photo that serves as the identity reference for future scenes. From there, they can generate new videos with prompt-based styling or use assets from the HeyGen library.
Avatar V supports output in 175 languages with automatic lip-sync adapted to the target language. Voice cloning is not required, but HeyGen says it improves realism. The company also recommends that users be expressive during recording, because the source clip strongly affects the quality of the generated result.
Now serving as the base layer across HeyGen’s platform
HeyGen says Avatar V now acts as the foundation for the rest of its platform features. It has also been integrated with Seedance 2.0 for more cinematic video generation. Access is available through the company’s paid subscription plans, alongside templates, translation tools, and studio features.
On its launch page, HeyGen describes the model with a simple standard: the output should be good enough that a user would be willing to put their own name on it. That framing puts the focus on sustained identity quality over flashy short demos, which is where many earlier avatar products tended to break down.

