OmniHuman
ByteDance's model that turns a single image plus an audio track into a lip-synced, full-body talking avatar video.
Overview
OmniHuman is ByteDance's digital-human model that generates realistic talking and singing avatar videos from a single reference image plus an audio track, with full-body motion, gesture, and lip-sync. OmniHuman-1.5 (2026) supports HD output, clips up to about 30 seconds, optional text guidance for camera and gesture, and even dual-person audio. It powers avatar and spokesperson features in ByteDance's Dreamina and CapCut apps and is available through APIs such as fal. It is intended for consented avatar, spokesperson, and localization use; because face-reenactment tools can be misused for non-consensual deepfakes, ByteDance gates access and users must have rights to the likenesses and voices they animate.
Pricing
Pricing shown for reference only. These figures reflect RECATOOLS research as of 4 Sep 2026 and may be out of date or incomplete. This is not financial or purchasing advice — always confirm the current price on the provider’s official website before making any decision.
What you can produce with OmniHuman
- A full-body, gesture-aware, lip-synced video from a single photograph plus an audio track
- Motion noticeably more natural than earlier talking-head tools, credited to its "omni-conditions" training
- Consented spokesperson, presenter and localisation clips — the use the gated access is explicitly framed around
- ⚠️ Rights you must actually hold, in the likeness AND the voice you animate. This is not paperwork: coverage called it possibly the most realistic deepfake algorithm yet, and commentators tie it directly to non-consensual deepfake pornography, election-cycle misinformation and impersonation fraud, with ByteDance conceding that regulating the technology remains an open problem
- Short clips only — roughly 30 seconds in the 1.5 release — so it suits avatar and localisation snippets rather than long-form production
- Access mostly through ByteDance's own Dreamina and CapCut apps, region- and quota-limited, or paid APIs such as fal. It is closed, with no self-hosting
What this is for: Turning one portrait and an audio track into a lip-synced, full-body talking or singing avatar video for consented spokesperson, presenter, and localization content.
Who this is for: Marketers, educators, and creators producing avatar videos who hold the rights to the person's likeness and voice.
Availability: Available in ByteDance's Dreamina and CapCut apps and via APIs (e.g. fal); freemium with paid credits. Closed/proprietary; access is gated and misuse for non-consensual likenesses is prohibited.
What people say
OmniHuman's reception is inseparable from the deepfake conversation, and honesty requires leading with it. When ByteDance published the work, coverage from TechSpot, AOL, and others called it possibly "the most realistic deepfake algorithm yet," and Northwestern's Matt Groh said realism had "reached a whole new level." Technically the results are striking—a single photo plus an audio track yields full-body, gesture-aware, lip-synced video—and reviewers credit the "omni-conditions" training for motion far more natural than earlier talking-head tools.
The dominant criticism is misuse risk, and it is serious rather than hypothetical. Commentators tie OmniHuman directly to real harms—non-consensual deepfake pornography, election-cycle misinformation, and impersonation fraud—and note that ByteDance has said it is building control mechanisms while conceding that regulating this technology remains an open problem. That's why access is gated and framed around consented spokesperson, presenter, and localization use, with users required to hold rights to any likeness and voice they animate.
More practical caveats: it's closed and reaches most users only through ByteDance's own Dreamina and CapCut apps (region- and quota-limited) or paid APIs like fal, and outputs are capped at short clips (roughly 30 seconds in the 1.5 release), so it suits avatar and localization snippets rather than long-form production.
Summary of public user & expert reviews, compiled by RECATOOLS.
About this listing
This entry was compiled from publicly available data including OmniHuman's official website, press releases, documentation, and reputable third-party publications. RECATOOLS is not affiliated with OmniHuman unless explicitly stated.
Third-party AI tools update their pricing, features, availability, and policies frequently. Information here may be outdated by the time you read this — we make reasonable efforts to keep listings current, but cannot guarantee absolute accuracy.
For the latest details, please refer to OmniHuman directly →
Spotted something out of date? Suggest an update →
More in Video & Audio