Omniverse Audio2Face

NVIDIAPage's Omniverse Audio2Face is an AI powered application that simplifies animation of a 3D character to match any voice-over track. I...

Last verified:

Visit Omniverse Audio2Face

What is Omniverse Audio2Face?

NVIDIA Omniverse Audio2Face is an AI-powered application that instantly creates expressive facial animation from just an audio source using generative AI. It is a foundation application for animating 3D characters' facial characteristics to match any voice-over track, whether for games, films, real-time digital assistants, or personal use. The app is built on Universal Scene Description (OpenUSD) and can be used for interactive real-time applications or as a traditional facial animation authoring tool.

Audio2Face comes preloaded with

Omniverse Audio2Face pricing

Pricing model: Free

Audio2Face is totally free to use, like all of NVIDIA Omniverse. It is available in open beta with no cost. Enterprise support is available as an optional paid option - users need to contact Enterprise sales for support pricing. There is no separate paid plan for the application itself; the free tier includes full access to all Audio2Face features.

Omniverse Audio2Face pros

  • Generates expressive facial animation from just audio using generative AI
  • Real-time facial animation driven by pre-trained Deep Neural Network
  • Preloaded with Digital Mark 3D character for instant starting
  • Live mode supports microphone input for real-time animation
  • Retargets to any 3D human or human-esque face, realistic or stylized
  • Supports multiple instances with multiple characters in same scene
  • Automatically animates emotions matching selected emotional range
  • AI infers emotion directly from audio clip automatically
  • Animates eyes, mouth, tongue, and head motion
  • Processes any language with continual language updates
  • Export-import support with Blendshapes for Blender and Unreal Engine
  • Blendshape conversion and blendweight export options available
  • Batch output multiple animation files from multiple audio sources
  • OpenUSD-based for integration with Omniverse pipeline
  • Dial up or down facial expression intensity per character
  • Swap human or animal characters on the fly with few clicks

Omniverse Audio2Face cons

  • Requires RTX GPU with minimum 8 GB VRAM
  • Only supports Windows 10 (1903+) and Ubuntu Linux 20.04+
  • Needs Omniverse Nucleus installed
  • 16 GB RAM minimum, 32 GB recommended
  • 250 GB SSD storage minimum required
  • look may occur without additional eye animation
  • Head mesh must be broken into individual mesh components
  • UI can become unusable if not topmost window on some setups

Frequently asked questions about Omniverse Audio2Face

What is Omniverse Audio2Face?

Omniverse Audio2Face is an AI-powered application that generates expressive facial animation and lip sync from just an audio source using generative AI. It is a foundation application for animating 3D characters' facial characteristics to match any voice-over track for games, films, real-time digital assistants, or personal use.

How does Audio2Face work?

Audio2Face feeds audio input into a pre-trained Deep Neural Network. The output of the network drives the facial animation of 3D characters in real-time by moving the 3D vertices of your character mesh to create expressions that match the audio.

Can I use my own 3D character with Audio2Face?

Yes, Audio2Face lets you retarget animations to any 3D human or human-esque face, whether realistic or stylized. Character Setup and Transfer provides the means to re-target the performance of the entire face including Eyes, Teeth, and Tongue from the default model to your own characters.

Does Audio2Face support live microphone input?

Yes, Audio2Face includes a Live mode that allows you to use a microphone to drive the application in real-time, generating facial animations live as you speak.

What languages does Audio2Face support?

Audio2Face can process any language easily. NVIDIA is continually updating the application with more and more languages.

Can I run multiple characters at the same time?

Yes, you can run multiple instances of Audio2Face with as many characters in a scene as you like. All characters can be animated from the same audio track or different audio tracks simultaneously.

What export formats does Audio2Face support?

Audio2Face supports export as USD Cache, blendshape conversion, and blendweight export. It now supports export-import with Blendshapes for Blender and Epic Games Unreal Engine through Omniverse Connectors.

What are the system requirements for Audio2Face?

Minimum requirements include Windows 10 (1903+) or Ubuntu Linux 20.04+, any RTX GPU with 8 GB VRAM, 16 GB RAM, and 250 GB SSD. Recommended specs are Intel Core i7 13th Series or AMD Ryzen 7 7th Series CPU, 32 GB RAM, 500 GB SSD, and GeForce RTX 4070 Ti or RTX A4500 or higher GPU.

How do I control emotions in Audio2Face?

Audio2Face gives you the ability to choose and animate your character's emotions instantly. The AI network automatically manipulates the face, eyes, mouth, tongue, and head motion to match your selected emotional range and customized intensity level, or it can automatically infer emotion directly from the audio clip.

Is Audio2Face free to use?

Yes, Audio2Face is totally free to use like all of NVIDIA Omniverse. It is available in open beta with no cost. Enterprise support is available as an optional paid service - contact Enterprise sales for support pricing.

Categories

Use cases

Browse all AI tools on NeedAnAI