IP Library Granted Patent US 12,367,628
Granted Patent B2
US 12,367,628 · App. 17/903,585 · Granted Jul 22, 2025

Speech-driven animation using one or more neural networks

Inventors: Siddharth Gururani (Santa Clara, CA); Arun Mallya (Mountain View, CA); Ting-Chun Wang (Santa Clara, CA); Jose Rafael Valle da Costa (Berkley, CA); Ming-Yu Liu (San Jose, CA)
Assignee: NVIDIA Corporation
G06T13/205G06V10/82G06V40/171
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,628
App. No.
17/903,585
Granted
Jul 22, 2025
Kind
B2
Abstract

Apparatuses, systems, and techniques are presented to generate digital content. In at least one embodiment, one or more neural networks are used to generate video information based at least in part upon voice information and a combination of image features and facial landmarks corresponding to one or more images of a person.

Claims (42)

1. A processor, comprising:

one or more circuits to;

predict facial landmarks of a person uttering a speech based, at least in part, upon voice information of the speech and one or more images of the person;

predict keypoints based, at least in part, upon the predicted facial landmarks and the voice information; and

generate, using one or more neural networks, video information of the person uttering the speech based, at least in part, upon the predicted keypoints.

2. The processor of claim 1 , wherein the one or more circuits are further to use a first neural network to determine a set of mouth landmarks and a second neural network to determine a set of face landmarks for the person corresponding to the voice information, wherein the predicted facial landmarks include both the set of face landmarks and the set of mouth landmarks.

3. The processor of claim 2 , wherein the one or more circuits are further to use a keypoint determination neural network to predict the keypoints for the person corresponding to the voice information based, at least in part, upon the predicted facial landmarks.

4. The processor of claim 3 , wherein the one or more circuits are further to use a generative neural network to generate the video information based, at least in part, upon the voice information and the predicted keypoints, and further upon pose information predicted by a pose generation network.

5. The processor of claim 1 , wherein the one or more circuits are further to generate a set of frontalized landmarks normalized for a face and mouth of the person to provide as input, along with a feature representation of the voice information, to at least one landmark prediction neural network.

6. The processor of claim 1 , wherein the one or more circuits are further to use at least one emotion prediction module to predict one or more emotional states for individual frames of the video information and generate the video information to reflect the predicted emotional states.

7. A system, comprising:

one or more processors to:

predict facial landmarks of a person uttering a speech based, at least in part, upon voice information of the speech and one or more images of the person;

predict keypoints based, at least in part, upon the predicted facial landmarks and the voice information; and

generate, using one or more neural networks, video information of the person uttering the speech based, at least in part, upon the predicted keypoints.

8. The system of claim 7 , wherein the one or more processors are further to use a first neural network to determine a set of mouth landmarks and a second neural network to determine a set of face landmarks for the person corresponding to the voice information, wherein the predicted facial landmarks include both the set of face landmarks and the set of mouth landmarks.

9. The system of claim 8 , wherein the one or more processors are further to use a keypoint determination neural network to predict the keypoints for the person corresponding to the voice information based, at least in part, upon the predicted facial landmarks.

10. The system of claim 9 , wherein the one or more processors are further to use a generative neural network to generate the video information based, at least in part, upon the voice information and the predicted keypoints, and further upon pose information predicted by a pose generation network.

11. The system of claim 7 , wherein the one or more processors are further to generate a set of frontalized landmarks normalized for a face and mouth of the person to provide as input, along with a feature representation of the voice information, to at least one landmark prediction neural network.

12. The system of claim 7 , wherein the one or more processors are further to use at least one emotion prediction module to predict one or more emotional states for individual frames of the video information and generate the video information to reflect the predicted emotional states.

13. A method, comprising:

predicting facial landmarks of a person uttering a speech based, at least in part, upon voice information of the speech and one or more images of the person;

predicting keypoints based, at least in part, upon the predicted facial landmarks and the voice information; and

generating, using one or more neural networks, video information of the person uttering the speech based, at least in part, upon the predicted keypoints.

14. The method of claim 13 , further comprising: using a first neural network to determine a set of mouth landmarks and a second neural network to determine a set of face landmarks for the person corresponding to the voice information, wherein the predicted facial landmarks include both the set of face landmarks and the set of mouth landmarks.

15. The method of claim 14 , further comprising: using a keypoint determination neural network to predict the keypoints for the person corresponding to the voice information based, at least in part, upon the predicted facial landmarks.

16. The method of claim 15 , further comprising: using a generative neural network to generate the video information based, at least in part, upon the voice information and the predicted keypoints, and further upon pose information predicted by a pose generation network.

17. The method of claim 13 , further comprising:

generating a set of frontalized landmarks normalized for a face and mouth of the person to provide as input, along with a feature representation of the voice information, to at least one landmark prediction neural network.

18. The method of claim 13 , further comprising:

using at least one emotion prediction module to predict one or more emotional states for individual frames of the video information and generate the video information to reflect the predicted emotional states.

19. A video generation system, comprising:

one or more processors to:

predict facial landmarks of a person uttering a speech based, at least in part, upon voice information of the speech and one or more images of the person;

predict keypoints based, at least in part, upon the predicted facial landmarks and the voice information; and

generate, using one or more neural networks, video information of the person uttering the speech based, at least in part, upon the predicted keypoints; and

memory for storing network parameters for the one or more first neural networks.

20. The video generation system of claim 19 , wherein the one or more processors are further to use a first neural network to determine a set of mouth landmarks and a second neural network to determine a set of face landmarks for the person corresponding to the voice information, wherein the predicted facial landmarks include both the set of face landmarks and the set of mouth landmarks.

21. The video generation system of claim 20 , wherein the one or more processors are further to use a keypoint determination neural network to predict the keypoints for the person corresponding to the voice information based, at least in part, upon the predicted facial landmarks.

22. The video generation system of claim 21 , wherein the one or more processors are further to use a generative neural network to generate the video information based, at least in part, upon the voice information and the predicted keypoints, and further upon pose information predicted by a pose generation network.

23. The video generation system of claim 19 , wherein the one or more processors are further to generate a set of frontalized landmarks normalized for a face and mouth of the person to provide as input, along with a feature representation of the voice information, to at least one landmark prediction neural network.

24. The video generation system of claim 19 , wherein the one or more processors are further to use at least one emotion prediction module to predict one or more emotional states for individual frames of the video information and generate the video information to reflect the predicted emotional states.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: GURURANI, SIDDHARTH; MALLYA, ARUN; WANG, TING-CHUN; VALLE DA COSTA, JOSE RAFAEL; LIU, MING-YU
To: NVIDIA CORPORATION
Reel/Frame 065153/0609 →
Continuity (1)
Related Publication 20240144568A1 · May 2, 2024
References Cited (18)
US 20100082345A1 · Wang · 2010 [cited by examiner]
US 20210056292A1 · Mao · 2021 [cited by examiner]
US 20210202090A1 · O'Donovan · 2021 [cited by examiner]
US 20210390748A1 · Liao · 2021 [cited by examiner]
US 20220103860A1 · Demyanov · 2022 [cited by examiner]
US 20220392131A1 · Li · 2022 [cited by examiner]
US 20230049729A1 · Berlin · 2023 [cited by examiner]
US 20230171380A1 · Oz · 2023 [cited by examiner]
US 20240013463A1 · Assouline · 2024 [cited by examiner]
Jiaqi Hao, Shiguang Liu, and Qing Xu. 2021. Controlling Eye Blink for Talking Face Generation via Eye Conversion. In SIGGRAPH Asia 2021 Technical Communications (SA '21). Association for Computing Machinery, New York, N… [cited by examiner]
Prateek Manocha and Prithwijit Guha, “Facial Keypoint Sequence Generation from Audio,” https://arxiv.org/abs/2011.01114 (Year: 2020). [cited by examiner]
IEEE “IEEE Standard for Floating-Point Arithmetric”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008. [cited by applicant]
International Electrotechnical Commission, “Functional safety of electrical/electronic/programmable electronic safety-related systems,” IEC Standard 61508-1, Apr. 2014, 23 pages. [cited by applicant]
International Organization for Standardization, “Road vehicles—Functional safety,” ISO Standard 26262, https://www.iso.org/obp/ui/#iso:std:iso:26262:-1:ed-1:v1:en, Nov. 11, 2011, 35 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Standard No. J3016-201806, dated Jun. … [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Wang et al., “Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion,” Jul. 20, 2021, 8 Pages. [cited by applicant]
Zhou et al., “MakeItTalk: Speaker-Aware Talking-Head Animation,” Oct. 7, 2020, 15 pages. [cited by applicant]