IP Library › Granted Patent US 12,651,459
Granted Patent B2
US 12,651,459 · App. 17/326,091 · Granted Jun 9, 2026

Synthesizing video from audio using one or more neural networks

Inventors: Ming-Yu Liu (San Jose, CA); Ting-Chun Wang (Santa Clara, CA); Arun Mallya (San Jose, CA)
Assignee: NVIDIA Corporation
G06V20/46G06N3/04G06V40/169G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,459
App. No.
17/326,091
Granted
Jun 9, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques are presented to reduce an amount of data to be transmitted for media content. In at least one embodiment, one or more neural networks are used to generate video and audio information corresponding to one or more people based, at least in part, on at least one image and voice information corresponding to the one or more people.

Claims (63)

1 . One or more processors, comprising:

circuitry to:

receive one or more transmissions of speech data representing speech of one or more participants of a live transmission;

receive a first image of the one or more participants;

receive an indication of a change to one or more characteristics of the one or more participants in the first image;

receive a second image of the one or more participants, that has been transmitted in response to the indication of the change;

generate, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and

cause the one or more video frames representing the one or more participants of the live transmission to be presented.

2 . The one or more processors of claim 1 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.

3 . The one or more processors of claim 1 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.

4 . The one or more processors of claim 1 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by the one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.

5 . The one or more processors of claim 1 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.

6 . The one or more processors of claim 1 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the live transmission.

7 . A system comprising:

one or more processors to:

receive one or more transmissions of speech data representing speech of one or more participants of a live transmission;

receive a first image of the one or more participants;

receive an indication of a change to one or more characteristics of the one or more participants in the first image;

receive a second image of the one or more participants, that has been transmitted in response to the indication of the change;

generate, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and

cause the one or more video frames representing the one or more participants of the live transmission to be presented.

8 . The system of claim 7 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.

9 . The system of claim 7 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.

10 . The system of claim 7 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by the one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.

11 . The system of claim 7 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.

12 . The system of claim 7 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the live transmission of speech data.

13 . A method comprising:

receiving one or more transmissions of speech data representing speech of one or more participants of a live transmission;

receiving a first image of the one or more participants;

receiving an indication of a change to one or more characteristics of the one or more participants in the first image;

receiving a second image of the one or more participants, that has been transmitted in response to the indication of the change;

generating, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and

causing the one or more video frames representing the one or more participants of the live transmission to be presented.

14 . The method of claim 13 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.

15 . The method of claim 13 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.

16 . The method of claim 13 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.

17 . The method of claim 13 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of the speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.

18 . The method of claim 13 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the live transmission of the speech data.

19 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

receive one or more transmissions of speech data representing speech of one or more participants of a live transmission;

receive a first image of the one or more participants;

receive an indication of a change to one or more characteristics of the one or more participants in the first image;

receive a second image of the one or more participants, that has been transmitted in response to the indication of the change;

generate, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and

cause the one or more video frames representing the one or more participants of the live transmission to be presented.

20 . The non-transitory machine-readable medium of claim 19 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.

21 . The non-transitory machine-readable medium of claim 19 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.

22 . The non-transitory machine-readable medium of claim 19 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by the one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.

23 . The non-transitory machine-readable medium of claim 19 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.

24 . The non-transitory machine-readable medium of claim 19 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the live transmission.

25 . A video communication system, comprising:

one or more processors to receive one or more transmissions of speech data representing speech of one or more participants of a live transmission;

receive a first image of the one or more participants;

receive an indication of a change to one or more characteristics of the one or more participants since the first image;

receive a second image of the one or more participants, that has been transmitted in response to the indication of the change;

generate, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and

cause the one or more video frames representing the one or more participants of the live transmission to be presented; and

store one or more network parameters for the one or more neural networks in one or more memories.

26 . The video communication system of claim 25 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.

27 . The video communication system of claim 25 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.

28 . The video communication system of claim 25 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by the one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.

29 . The video communication system of claim 25 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.

30 . The video communication system of claim 25 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the one or more received transmissions of speech data.

Assignments (2)
CONFIRMATORY ASSIGNMENT Recorded Mar 12, 2026
From: LIU, MING-YU; WANG, TING-CHUN; MALLYA, ARUN
To: NVIDIA CORPORATION
Reel/Frame 075109/0398 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2021
From: LIU, MING-YU; WANG, TING; MALLYA, ARUN
To: NVIDIA CORPORATION
Reel/Frame 056332/0516 →
Continuity (1)
Related Publication 20220374637A1 · Nov 24, 2022
References Cited (58)
US 11144764B1 · Sharma · 2021 [cited by examiner]
US 11295501B1 · Biswas · 2022 [cited by examiner]
US 11455790B2 · Karras et al. · 2022 [cited by applicant]
US 11568227B1 · Ko · 2023 [cited by examiner]
US 11580395B2 · Karras et al. · 2023 [cited by applicant]
US 20210021949A1 · Sridharan et al. · 2021 [cited by applicant]
US 20210133990A1 · Eckart et al. · 2021 [cited by applicant]
US 20210329306A1 · Liu et al. · 2021 [cited by applicant]
US 20220171960A1 · Nelson · 2022 [cited by examiner]
US 20220351439A1 · Chae · 2022 [cited by examiner]
US 20220358703A1 · Chae · 2022 [cited by examiner]
US 20220366927A1 · Pishehvar · 2022 [cited by examiner]
US 20220375224A1 · Chae · 2022 [cited by examiner]
US 20220398793A1 · Chae · 2022 [cited by examiner]
US 20220399025A1 · Chae · 2022 [cited by examiner]
CN 106559636A · 2017 [cited by applicant]
CN 108447474A · 2018 [cited by applicant]
CN 108831463A · 2018 [cited by applicant]
CN 109377539A · 2019 [cited by applicant]
CN 110677598A · 2020 [cited by examiner]
CN 110751708A · 2020 [cited by applicant]
CN 110866968A · 2020 [cited by applicant]
CN 111225237A · 2020 [cited by applicant]
CN 111243626A · 2020 [cited by applicant]
CN 112102329A · 2020 [cited by applicant]
CN 112750185A · 2021 [cited by applicant]
CN 112752118A · 2021 [cited by applicant]
EP 3996047A1 · 2022 [cited by applicant]
JP 2021009669A · 2021 [cited by applicant]
KR 20060090687A · 2006 [cited by applicant]
KR 20200112647A · 2020 [cited by applicant]
KR 20200145700A · 2020 [cited by applicant]
WO 2017137948A1 · 2017 [cited by applicant]
WO 2020256475A1 · 2020 [cited by applicant]
WO WO2020256472A1 · 2020 [cited by examiner]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Liao et al., “Speech2Video Synthesis with 3D Skeleton Regularization and Expressive Body Poses,” ACCV 2020, 16 pages. [cited by applicant]
Siarohin et al., “First Order Motion Model for Image Animation,” Neural Information Processing Systems, 2019, 11 pages. [cited by applicant]
Thies et al., “Neural Voice Puppetry: Audio-driven Facial Reenactment,” ECCV, 2020, 12 pages. [cited by applicant]
Wang et al., “One-Shot Free-View Neural Talking-Head Synthesis for Video Conferencing,” CVPR, 2021, 11 pages. [cited by applicant]
Wen et al., “Photorealistic Audio-driven Video Portraits,” retrieved from https://richardt.name/publications/audio-dvp/, 2020, 10 pages. [cited by applicant]
Yao et al., “Iterative Text-based Editing of Talking heads Using Neural Retargeting,” Nov. 21, 2020, 14 pages. [cited by applicant]
Zhou et al., “MakeItTalk: Speaker-Aware Talking-Head Animation,” Oct. 7, 2020, 15 pages. [cited by applicant]
United Kingdom Combined Search and Examination Report for Application No. GB2207450.4, mailed Nov. 25, 2022, 6 pages. [cited by applicant]
Vougioukas et al., “Realistic Speech-Driven Facial Animation with GANs,” International Journal of Computer Vision, 2020, 16 pages. [cited by applicant]
Office Action for Chinese Application No. 202210530954.8, mailed Jan. 6, 2024, 9 pages. [cited by applicant]
Karras et al., “Audio-Driven Facial Animation by Joint End-to-End Learning of Pose and Emotion,” ACM Transactions on Graphics vol. 36, No. 4, Jul. 30, 2017, 12 Pages. [cited by applicant]
Nvidia, “Omniverse Audio2Face,” Retrieved from, https://www.nvidia.com/en-us/omniverse/apps/audio2face/, undated, 13 Pages. [cited by applicant]
Prenger et al., “WaveGlow: A Flow-based Generative Network for Speech Synthesis,” Oct. 31, 2018, 5 Pages. [cited by applicant]
Valle et al., Flowtron: an Autoregressive Flow-based Generative Network for Text-to-Speech Synthesis, Jul. 16, 2020, 10 Pages. [cited by applicant]
Valle et al., “Mellotron: Multispeaker expressive voice synthesis by conditioning on rhythm, pitch and global style tokens,” Oct. 26, 2019, 5 Pages. [cited by applicant]
Wang et al., “Few-Shot Video-to-Video Synthesis,” Advances in Neural Information Processing Systems, Oct. 28, 2019, 14 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2207450.4, mailed Dec. 5, 2023, 3 pages. [cited by applicant]
Office Action for Chinese Application No. 202210530954.8, mailed Nov. 25, 2024, 20 pages. [cited by applicant]
Decision of Rejection for Chinese Application No. 202210530954.8, mailed Feb. 19, 2025, 15 pages. [cited by applicant]
Combined Search and Examination Report for United Kingdom Application No. GB2418263.6, mailed May 2, 2025, 6 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2207450.4, mailed Jul. 3, 2024, 3 pages. [cited by applicant]
Office Action for Chinese Application No. 202210530954.8, mailed Jun. 21, 2024, 14 pages. [cited by applicant]