IP Library › Patent Application 18527668
Patent Application
App. No. 18/527,668

TWO-STAGE FRAMEWORK FOR ZERO-SHOT IDENTITY-AGNOSTIC TALKING-HEAD GENERATION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/527,668
Abstract

Methods, systems, apparatuses, devices, and computer program products are described. A system may input a first audio stream (e.g., audio recording) and a corresponding text sting into a machine learning model. The first audio stream and the text string may correspond to a first identity (e.g., person). Based on an output of the machine learning model, the system may generate a second audio stream associated with a second identity and mimics the first audio steam. For example, the second audio stream may be a generated recording of the second identity speaking the first text string. In addition, the system may generate a video depicting the second identity speaking the first text string (e.g., the second audio stream) based on combining the second audio stream with some image or previous video of the second identity. For example, the system may generate the video based on generating a head motion sequence.

Claims (50)

1 . A method for data generation, comprising:

inputting a first audio stream and a first text string corresponding to the first audio stream into a machine learning model, wherein the first audio stream and the first text string correspond to a first identity;

generating a second audio stream based at least in part on an output of the machine learning model, wherein the second audio stream is associated with a second identity and mimics the first audio stream; and

generating a video that displays the second identity speaking the first text string based at least in part on combining the second audio stream with a visual medium associated with the second identity.

2 . The method of claim 1 , further comprising:

training the machine learning model based at least in part on a set of audio streams and a set of text strings, wherein the set of audio streams and the set of text strings correspond to a plurality of identifiers.

3 . The method of claim 1 , further comprising:

identifying a first set of features associated with the first audio stream;

identifying a second set of features associated with the first text string; and

generating a head motion sequence corresponding to the second identity speaking the first text string based at least in part on the first set of features and the second set of features.

4 . The method of claim 3 , wherein generating the video comprises:

generating the video based at least in part on the generated head motion sequence.

5 . The method of claim 1 , further comprising:

identifying a set of characteristics associated with the visual medium, wherein the set of characteristics include one or more geometric parameters and one or more appearance characteristics associated with a head motion of the second identity.

6 . The method of claim 5 , wherein generating the video comprises:

generating the video based at least in part on combining the set of characteristics with the first audio stream.

7 . An apparatus for data generation, comprising:

one or more memories storing processor-executable code; and

one or more processors coupled with the one or more memories and individually or collectively operable to execute the code to cause the apparatus to:

input a first audio stream and a first text string corresponding to the first audio stream into a machine learning model, wherein the first audio stream and the first text string correspond to a first identity;

generate a second audio stream based at least in part on an output of the machine learning model, wherein the second audio stream is associated with a second identity and mimics the first audio stream; and

generate a video that displays the second identity speaking the first text string based at least in part on combining the second audio stream with a visual medium associated with the second identity.

8 . The apparatus of claim 7 , wherein the one or more processors are individually or collectively further operable to execute the code to cause the apparatus to:

train the machine learning model based at least in part on a set of audio streams and a set of text strings, wherein the set of audio streams and the set of text strings correspond to a plurality of identifiers.

9 . The apparatus of claim 7 , wherein the one or more processors are individually or collectively further operable to execute the code to cause the apparatus to:

identify a first set of features associated with the first audio stream;

identify a second set of features associated with the first text string; and

generate a head motion sequence corresponding to the second identity speaking the first text string based at least in part on the first set of features and the second set of features.

10 . The apparatus of claim 9 , wherein, to generate the video, the one or more processors are individually or collectively operable to execute the code to cause the apparatus to:

generate the video based at least in part on the generated head motion sequence.

11 . The apparatus of claim 7 , wherein the one or more processors are individually or collectively further operable to execute the code to cause the apparatus to:

identify a set of characteristics associated with the visual medium, wherein the set of characteristics include one or more geometric parameters and one or more appearance characteristics associated with a head motion of the second identity.

12 . The apparatus of claim 11 , wherein, to generate the video, the one or more processors are individually or collectively operable to execute the code to cause the apparatus to:

generate the video based at least in part on combining the set of characteristics with the first audio stream.

13 . A non-transitory computer-readable medium storing code for data generation, the code comprising instructions executable by one or more processors to:

input a first audio stream and a first text string corresponding to the first audio stream into a machine learning model, wherein the first audio stream and the first text string correspond to a first identity;

generate a second audio stream based at least in part on an output of the machine learning model, wherein the second audio stream is associated with a second identity and mimics the first audio stream; and

generate a video that displays the second identity speaking the first text string based at least in part on combining the second audio stream with a visual medium associated with the second identity.

14 . The non-transitory computer-readable medium of claim 13 , wherein the instructions are further executable by the one or more processors to:

train the machine learning model based at least in part on a set of audio streams and a set of text strings, wherein the set of audio streams and the set of text strings correspond to a plurality of identifiers.

15 . The non-transitory computer-readable medium of claim 13 , wherein the instructions are further executable by the one or more processors to:

identify a first set of features associated with the first audio stream;

identify a second set of features associated with the first text string; and

generate a head motion sequence corresponding to the second identity speaking the first text string based at least in part on the first set of features and the second set of features.

16 . The non-transitory computer-readable medium of claim 15 , wherein the instructions to generate the video are executable by the one or more processors to:

generate the video based at least in part on the generated head motion sequence.

17 . The non-transitory computer-readable medium of claim 13 , wherein the instructions are further executable by the one or more processors to:

identify a set of characteristics associated with the visual medium, wherein the set of characteristics include one or more geometric parameters and one or more appearance characteristics associated with a head motion of the second identity.

18 . The non-transitory computer-readable medium of claim 17 , wherein the instructions to generate the video are executable by the one or more processors to:

generate the video based at least in part on combining the set of characteristics with the first audio stream.

Assignments (2)
CHANGE OF NAME Recorded Aug 4, 2026
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 076118/0548 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2023
From: WANG, ZHICHAO; LUNDGAARD, KELD; DAI, MENGYU
To: SALESFORCE.COM, INC.
Reel/Frame 065748/0799 →