IP Library Granted Patent US 11,968,433
Granted Patent B2
US 11,968,433 · App. 17/722,258 · Granted Apr 23, 2024

Systems and methods for generating synthetic videos based on audio contents

Inventor: Jing Zhao (Beijing, CN)
Assignee: REALSEE (BEIJING) TECHNOLOGY CO., LTD.
H04N21/8547G06N3/045G06V20/46G06V40/165G11B27/031G11B27/10H04N21/8456
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,968,433
App. No.
17/722,258
Granted
Apr 23, 2024
Kind
B2
Abstract

Systems and methods for generating a synthetic video based on an audio are provided. An exemplary system may include a memory storing computer-readable instructions and at least one processor. The processor may execute the computer-readable instructions to perform operations. The operations may include receiving a reference video including a motion picture of a human face and receiving the audio including a speech. The operations may also include generating a synthetic motion picture of the human face based on the reference video and the audio. The synthetic motion picture of the human face may include a motion of a mouth of the human face presenting the speech. The motion of the mouth may match a content of the speech. The operations may further include generating the synthetic video based on the synthetic motion picture of the human face.

Claims (58)

1. A system for generating a synthetic video based on an audio, comprising:

a memory storing computer-readable instructions; and

at least one processor communicatively coupled to the memory, wherein the computer-readable instructions, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

receiving a reference video comprising a motion picture of a human face;

receiving the audio comprising a speech describing a real property, the audio being recorded at a different occasion from the reference video;

generating a synthetic motion picture of the human face based on the reference video and the audio, wherein the synthetic motion picture of the human face comprises a sequence of synthetic images simulating a motion of a mouth of the human face, each synthetic image temporally corresponding to a content of an audio segment of the received audio, wherein the synthetic image is generated based on an image of the motion picture of the human face and temporally-corresponding time-domain signals extracted from the received audio to match a content of the speech describing the real property; and

generating the synthetic video showing a person describing the real property based on the synthetic motion picture of the human face.

2. The system of claim 1 , wherein the operations comprise:

conditioning the reference video to match a time duration of the audio.

3. The system of claim 2 , wherein conditioning the reference video to match the time duration of the audio comprises:

truncating the reference video when the reference video is longer in duration than the audio; and

extending the reference video with a duplication of at least a portion of the reference video when the reference video is shorter in duration than the audio.

4. The system of claim 1 , wherein the operations comprise:

extracting, from the reference video, a plurality of frames each containing an image of the human face;

replacing, in each of the plurality of frames, a portion of the image corresponding to the mouth of the human face with a synthetic image corresponding to that frame; and

generating the synthetic motion picture of the human face by combining the plurality of frames each containing the respective synthetic image.

5. The system of claim 4 , wherein the operations comprise:

dividing the audio into a plurality of audio segments based on the plurality of frames, each of the plurality of audio segments corresponding to one of the plurality of frames; and

generating, for each of the plurality of frames, the synthetic image corresponding to that frame based on the audio segment corresponding to that frame, the synthetic image comprising a shape of the mouth matching a content of the audio segment.

6. The system of claim 5 , wherein the operations comprise:

processing, for each of the plurality of frames, the image of the human face corresponding to that frame and the audio segment corresponding to that frame using a neural network generator to generate the synthetic image corresponding to that frame.

7. The system of claim 6 , wherein:

the neural network generator is trained in a generative adversarial network comprising the neural network generator and a discriminator using a plurality of training samples, each of the plurality of training samples comprising a training audio segment and a corresponding training human face image containing a mouth shape matching a content of the training audio segment.

8. The system of claim 4 , wherein the operations comprise:

performing facial recognition to the reference video to extract the plurality of frames.

9. The system of claim 1 , wherein the operations comprise:

generating the synthetic video by combining the synthetic motion picture of the human face and the audio.

10. The system of claim 1 , wherein the speech comprises a presentation of a real estate property.

11. A method for generating a synthetic video based on an audio, the method comprising:

receiving a reference video comprising a motion picture of a human face;

receiving the audio comprising a speech describing a real property, the audio being recorded at a different occasion from the reference video;

generating a synthetic motion picture of the human face based on the reference video and the audio, wherein the synthetic motion picture of the human face comprises a sequence of synthetic images simulating a motion of a mouth of the human face, each synthetic image temporally corresponding to a content of an audio segment of the received audio, wherein the synthetic image is generated based on an image of the motion picture of the human face and temporally-corresponding time-domain signals extracted from the received audio to match a content of the speech describing the real property; and

generating the synthetic video showing a person describing the real property based on the synthetic motion picture of the human face.

12. The method of claim 11 , comprising:

conditioning the reference video to match a time duration of the audio.

13. The method of claim 12 , wherein conditioning the reference video to match the time duration of the audio comprises:

truncating the reference video when the reference video is longer in duration than the audio; and

extending the reference video with a duplication of at least a portion of the reference video when the reference video is shorter in duration than the audio.

14. The method of claim 11 , comprising:

extracting, from the reference video, a plurality of frames each containing an image of the human face;

replacing, in each of the plurality of frames, a portion of the image corresponding to the mouth of the human face with a synthetic image corresponding to that frame; and

generating the synthetic motion picture of the human face by combining the plurality of frames each containing the respective synthetic image.

15. The method of claim 14 , comprising:

dividing the audio into a plurality of audio segments based on the plurality of frames, each of the plurality of audio segments corresponding to one of the plurality of frames; and

generating, for each of the plurality of frames, the synthetic image corresponding to that frame based on the audio segment corresponding to that frame, the synthetic image comprising a shape of the mouth matching a content of the audio segment.

16. The method of claim 15 , comprising:

processing, for each of the plurality of frames, the image of the human face corresponding to that frame and the audio segment corresponding to that frame using a neural network generator to generate the synthetic image corresponding to that frame.

17. The method of claim 16 , wherein:

the neural network generator is trained in a generative adversarial network comprising the neural network generator and a discriminator using a plurality of training samples, each of the plurality of training samples comprising a training audio segment and a corresponding training human face image containing a mouth shape matching a content of the training audio segment.

18. The method of claim 14 , comprising:

performing facial recognition to the reference video to extract the plurality of frames.

19. The method of claim 11 , comprising:

generating the synthetic video by combining the synthetic motion picture of the human face and the audio.

20. A non-transitory computer-readable medium storing computer-readable instructions, wherein the computer-readable instructions, when executed by at least one processor, cause the at least one processor to perform a method for generating a synthetic video based on an audio, the method comprising:

receiving a reference video comprising a motion picture of a human face;

receiving the audio comprising a speech describing a real property, the audio being recorded at a different occasion from the reference video;

generating a synthetic motion picture of the human face based on the reference video and the audio, wherein the synthetic motion picture of the human face comprises a sequence of synthetic images simulating a motion of a mouth of the human face, each synthetic image temporally corresponding to a content of an audio segment of the received audio, wherein the synthetic image is generated based on an image of the motion picture of the human face and temporally-corresponding time-domain signals extracted from the received audio to match a content of the speech describing the real property; and

generating the synthetic video showing a person describing the real property based on the synthetic motion picture of the human face.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2022
From: ZHAO, JING
To: REALSEE (BEIJING) TECHNOLOGY CO., LTD.
Reel/Frame 059617/0110 →
Priority Claims (1)
CN 202110437420.6 · Apr 22, 2021 · national
Continuity (1)
Related Publication 20220345796A1 · Oct 27, 2022