IP Library › Granted Patent US 12,347,197
Granted Patent B2
US 12,347,197 · App. 17/762,926 · Granted Jul 1, 2025

Device and method for generating speech video along with landmark

Inventor: Gyeongsu Chae (Seoul, KR)
Assignee: DEEPBRAIN AI INC.
G06V20/46G06V10/82G06V40/171G10L15/02G10L25/30G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,347,197
App. No.
17/762,926
Granted
Jul 1, 2025
Kind
B2
Abstract

A speech video generation device according to an embodiment includes a first encoder, which receives an input of a person background image that is a video part in a speech video of a predetermined person, and extracts an image feature vector from the person background image, a second encoder, which receives an input of a speech audio signal that is an audio part in the speech video, and extracts a voice feature vector from the speech audio signal, a combining unit, which generates a combined vector by combining the image feature vector output from the first encoder and the voice feature vector output from the second encoder, a first decoder, which reconstructs the speech video of the person using the combined vector as an input, and a second decoder, which predicts a landmark of the speech video using the combined vector as an input.

Claims (51)

1. A speech video generation device that is a computing device having one or more processors and a memory which stores one or more programs executed by the one or more processors, the speech video generation device comprising:

a first encoder configured to receive an input of an image that is a video part of a speech video of a person and extract an image feature vector from the image;

a second encoder configured to receive an input of a speech audio signal of the speech video and extract a voice feature vector from the speech audio signal;

a combining unit configured to generate a combined vector by combining the image feature vector extracted by the first encoder and the voice feature vector extracted by the second encoder;

a first decoder configured to reconstruct the speech video of the person using the combined vector as an input; and

a second decoder configured to predict a landmark of the speech video using the combined vector as an input,

wherein a portion related to a speech of the person is hidden by a mask in the image,

wherein the speech audio signal is an audio part of the same section as the image of the speech video of the person, wherein the first decoder is a machine learning model trained to reconstruct the portion hidden by the mask in the image based on the voice feature vector.

2. The speech video generation device of claim 1 , wherein when the image of the person is input to the first encoder, and the speech audio signal not related to the image is input to the second encoder;

the combining unit generates the combined vector by combining the image feature vector output from the first encoder and the voice feature vector output from the second encoder;

the first decoder receives the combined vector to generate the speech video of the person by reconstructing, based on the speech audio signal not related to the image, the portion related to the speech in the image; and

the second decoder predicts and outputs the landmark of the speech video.

3. The speech video generation device of claim 1 , wherein the second decoder comprises:

an extraction module trained to extract a feature vector from the input combined vector; and

a prediction module trained to predict landmark coordinates of the speech video based on the feature vector extracted by the extraction module.

4. The speech video generation device of claim 3 , wherein an objective function L prediction of the second decoder is expressed as an equation below:

L prediction =∥K−G ( I ;θ)∥  [Equation]

where K: Labeled landmark coordinates of a speech video;

G: Neural network constituting the second decoder;

θ: Parameter of the neural network constituting the second decoder;

I: Combined vector;

G(I;θ): Landmark coordinates predicted by the second decoder; and

∥K−G(I;θ)∥: Function for deriving a difference between the labeled landmark coordinates and the predicted landmark coordinates of the speech video.

5. The speech video generation device of claim 1 , wherein the second decoder comprises:

an extraction module trained to extract a feature tensor from the input combined vector; and

a prediction module trained to predict a landmark image based on the feature tensor extracted by the extraction module.

6. The speech video generation device of claim 5 , wherein the landmark image is an image indicating, by a probability value, whether each pixel corresponds to a landmark in an image space corresponding to the speech video.

7. The speech video generation device of claim 5 , wherein an objective function L prediction of the second decoder is expressed as an equation below:

L prediction =−Σ{p target ( x i ,y i )log( p ( x i ,y i ))+(1− p target ( x i ,y i )log(1− p ( x i ,y i ))}  [(Equation)

where p(x i , y i ): Probability value pertaining to whether a pixel (x i , y i ) is a landmark where p(x i , y i )=probability distribution (P(F(x i , y i );δ));

P: Neural network constituting the second decoder;

F(x i , y i ): Feature tensor of the pixel (x i , y i );

δ: Parameter of the neural network constituting the second decoder; and

p target (x i , y i ): Labeled landmark indication value of the pixel (x i , y i ) of a speech video.

8. A speech video generation device that is a computing device having one or more processors and a memory which stores one or more programs executed by the one or more processors, the speech video generation device comprising:

a first encoder, which receives an input of an image that is a video part in a speech video of a person, and extracts an image feature vector from the image;

a second encoder, which receives an input of a speech audio signal that is an audio part in the speech video, and extracts a voice feature vector from the speech audio signal;

a combining unit, which generates a combined vector by combining the image feature vector output from the first encoder and the voice feature vector output from the second encoder;

a decoder, which uses the combined vector as an input, and performs deconvolution and up-sampling on the combined vector;

a first output layer, which is connected to the decoder and outputs a reconstructed speech video of the person based on up-sampled data; and

a second output layer, which is connected to the decoder and outputs a predicted landmark of the speech video based on the up-sampled data,

wherein a portion related to a speech of the person is hidden by a mask in the image,

wherein the speech audio signal is an audio part of the same section as the image of the speech video of the person, wherein the decoder is a machine learning model trained to reconstruct the portion hidden by the mask in the image based on the voice feature vector.

9. A speech video generation method performed by a computing device having one or more processors and a memory which stores one or more programs executed by the one or more processors, the speech video generation method comprising:

receiving an input of an image that is a video part of a speech video of a person, and extracting an image feature vector from the image by a first encoder;

receiving an input of a speech audio signal that is an audio part of the speech video, and extracting a voice feature vector from the speech audio signal by a second encoder;

generating a combined vector by combining the image feature vector extracted by the first encoder and the voice feature vector extracted by the second encoder;

reconstructing, by a first decoder, the speech video of the person by using the combined vector as an input; and

predicting, by a second decoder, a landmark of the speech video by using the combined vector as an input,

wherein a portion related to a speech of the person is hidden by a mask in the image,

wherein the speech audio signal is an audio part of the same section as the image of the speech video of the person, wherein the first decoder is a machine learning model trained to reconstruct the portion hidden by the mask in the image based on the voice feature vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2022
From: CHAE, GYEONGSU
To: DEEPBRAIN AI INC.
Reel/Frame 059376/0860 →
Priority Claims (1)
KR 10-2020-0109173 · Aug 28, 2020 · national
Continuity (1)
Related Publication 20220375224A1 · Nov 24, 2022
References Cited (19)
US 10972682B1 · Muenster · 2021 [cited by examiner]
US 20040235531A1 · Anzawa · 2004 [cited by examiner]
US 20180174348A1 · Bhat et al. · 2018 [cited by applicant]
US 20210027511A1 · Shang · 2021 [cited by examiner]
US 20210056348A1 · Berlin · 2021 [cited by examiner]
US 20210248801A1 · Li · 2021 [cited by examiner]
KR 1020060090687A · 2006 [cited by applicant]
KR 1020140037410A · 2014 [cited by applicant]
KR 1020200043660A · 2020 [cited by applicant]
International Search Report for PCT/KR2020/018372 mailed on May 24, 2021. [cited by applicant]
Konstantinos Vougioukas et al. “Realistic Speech-Driven Facial Animation with GANs”,arXiv:1906.06337. 2019 [Retreived on Apr. 27, 2021]. Retrieved form <URL: https://arxiv.org/abs/1906.06337>. [cited by applicant]
Sinha, Sanjana et al. “Identity-Preserving Realistic Talking Face Generation”, arXiv:2005.12318. 2020 [Retrieved on Apr. 27, 2021]. Retrieved form <URL: https;//arxiv.org/abs/2005.12318>. [cited by applicant]
Office action issued on Jan. 12, 2022 from Korean Patent Office in a counterpart Korean Patent Application No. 10-2020-0109173 (all the cited references are listed in this IDS.) (English translation is also submitted he… [cited by applicant]
Office action issued on Jul. 27, 2022 from Korean Patent Office in a counterpart Korean Patent Application No. 10-2020-0109173 (all the cited references are listed in this IDS.) (English translation is also submitted he… [cited by applicant]
Lezi Wang et al, “A coupled encoder decoder network for joint face detection and landmark localization”, Image and Vision Computing 87, 2019, pp. 37-46. [cited by applicant]
Hao Zhu et al., “Arbitrary talking face generation via attentional audio-visual coherence learning”, 2020, arXiv: 1812.06589v2, [cs.CV]. [cited by applicant]
Linsen Song et al., “Geometry Aware Face Completion and Editing”, Association for the Advancement of Artificial Intelligence, Feb. 2019, pp. 2506-2513. [cited by applicant]
Lele Chen et al., “Hierarchical Cross-Modal Talking Face Generation with Dynamic Pixel-Wise Loss”, the Procs. of the IEEE / CVF Conference on CVPR, 2019, pp. 7832-7841. [cited by applicant]
Alexandros Koumparoulis et al., “Audio-Assisted Image Inpainting for Talking Faces”, Icassp 2020, Apr. 2020, pp. 7664-7668. [cited by applicant]