IP Library › Granted Patent US 12,236,943
Granted Patent B2
US 12,236,943 · App. 17/764,651 · Granted Feb 25, 2025

Apparatus and method for generating lip sync image

Inventors: Guem Buel Hwang (Seoul, KR); Gyeong Su Chae (Seoul, KR)
Assignee: DEEPBRAIN AI INC.
G10L15/16G10L21/10G10L15/25G10L2021/105
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,943
App. No.
17/764,651
Granted
Feb 25, 2025
Kind
B2
Abstract

An apparatus for generating a lip sync image according to disclosed embodiment has one or more processors and a memory which stores one or more programs executed by the one or more processors. The apparatus includes a first artificial neural network model configured to generate an utterance match synthesis image by using a person background image and an utterance match audio signal corresponding to the person background image as an input, and generate an utterance mismatch synthesis image by using the person background image and an utterance mismatch audio signal not corresponding to the person background image as an input, and a second artificial neural network model configured to output classification values for an input pair in which an image and a voice match and an input pair in which an image and a voice do not match by using the input pairs as an input.

Claims (40)

1. An apparatus for generating a lip sync image having one or more processors and a memory which stores one or more programs executed by the one or more processors, the apparatus comprising:

a first artificial neural network model configured to generate an utterance match synthesis image by using a person background image and an utterance match audio signal as an input, and generate an utterance mismatch synthesis image by using the person background image and an utterance mismatch audio signal as an input; and

a second artificial neural network model configured to output classification values for an input pair in which an image and a voice match and an input pair in which an image and a voice do not match by using the input pairs as an input,

wherein the utterance match audio signal is a voice signal which matches a figure in which the corresponding person utters in the person background image,

wherein the utterance mismatch audio signal is a voice signal which does not match the figure in which the corresponding person utters in the person background image,

wherein the second artificial neural network model is trained to classify the input pair in which an image and a voice match as True, and to classify the input pair in which an image and a voice do not match as False,

wherein the first artificial neural network model is configured to receive the utterance mismatch synthesis image generated by the first artificial neural network model and the utterance mismatch audio signal used as the input when generating the utterance mismatch synthesis image and classify the utterance mismatch synthesis image and the utterance mismatch audio signal as True, and propagate a generative adversarial error to the first artificial neural network model through an adversarial learning method.

2. The apparatus of claim 1 , wherein the person background image is an image in which a portion associated with an utterance of a person is masked.

3. The apparatus of claim 1 , wherein the first artificial neural network model comprises:

a first encoder configured to use the person background image as an input, and extract an image feature vector from the input person background image;

a second encoder configured to use the utterance match audio signal corresponding to the person background image as an input, and extract a voice feature vector from the input utterance match audio signal;

a combiner configured to generate a combined vector by combining the image feature vector and the voice feature vector; and

a decoder configured to use the combined vector as an input, and generate the utterance match synthesis image based on the combined vector.

4. The apparatus of claim 3 , wherein an objective function L reconstruction for the generation of the utterance match synthesis image of the first artificial neural network model is represented by the following equation:

L reconstruction =∥I i −Î ii ∥

where I i is Original utterance image;

Î ii is Utterance match synthesis image; and

∥A−B∥ is Function for obtaining difference between A and B.

5. The apparatus of claim 4 , wherein an objective function L discriminator of the second artificial neural network model is represented by the following equation:

L discriminator =log(1− D ( I i ,A i ))+log( D ( I i ,A i ))

where D is Neural network of the second artificial neural network model;

(I i , A i ) is Input pair in which an image and a voice match (i-th image and i-th voice); and

(I i , A j ) is Input pair in which an image and a voice do not match (i-th image and j-th voice).

6. The apparatus of claim 5 , wherein an adversarial objective function L adversarial for the generation of the utterance mismatch synthesis image of the first artificial neural network model is represented by the following equation:

L adversarial =−log( D ( G ( M i * I i ,A i ), A j ))

where G is Neural network constituting the first artificial neural network model;

M i *I i is Person background image in which portion associated with utterance is masked (M i : mask);

G (M i *I i , A j ) is Utterance mismatch synthesis image generated by the first artificial neural network model; and

A j is Utterance mismatch audio signal not corresponding to person background image.

7. The apparatus of claim 6 , wherein a final objective function L T for the generation of the utterance match synthesis image and the utterance mismatch synthesis image of first artificial neural network model is represented by the following equation:

L T =L reconstruction +λL adversarial

where λ is Weight.

8. A method for generating a lip sync image performed by a computing device having one or more processors and a memory which stores one or more programs executed by the one or more processors, the method comprising:

generating, in a first artificial neural network model, an utterance match synthesis image by using a person background image and an utterance match audio signal as an input;

generating, in a first artificial neural network model, an utterance mismatch synthesis image by using the person background image and an utterance mismatch audio signal as an input; and

outputting, in a second artificial neural network model, classification values for an input pair in which an image and a voice match and an input pair in which an image and a voice do not match by using the input pairs as an input,

wherein the utterance match audio signal is a voice signal which matches a figure in which the corresponding person utters in the person background image,

wherein the utterance mismatch audio signal is a voice signal which does not match the figure in which the corresponding person utters in the person background image,

wherein the second artificial neural network model is trained to classify the input pair in which an image and a voice match as True, and to classify the input pair in which an image and a voice do not match as False,

wherein the first artificial neural network model is configured to receive the utterance mismatch synthesis image generated by the first artificial neural network model and the utterance mismatch audio signal used as the input when generating the utterance mismatch synthesis image and classify the utterance mismatch synthesis image and the utterance mismatch audio signal as True, and propagate a generative adversarial error to the first artificial neural network model through an adversarial learning method.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2022
From: HWANG, GUEM BUEL; CHAE, GYEONG SU
To: DEEPBRAIN AI INC.
Reel/Frame 059422/0875 →
Priority Claims (1)
KR 10-2021-0003375 · Jan 11, 2021 · national
Continuity (1)
Related Publication 20230178072A1 · Jun 8, 2023
References Cited (13)
US 20170243387A1 · Li et al. · 2017 [cited by applicant]
US 20180039990A1 · Lindemann · 2018 [cited by examiner]
US 20190206441A1 · De la Torre · 2019 [cited by examiner]
US 20210248801A1 · Li · 2021 [cited by examiner]
KR 1020190114150A · 2019 [cited by applicant]
KR 1020200112647A · 2020 [cited by applicant]
Jung-Woo Lee et al, “Korean Lip-Sync using Conditional GAN” The Magazine of the IEIE; Aug. 2020, pp. 1241-1243, URL <http://www.dbpia.co.kr/journal/articleDetail?nodeId=NODE10448132> (English translation of abstract is … [cited by applicant]
Jae Hyun Lee et al, “Avatar's Lip Synchronization in Talking Involved Virtual Reality” Korea Computer Graphics Society; vol. 26, No. 4, p. 9-15 DOI: 10.15701/kcgs.2020.26.4.9 (English translation of abstract is included… [cited by applicant]
Office action issued on Apr. 29, 2022 from Korean Patent Office in a counterpart Korean Patent Application No. 10-2021-0003375 (all the cited references are listed in this IDS.). [cited by applicant]
Konstantinos Vougioukas et al., “Realistic Speech-Driven Facial Animation with GAN”, International Journal of Computer Vision, 2020 vol. 128, pp. 1398-1413, https://doi.org/10.1007/s11263-019-01251-8. [cited by applicant]
Ruobing Zheng et al., “Photorealistic Lip Sync with Adversarial Temporal Convolutional Networks”, arXiv: 2002.08700v1 [cs.CV] Feb. 20, 2020. [cited by applicant]
Hao Zhu et al., “Arbitrary talking face generation via attentional audio-visual coherence learning” arXiv: 1812.06589x2, [cs.CV], and May 13, 2020. [cited by applicant]
K R Prajwal et al., “A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild” arXiv: 2008.10010v1 [cs.CV] Aug. 23, 2020. [cited by applicant]