IP Library › Granted Patent US 12,437,773
Granted Patent B2
US 12,437,773 · App. 17/779,693 · Granted Oct 7, 2025

Apparatus and method for generating speech synthesis image

Inventors: Gyeong Su Chae (Seoul, KR); Guem Buel Hwang (Seoul, KR)
Assignee: DEEPBRAIN AI INC.
G10L21/10G06T7/248G10L15/25G06T2207/20076G06T2207/20081G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,773
App. No.
17/779,693
Granted
Oct 7, 2025
Kind
B2
Abstract

An apparatus for generating a speech synthesis image according to a disclosed embodiment is an apparatus for generating a speech synthesis image based on machine learning, the apparatus including a first global geometric transformation predictor configured to be trained to receive each of a source image and a target image including the same person, and predict a global geometric transformation for a global motion of the person between the source image and the target image based on the source image and the target image, a local feature tensor predictor configured to be trained to predict a feature tensor for a local motion of the person based on input target image-related information, and an image generator configured to be trained to reconstruct the target image based on the global geometric transformation, the source image, and the feature tensor for the local motion.

Claims (35)

1. An apparatus for generating a speech synthesis image based on machine learning, the apparatus comprising:

a first global geometric transformation predictor configured to be trained to receive each of a source image and a target image including the same person, and predict a global geometric transformation for a global motion of the person between the source image and the target image, based on the source image and the target image;

a local feature tensor predictor configured to be trained to predict a feature tensor for a local motion of the person, based on input target image-related information; and

an image generator configured to be trained to reconstruct the target image, based on the global geometric transformation, the source image, and the feature tensor for the local motion,

wherein the global motion is a motion of the person with an amount greater than or equal to a preset threshold amount of motion,

wherein the first global geometric transformation predictor is further configured to:

extract a source image heat map based on the source image, the source image heat map being a probability distribution map in an image space indicating whether each pixel in the source image is a pixel related to the global motion of the person;

extract a target image heat map based on the target image, the target image heat map being a probability distribution map in the image space as to whether each pixel in the target image is a pixel related to the global motion of the person; and

calculate the global geometric transformation based on the source image heat map and the target image heat map.

2. The apparatus according to claim 1 , wherein

the local motion is a motion of a face when the person is speaking.

3. The apparatus according to claim 2 , wherein the first global geometric transformation predictor is further configured to extract a geometric transformation into the source image heat map from a preset reference probability distribution, based on the source image, extract a geometric transformation into the target image heat map from the preset reference probability distribution, based on the target image, and calculate the global geometric transformation, based on the geometric transformation into the source image heat map from the reference probability distribution and the geometric transformation into the target image heat map from the reference probability distribution.

4. The apparatus according to claim 2 , wherein the local feature tensor predictor includes a first local feature tensor predictor configured to be trained to predict a speech feature tensor for a local speech motion of the person, based on the input target image-related information, and

the local speech motion is a motion related to a speech of the local motion of the person.

5. The apparatus according to claim 4 , wherein the local feature tensor predictor further includes a second local feature tensor predictor configured to be trained to predict a non-speech feature tensor for a local non-speech motion of the person, based on the input target image-related information, and

the local non-speech motion is a motion not related to the speech of the local motion of the person.

6. The apparatus according to claim 5 , wherein the second local feature tensor predictor is trained to receive a target partial image including only a motion not related to speech of the person in the target image, and predict the non-speech feature tensor, based on the target partial image.

7. The apparatus according to claim 2 , further comprising an optical flow predictor configured to be trained to calculate an optical flow between the source image and the target image, based on the source image and the global geometric transformation,

wherein the image generator is trained to reconstruct the target image, based on the optical flow between the source image and the target image, the source image, and the feature tensor for the local motion.

8. The apparatus according to claim 1 , wherein the first global geometric transformation predictor is further configured to calculate a geometric transformation into any i-th (iϵ{1, 2, . . . , n}) (n is a natural number equal to or greater than 2) frame heat map in an image having n frames from a preset reference probability distribution when the image is input, and calculate a global geometric transformation between two adjacent frames in the image, based on the geometric transformation into the i-th frame heat map from the reference probability distribution.

9. The apparatus according to claim 8 , further comprising a second global geometric transformation predictor configured to receive sequential voice signals corresponding to the n frames, and to be trained to predict a global geometric transformation between two adjacent frames in the image from the sequential voice signals.

10. The apparatus according to claim 9 , wherein the second global geometric transformation predictor is further configured to adjust a parameter of an artificial neural network to minimize a difference between the global geometric transformation between the two adjacent frames which is predicted in the second global geometric transformation predictor and the global geometric transformation between the two adjacent frames which is calculated in the first global geometric transformation predictor.

11. The apparatus according to claim 10 , wherein in a test process for speech synthesis image generation,

the second global geometric transformation predictor is further configured to receive sequential voice signals of a person, calculate a global geometric transformation between two adjacent frames in an image corresponding to the sequential voice signals from the sequential voice signals, and calculate a global geometric transformation between a preset target frame and a preset start frame, based on the global geometric transformation between the two adjacent frames,

the local feature tensor predictor is further configured to predict the feature tensor for the local motion of the person, based on the input target image-related information, and

the image generator is further configured to reconstruct the target frame, based on the global geometric transformation, the source image, and the feature tensor for the local motion.

12. A method for generating a speech synthesis image, based on machine learning that is performed in a computing device including one or more processors and a memory storing one or more programs executed by the one or more processors, the method comprising:

training a first global geometric transformation predictor to receive each of a source image and a target image including the same person, and predict a global geometric transformation for a global motion of the person between the source image and the target image, based on the source image and the target image;

training a local feature tensor predictor to predict a feature tensor for a local motion of the person, based on input target image-related information; and

training an image generator to reconstruct the target image, based on the global geometric transformation, the source image, and the feature tensor for the local motion,

wherein the global motion is a motion of the person with an amount greater than or equal to a preset threshold amount of motion,

wherein the training of the first global geometric transformation predictor comprises:

extracting a source image heat map based on the source image, the source image heat map being a probability distribution map in an image space indicating whether each pixel in the source image is a pixel related to the global motion of the person;

extracting a target image heat map based on the target image, the target image heat map being a probability distribution map in the image space as to whether each pixel in the target image is a pixel related to the global motion of the person; and

calculating the global geometric transformation based on the source image heat map and the target image heat map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2022
From: CHAE, GYEONG SU; HWANG, GUEM BUEL
To: DEEPBRAIN AI INC.
Reel/Frame 060011/0605 →
Priority Claims (1)
KR 10-2022-0019075 · Feb 14, 2022 · national
Continuity (1)
Related Publication 20240303830A1 · Sep 12, 2024
References Cited (16)
US 11295501B1 · Biswas · 2022 [cited by examiner]
US 20170262736A1 · Yu · 2017 [cited by examiner]
US 20200393943A1 · Sowden · 2020 [cited by applicant]
US 20210271862A1 · Ji · 2021 [cited by examiner]
US 20220051000A1 · Zhou · 2022 [cited by examiner]
US 20220417590A1 · Jiao · 2022 [cited by examiner]
US 20230368580A1 · Miyamoto · 2023 [cited by examiner]
KR 1020200145700A · 2020 [cited by applicant]
WO WO0231772A2 · 2002 [cited by applicant]
Rui Huang, et al., “Beyond Face Rotation: Global and Local Perception GAN for Photorealistic and Identity Preserving Frontal View Synthesis”, IEEE International Conference on Computer Vision, 2017, pp. 1-11 (Year: 2017). [cited by examiner]
T. Wang, et. al., “One-shot free-view neural talking-head synthesis for video conferencing”, in: IEEE CVPR, 2021, pp. 10039-10049 (Year: 2021). [cited by examiner]
K. Liu, et al., Towards Disentangling Latent Space for Unsupervised Semantic Face Editing, IEEE Transactions on Image Processing, vol. 31, 2022, Jan. 28, 2022, pp. 1475-1489 (Year: 2022). [cited by examiner]
E. Richardson et. al., “Encoding in style: a STYLEGAN encoder for image-to-image translation”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20-25, 2021, pp. 2287-2296 (Year: 2021). [cited by examiner]
Yurui Ren, et al., “Deep Spatial Transformation for Pose-Guided Person Image Generation and Animation”, IEEE Transactions on Image Processing, vol. 29, Aug. 27, 2020, pp. 8622-8635 (Year: 2020). [cited by examiner]
Lijuan Wang et al., “High Quality Lip-Sync Animation for 3D Photo-Realistic Talking Head”, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4529-4532, 2012. [cited by applicant]
M. Wang et al., “Facial Expression Synthesisusing a Global-Local Multilinear Framework”, Computer Graphics Forum, vol. 39, No. 2, pp. 235-245, 2020. [cited by applicant]