IP Library Granted Patent US 12,299,792
Granted Patent B2
US 12,299,792 · App. 17/866,114 · Granted May 13, 2025

Electronic device for generating mouth shape and method for operating thereof

Inventors: Jooyoung Kim (Suwon-si, KR); Jihyeok Yoon (Suwon-si, KR); Sunyoung Jung (Suwon-si, KR); Junsik Jeong (Suwon-si, KR); Haegang Jeong (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T13/00G06T3/4053G06T7/11G06T11/00G10L13/02G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,792
App. No.
17/866,114
Granted
May 13, 2025
Kind
B2
Abstract

An electronic device includes at least one processor, and at least one memory storing instructions executable by the at least one processor and operatively connected to the at least one processor, where the at least one processor is configured to acquire voice data to be synthesized with at least one first image, generate a plurality of mouth shape candidates by using the voice data, select a mouth shape candidate among the plurality of mouth shape candidates, generate at least one second image based on the selected mouth shape candidate and at least a portion of each of the at least one first image, and generate at least one third image by applying at least one super-resolution model to the at least one second image.

Claims (70)

1. An electronic device comprising:

at least one processor; and

at least one memory storing instructions executable by the at least one processor and operatively connected to the at least one processor,

wherein the instructions, when executed by the at least one processor, cause the electronic device to:

acquire voice data to be synthesized with first images,

generate a plurality of mouth shape candidates for each of the first images by using the voice data,

provide a user interface capable of selecting one of the plurality of mouth shape candidates, wherein the user interface provides a function of outputting a voice corresponding to the voice data while reproducing at least some of the plurality of mouth shape candidates or at least some areas of face areas including each of at least some of the plurality of mouth shape candidates,

based on obtaining an input with respect to the user interface, select a mouth shape candidate among the plurality of mouth shape candidates for each of the first images to obtain a plurality of selected mouth shape candidates corresponding to the first images,

generate second images based on the plurality of selected mouth shape candidates and at least a portion of each of the first images, and

generate third images by applying at least one super-resolution model to the generated second images.

2. The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor, cause the electronic device to generate at least one processed voice data by performing at least one analog processing on the voice data, input the voice data and the at least one processed voice data into a mouth shape generation model, and generate a plurality of output values from the mouth shape generation model as the plurality of mouth shape candidates.

3. The electronic device of claim 2 , wherein the at least one analog processing includes at least one of an increase in an amplitude of the voice data, a decrease in the amplitude of the voice data, an increase of a reproduction speed of the voice data, a decrease in the reproduction speed of the voice data, addition of first noise to the voice data, suppression of second noise from the voice data, separation of a first background sound from the voice data, or addition of a second background sound to the voice data.

4. The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor, cause the electronic device to input the voice data into a plurality of mouth shape generation models and generate a plurality of output values from the plurality of mouth shape generation models as the plurality of mouth shape candidates.

5. The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor, cause the electronic device to provide a user interface capable of selecting a time section in which the voice data is synthesized and identify the first images corresponding to the time section in which the voice data is synthesized, based on a user input made through the user interface.

6. The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor, cause the electronic device to input each of the plurality of mouth shape candidates into an assessment model and select one of the plurality of mouth shape candidates, based on a plurality of scores output from the assessment model.

7. The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor, cause the electronic device to generate a plurality of completely synthesized images by synthesizing at least some of the third images with the first images, respectively.

8. The electronic device of claim 7 , wherein the instructions, when executed by the at least one processor, cause the electronic device to perform segmentation on each of the third images, identify at least one area in each of the third images, based on a result of the segmentation, and generate the plurality of completely synthesized images by synthesizing the at least one area with the first images, respectively.

9. The electronic device of claim 1 , wherein the instructions, when executed by the at least one processor, cause the electronic device to generate the voice data, based on a voice received through a microphone of the electronic device or convert text into the voice data.

10. A method of operating an electronic device, the method comprising:

acquiring voice data to be synthesized with first images,

generating a plurality of mouth shape candidates for each of the first images by using the voice data,

providing a user interface capable of selecting one of the plurality of mouth shape candidates, wherein the user interface provides a function of outputting a voice corresponding to the voice data while reproducing at least some of the plurality of mouth shape candidates or at least some areas of face areas including each of at least some of the plurality of mouth shape candidates,

based on obtaining an input with respect to the user interface, selecting a mouth shape candidate among the plurality of mouth shape candidates for each of the first images to obtain a plurality of selected mouth shape candidates corresponding to the first images,

generating second images based on the plurality of selected mouth shape candidates and at least a portion of each of the first images, and

generating third images by applying at least one super-resolution model to the second images.

11. The method of claim 10 , wherein the generating of the plurality of mouth shape candidates comprises:

generating at least one processed voice data by performing at least one analog processing on the voice data,

inputting the voice data and the at least one processed voice data into a mouth shape generation model, and

generating a plurality of output values from the mouth shape generation model as the plurality of mouth shape candidates.

12. The method of claim 11 , wherein the analog processing comprises at least one of an increase in an amplitude of the voice data, a decrease in the amplitude of the voice data, an increase of a reproduction speed of the voice data, a decrease in the reproduction speed of the voice data, addition of first noise to the voice data, suppression of second noise from the voice data, separation of a first background sound from the voice data, or addition of a second background sound to the voice data.

13. The method of claim 10 , wherein the generating of the plurality of mouth shape candidates comprises inputting the voice data into a plurality of mouth shape generation models and generating a plurality of output values from the plurality of mouth shape generation models as the plurality of mouth shape candidates.

14. A non-transitory computer-readable storage medium storing at least one instruction, the at least one instruction causing at least one processor to:

acquire voice data to be synthesized with first images,

generate a plurality of mouth shape candidates for each of the first images by using the voice data,

provide a user interface capable of selecting one of the plurality of mouth shape candidates, wherein the user interface provides a function of outputting a voice corresponding to the voice data while reproducing at least some of the plurality of mouth shape candidates or at least some areas of face areas including each of at least some of the plurality of mouth shape candidates,

based on obtaining an input with respect to the user interface, select a mouth shape candidate among the plurality of mouth shape candidates for each of the first images to obtain a plurality of selected mouth shape candidates corresponding to the first images,

generate second images based on the plurality of selected mouth shape candidates and at least a portion of each of the first images, and

generate third images by applying at least one super-resolution model to the second images.

15. An electronic device comprising:

at least one processor; and

at least one memory storing instructions executable by the at least one processor and operatively connected to the at least one processor,

wherein the instructions, when executed by the at least one processor, cause the electronic device to:

acquire voice data to be synthesized with at least one first image,

generate a plurality of mouth shape candidates by using the voice data,

select a mouth shape candidate among the plurality of mouth shape candidates,

generate at least one second image based on the selected mouth shape candidate and at least a portion of each of the at least one first image, and

generate at least one third image by applying at least one super-resolution model to the at least one second image,

wherein the at least one processor is further configured to divide each of the at least one second image into at least one first area corresponding to a mouth and at least one second area including a remaining area except for the first area, generate at least one first high-resolution area by applying the at least one first area to a first super-resolution model specialized for the at least one first area, generate at least one second high-resolution area by applying the at least one second area to a second super-resolution model specialized for the at least one second area, and generate the at least one third image by synthesizing the at least one first high-resolution area with the at least one second high-resolution area, respectively.

16. A method of operating an electronic device, the method comprising:

acquiring voice data to be synthesized with at least one first image,

generating a plurality of mouth shape candidates by using the voice data,

selecting a mouth shape candidate among the plurality of mouth shape candidates,

generating at least one second image based on the selected mouth shape candidate and at least a portion of each of the at least one first image, and

generating at least one third image by applying at least one super-resolution model to the at least one second image,

wherein the generating of the at least one third image comprises:

dividing each of the at least one second image into at least one first area corresponding to a mouth and at least one second area including a remaining area except for the first area,

generating at least one first high-resolution area by applying the at least one first area to a first super-resolution model specialized for the at least one first area,

generating at least one second high-resolution area by applying the at least one second area to a second super-resolution model specialized for the at least one second area, and

generating the at least one third image by synthesizing the at least one first high-resolution area with the at least one second high-resolution area, respectively.

17. A non-transitory computer-readable storage medium storing at least one instruction, the at least one instruction causing at least one processor to:

acquire voice data to be synthesized with at least one first image,

generate a plurality of mouth shape candidates by using the voice data,

select a mouth shape candidate among the plurality of mouth shape candidates,

generate at least one second image based on the selected mouth shape candidate and at least a portion of each of the at least one first image, and

generate at least one third image by applying at least one super-resolution model to the at least one second image,

wherein the at least one instruction further causes the at least one processor to:

divide each of the at least one second image into at least one first area corresponding to a mouth and at least one second area including a remaining area except for the first area,

generate at least one first high-resolution area by applying the at least one first area to a first super-resolution model specialized for the at least one first area,

generate at least one second high-resolution area by applying the at least one second area to a second super-resolution model specialized for the at least one second area, and

generate the at least one third image by synthesizing the at least one first high-resolution area with the at least one second high-resolution area, respectively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2022
From: KIM, JOOYOUNG; YOON, JIHYEOK; JUNG, SUNYOUNG; JEONG, JUNSIK; JEONG, HAEGANG
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 060525/0046 →
Priority Claims (2)
KR 10-2021-0093812 · Jul 16, 2021 · national
KR 10-2021-0115930 · Aug 31, 2021 · national
Continuity (2)
Continuation PCTKR2022009589 · Jul 4, 2022
Related Publication 20230014604A1 · Jan 19, 2023
References Cited (52)
US 6539354B1 · Sutton et al. · 2003 [cited by applicant]
US 7168953B1 · Poggio · 2007 [cited by examiner]
US 20010050689A1 · Park · 2001 [cited by examiner]
US 20020114534A1 · Nakamura et al. · 2002 [cited by applicant]
US 20060012601A1 · Francini et al. · 2006 [cited by applicant]
US 20060164440A1 · Sullivan · 2006 [cited by examiner]
US 20090240734A1 · Lloyd-Jones · 2009 [cited by examiner]
US 20100171759A1 · Nickolov · 2010 [cited by examiner]
US 20110224978A1 · Sawada · 2011 [cited by examiner]
US 20130195428A1 · Marks et al. · 2013 [cited by applicant]
US 20130343647A1 · Aoki · 2013 [cited by applicant]
US 20150195397A1 · Rice · 2015 [cited by examiner]
US 20150242394A1 · Kim · 2015 [cited by applicant]
US 20160267686A1 · Ohta · 2016 [cited by applicant]
US 20170287481A1 · Bhat et al. · 2017 [cited by applicant]
US 20170345201A1 · Lin et al. · 2017 [cited by applicant]
US 20180232561A1 · Zheng et al. · 2018 [cited by applicant]
US 20190244623A1 · Hall et al. · 2019 [cited by applicant]
US 20190332851A1 · Han et al. · 2019 [cited by applicant]
US 20190392625A1 · Wang et al. · 2019 [cited by applicant]
US 20200058288A1 · Lin · 2020 [cited by examiner]
US 20200234480A1 · Volkov · 2020 [cited by examiner]
US 20200296535A1 · Tsingos et al. · 2020 [cited by applicant]
US 20210005180A1 · Kim · 2021 [cited by applicant]
US 20210150793A1 · Stratton · 2021 [cited by examiner]
US 20210312685A1 · Guo · 2021 [cited by examiner]
US 20210390748A1 · Liao · 2021 [cited by examiner]
CN 111277912A · 2020 [cited by applicant]
JP 0644365A · 1994 [cited by applicant]
JP 6348811A · 1994 [cited by applicant]
JP 2000348173A · 2000 [cited by applicant]
JP 2002170112A · 2002 [cited by applicant]
JP 2016173790A · 2016 [cited by applicant]
JP 2019175210A · 2019 [cited by applicant]
KR 19980058764A · 1998 [cited by applicant]
KR 19990034391A · 1999 [cited by applicant]
KR 20000009490A · 2000 [cited by applicant]
KR 1020050009504A · 2005 [cited by applicant]
KR 1020140037410A · 2014 [cited by applicant]
KR 1020160110039A · 2016 [cited by applicant]
KR 101827168B1 · 2018 [cited by applicant]
KR 1020180109634A · 2018 [cited by applicant]
KR 1020190111278A · 2019 [cited by applicant]
KR 1020190114150A · 2019 [cited by applicant]
KR 102251781B1 · 2021 [cited by applicant]
WO 2009114488A1 · 2009 [cited by applicant]
WO 2012120697A1 · 2012 [cited by applicant]
WO 2017190646A1 · 2017 [cited by applicant]
Suwajanakorn et al., “Synthesizing obama: learning lip sync from audio,” ACM Transactions on Graphics (ToG) 36.4 (2017): 1-13 (Year: 2017). [cited by examiner]
International Search Report (PCT/ISA/210) and Written Opinion (PCT/ISA/237) dated Oct. 5, 2022 issued by the International Searching Authority in International Application No. PCT/KR2022/009589. [cited by applicant]
Search Report dated Jul. 3, 2024, issued by the European Patent Office in European Application No. 22842351.3. [cited by applicant]
Communication dated Sep. 18, 2024, issued by the European Patent Office in counterpart European Application No. 22842351.3. [cited by applicant]