IP Library Granted Patent US 11,568,645
Granted Patent B2
US 11,568,645 · App. 16/823,752 · Granted Jan 31, 2023

Electronic device and controlling method thereof

Inventors: Victor Sergeevich Lempitsky (Moscow, RU); Aliaksandra Petrovna Shysheya (Moscow, RU); Egor Olegovich Zakharov (Moscow, RU); Egor Andreevich Burkov (Moscow, RU)
Assignee: Samsung Electronics Co., Ltd.
G06V20/46G06F16/70G06N3/08G06V20/41G06V40/00G06V40/168G06V40/169G06V40/172G06V40/179G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,568,645
App. No.
16/823,752
Granted
Jan 31, 2023
Kind
B2
Abstract

An electronic device and a controlling method thereof are provided. A controlling method of an electronic device according to the disclosure includes: performing first learning for a neural network model for acquiring a video sequence including a talking head of a random user based on a plurality of learning video sequences including talking heads of a plurality of users, performing second learning for fine-tuning the neural network model based on at least one image including a talking head of a first user different from the plurality of users and first landmark information included in the at least one image, and acquiring a first video sequence including the talking head of the first user based on the at least one image and pre-stored second landmark information using the neural network model for which the first learning and the second learning were performed.

Claims (82)

1. A method of controlling an electronic device, comprising:

performing first learning for a neural network model for acquiring a video sequence including a talking head of a random user based on a plurality of learning video sequences including talking heads of a plurality of users;

performing second learning for fine-tuning the neural network model based on at least one image including a talking head of a first user different from the plurality of users and first landmark information included in the at least one image; and

acquiring a first video sequence including the talking head of the first user based on the at least one image and pre-stored second landmark information using the neural network model for which the first learning and the second learning were performed,

wherein the performing second learning comprises:

acquiring the at least one image;

acquiring the first landmark information based on the at least one image;

acquiring a first embedding vector including information related to an identity of the first user by inputting the at least one image and the first landmark information into an embedder of the neural network model for which the first learning was performed; and

fine-tuning a parameter set of a generator of the neural network model for which the first learning was performed to be matched with the at least one image based on the first embedding vector.

2. The method of claim 1 ,

wherein the acquiring a first video sequence comprises:

acquiring the first video sequence by inputting the second landmark information and the first embedding vector into the generator.

3. The method of claim 1 ,

wherein the first landmark information and the second landmark information include information on a head pose and information on a mimics descriptor.

4. The method of claim 1 ,

wherein the embedder and the generator include a convolutional network, and

based on the generator being instantiated, acquiring normalization coefficients inside the instantiated generator based on the first embedding vector acquired by the embedder.

5. The method of claim 1 ,

wherein the performing first learning comprises:

acquiring at least one learning image from a learning video sequence including a talking head of a second user among the plurality of learning video sequences;

acquiring third landmark information for the second user based on the at least one learning image;

acquiring a second embedding vector including information related to an identity of the second user by inputting the at least one learning image and the third landmark information into the embedder;

instantiating the generator based on the parameter set of the generator and the second embedding vector;

acquiring a second video sequence including the talking head of the second user by inputting the second landmark information and the second embedding vector into the generator; and

updating the parameter set of the neural network model based on a degree of similarity between the second video sequence and the learning video sequence.

6. The method of claim 5 ,

wherein the performing first learning further comprises:

acquiring a realism score for the second video sequence through a discriminator of the neural network model;

updating the parameter set of the generator and the parameter set of the embedder based on the realism score; and

updating the parameter set of the discriminator.

7. The method of claim 6 ,

wherein the discriminator includes a projection discriminator acquiring the realism score based on a third embedding vector different from the first embedding vector and the second embedding vector.

8. The method of claim 7 ,

wherein, based on the first learning being performed, penalizing a difference between the second embedding vector and the third embedding vector, and

initializing the third embedding vector based on the first embedding vector at the start of the second learning.

9. The method of claim 1 ,

wherein the at least one image includes from 1 to 32 images.

10. An electronic device comprising:

a memory storing at least one instruction; and

a processor configured to execute the at least one instruction,

wherein the processor is configured to:

perform first learning for a neural network model for acquiring a video sequence including a talking head of a random user based on a plurality of learning video sequences including talking heads of a plurality of users,

perform second learning for fine-tuning the neural network model based on at least one image including a talking head of a first user different from the plurality of users and first landmark information included in the at least one image, and

acquire a first video sequence including the talking head of the first user based on the at least one image and pre-stored second landmark information using the neural network model for which the first learning and the second learning were performed,

wherein the perform second learning comprises:

acquire the at least one image,

acquire the first landmark information based on the at least one image,

acquire a first embedding vector including information related to an identity of the first user by inputting the at least one image and the first landmark information into an embedder of the neural network model for which the first learning was performed, and

fine-tune a parameter set of a generator of the neural network model for which the first learning was performed to be matched with the at least one image based on the first embedding vector.

11. The electronic device of claim 10 ,

wherein the processor is configured to:

acquire the first video sequence by inputting the second landmark information and the first embedding vector into the generator.

12. The electronic device of claim 10 ,

wherein the first landmark information and the second landmark information include information on a head pose and information on a mimics descriptor.

13. The electronic device of claim 10 ,

wherein the embedder and the generator include a convolutional network, and

based on the generator being instantiated, acquire normalization coefficients inside the instantiated generator based on the first embedding vector acquired by the embedder.

14. The electronic device of claim 10 , wherein the processor is configured to:

acquire at least one learning image from a learning video sequence including a talking head of a second user among the plurality of learning video sequences,

acquire third landmark information for the second user based on the at least one learning image,

acquire a second embedding vector including information related to an identity of the second user by inputting the at least one learning image and the third landmark information into the embedder,

instantiate the generator based on the parameter set of the generator and the second embedding vector,

acquire a second video sequence including the talking head of the second user by inputting the second landmark information and the second embedding vector into the generator, and

update the parameter set of the neural network model based on a degree of similarity between the second video sequence and the learning video sequence.

15. The electronic device of claim 14 ,

wherein the processor is configured to:

acquire a realism score for the second video sequence through a discriminator of the neural network model, and update the parameter set of the generator and the parameter set of the embedder based on the realism score, and

update the parameter set of the discriminator.

16. The electronic device of claim 15 ,

wherein the discriminator includes a projection discriminator acquiring the realism score based on a third embedding vector different from the first embedding vector and the second embedding vector.

17. The electronic device of claim 16 ,

wherein, based on the first learning being performed, penalizing a difference between the second embedding vector and the third embedding vector, and

the third embedding vector is initialized based on the first embedding vector at the start of the second learning.

18. A non-transitory computer readable recording medium having recorded thereon a program which, when executed by a processor of an electronic device causes the electronic device to perform operations comprising:

performing first learning for a neural network model for acquiring a video sequence including a talking head of a random user based on a plurality of learning video sequences including talking heads of a plurality of users;

performing second learning for fine-tuning the neural network model based on at least one image including a talking head of a first user different from the plurality of users and first landmark information included in the at least one image; and

acquiring a first video sequence including the talking head of the first user based on the at least one image and pre-stored second landmark information using the neural network model for which the first learning and the second learning were performed,

wherein the performing second learning comprises:

acquiring the at least one image;

acquiring the first landmark information based on the at least one image;

acquiring a first embedding vector including information related to an identity of the first user by inputting the at least one image and the first landmark information into an embedder of the neural network model for which the first learning was performed; and

fine-tuning a parameter set of a generator of the neural network model for which the first learning was performed to be matched with the at least one image based on the first embedding vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2020
From: LEMPITSKY, VICTOR SERGEEVICH; SHYSHEYA, ALIAKSANDRA PETROVNA; ZAKHAROV, EGOR OLEGOVICH; BURKOV, EGOR ANDREEVICH
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 052191/0252 →
Priority Claims (3)
RU RU2019108227 · Mar 21, 2019 · national
RU RU2019125940 · Aug 16, 2019 · national
KR 10-2020-0011360 · Jan 30, 2020 · national
Continuity (1)
Related Publication 20200302184A1 · Sep 24, 2020
Cited By (1)
US 12,307,749