IP Library › Granted Patent US 11,398,062
Granted Patent B2
US 11,398,062 · App. 16/957,363 · Granted Jul 26, 2022

Face synthesis

Inventors: Dong Chen (Beijing, CN); Fang Wen (Beijing, CN); Gang Hua (Sammamish, WA); Jianmin Bao (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06T11/00G06K9/6256G06V40/168G06V40/172
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,398,062
App. No.
16/957,363
Granted
Jul 26, 2022
Kind
B2
Abstract

In accordance with implementations of the subject matter described herein, there is provided a solution for face synthesis. In this solution, a first image about a face of a first user and a second image about a face of a second user are obtained. A first feature characterizing an identity of the first user is extracted from the first image, and a second feature characterizing a plurality of attributes of the second image is extracted from the second image, where the plurality of attributes do not include the identity of the second user. Then, a third image about a face of the first user is generated based on the first and second features, the third image reflecting the identity of the first user and the plurality of attributes of the second image.

Claims (42)

1. An electronic device, comprising:

a processing unit;

a memory coupled to the processing unit and storing instructions for execution by the processing unit, the instructions, when executed by the processing unit, causing the device to perform operations comprising:

obtaining a first image about a face of a first user and a second image about a face of a second user;

extracting a first feature from the first image, the first feature characterizing a first identity of the first user;

extracting a second feature from the second image, the second feature characterizing a plurality of attributes of the second image other than a second identity of the second user; and

generating, a third image about a face of the first user based on the first and the second features using a learning network previously trained by minimizing a first reconstruction loss between the second image and the third image and preserving one or more of the plurality of attributes of the second image while minimizing a second reconstruction loss between the first image and the third image, the third image reflecting the first identity of the first user and the plurality of attributes of the second image.

2. The device according to claim 1 , wherein extracting the first feature comprises:

extracting the first feature from the first image using a first sub-network in the learning network for face synthesis, the first feature being extracted from at least one layer of the first sub-network.

3. The device according to claim 2 , wherein extracting the second feature comprises:

extracting the second feature from the second image using a second sub-network in the learning network, the second feature being extracted from at least one layer of the second sub-network.

4. The device according to claim 3 , wherein generating the third image comprises:

generating the third image based on the first and second features using a third sub-network in the learning network, outputs of the first and second sub-networks being coupled to an input of the third sub-network.

5. An electronic device, comprising:

a processing unit;

a memory coupled to the processing unit and storing instructions for execution by the processing unit, the instructions, when executed by the processing unit, causing the device to perform operations comprising:

obtaining a first image about a face of a first user and a second image about a face of a second user, the first image being labeled with a first identity of the first user; and

training a learning network for face synthesis based on the first and second images, such that the learning network:

extracts a first feature from the first image, the first feature characterizing the first identity of the first user;

extracts a second feature from the second image, the second feature characterizing a plurality of attributes of the second image other than a second identity of the second user; and

generates a third image about a face of the first user based on the first and second features using the learning network previously trained by minimizing a first reconstruction loss between the second image and the third image and preserving one or more of the plurality of attributes of the second image while minimizing a second reconstruction loss between the first image and the third image, the third image reflecting the first identity of the first user and the plurality of attributes of the second image.

6. The device according to claim 5 , wherein the learning network includes a first sub-network, and training the learning network comprises:

training the first sub-network such that the first sub-network extracts the first feature from the first image.

7. The device according to claim 6 , wherein the learning network includes a second sub-network, and training the learning network comprises:

training the second sub-network such that the second sub-network extracts the second feature from the second image.

8. The device according to claim 7 , wherein the learning network includes a third sub-network, outputs of the first and the second sub-networks being coupled to an input of the third sub-network, and training the learning network comprises:

training the third sub-network such that the third sub-network generates the third image based on the first and second features.

9. The device according to claim 8 , wherein the learning network includes a fourth sub-network, an input of the first sub-network and an output of the third sub-network being coupled to an input of the fourth sub-network, and training the learning network comprises:

training the fourth sub-network such that the fourth sub-network classifies the first image and the third image as about a same user.

10. The device according to claim 9 , wherein the learning network includes a fifth sub-network, an output of the third sub-network and an input of the second sub-network being coupled to an input of the fifth sub-network, and training the learning network comprises:

training the fifth sub-network such that the fifth sub-network classifies the second image as an original image and the third image as a synthesized image.

11. A computer-implemented method, comprising:

obtaining a first image about a face of a first user and a second image about a face of a second user;

extracting a first feature from the first image, the first feature characterizing a first identity of the first user;

extracting a second feature from the second image, the second feature characterizing a plurality of attributes of the second image other than a second identity of the second user; and

generating a third image about a face of the first user based on the first and second features using a learning network previously trained by minimizing a first reconstruction loss between the second image and the third image and preserving one or more of the plurality of attributes of the second image while minimizing a second reconstruction loss between the first image and the third image, the third image reflecting the first identity of the first user and the plurality of attributes of the second image.

12. The method according to claim 11 , wherein extracting the first feature comprises:

extracting the first feature from the first image using a first sub-network in the learning network for face synthesis, the first feature being extracted from at least one layer of the first sub-network.

13. The method according to claim 12 , wherein extracting the second feature comprises:

extracting the second feature from the second image using a second sub-network in the learning network, the second feature being extracted from at least one layer of the second sub-network.

14. The method according to claim 13 , wherein generating the third image comprises:

generating the third image based on the first and second features using a third sub-network in the learning network, outputs of the first and second sub-networks being coupled to an input of the third sub-network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2020
From: CHEN, DONG; WEN, FANG; HUA, GANG; BAO, JIANMIN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 053295/0737 →
Priority Claims (1)
CN 201810082732.8 · Jan 29, 2018 · national
Continuity (1)
Related Publication 20200334867A1 · Oct 22, 2020