IP Library › Granted Patent US 11,494,959
Granted Patent B2
US 11,494,959 · App. 17/344,114 · Granted Nov 8, 2022

Method and apparatus with generation of transformed image

Inventors: Minsu Ko (Suwon-si, KR); Sungjoo Suh (Seongnam-si, KR); Young Chun Ahn (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06T11/60G06N3/0472G06T5/50G06T2207/20081G06T2207/20084G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,494,959
App. No.
17/344,114
Granted
Nov 8, 2022
Kind
B2
Abstract

A method with generation of a transformed image includes: receiving an input image; extracting, from the input image, coefficients corresponding to semantic elements of the input image; selecting at least one first target coefficient, among the coefficients, corresponding to at least one target semantic element that is to be changed among the semantic elements of the input image; changing the at least one first target coefficient; and generating a transformed image from the input image by applying the coefficients, including the changed at least one first target coefficient, to basis vectors used to represent the semantic elements of the input image in an embedding space of a neural network, the basis vectors corresponding to the semantic elements of the input image.

Claims (89)

1. A method with generation of a transformed image, the method comprising:

receiving an input image;

extracting, from the input image, coefficients corresponding to semantic elements of the input image;

selecting at least one first target coefficient, among the coefficients, corresponding to at least one target semantic element that is to be changed among the semantic elements of the input image;

changing the at least one first target coefficient, including adjusting a first weight of the at least one first target coefficient and a second weight of at least one second target coefficient and determining the at least one first target coefficient based on the at least one first target coefficient reflecting the adjusted first weight and the at least one second target coefficient reflecting the adjusted second weight; and

generating a transformed image from the input image by applying the coefficients, including the changed at least one first target coefficient, to basis vectors used to represent the semantic elements of the input image in an embedding space of a neural network, the basis vectors corresponding to the semantic elements of the input image.

2. The method of claim 1 , wherein the neural network comprises:

the basis vectors;

an encoder configured to estimate coefficients of the basis vectors to map the input image to the embedding space through the basis vectors; and

a generator configured to generate the transformed image by reflecting the coefficients corresponding to the semantic elements of the input image to the basis vectors.

3. The method of claim 2 , wherein an input space of the generator, and an output space to which a vector mapped using the encoder is projected from the input image based on the basis vectors are shared as the embedding space.

4. The method of claim 1 , wherein the selecting of the at least one first target coefficient comprises selecting the at least one first target coefficient using a matching table between the semantic elements and the coefficients.

5. The method of claim 1 , wherein the changing of the at least one first target coefficient comprises:

determining either one or both of a direction of a change in the at least one first target coefficient and a degree of the change in the at least one first target coefficient; and

changing the at least one first target coefficient based on either one or both of the direction of the change and the degree of the change.

6. The method of claim 1 , wherein the semantic elements comprise any one or any combination of any two or more of:

first elements including a gender, an age, a race, a pose, a facial expression, and a size of an object included in the input image;

second elements including a color, a shape, a style, a pose, a size, and a position of each component of the object; and

third elements including glasses, a hat, accessories, and clothes added to the object.

7. The method of claim 1 , further comprising:

receiving a second input image for changing the at least one first target coefficient; and

selecting, from the second input image, the at least one second target coefficient corresponding to at least one target semantic element among semantic elements of the second input image,

wherein the changing of the at least one first target coefficient comprises determining the at least one first target coefficient based on the at least one second target coefficient.

8. The method of claim 7 , wherein the determining of the at least one first target coefficient based on the at least one second target coefficient comprises either one or both of:

determining the at least one first target coefficient by combining the at least one first target coefficient and the at least one second target coefficient; and

determining the at least one first target coefficient by swapping the at least one first target coefficient with the at least one second target coefficient.

9. The method of claim 1 , wherein the changing of the at least one first target coefficient comprises:

determining either one or both of a direction of a change in the at least one first target coefficient and a degree of the change in the at least one first target coefficient; and

determining the at least one first target coefficient by combining the at least one first target coefficient reflecting the adjusted first weight and the at least one second target coefficient reflecting the adjusted second weight.

10. A method with generation of a transformed image, the method comprising:

receiving an input image;

extracting, from the input image, coefficients corresponding to semantic elements of the input image;

selecting at least one first target coefficient, among the coefficients, corresponding to at least one target semantic element that is to be changed among the semantic elements of the input image;

changing the at least one first target coefficient; and

generating a transformed image from the input image by applying the coefficients, including the changed at least one first target coefficient, to basis vectors used to represent the semantic elements of the input image in an embedding space of a neural network, the basis vectors corresponding to the semantic elements of the input image,

wherein the changing includes determining the at least one first target coefficient based on at least one second target coefficient, further comprising:

selecting, from a second input image, the at least one second target coefficient corresponding to at least one target semantic element among semantic elements of the second input image;

determining either one or both of a direction of a change in the at least one first target coefficient and a degree of the change in the at least one first target coefficient;

adjusting a first weight of the at least one first target coefficient and a second weight of the at least one second target coefficient based on either one or both of the direction of the change and the degree of the change; and

determining the at least one first target coefficient by combining the at least one first target coefficient reflecting the adjusted first weight and the at least one second target coefficient reflecting the adjusted second weight.

11. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 .

12. An apparatus with generation of a transformed image, the apparatus comprising:

a communication interface configured to receive an input image; and

a processor configured to:

extract, from the input image, coefficients corresponding to semantic elements of the input image;

select at least one first target coefficient, among the coefficients, corresponding to at least one target semantic element that is to be changed among the semantic elements of the input image;

change the at least one first target coefficient, including adjusting a first weight of the at least one first target coefficient and a second weight of at least one second target coefficient and determining the at least one first target coefficient based on the at least one first target coefficient reflecting the adjusted first weight and the at least one second target coefficient reflecting the adjusted second weight; and

generate a transformed image from the input image by applying the coefficients, including the changed at least one first target coefficient, to basis vectors used to represent the semantic elements of the input image in an embedding space of a neural network, the basis vectors corresponding to the semantic elements of the input image.

13. The apparatus of claim 12 , wherein the neural network comprises:

the basis vectors;

an encoder configured to estimate coefficients of the basis vectors to map the input image to the embedding space through the basis vectors; and

a generator configured to generate the transformed image by reflecting the coefficients corresponding to the semantic elements of the input image to the basis vectors.

14. The apparatus of claim 13 , wherein an input space of the generator, and an output space to which a vector mapped using the encoder is projected from the input image based on the basis vectors are shared as the embedding space.

15. The apparatus of claim 12 , wherein the processor is further configured to select the at least one first target coefficient using a matching table between the semantic elements and the coefficients.

16. The apparatus of claim 12 , wherein the processor is further configured to:

determine either one or both of a direction of a change in the at least one first target coefficient and a degree of the change in the at least one first target coefficient; and

change the at least one first target coefficient based on either one or both of the direction of the change and the degree of the change.

17. The apparatus of claim 12 , wherein the semantic elements comprise any one or any combination of any two or more of:

first elements including a gender, an age, a race, a pose, a facial expression, and a size of an object included in the input image;

second elements including a color, a shape, a style, a pose, a size, and a position of each component of the object; and

third elements including glasses, a hat, accessories, and clothes added to the object.

18. The apparatus of claim 12 , wherein the communication interface is further configured to receive a second input image for changing the at least one first target coefficient, and

wherein the processor is further configured to:

select, from the second input image, the at least one second target coefficient corresponding to at least one target semantic element among semantic elements of the second input image; and

determine the at least one first target coefficient based on the at least one second target coefficient.

19. The apparatus of claim 18 , wherein the processor is further configured to determine the at least one first target coefficient by combining the at least one first target coefficient and the at least one second target coefficient, or by swapping the at least one first target coefficient with the at least one second target coefficient.

20. The apparatus of claim 18 , wherein the processor is further configured to:

determine either one or both of a direction of a change in the at least one first target coefficient and a degree of the change in the at least one first target coefficient;

and

determine the at least one first target coefficient by combining the at least one first target coefficient reflecting the adjusted first weight and the at least one second target coefficient reflecting the adjusted second weight.

21. The apparatus of claim 12 , further comprising:

a display configured to display the transformed image.

22. The apparatus of claim 12 , wherein the input image comprises an enrollment image for user authentication, and

wherein the apparatus is configured to perform the user authentication by comparing an authentication image of a user to the transformed image.

23. An apparatus with image transformation, comprising:

at least one processor configured to implement a neural network, the neural network comprising:

an embedding space;

basis vectors representing semantic elements of an input image inputted to the apparatus;

an encoder configured to estimate coefficients of the basis vectors to map the input image to the embedding space through the basis vectors; and

a generator configured to generate a transformed image by:

changing at least one first target coefficient, among the coefficients, corresponding to at least one target semantic element that is to be changed among the semantic elements of the input image, including adjusting a first weight of the at least one first target coefficient and a second weight of at least one second target coefficient and determining the at least one first target coefficient based on the at least one first target coefficient reflecting the adjusted first weight and the at least one second target coefficient reflecting the adjusted second weight; and

applying the coefficients, including the changed at least one first target coefficient, to the basis vectors.

24. The apparatus of claim 23 , wherein the embedding space is configured as an input space of the generator and an output space of the encoder.

25. The apparatus of claim 23 , wherein the semantic elements comprise facial and hair appearance attributes of a person.

26. The apparatus of claim 23 , wherein the generator is further configured to:

select, from a second input image, the at least one second target coefficient corresponding to at least one target semantic element among semantic elements of the second input image; and

determine the at least one first target semantic element by performing either one of:

combining the at least one first target coefficient and the at least one second target coefficient; and

swapping the at least one first target coefficient with the at least one second target coefficient.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 10, 2021
From: KO, MINSU; SUH, SUNGJOO; AHN, YOUNG CHUN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 056499/0985 →
Priority Claims (2)
KR 10-2020-0150747 · Nov 12, 2020 · national
KR 10-2020-0181653 · Dec 23, 2020 · national
Continuity (1)
Related Publication 20220148244A1 · May 12, 2022