IP Library Granted Patent US 12,327,280
Granted Patent B2
US 12,327,280 · App. 18/067,597 · Granted Jun 10, 2025

Method and system for outfit simulation using layer mask

Inventors: Benjamin James Biggs (San Francisco, CA); Philip Pinette (Seattle, WA); Charu Kothari (San Jose, CA); Caitlin Isaac Cagampan (Pasadena, CA); Gerard Guy Medioni (Los Angeles, CA); Achal Dushyant Dave (San Francisco, CA); Scott Chenghui Sun (Milpitas, CA)
Assignee: AMAZON TECHNOLOGIES, INC.
G06Q30/0643G06Q30/0629G06T11/00G06T19/20G06V10/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,327,280
App. No.
18/067,597
Granted
Jun 10, 2025
Kind
B2
Abstract

A method includes generating a virtual model of a human body based at least in part on a selected image of the human body and generating a segment of an article of clothing based at least in part on a selected image of the article of clothing. The method also includes generating a layer mask indicating whether a plurality of output pixels of an output image should be produced according to the image of the human body, the image of the shirt, or the image of the pair of pants and producing the plurality of output pixels of the output image according to the layer mask. The output image shows the article of clothing on the human body in the selected image of the human body.

Claims (80)

1. A method comprising:

receiving, from a server, first two-dimensional image data representing a human body;

generating, using a first neural network, a three-dimensional virtual model of the human body based at least in part on the first two-dimensional image data representing the human body;

receiving, from a first database, second two-dimensional image data representing a first article of clothing and third two-dimensional image data representing a second article of clothing;

generating, using a second neural network, segment features of the first article of clothing based at least in part on the second two-dimensional image data representing the first article of clothing;

determining, using the second neural network, a first clothing type based on the segment features of the first article of clothing, wherein the first clothing type includes data relating to where the first article of clothing is worn;

generating, using the second neural network, segment features of the second article of clothing based at least in part on the third two-dimensional image data representing the second article of clothing;

determining, using the second neural network, a second clothing type based on the segment features of the second article of clothing, wherein the second clothing type includes data relating to where the second article of clothing is worn;

reposing, using the second neural network, the segment features of the first article of clothing based at least in part on the three-dimensional virtual model of the human body and the first clothing type;

reposing, using the second neural network, the segment features of the second article of clothing based at least in part on the three-dimensional virtual model of the human body and the second clothing type;

determining, using the second neural network and based on the first clothing type and the second clothing type, that the reposed segment features of the first article of clothing and the reposed segment features of the second article of clothing overlap and define an overlapping region;

determining, using the second neural network, clothing positioning comprising positioning of the first article of clothing and the second article of clothing, wherein the clothing positioning defines which pixels of the first article of clothing and which pixels the second article of clothing will be used as output pixels for the overlapping region;

generating, using the second neural network, based at least in part on the three-dimensional virtual model of the human body, the reposed segment features of the first article of clothing, the reposed segment features of the second article of clothing, and the clothing positioning, a layer mask indicating whether a plurality of output pixels of an output image should be produced according to the first two-dimensional image data representing the human body, the second two-dimensional image data representing the first article of clothing, or the third-two-dimensional image data representing the second article of clothing;

generating, using a third neural network, texture features of the human body based at least in part on the first two-dimensional image data representing the human body;

generating, using the third neural network, texture features of the first article of clothing based at least in part on the second two-dimensional image data representing the first article of clothing;

generating, using the third neural network, texture features of the second article of clothing based at least in part on the third two-dimensional image data representing the second article of clothing;

reposing, using the third neural network, the texture features of the first article of clothing based at least in part on the three-dimensional virtual model of the human body;

reposing, using the third neural network, the texture features of the second article of clothing based at least in part on the three-dimensional virtual model of the human body;

sampling, using the third neural network, the texture features of the human body, the reposed texture features of the first article of clothing, and the reposed texture features of the second article of clothing according to the layer mask; and

producing, using a fourth neural network, the plurality of output pixels of the output image based at least in part on the sampled texture features of the human body, the sampled reposed texture features of the first article of clothing, and the sampled reposed texture features of the second article of clothing, wherein the output image shows the first article of clothing and the second article of clothing on the three-dimensional virtual model of the human body in the first two-dimensional image data representing the human body.

2. The method of claim 1 , wherein generating the layer mask comprises:

generating, using the second neural network, a segment of the first article of clothing based at least in part on the reposed segment features of the first article of clothing; and

generating, using the second neural network, a segment of the second article of clothing based at least in part on the reposed segment features of the second article of clothing, wherein the layer mask is generated based at least in part on the segment of the first article of clothing and the segment of the second article of clothing.

3. The method of claim 1 , wherein reposing the texture features of the first article of clothing comprises resizing or reorienting the texture features of the first article of clothing to fit onto the three-dimensional virtual model of the human body.

4. The method of claim 1 , wherein producing the plurality of output pixels of the output image comprises altering colors of the sampled texture features of the human body, the sampled reposed texture features of the first article of clothing, and the sampled reposed texture features of the second article of clothing.

5. A method comprising:

receiving, from a first server, first two-dimensional image data representing a human body;

generating, using at least one neural network executed by one or more processors, a three-dimensional virtual model of the human body based at least in part on the first two-dimensional image data representing the human body;

receiving, from a second server, second two-dimensional image data representing a first article of clothing and third two-dimensional image data representing a second article of clothing;

generating, using the at least one neural network executed by the one or more processors, segment features of the first article of clothing based at least in part on the second two-dimensional image data representing the first article of clothing;

determining, using the at least one neural network executed by the one or more processors, a first clothing type based on the segment features of the first article of clothing, wherein the first clothing type includes data relating to where the first article of clothing is worn;

generating, using the at least one neural network executed by the one or more processors, segment features of the second article of clothing based at least in part on the third two-dimensional image data representing the second article of clothing;

determining, using the at least one neural network executed by the one or more processors, a second clothing type based on the segment features of the second article of clothing, wherein the second clothing type includes data relating to where the second article of clothing is worn;

reposing, using the at least one neural network executed by the one or more processors, the segment features of the first article of clothing based at least in part on the three-dimensional virtual model of the human body;

reposing, using the at least one neural network executed by the one or more processors, the segment features of the second article of clothing based at least in part on the three-dimensional virtual model of the human body;

determining, using the at least one neural network executed by the one or more processors and based on the first clothing type and the second clothing type, that segment features of the first article of clothing and the segment features of the second article of clothing overlap and define an overlapping region;

determining, using the at least one neural network executed by the one or more processors, clothing positioning comprising positioning of the first article of clothing and the second article of clothing, wherein the clothing positioning defines which pixels of the first article of clothing and which pixels the second article of clothing will be used as output pixels for the overlapping region;

generating, using the at least one neural network executed by the one or more processors, based at least in part on the three-dimensional virtual model of the human body, the segment features of the first article of clothing, the segment features of the second article of clothing, and the clothing positioning, a layer mask indicating whether a plurality of output pixels of an output image should be produced according to the first two-dimensional image data representing the human body, the second two-dimensional image data representing the first article of clothing, or the third two-dimensional image data representing the second article of clothing;

generating, using the at least one neural network executed by the one or more processors, texture features of the human body based at least in part on the first two-dimensional image data representing the human body;

generating, using the at least one neural network executed by the one or more processors, texture features of the first article of clothing based at least in part on the second two-dimensional image data representing the first article of clothing;

generating, using the at least one neural network executed by the one or more processors, texture features of the second article of clothing based at least in part on the third two-dimensional image data representing the second article of clothing;

reposing, using the at least one neural network executed by the one or more processors, the texture features of the first article of clothing based at least in part on the three-dimensional virtual model of the human body;

reposing, using the at least one neural network executed by the one or more processors, the texture features of the second article of clothing based at least in part on the three-dimensional virtual model of the human body;

sampling, using the at least one neural network executed by the one or more processors, the texture features of the human body, the reposed texture features of the first article of clothing, and the reposed texture features of the second article of clothing according to the layer mask; and

producing, using the at least one neural network executed by the one or more processors, the plurality of output pixels of the output image based at least in part on the sampled texture features of the human body, the sampled reposed texture features of the first article of clothing, and the sampled reposed texture features of the second article of clothing according to the layer mask, wherein the output image shows the first article of clothing and the second article of clothing on the three-dimensional virtual model of the human body in the first two-dimensional image data representing the human body.

6. The method of claim 5 , further comprising:

generating, using the at least one neural network executed by the one or more processors, a segment of the first article of clothing based at least in part on the reposed segment features of the first article of clothing; and

generating, using the at least one neural network executed by the one or more processors, a segment of the second article of clothing based at least in part on the reposed segment features of the second article of clothing.

7. The method of claim 6 , wherein the layer mask is generated based at least in part on the segment of the first article of clothing and the segment of the second article of clothing.

8. The method of claim 5 , wherein reposing the segment features of the first article of clothing comprises resizing or reorienting the segment features of the first article of clothing to fit onto the three-dimensional virtual model of the human body.

9. The method of claim 5 , wherein generating the output image comprises altering colors of the sampled texture features of the human body, the sampled reposed texture features of the first article of clothing, and the sampled reposed texture features of the second article of clothing.

10. A system comprising:

a memory;

at least one neural network executed by one or more processors; and

one or more processors communicatively coupled to the memory, the one or more processors configured to:

receive, from a first server, first two dimensional image data representing a human body;

generate, using the at least one neural network executed by the one or more processors, a three-dimensional virtual model of the human body based at least in part on the first two-dimensional image data representing the human body;

receive, from a second server, second two-dimensional image data representing a first article of clothing and third image data representing a second article of clothing;

generate, using the at least one neural network executed by the one or more processors, segment features of the first article of clothing based at least in part on the second two-dimensional image data representing the first article of clothing;

determine, using the at least one neural network executed by the one or more processors, a first clothing type based on the segment features of the first article of clothing, wherein the first clothing type includes data relating to where the first article of clothing is worn;

generate, using the at least one neural network executed by the one or more processors, segment features of the second article of clothing based at least in part on the second two dimensional image data representing the second article of clothing;

determine, using the at least one neural network executed by the one or more processors, a second clothing type based on the segment features of the second article of clothing, wherein the second clothing type includes data relating to where the second article of clothing is worn;

repose, using the at least one neural network executed by the one or more processors, the segment features of the first article of clothing based at least in part on the three-dimensional virtual model of the human body;

repose, using the at least one neural network executed by the one or more processors, the segment features of the second article of clothing based at least in part on the three-dimensional virtual model of the human body;

determine, using the at least one neural network executed by the one or more processors and based on the first clothing type and the second clothing type, that segment features of the first article of clothing and the segment features of the second article of clothing overlap and define an overlapping region;

determine, using the at least one neural network executed by the one or more processors, clothing positioning comprising positioning of the first article of clothing and the second article of clothing, wherein the clothing positioning defines which pixels of the first article of clothing and which pixels the second article of clothing will be used as output pixels for the overlapping region;

generate, using the at least one neural network executed by the one or more processors, based at least in part on the three-dimensional virtual model of the human body, the segment features of the first article of clothing, the segment features of the second article of clothing, and the clothing positioning, a layer mask indicating whether a plurality of output pixels of an output image should be produced according to the first two-dimensional image data representing the human body, the second two-dimensional image data representing the first article of clothing, or the third two-dimensional image data representing the second article of clothing;

generate, using the at least one neural network executed by the one or more processors, texture features of the human body based at least in part on the first two-dimensional image data representing the human body;

generate, using the at least one neural network executed by the one or more processors, texture features of the first article of clothing based at least in part on the second two-dimensional image data representing the first article of clothing;

generate, using the at least one neural network executed by the one or more processors, texture features of the second article of clothing based at least in part on the third two-dimensional image data representing the second article of clothing;

repose, using the at least one neural network executed by the one or more processors, the texture features of the first article of clothing based at least in part on the three-dimensional virtual model of the human body;

repose, using the at least one neural network executed by the one or more processors, the texture features of the second article of clothing based at least in part on the three-dimensional virtual model of the human body;

sample, using the at least one neural network executed by the one or more processors, the texture features of the human body, the reposed texture features of the first article of clothing, and the reposed texture features of the second article of clothing according to the layer mask;

produce, using the at least one neural network executed by the one or more processors, the plurality of output pixels of the output image according to the layer mask based at least in part on the sampled texture features of the human body, the sampled reposed texture features of the first article of clothing, and the sampled reposed texture features of the second article of clothing, wherein the output image shows the first article of clothing and the second article of clothing on the three-dimensional virtual model of the human body in the first two-dimensional image data representing the human body.

11. The system of claim 10 , wherein the one or more processors are further configured to:

generate, using the at least one neural network executed by the one or more processors, a segment of the first article of clothing based at least in part on the reposed segment features of the first article of clothing; and

generate, using the at least one neural network executed by the one or more processors, a segment of the second article of clothing based at least in part on the reposed segment features of the second article of clothing.

12. The system of claim 11 , wherein the layer mask is generated based at least in part on the segment of the first article of clothing and the segment of the second article of clothing.

13. The system of claim 10 , wherein reposing the segment features of the first article of clothing comprises resizing or reorienting the segment features of the first article of clothing to fit onto the three-dimensional virtual model of the human body.

14. The system of claim 10 , wherein generating the output image comprises altering colors of the sampled texture features of the human body, the sampled reposed texture features of the first article of clothing, and the sampled reposed texture features of the second article of clothing.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2023
From: BIGGS, BENJAMIN JAMES; PINETTE, PHILIP
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 065236/0653 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2022
From: KOTHARI, CHARU; CAGAMPAN, CAITLIN ISAAC; MEDIONI, GERARD GUY; DAVE, ACHAL DUSHYANT; SUN, SCOTT CHENGHUI
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 062131/0622 →
Continuity (1)
Related Publication 20240202809A1 · Jun 20, 2024
References Cited (11)
US 11200689B1 · Smith · 2021 [cited by examiner]
US 12079961B2 · Song · 2024 [cited by examiner]
US 12100156B2 · Dudovitch · 2024 [cited by examiner]
US 20200151807A1 · Zhou · 2020 [cited by examiner]
US 20210142539A1 · Ayush · 2021 [cited by examiner]
US 20210241531A1 · Lee et al. · 2021 [cited by applicant]
US 20220318892A1 · Lee · 2022 [cited by examiner]
Aiyu Cui, Daniel McKee, Svetlana Lazebnik, “Dressing in Order: Recurrent Person Image Generation for Pose Transfer, Virtual Try-On and Outfit Editing,” Oct. 18, 2022, University of Illinois at Urbana-Champaign, pp. 1-19… [cited by examiner]
Demo/Looklet Dressing Room, retrieved Dec. 15, 2022, <https://dressing-room.looklet.com/>. [cited by applicant]
Grigorev et al.; Coordinate-based Texture Inpainting for Pose-Guided Human Image Generation; 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE; Jun. 15, 2019; pp. 12127-12136. [cited by applicant]
Pomocka; International Search Report of PCT/US2023/081339; Apr. 15, 2024; 6 pgs. [cited by applicant]