IP Library Granted Patent US 12,437,367
Granted Patent B2
US 12,437,367 · App. 18/068,383 · Granted Oct 7, 2025

Real-time try-on using body landmarks

Inventors: Avihay Assouline (Tel Aviv, IL); Nir Malbin (Shoham, IL); Iason Kokkinos (London, GB); Riza Alp Guler (London, GB); Himmy Tam (London, GB); Mohammad Rami Koujan (London, GB)
Assignee: SNAP INC.
G06T5/50G06T7/70G06V10/761G06V10/82G06T2207/20221G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,367
App. No.
18/068,383
Granted
Oct 7, 2025
Kind
B2
Abstract

Methods and systems are disclosed for transferring garments from one real-world object to another in real time using body landmarks. The system receives a first image that includes a depiction of a first person wearing a fashion item in a first pose. The system obtains a second image that includes a depiction of a second person in a second pose and generates a first set of body landmarks corresponding the first person in the first pose and a second set of body landmarks corresponding the second person wearing in the first pose. The system computes a deviation between the first set of body landmarks and the second set of body landmarks. The system generates a new image that depicts the second person wearing the fashion item worn by the first person based on the deviation between the first set of body landmarks and the second set of body landmarks.

Claims (60)

1. A method comprising:

receiving, by one or more processors, a first image that includes a depiction of a first person wearing a fashion item in a first pose;

obtaining a second image that includes a depiction of a second person in a second pose;

generating a first set of body landmarks corresponding the first person in the first pose and a second set of body landmarks corresponding the second person in the second pose;

computing a deviation between the first set of body landmarks and the second set of body landmarks;

determining, based on the first set of body landmarks and the second set of body landmarks, one or more adjustments to the first pose of the first person, depicted in the first image that has been captured by an image capture device, that correspond to the second pose of the second person and modifications to one or more visual parameters of the fashion item depicted in the first image as being worn by the first person that correspond to the second pose of the second person to enable placement of the fashion item worn by the first person on the second person; and

generating a new image that depicts the second person wearing the fashion item worn by the first person based on the deviation between the first set of body landmarks and the second set of body landmarks and the determined one or more adjustments and modifications.

2. The method of claim 1 , further comprising:

modifying the first set of body landmarks associated with the first person to match the second set of body landmarks associated with the second person based on the deviation;

applying a fitting model to the first and second sets of body landmarks to adjust the one or more visual parameters of the fashion item corresponding to the modified first set of body landmarks; and

generating an intermediate image by the fitting model depicting the fashion item with the adjusted one or more visual parameters overlaid on the second person depicted in the second image.

3. The method of claim 2 , further comprising:

applying a generative machine learning model to the intermediate image to render the new image, the generative machine learning model being configured to blend sets of pixels corresponding to one or more gaps or occlusions that appear in the intermediate image and adjust for differences in lighting conditions and skin tones of users depicted in images.

4. The method of claim 2 , wherein the fitting model comprises a parametrized machine learning model.

5. The method of claim 2 , wherein the fitting model comprises a non-learned model.

6. The method of claim 2 , further comprising feeding an output of the fitting model to a first machine learning model used to generate the first and second sets of body landmarks.

7. The method of claim 6 , wherein generating the first and second sets of body landmarks comprises:

applying a body landmarks model comprising the first machine learning model to the first image; and

applying the body landmarks model comprising the first machine learning model to the second image.

8. The method of claim 6 , wherein the new image is generated by a second machine learning model.

9. The method of claim 8 , further comprising training the first and second machine learning models by iterating through a sequence of training operations comprising:

receiving a first training image that depicts a training person in a first training pose and wearing a training fashion item;

receiving a training video that depicts the training person in a second training pose;

applying the first machine learning model to the first training image and a given frame of the training video to generate first and second sets of estimated body landmarks associated with the training person;

computing fit between the first and second sets of estimated body landmarks; and

applying the computed fit between the first and second sets of estimated body landmarks associated with the training person to the second machine learning model to generate a depiction of the training person in the second training pose wearing the training fashion item.

10. The method of claim 9 , further comprising:

computing a deviation between the generated depiction of the training person in the second training pose wearing the training fashion item and the given frame of the training video; and

updating one or more parameters of the first and second machine learning models based on the computed deviation.

11. The method of claim 10 , wherein the second machine learning model comprises a neural network comprising a generative adversarial network (GAN).

12. The method of claim 1 , further comprising:

extracting the second image from a real-time video feed captured by a camera of a client device.

13. The method of claim 12 , wherein the real-time video feed is captured using a rear-facing camera of the client device.

14. The method of claim 1 , wherein a video comprising the new image is generated in real-time as the second image is being captured.

15. A system comprising:

at least one processor; and

a memory component having instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

receiving a first image that includes a depiction of a first person wearing a fashion item in a first pose;

obtaining a second image that includes a depiction of a second person in a second pose;

generating a first set of body landmarks corresponding the first person in the first pose and a second set of body landmarks corresponding the second person wearing in the first pose;

computing a deviation between the first set of body landmarks and the second set of body landmarks;

determining, based on the first set of body landmarks and the second set of body landmarks, one or more adjustments to the first pose of the first person, depicted in the first image that has been captured by an image capture device, that correspond to the second pose of the second person and modifications to one or more visual parameters of the fashion item depicted in the first image as being worn by the first person that correspond to the second pose of the second person to enable placement of the fashion item worn by the first person on the second person; and

generating a new image that depicts the second person wearing the fashion item worn by the first person based on the deviation between the first set of body landmarks and the second set of body landmarks and the determined one or more adjustments and modifications.

16. The system of claim 15 , the operations further comprising:

modifying the first set of body landmarks associated with the first person to match the second set of body landmarks associated with the second person based on the deviation;

applying a fitting model to the first and second sets of body landmarks to adjust the one or more visual parameters of the fashion item corresponding to the modified first set of body landmarks; and

generating an intermediate image by the fitting model depicting the fashion item with the adjusted one or more visual parameters overlaid on the second person depicted in the second image.

17. The system of claim 16 , the operations further comprising:

applying a generative machine learning model to the intermediate image to render the new image, the generative machine learning model being configured to blend sets of pixels corresponding to one or more gaps or occlusions that appear in the intermediate image and adjust for differences in lighting conditions and skin tones of users depicted in images.

18. The system of claim 16 , the operations further comprising:

capturing the first image that includes the depiction of the first person by a first camera of a user device that points towards a first direction; and

capturing the second image that includes the depiction of the second person by a second camera of the user device that points towards a second direction that is opposite the first direction of the first camera used to capture the first image.

19. The system of claim 16 , wherein the fitting model comprises a non-learned model.

20. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

receiving a first image that includes a depiction of a first person wearing a fashion item in a first pose;

obtaining a second image that includes a depiction of a second person in a second pose;

generating a first set of body landmarks corresponding the first person in the first pose and a second set of body landmarks corresponding the second person wearing in the first pose;

computing a deviation between the first set of body landmarks and the second set of body landmarks;

determining, based on the first set of body landmarks and the second set of body landmarks, one or more adjustments to the first pose of the first person, depicted in the first image that has been captured by an image capture device, that correspond to the second pose of the second person and modifications to one or more visual parameters of the fashion item depicted in the first image as being worn by the first person that correspond to the second pose of the second person to enable placement of the fashion item worn by the first person on the second person; and

generating a new image that depicts the second person wearing the fashion item worn by the first person based on the deviation between the first set of body landmarks and the second set of body landmarks and the determined one or more adjustments and modifications.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2022
From: ASSOULINE, AVIHAY; MALBIN, NIR; KOKKINOS, IASON; GULER, RIZA ALP; TAM, HIMMY; RAMI KOUJAN, MOHAMMAD
To: SNAP INC.
Reel/Frame 062148/0020 →
Priority Claims (1)
GR 20220100947 · Nov 16, 2022 · national
Continuity (1)
Related Publication 20240161242A1 · May 16, 2024
References Cited (28)
US 20050154487A1 · Wang · 2005 [cited by applicant]
US 20140126769A1 · Reitmayr et al. · 2014 [cited by applicant]
US 20140270357A1 · Hampiholi · 2014 [cited by examiner]
US 20150279098A1 · Kim et al. · 2015 [cited by applicant]
US 20190371080A1 · Sminchisescu · 2019 [cited by examiner]
US 20210133919A1 · Ayush · 2021 [cited by examiner]
US 20210275925A1 · Kolen · 2021 [cited by examiner]
US 20240290043A1 · Zhou et al. · 2024 [cited by applicant]
KR 102381566B1 · 2022 [cited by applicant]
KR 20220053739A · 2022 [cited by applicant]
KR 20230007255A · 2023 [cited by applicant]
WO WO2024107634A1 · 2024 [cited by applicant]
WO WO2024177859A1 · 2024 [cited by applicant]
Roy, Debapriya, “LGVTON: a landmark guided approach for model to person virtual try-on” (Jan. 8, 2022), Multimedia Tools and Applications, Kluwer Academic Publishers, Boston, US, vol. 81, No. 4, (Jan. 8, 2022), 37 pgs. [cited by examiner]
Cao , Z., “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields” (Year 2017), Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (Year 2017), 9 pgs. [cited by examiner]
“International Application Serial No. PCT US2023 079493, International Search Report mailed Mar. 25, 2024”, 4 pgs. [cited by applicant]
“International Application Serial No. PCT US2023 079493, Written Opinion mailed Mar. 25, 2024”, 7 pgs. [cited by applicant]
Cao, Z., “Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (2017), 9 pgs. [cited by applicant]
Roy, Debapriya, “LGVTON: a landmark guided approach for model to person virtual try-on”, Multimedia Tools and Applications, Kluwer Academic Publishers, Boston, US, vol. 81, No. 4, (Jan. 8, 2022), 37 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/015779, International Search Report mailed Jun. 18, 2024”, 4 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/015779, Written Opinion mailed Jun. 18, 2024”, 4 pgs. [cited by applicant]
“U.S. Appl. No. 18/135,599, Examiner Interview Summary mailed Mar. 20, 2025”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 18/135,599, Non Final Office Action mailed Feb. 20, 2025”, 24 pgs. [cited by applicant]
“U.S. Appl. No. 18/135,599, Notice of Allowance mailed Apr. 16, 2025”, 8 pgs. [cited by applicant]
“U.S. Appl. No. 18/135,599, Response filed Apr. 2, 2025 to Non Final Office Action mailed Feb. 20, 2025”, 14 pgs. [cited by applicant]
Brownridge, Andrew, et al., “Body Scanning for Avatar Production and Animation”, (2014), 13 pgs. [cited by applicant]
Su, Zhaoqi, et al., “MulayCap: Multi-layer Human Performance Capture Using A Monocular Video Camera”, arXiv:2004.05815v3 [cs.CV], (Oct. 1, 2020), 18 pgs. [cited by applicant]
Yuan, Miaolong, et al., “A Mixed Reality Virtual Clothes Try-On System”, IEEE Transactions On Multimedia, vol. 15, No. 8, (Dec. 2013), 1958-1968. [cited by applicant]
Cited By (1)
US 12,682,461