IP Library Granted Patent US 12682508
Granted Patent B2
US 12682508 · App. 18/425,960 · Granted Jul 14, 2026

N-flows estimation for garment transfer

Inventors: Avihay Assouline (Tel Aviv, IL); Amir Fruchtman (Holon, IL); Riza Alp Guler (London, GB); Jonathan Heimann (Herzliya, IL); Nir Malbin (Shoham, IL)
Assignee: SNAP INC.
G06T11/00G06T3/40G06T7/248G06T7/73G06T2207/20081G06T2207/20084G06T2207/20221G06T2207/30196G06T2210/16G06T2210/36
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682508
App. No.
18/425,960
Granted
Jul 14, 2026
Kind
B2
Abstract

Methods and systems are disclosed for using machine learning models to perform pixel-based deformation of fashion items using N-flows estimation. The methods and systems receive one or more images depicting a first person in a first pose and receive a source image depicting a target fashion item worn on a portion of a body of a second person in a second pose, the target fashion item comprising a plurality of parts. The methods and systems process, using one or more machine learning models, the one or more images together with the source image to generate a set of data comprising a plurality of flow fields each associated with a different one of the plurality of parts of the target fashion item. The methods and systems modify a portion of the one or more images to overlay the target fashion item on the first person in the first pose.

Claims (69)

1 . A method comprising:

receiving, by one or more processors, one or more images depicting a first person in a first pose;

receiving a source image depicting a target fashion item worn on a portion of a body of a second person in a second pose, the target fashion item comprising a plurality of parts of an individual garment worn on an individual body part of the second portion;

processing, using one or more machine learning models, the one or more images together with the source image to generate a set of data comprising a plurality of flow fields each associated with a different one of the plurality of parts of the individual garment worn on the individual body part of the second person, each of the plurality of flow fields indicating existence and location of each pixel of the one or more images in the source image, the set of data comprising source part masks and flow selection information;

sampling the source image using the plurality of flow fields to select one or more pixels from the source image to replace one or more target pixels in the one or more images, the sampling comprising:

selecting an individual pixel from a first flow field of the plurality of flow fields;

determining, based on the flow selection information, a first part of the plurality of parts that is associated with the individual pixel;

obtaining a first source part mask of the source part masks that corresponds to the first part; and

determining whether the individual pixel is included in the first source part mask, the one or more pixels being selected from the source image to replace the one or more target pixels in the one or more images based on determining whether the individual pixel is included in the first source part mask; and

modifying, based on the set of data generated by the one or more machine learning models, a portion of the one or more images to overlay the target fashion item on the first person in the first pose.

2 . The method of claim 1 , wherein the source part masks specify correspondence between the plurality of parts of the target fashion item and each pixel in the source image.

3 . The method of claim 1 , wherein the flow selection information specifies a correspondence between the plurality of parts of the target fashion item and each pixel in the one or more images.

4 . The method of claim 1 , further comprising:

in response to determining that the individual pixel is included in the first source part mask, replacing the individual pixel in the one or more images based on a corresponding pixel of the one or more pixels in the first source part mask.

5 . The method of claim 1 , wherein sampling the source image comprises:

selecting an additional individual pixel from the second flow field of the plurality of flow fields;

determining, based on the flow selection information, a second part of the plurality of parts that is associated with the additional individual pixel;

obtaining a second source part mask of the source part masks that corresponds to the second part; and

determining whether the additional individual pixel is included in the second source part mask, the one or more pixels being selected from the source image to replace the one or more target pixels in the one or more images based on determining whether the additional individual pixel is included in the second source part mask.

6 . The method of claim 1 , wherein the plurality of parts comprise upper arms, lower arms, left arm, right arm, and torso.

7 . The method of claim 1 , wherein the first flow field comprises a first image resolution, wherein the modified portion of the one or more images comprising the target fashion item comprises a second image resolution that is greater than the first image resolution.

8 . The method of claim 7 , wherein an image resolution of each of the plurality of flow fields is lower than the modified portion of the one or more images.

9 . The method of claim 7 , wherein the first image resolution is 33×33 pixels.

10 . The method of claim 1 , further comprising:

applying a pose estimation machine learning model to the one or more images to generate first pose estimation information representing the first pose of the first person; and

applying the pose estimation machine learning model to the source image to generate second pose estimation information representing the second pose of the second person.

11 . The method of claim 10 , further comprising:

processing, by a flow estimation machine learning model, the first pose estimation information and the second pose estimation information, to generate the set of data.

12 . The method of claim 1 , further comprising:

sampling one or more pixels of the source image based on one of the flow fields to extract and adjust a pose of the target fashion item to match the first pose of the first person.

13 . The method of claim 1 , wherein the one or more machine learning models are trained by performing training operations comprising:

accessing training data comprising a first training image depicting a first training object in a first training pose, a second training image depicting a second training object in a second training pose, and a set of ground truth flow fields and a set of ground truth segmentation masks;

analyzing, using the one or more machine learning models, the first and second training images to estimate a first flow field for a first body part using a first ground truth segmentation mask of the set of ground truth segmentation masks corresponding to the first body part;

analyzing, using the one or more machine learning models, the first and second training images to estimate a second flow field for a second body part using a second ground truth segmentation mask of the set of ground truth segmentation masks corresponding to the second body part;

computing a loss based on a deviation between the estimated first and second flow fields for the first and second training images and the set of ground flow fields; and

updating one or more parameters of the one or more machine learning models based on the computed loss.

14 . The method of claim 13 , further comprising generating the first and second training images by performing operations comprising:

accessing a first set of images depicting one or more objects;

accessing a second set of images depicting one or more fashion items;

generating the set of ground truth segmentation masks that segment individual parts of the one or more fashion items;

computing a first ground truth flow field indicating existence and location in the second set of images of a first part of the one or more objects depicted in the first set of images;

computing a second ground truth flow field indicating existence and location in the second set of images of a second part of the one or more objects depicted in the first set of images; and

associating the first and second ground truth flow fields and the set of ground truth segmentation masks with the first and second sets of images.

15 . A system comprising:

at least one processor; and

at least one memory component having instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

receiving one or more images depicting a first person in a first pose;

receiving a source image depicting a target fashion item worn on a portion of a body of a second person in a second pose, the target fashion item comprising a plurality of parts of an individual garment worn on an individual body part of the second portion;

processing, using one or more machine learning models, the one or more images together with the source image to generate a set of data comprising a plurality of flow fields each associated with a different one of the plurality of parts of the individual garment worn on the individual body part of the second person, each of the plurality of flow fields indicating existence and location of each pixel of the one or more images in the source image, the set of data comprising source part masks and flow selection information;

sampling the source image using the plurality of flow fields to select one or more pixels from the source image to replace one or more target pixels in the one or more images, the sampling comprising:

selecting an individual pixel from a first flow field of the plurality of flow fields;

determining, based on the flow selection information, a first part of the plurality of parts that is associated with the individual pixel;

obtaining a first source part mask of the source part masks that corresponds to the first part; and

determining whether the individual pixel is included in the first source part mask, the one or more pixels being selected from the source image to replace the one or more target pixels in the one or more images based on determining whether the individual pixel is included in the first source part mask; and

modifying, based on the set of data generated by the one or more machine learning models, a portion of the one or more images to overlay the target fashion item on the first person in the first pose.

16 . A non-transitory computer-readable storage medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising:

receiving one or more images depicting a first person in a first pose;

receiving a source image depicting a target fashion item worn on a portion of a body of a second person in a second pose, the target fashion item comprising a plurality of parts of an individual garment worn on an individual body part of the second portion;

processing, using one or more machine learning models, the one or more images together with the source image to generate a set of data comprising a plurality of flow fields each associated with a different one of the plurality of parts of the individual garment worn on the individual body part of the second person, each of the plurality of flow fields indicating existence and location of each pixel of the one or more images in the source image, the set of data comprising source part masks and flow selection information;

sampling the source image using the plurality of flow fields to select one or more pixels from the source image to replace one or more target pixels in the one or more images, the sampling comprising:

selecting an individual pixel from a first flow field of the plurality of flow fields;

determining, based on the flow selection information, a first part of the plurality of parts that is associated with the individual pixel;

obtaining a first source part mask of the source part masks that corresponds to the first part; and

determining whether the individual pixel is included in the first source part mask, the one or more pixels being selected from the source image to replace the one or more target pixels in the one or more images based on determining whether the individual pixel is included in the first source part mask; and

modifying, based on the set of data generated by the one or more machine learning models, a portion of the one or more images to overlay the target fashion item on the first person in the first pose.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein the source part masks specify correspondence between the plurality of parts of the target fashion item and each pixel in the source image.

18 . The non-transitory computer-readable storage medium of claim 16 , wherein the flow selection information specifies a correspondence between the plurality of parts of the target fashion item and each pixel in the one or more images.

19 . The non-transitory computer-readable storage medium of claim 16 , wherein the plurality of parts comprise upper arms, lower arms, left arm, right arm, and torso.

20 . The non-transitory computer-readable storage medium of claim 16 , wherein the first flow field comprises a first image resolution, wherein the modified portion of the one or more images comprising the target fashion item comprises a second image resolution that is greater than the first image resolution.