IP Library Granted Patent US 12664701
Granted Patent B1
US 12664701 · App. 18/121,418 · Granted Jun 23, 2026

Few-shot item inpainting in input images

Inventors: Mehmet Saygin Seyfioglu (Seattle, WA); Karim Bouyarmane (Seattle, WA); Suren Kumar (Seattle, WA); Amirhossein Tavanaei (Mason, OH); Ismail Baha Tutar (Seattle, WA)
Assignee: AMAZON TECHNOLOGIES, INC.
G06T11/60G06F40/126G06F40/284G06F40/40G06T5/77G06T2200/24G06T2207/20081G06T2207/20084G06T2207/20092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664701
App. No.
18/121,418
Granted
Jun 23, 2026
Kind
B1
Abstract

Techniques are generally described for item inpainting in images using a small number of reference images without 3D models. In various examples, a first selection of a first item may be received. In some further examples, a first set of weights learned for a generative latent representation model fine-tuned using at least one image of the first item may be determined. In some cases, first user-input image data representing a target environment may be received. In various examples, the generative latent representation model may generate, using the first set of weights, first output image data representing a representation of the first item within the target environment.

Claims (62)

1 . A computer-implemented method, comprising:

receiving a selection of a first item on a graphical user interface;

determining a first set of weights associated with the first item, wherein the first set of weights are learned by fine-tuning an encoder of a pre-trained latent diffusion model using randomly-masked images of the first item;

loading the first set of weights into first non-transitory computer-readable memory associated with the pre-trained latent diffusion model;

receiving a first user-input image depicting a target environment;

receiving user-defined mask data identifying an area within the first user-input image;

inputting the first user-input image and the user-defined mask data into the pre-trained latent diffusion model, wherein the pre-trained latent diffusion model is loaded with the first set of weights;

generating, by the pre-trained latent diffusion model, a first output image depicting a representation of the first item within the target environment; and

displaying, on the graphical user interface, the first output image.

2 . The computer-implemented method of claim 1 , further comprising:

determining a second set of weights associated with the first item, wherein the second set of weights are learned by fine-tuning a text encoder of the pre-trained latent diffusion model using a text input describing a type of the first item and first token data uniquely identifying the first item; and

in response to receiving the selection of the first item, loading the second set of weights into second non-transitory computer-readable memory associated with a text encoder of the pre-trained latent diffusion model, wherein the user-defined mask data comprises first user-input text describing an object appearing in the first user-input image to be replaced by the first item.

3 . The computer-implemented method of claim 2 , further comprising:

receiving, in a field of the graphical user interface, second user-input text describing a desired condition of an appearance of the first item; and

generating, by the pre-trained latent diffusion model, a second output image depicting a second representation of the first item within the target environment, wherein the second representation corresponds to the desired condition of the appearance of the first item.

4 . A method comprising:

receiving, by a first graphical user interface, a first selection of a first item;

determining, based at least in part on the first selection of the first item, a first set of weights learned for a generative machine learning model fine-tuned using at least one image of the first item;

receiving, by the first graphical user interface, first user-input image data representing a target environment; and

generating, by the generative machine learning model using the first set of weights, first output image data representing a representation of the first item within the target environment.

5 . The method of claim 4 , further comprising:

receiving, by the first graphical user interface, first mask data identifying a location in the target environment for replacement by the representation of the first item, wherein the generative machine learning model generates the first output image data further based at least in part on the first mask data.

6 . The method of claim 4 , further comprising:

receiving, by the first graphical user interface, first text data identifying an object represented in the first user-input image data for replacement by the representation of the first item, wherein the generative machine learning model generates the first output image data further based at least in part on the first text data.

7 . The method of claim 4 , wherein the generative machine learning model is a pre-trained generative latent representation model, the method further comprising:

generating a plurality of masked images of the first item; and

determining the first set of weights associated with the first item based at least in part on the plurality of masked images of the first item.

8 . The method of claim 4 , further comprising:

storing the first set of weights in non-transitory computer-readable memory in association with first identifier data identifying the first item; and

loading the first set of weights based at least in part on the first selection of the first item.

9 . The method of claim 4 , wherein the generative machine learning model is a pre-trained Stable Diffusion model, the method further comprising:

determining the first set of weights for a UNet of the pre-trained Stable Diffusion model while maintaining weights of a variational autoencoder of the pre-trained Stable Diffusion model.

10 . The method of claim 4 , further comprising:

receiving, in a field of the first graphical user interface, user-input text describing a first condition of an appearance of the first item; and

generating, by the generative machine learning model, a second output image data representing a second representation of the first item within the target environment, wherein the second representation corresponds to the first condition of the appearance of the first item.

11 . The method of claim 4 , further comprising:

determining, based at least in part on the first selection of the first item, a second set of weights learned for a text encoder of the generative machine learning model, wherein the second set of weights learned for the text encoder are learned based at least in part on the at least one image of the first item, first text data describing a type of the first item, a first identifier data uniquely identifying the first item.

12 . The method of claim 4 , wherein the generative machine learning model is pre-trained and wherein the first set of weights are learned by updating pre-trained weights of a UNet model of the generative machine learning model using training images consisting of one or more images of the first item.

13 . A system comprising:

at least one processor; and

non-transitory computer-readable memory storing instructions that, when executed by the at least one processor, are effective to:

receive, at a first graphical user interface, a first selection of a first item;

determine, based at least in part on the first selection of the first item, a first set of weights learned for a generative machine learning model fine-tuned using at least one image of the first item;

receive, by the first graphical user interface, first user-input image data representing a target environment; and

generate, by the generative machine learning model using the first set of weights, first output image data representing a representation of the first item within the target environment.

14 . The system of claim 13 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

receive, by the first graphical user interface, first mask data identifying a location in the target environment for replacement by the representation of the first item, wherein the generative machine learning model generates the first output image data further based at least in part on the first mask data.

15 . The system of claim 13 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

receive, by the first graphical user interface, first text data identifying an object represented in the first user-input image data for replacement by the representation of the first item, wherein the generative machine learning model generates the first output image data further based at least in part on the first text data.

16 . The system of claim 13 , wherein the generative machine learning model is a pre-trained generative machine learning model, the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

generate a plurality of masked images of the first item; and

determine the first set of weights associated with the first item based at least in part on the plurality of masked images of the first item.

17 . The system of claim 13 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

store the first set of weights in non-transitory computer-readable memory in association with first identifier data identifying the first item; and

load the first set of weights based at least in part on the first selection of the first item.

18 . The system of claim 13 , wherein the generative machine learning model is a pre-trained Stable Diffusion model, the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

determine the first set of weights for a UNet of the pre-trained Stable Diffusion model while maintaining weights of a variational autoencoder of the pre-trained Stable Diffusion model.

19 . The system of claim 13 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

receive, in a field of the first graphical user interface, user-input text describing a first condition of an appearance of the first item; and

generate, by the generative machine learning model, a second output image data representing a second representation of the first item within the target environment, wherein the second representation corresponds to the first condition of the appearance of the first item.

20 . The system of claim 13 , the non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to:

determine, based at least in part on the first selection of the first item, a second set of weights learned for a text encoder of the generative machine learning model, wherein the second set of weights learned for the text encoder are learned based at least in part on the at least one image of the first item, first text data describing a type of the first item, a first identifier data uniquely identifying the first item.