Image inpainting using local content preservation
Methods, non-transitory computer readable media, apparatuses, and systems for image inpainting include obtaining, via a user interface, an input image and a local content preservation value and receiving a content generation selection. The content generation selection applies the local content preservation value to at least one pixel of the input image. An image generation model then generates an output image based on the input image and the content generation selection. The output image includes synthetic content in a region specified by the content generation selection, and a degree of adherence of the synthetic content to the input image is based on the local content preservation value.
1 . A method for image generation, comprising:
obtaining, via a user interface, a text prompt, an input image, and a local content preservation value;
receiving a content generation selection, wherein the content generation selection applies the local content preservation value to at least one pixel of the input image; and
generating, using an image generation process guided by the text prompt and performed by an image generation model, an output image based on the input image and the content generation selection, wherein the output image includes synthetic content in a region specified by the content generation selection, and wherein a weight of an effect of the text prompt on the image generation process corresponds to the local content preservation value.
2 . The method of claim 1 , further comprising:
receiving, via a brush size input element, a brush size input, wherein the content generation selection is based at least in part on the brush size input.
3 . The method of claim 1 , further comprising:
receiving, via a brush hardness input element, a brush hardness input, wherein the content generation selection is based at least in part on the brush hardness input.
4 . The method of claim 1 , further comprising:
receiving, via a local content preservation input element, an additional local content preservation value from a user; and
receiving an additional content generation selection from the user, wherein the additional content generation selection applies the additional local content preservation value to at least one additional pixel of the input image, and wherein the output image is further based on the additional content generation selection.
5 . The method of claim 1 , further comprising:
receiving, via a global content preservation input element, a global content preservation value, wherein the output image is further based on the global content preservation value.
6 . The method of claim 5 , further comprising:
combining the local content preservation value and the global content preservation value to obtain a combined content preservation value for the at least one pixel, wherein the output image is based on the combined content preservation value.
7 . The method of claim 1 , further comprising:
receiving, via a match shape input element, a match shape input, wherein a closeness of a match between the synthetic content of the output image and the content generation selection is based on the match shape input.
8 . The method of claim 1 , further comprising:
receiving, via a guidance strength input element, a guidance strength input, wherein a closeness of a match between the synthetic content of the output image and the text prompt is based on the guidance strength input.
9 . The method of claim 1 , further comprising:
receiving, via an insert content input element, an insert content input, wherein the output image is based on the insert content input.
10 . The method of claim 1 , further comprising:
receiving, via a remove content input element, a remove content input, wherein the output image is based on the remove content input.
11 . The method of claim 1 , further comprising:
saving the content generation selection as metadata in an image layer of the output image.
12 . The method of claim 1 , further comprising:
generating a plurality of variations of the output image;
displaying a preview of each of the plurality of variations of the output image; and
receiving a selection input indicating the output image from among the plurality of variations.
13 . The method of claim 1 , further comprising:
displaying a content generation selection overlay depicting the content generation selection overlapping the input image, wherein a transparency value of the content generation selection overlay is based on the local content preservation value.
14 . The method of claim 1 , further comprising:
displaying the output image including the synthetic content at a location of the content generation selection.
15 . A non-transitory computer readable medium storing instructions that, when executed by a processor, cause the processor to:
obtain an input image;
receive, via text input element, a text prompt;
receive, via a local content preservation input element, a local content preservation value from a user;
receive, via a content generation brush tool, a content generation selection from the user, wherein the content generation selection applies the local content preservation value to at least one pixel of the input image; and
generate, using an image generation process guided by the text prompt and performed by an image generation model, an output image based on the input image, wherein the output image includes content based on the text prompt in a region specified by the content generation selection, and wherein a weight of an effect of the text prompt on the image generation process corresponds to the local content preservation value.
16 . A system for image generation, comprising:
one or more processors;
one or more memory components coupled with the one or more processors;
a user interface displaying:
a text input element configured to receive a text prompt,
a local content preservation input element configured to receive a local content preservation value from a user, and
a content generation brush tool configured to receive a content generation selection from the user, wherein the content generation selection applies the local content preservation value to at least one pixel of an input image; and
an image generation model configured to perform an image generation process to generate an output image based on the input image, wherein the output image includes content based on the text prompt in a region specified by the content generation selection, and wherein a weight of an effect of the text prompt on the image generation process corresponds to the local content preservation value.
17 . The system of claim 16 , the user interface further displaying:
a brush size input element configured to receive a brush size input, wherein the content generation selection is based at least in part on the brush size input.
18 . The system of claim 16 , the user interface further displaying:
a brush hardness input element configured to receive a brush hardness input, wherein the content generation selection is based at least in part on the brush hardness input.
19 . The system of claim 16 , wherein:
the local content preservation input element is further configured to receive an additional local content preservation value from the user; and
the content generation brush tool is further configured to receive an additional content generation selection from the user, wherein the additional content generation selection applies the additional local content preservation value to at least one additional pixel of the input image, and wherein the output image is further based on the additional content generation selection.