Generating color-edited digital images utilizing a content aware diffusion neural network
This disclosure describes one or more implementations of systems, non-transitory computer-readable media, and methods that trains (and utilizes) an image color editing diffusion neural network to generate a color edited digital image(s) for a digital image. In particular, in one or more implementations, the disclosed systems identify a digital image depicting content in a first color style. Moreover, the disclosed systems generate, from the digital image utilizing an image color editing diffusion neural network, a color-edited digital image depicting the content in a second color style different from the first color style. Further, the disclosed systems provide, for display within a graphical user interface, the color-edited digital image.
1 . A method comprising:
identifying a digital image depicting content in a first color style;
identifying a noise representation corresponding to a second color style from an additional digital image different than the digital image;
generating a denoised image representation corresponding to the second color style utilizing an image color editing diffusion neural network to remove noise from the noise representation based on conditioning the image color editing diffusion neural network on the digital image depicting content in the first color style;
generating, utilizing the image color editing diffusion neural network, a color-edited digital image depicting the content in the second color style based on the denoised image representation, wherein the second color style is different from the first color style; and
providing, for display within a graphical user interface, the color-edited digital image.
2 . The method of claim 1 , further comprising:
providing, to a graphical user interface, selectable options for target styles; and
in response to a selection of a selectable option comprising the second color style, identifying the noise representation from the additional digital image different than the digital image.
3 . The method of claim 1 , further comprising:
identifying an additional noise representation corresponding to a third color style; and
utilizing the additional noise representation and the digital image with the image color editing diffusion neural network to generate an additional color-edited digital image depicting the content in the third color style.
4 . The method of claim 3 , further comprising providing, for display within the graphical user interface, the color-edited digital image and the additional color-edited digital image.
5 . The method of claim 1 , wherein identifying the noise representation corresponding to the second color style comprises:
generating, utilizing an invertible diffusion encoder, the noise representation from the additional digital image.
6 . The method of claim 1 , wherein identifying the noise representation corresponding to the second color style comprises:
identifying a raw digital image and an edited digital image generated from the raw digital image, wherein the edited digital image is the additional digital image that is different than the digital image; and
generating, utilizing an invertible diffusion encoder from the raw digital image and the edited digital image, the noise representation, wherein the noise representation reflects an editing direction between the raw digital image and the edited digital image.
7 . The method of claim 1 , wherein the image color editing diffusion neural network is trained to generate, from a training digital image, a denoised color-edited image conditioned utilizing a color space restoration task.
8 . The method of claim 1 , further comprising:
receiving an input text prompt, from a client device, indicating a request to edit the digital image from the first color style to the second color style;
generating, utilizing a vision language model, an encoding of the input text prompt; and
generating the color-edited digital image utilizing the image color editing diffusion neural network from the digital image and the input text prompt.
9 . A non-transitory computer-readable medium storing executable instructions which, when executed by a computing device, cause the computing device to perform operations comprising:
identifying a digital image depicting content in a first color style;
identifying a noise representation corresponding to a second color style from an additional digital image different than the digital image;
generating a denoised image representation corresponding to the second color style utilizing an image color editing diffusion neural network to remove noise from the noise representation based on conditioning the image color editing diffusion neural network on the digital image depicting content in the first color style;
generating, utilizing the image color editing diffusion neural network, a color-edited digital image depicting the content in the second color style based on the denoised image representation, wherein the second color style is different from the first color style; and
providing, for display within a graphical user interface, the color-edited digital image.
10 . The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise:
providing, to a graphical user interface, selectable options for target styles; and
in response to a selection of a selectable option comprising the second color style, identifying the noise representation from the additional digital image different than the digital image.
11 . The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise:
identifying a plurality of noise representations corresponding to a plurality of color styles; and
generating, utilizing the image color editing diffusion neural network with the digital image and the plurality of noise representations, color-edited digital images depicting the content of the digital image in the plurality of color styles.
12 . The non-transitory computer-readable medium of claim 9 , wherein the image color editing diffusion neural network is trained to generate, from a training digital image, a denoised color-edited image conditioned utilizing a color space restoration task.
13 . The non-transitory computer-readable medium of claim 9 , wherein the operations further comprise:
receiving an input text prompt comprising the second color style; and
generating, utilizing the image color editing diffusion neural network, the color-edited digital image from the input text prompt.
14 . A system comprising:
a memory component; and
one or more processing devices coupled to the memory component, the one or more processing devices to perform operations comprising:
identifying a digital image depicting content in a first color style;
identifying a noise representation corresponding to a second color style from an additional digital image different than the digital image;
generating a denoised image representation corresponding to the second color style utilizing an image color editing diffusion neural network to remove noise from the noise representation based on conditioning the image color editing diffusion neural network on the digital image depicting content in the first color style;
generating, utilizing the image color editing diffusion neural network, a color-edited digital image depicting the content in the second color style based on the denoised image representation, wherein the second color style is different from the first color style; and
providing, for display within a graphical user interface, the color-edited digital image.
15 . The system of claim 14 , wherein the operations further comprise:
providing, to a graphical user interface, selectable options for target styles; and
in response to a selection of a selectable option comprising the second color style, identifying the noise representation from the additional digital image different than the digital image.
16 . The system of claim 14 , wherein the operations further comprise:
identifying an additional noise representation corresponding to a third color style; and
utilizing the additional noise representation and the digital image with the image color editing diffusion neural network to generate an additional color-edited digital image depicting the content in the third color style.
17 . The system of claim 16 , wherein the operations further comprise providing, for display within the graphical user interface, the color-edited digital image and the additional color-edited digital image.
18 . The system of claim 14 , wherein identifying the noise representation corresponding to the second color style comprises generating, utilizing an invertible diffusion encoder, the noise representation from the additional digital image.
19 . The system of claim 14 , wherein identifying the noise representation corresponding to the second color style comprises:
identifying a raw digital image and an edited digital image generated from the raw digital image, wherein the edited digital image is the additional digital image that is different than the digital image; and
generating, utilizing an invertible diffusion encoder from the raw digital image and the edited digital image, the noise representation, wherein the noise representation reflects an editing direction between the raw digital image and the edited digital image.
20 . The system of claim 14 , wherein the image color editing diffusion neural network is trained to generate, from a training digital image, a denoised color-edited image conditioned utilizing a color space restoration task.