AI-based shape-adaptive consistent visual effect generation
A data processing system implements constructing a first prompt including a font mask of a reference character (RC) and a style prompt, sending the first prompt to a text2image model to iteratively generate salient content and concentrate the salient content within the font mask of RC as a first image of RC; concatenating two of the first images as a second image; generating a combined font mask of the font mask of RC and a font mask of a target character (TC); constructing a second prompt including the combined font mask and the second image, sending the second prompt to the model to iteratively generate salient content and in-paint the salient content within a half of the combined font mask as a third image of RC and TC; cropping a styled TC image from the third image using the font mask of TC; providing the styled TC image to a client device.
1 . A data processing system comprising:
a processor, and
a machine-readable storage medium storing executable instructions which, when executed by the processor, cause the processor alone or in combination with other processors to perform the following operations:
receiving, at a client device, a style prompt including at least one visual object to be visible in a character in a desired style;
constructing, via a prompt construction unit, a first prompt by appending a font mask of a reference character and the style prompt to a first instruction string, the first instruction string including instructions to a first text-to-image model to iteratively generate salient content based on the style prompt and concentrate the salient content within the font mask of the reference character as a first image of the reference character;
providing as an input the first prompt to the first text-to-image model and receiving as an output the first image from the first text-to-image model;
duplicating the first image and concatenating the first images as a second image;
generating a combined font mask of the font mask of the reference character and a font mask of a target character, wherein the target character is different from the reference character;
constructing, via the prompt construction unit, a second prompt by appending the combined font mask and the second image to a second instruction string, the second instruction string including instructions to the first text-to-image model to iteratively generate salient content based on the second image and in-paint the salient content within a half of the combined font mask as a third image of the reference character and the target character;
providing as an input the second prompt to the first text-to-image model and receiving as an output the third image from the first text-to-image model;
cropping a styled target character image from the third image using the font mask of the target character;
providing the styled target character image to the client device; and
causing a user interface of the client device to display the styled target character image.
2 . The data processing system of claim 1 , wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
cropping a styled reference character image from the first image using the font mask of the reference character;
constructing, via the prompt construction unit, a third prompt by appending the styled reference character image and the style prompt to a third instruction string, the third instruction string including instructions to the text-to-image model to iteratively regenerate edge details of the styled reference character image as a refined image of the reference character; and
providing as an input the third prompt to the text-to-image model and receiving as an output the refined image of the reference character from the text-to-image model.
3 . The data processing system of claim 2 , wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
concatenating the refined image of the reference character and the styled target character image as a fourth image;
constructing, via the prompt construction unit, a fourth prompt by appending the fourth image to a fourth instruction string, the fourth instruction string including instructions to the text-to-image model to iteratively in-paint edge details of the styled target character image to a half of the fourth image including the styled target character image as a refined image of the reference character and the target character;
providing as an input the fourth prompt to the text-to-image model and receiving as an output the refined image of the target character and the refined font mask of the target character from the text-to-image model;
cropping a refined image of the target character from the refined image of the reference character and target character;
constructing, via the prompt construction unit, a fifth prompt by appending the refined image of the target character to a fifth instruction string, the fifth instruction string including instructions to a vision generative model to extract an alpha channel of the refined image of the target character as a refined mask of the target character;
providing as an input the fifth prompt to the vision generative model and receiving as an output the refined mask of the target character from the vision generative model;
providing the refined image of the target character and the refined font mask of the target character to the client device; and
causing the user interface of the client device to display the refined image of the target character and the refined font mask of the target character,
wherein the first image includes the at least one visual object visible in the reference character in the desired style, and wherein the third image includes the first image and an image of the target character with the at least one visual object visible in the desired style.
4 . The data processing system of claim 2 , wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
constructing, via the prompt construction unit, a sixth prompt by appending image captions to a sixth instruction string, the sixth instruction string including instructions to a second text-to-image model to create a plurality of text prompts based on the image captions, to iteratively (1) generate based on a respective one of the text prompts a respective image differentiating a foreground from a background, and (2) segregate the foreground from the respective image as a respective irregular-shaped canvas mask and a respective irregular-shaped image, and to generate a dataset of triplet instances, each instance consisting of the respective irregular-shaped canvas mask, the respective irregular-shaped image, and the respective text prompt; and
providing as an input the sixth prompt to the second text-to-image model and receiving as an output the dataset of triplet instances from the second text-to-image model.
5 . The data processing system of claim 4 , wherein the second text-to-image model is DALL-E 3 or a diffusion model.
6 . The data processing system of claim 4 , wherein the first text-to-image model is a conditional diffusion model, and the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
training the conditional diffusion model based on the dataset of triplet instances.
7 . The data processing system of claim 3 , wherein the vision generative model is a conditional variational autoencoder (VAE), and the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
augmenting a decoder of the conditional VAE with an additional input channel and an additional output channel that facilitate mask conditioning and prediction.
8 . The data processing system of claim 7 , wherein the machine-readable storage medium further includes instructions configured to cause the processor alone or in combination with other processors to perform operations of:
prompting a segmentation model to generate original segmentation masks;
applying alpha mask augmentation on the original segmentation masks to generate augmented masks; and
training the conditional VAE using the augmented masks as input conditions.
9 . The data processing system of claim 8 , wherein the segmentation model is a prompt-based instance segmentation model or a U-net.
10 . The data processing system of claim 1 , wherein the style prompt is a text prompt or an image prompt.
11 . A method comprising:
receiving, at a client device, a style prompt including at least one visual object to be visible in a character in a desired style;
constructing, via a prompt construction unit, a first prompt by appending a font mask of a reference character and the style prompt to a first instruction string, the first instruction string including instructions to a first text-to-image model to iteratively generate salient content based on the style prompt and concentrate the salient content within the font mask of the reference character as a first image of the reference character;
providing as an input the first prompt to the first text-to-image model and receiving as an output the first image from the first text-to-image model;
duplicating the first image and concatenating the first images as a second image;
generating a combined font mask of the font mask of the reference character and a font mask of a target character, wherein the target character is different from the reference character;
constructing, via the prompt construction unit, a second prompt by appending the combined font mask and the second image to a second instruction string, the second instruction string including instructions to the first text-to-image model to iteratively generate salient content based on the second image and in-paint the salient content within a half of the combined font mask as a third image of the reference character and the target character;
providing as an input the second prompt to the first text-to-image model and receiving as an output the third image from the first text-to-image model;
cropping a styled target character image from the third image using the font mask of the target character;
providing the styled target character image to the client device; and
causing a user interface of the client device to display the styled target character image.
12 . The method of claim 11 , further comprising:
cropping a styled reference character image from the first image using the font mask of the reference character;
constructing, via the prompt construction unit, a third prompt by appending the styled reference character image and the style prompt to a third instruction string, the third instruction string including instructions to the text-to-image model to iteratively regenerate edge details of the styled reference character image as a refined image of the reference character; and
providing as an input the third prompt to the text-to-image model and receiving as an output the refined image of the reference character from the text-to-image model.
13 . The method of claim 12 , further comprising:
concatenating the refined image of the reference character and the styled target character image as a fourth image;
constructing, via the prompt construction unit, a fourth prompt by appending the fourth image to a fourth instruction string, the fourth instruction string including instructions to the text-to-image model to iteratively in-paint edge details of the styled target character image to a half of the fourth image including the styled target character image the as a refined image of the reference character and the target character;
providing as an input the fourth prompt to the text-to-image model and receiving as an output the refined image of the target character and the refined font mask of the target character from the text-to-image model;
cropping a refined image of the target character from the refined image of the reference character and target character;
constructing, via the prompt construction unit, a fifth prompt by appending the refined image of the target character to a fifth instruction string, the fifth instruction string including instructions to a vision generative model to extract an alpha channel of the refined image of the target character as a refined mask of the target character;
providing as an input the fifth prompt to the vision generative model and receiving as an output the refined mask of the target character from the vision generative model;
providing the refined image of the target character and the refined font mask of the target character to the client device; and
causing the user interface of the client device to display the refined image of the target character and the refined font mask of the target character.
14 . The method of claim 12 , further comprising:
constructing, via the prompt construction unit, a sixth prompt by appending image captions to a sixth instruction string, the sixth instruction string including instructions to a second text-to-image model to create a plurality of text prompts based on the image captions, to iteratively (1) generate based on a respective one of the text prompts a respective image differentiating a foreground from a background, and (2) segregate the foreground from the respective image as a respective irregular-shaped canvas mask and a respective irregular-shaped image, and to generate a dataset of triplet instances, each instance consisting of the respective irregular-shaped canvas mask, the respective irregular-shaped image, and the respective text prompt; and
providing as an input the sixth prompt to the second text-to-image model and receiving as an output the dataset of triplet instances from the second text-to-image model.
15 . The method of claim 14 , wherein the first text-to-image model is a conditional diffusion model, and the method further comprises:
training the conditional diffusion model based on the dataset of triplet instances.
16 . A non-transitory computer readable medium on which are stored instructions that, when executed, cause a programmable device to perform functions of:
receiving, at a client device, a style prompt including at least one visual object to be visible in a character in a desired style;
constructing, via a prompt construction unit, a first prompt by appending a font mask of a reference character and the style prompt to a first instruction string, the first instruction string including instructions to a first text-to-image model to iteratively generate salient content based on the style prompt and concentrate the salient content within the font mask of the reference character as a first image of the reference character;
providing as an input the first prompt to the first text-to-image model and receiving as an output the first image from the first text-to-image model;
duplicating the first image and concatenating the first images as a second image;
generating a combined font mask of the font mask of the reference character and a font mask of a target character, wherein the target character is different from the reference character;
constructing, via the prompt construction unit, a second prompt by appending the combined font mask and the second image to a second instruction string, the second instruction string including instructions to the first text-to-image model to iteratively generate salient content based on the second image and in-paint the salient content within a half of the combined font mask as a third image of the reference character and the target character;
providing as an input the second prompt to the first text-to-image model and receiving as an output the third image from the first text-to-image model;
cropping a styled target character image from the third image using the font mask of the target character;
providing the styled target character image to the client device; and
causing a user interface of the client device to display the styled target character image.
17 . The non-transitory computer readable medium of claim 16 , wherein the instructions when executed, further cause the programmable device to perform functions of:
cropping a styled reference character image from the first image using the font mask of the reference character;
constructing, via the prompt construction unit, a third prompt by appending the styled reference character image and the style prompt to a third instruction string, the third instruction string including instructions to the text-to-image model to iteratively regenerate edge details of the styled reference character image as a refined image of the reference character; and
providing as an input the third prompt to the text-to-image model and receiving as an output the refined image of the reference character from the text-to-image model.
18 . The non-transitory computer readable medium of claim 17 , wherein the instructions when executed, further cause the programmable device to perform functions of:
concatenating the refined image of the reference character and the styled target character image as a fourth image;
constructing, via the prompt construction unit, a fourth prompt by appending the fourth image to a fourth instruction string, the fourth instruction string including instructions to the text-to-image model to iteratively in-paint edge details of the styled target character image to a half of the fourth image including the styled target character image the as a refined image of the reference character and the target character;
providing as an input the fourth prompt to the text-to-image model and receiving as an output the refined image of the target character and the refined font mask of the target character from the text-to-image model;
cropping a refined image of the target character from the refined image of the reference character and target character;
constructing, via the prompt construction unit, a fifth prompt by appending the refined image of the target character to a fifth instruction string, the fifth instruction string including instructions to a vision generative model to extract an alpha channel of the refined image of the target character as a refined mask of the target character;
providing as an input the fifth prompt to the vision generative model and receiving as an output the refined mask of the target character from the vision generative model;
providing the refined image of the target character and the refined font mask of the target character to the client device; and
causing the user interface of the client device to display the refined image of the target character and the refined font mask of the target character.
19 . The non-transitory computer readable medium of claim 17 , wherein the instructions when executed, further cause the programmable device to perform functions of:
constructing, via the prompt construction unit, a sixth prompt by appending image captions to a sixth instruction string, the sixth instruction string including instructions to a second text-to-image model to create a plurality of text prompts based on the image captions, to iteratively (1) generate based on a respective one of the text prompts a respective image differentiating a foreground from a background, and (2) segregate the foreground from the respective image as a respective irregular-shaped canvas mask and a respective irregular-shaped image, and to generate a dataset of triplet instances, each instance consisting of the respective irregular-shaped canvas mask, the respective irregular-shaped image, and the respective text prompt; and
providing as an input the sixth prompt to the second text-to-image model and receiving as an output the dataset of triplet instances from the second text-to-image model.
20 . The non-transitory computer readable medium of claim 19 , wherein the first text-to-image model is a conditional diffusion model, and wherein the instructions when executed, further cause the programmable device to perform functions of:
training the conditional diffusion model based on the dataset of triplet instances.