Automatic generation of composite images
Certain aspects and features of this disclosure relate to automatic generation of composite images. For example, a method involves producing a representative image corresponding to a composite image based on a presentation context of input objects and segmenting the generated objects from the representative image to extract the generated objects from the representative image. The method also includes generating an inferred disposition of each of the generated objects and transforming each of the input objects to the inferred disposition of a corresponding generated object. The method can also include transmitting, storing, display, or rendering, in response to the transforming, the composite image of the input objects. Certain aspects and features also include computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
1 . A method comprising:
accessing a plurality of items from a stored catalog, each item including an image and a text description of the image;
determining, using at least one of image metadata or the text description, semantic content for each input image of a plurality of input images from the plurality of items, each including the image metadata and depicting an input object from the plurality of items;
generating a presentation context for a composite image based on the semantic content from the plurality of the images;
applying a text generative model to the image, the image metadata, and the text description to generate a text-based, contextually relevant prompt for a trained generative image model configured to provide contextual relevance for a commerce space;
prompting, using the text-based, contextually relevant prompt, the trained generative image model to produce a representative image corresponding to the composite image, the representative image including a plurality of generated objects, each generated object corresponding to one of a plurality of input objects;
segmenting, using an image mask module, the plurality of generated objects from the representative image to extract the plurality of generated objects from the representative image;
generating, using an image overlay module and the image metadata, an inferred disposition of each of the plurality of generated objects as extracted from the representative image;
transforming, using the image overlay module, each of the plurality of input objects to the inferred disposition of a corresponding generated object from the plurality of generated objects; and
rendering, in response to the transforming, the composite image of the plurality of input objects within the presentation context.
2 . The method of claim 1 , further comprising inpainting the composite image to fill any spaces in the composite image.
3 . The method of claim 1 , further comprising optimizing, using the inferred disposition, the text-based, contextually relevant prompt.
4 . The method of claim 3 , further comprising:
generating a vector corresponding to an embedding space that represents the presentation context;
sampling the embedding space around the vector; and
optimizing the text-based, contextually relevant prompt at least in part in response to the sampling of the embedding space.
5 . The method of claim 1 , wherein the inferred disposition comprises pose, scale, perspective, and rotation.
6 . The method of claim 1 , further comprising inferring, using an object inference module, a location corresponding to each of the plurality of input objects in a plurality of object images.
7 . The method of claim 6 , further comprising:
identifying each of the plurality of input objects based in part on the image metadata; and
inferring the location corresponding to each of the plurality of input objects based at least in part on the image metadata.
8 . A system comprising:
a memory component;
a processing device coupled to the memory component to perform operations of accessing a plurality of items from a stored catalog, each item including an image and a text description of the image, and causing a composite image of the plurality of items to be transmitted, stored, or displayed;
an object inference module configured to determine, using at least one of image metadata or the text description, semantic content for each input image of a plurality of input images, each including the image metadata and depicting an input object from the plurality of items, and generate a presentation context for the composite image based on the semantic content;
a text generative model configured to generate, based on the image, the image metadata, and the text description, a text-based, contextually relevant prompt for a trained generative image model configured to provide contextual relevance for a commerce space;
a trained generative image model configured to produce, in response to the text-based, contextually relevant prompt, and based on the presentation context, a representative image corresponding to the composite image based on the presentation context, the representative image including a plurality of generated objects, each generated object corresponding to one of the input objects;
an image mask module configured to segment the plurality of generated objects from the representative image to extract the plurality of generated objects from the representative image; and
an image overlay module configured to transform, using the image metadata, each of the input objects to an inferred disposition of a corresponding generated object from the plurality of generated objects to produce the composite image of the input objects within the presentation context.
9 . The system of claim 8 , further comprising an inpainting module configured to fill any spaces in the composite image.
10 . The system of claim 8 , wherein the object inference module is further configured to infer the input objects from the plurality of input images.
11 . The system of claim 8 , further comprising a prompt generation module configured to sample an embedding space around a vector that represents the presentation context and optimize the text-based, contextually relevant prompt to cause the trained generative image model to produce the representative image.
12 . The system of claim 8 , wherein the inferred disposition comprises pose, scale, perspective, and rotation.
13 . The system of claim 8 , wherein the input objects and the presentation context of the input objects are identified based at least in part on the image metadata.
14 . A non-transitory computer-readable medium storing executable instructions, which when executed by a processing device, cause the processing device to perform operations comprising:
accessing a plurality of items from a stored catalog, each item including an image and a text description of the image;
determining, using at least one of image metadata or the text description, semantic content for each input image of a plurality of input images from the plurality of items, each including the image metadata and depicting an input object from the plurality of items;
generating a presentation context for a composite image based on the semantic content from the plurality of the images;
applying a text generative model to the image, the image metadata, and the text description to generate a text-based, contextually relevant prompt for a trained generative image model configured to provide contextual relevance for a commerce space;
prompting, using the text-based, contextually relevant prompt, the trained generative image model to produce a representative image corresponding to the composite image, the representative image including a plurality of generated objects, a generated object corresponding to one of a plurality of input objects;
a step for transforming, based on the image metadata, each of the plurality of input objects to an inferred disposition of the generated object in the representative image to produce the composite image of the plurality of input objects; and
rendering the composite image of the plurality of input objects within the presentation context.
15 . The non-transitory computer-readable medium of claim 14 , wherein the instructions further cause the processing device to perform operations comprising inpainting the composite image to fill any spaces in the composite image.
16 . The non-transitory computer-readable medium of claim 14 , wherein the instructions further cause the processing device to perform operations further comprising optimizing, using the inferred disposition, the text-based, contextually relevant prompt.
17 . The non-transitory computer-readable medium of claim 16 , wherein the instructions further cause the processing device to perform operations comprising:
generating a vector corresponding to an embedding space that represents the presentation context;
sampling the embedding space around the vector; and
optimizing the text-based, contextually relevant prompt at least in part in response to the sampling of the embedding space.
18 . The non-transitory computer-readable medium of claim 14 , wherein the inferred disposition comprises pose, scale, perspective, and rotation.
19 . The non-transitory computer-readable medium of claim 14 , wherein the instructions further cause the processing device to perform operations comprising inferring a location corresponding to each of the plurality of input objects in each of the plurality of input images based at least in part on the image metadata.
20 . The non-transitory computer-readable medium of claim 19 , wherein the instructions further cause the processing device to perform operations comprising:
identifying each of the plurality of input objects based in part on the image metadata; and
inferring the location corresponding to each of the plurality of input objects based at least in part on the image metadata.