Aspect ratio conversion for automated image generation
Examples disclosed herein describe aspect ratio conversion techniques for automated image generation. An image generation request comprising a prompt is received from a user device. A processor-implemented automated image generator may generate a first image based on the prompt. The first image has a first aspect ratio. According to some examples, a region of interest is determined in the first image, based on a prompt alignment indicator for the region of interest. The first image is then processed to obtain a second image. The processing includes an automatic cropping operation directed at the region of interest. The second image has a second aspect ratio that is different from the first aspect ratio. The second image is caused to be presented on the user device.
1 . A method comprising:
receiving, from a user device, an image generation request comprising a prompt;
responsive to receiving the image generation request, generating, by a processor-implemented automated image generator and based on the prompt, a first image having a first aspect ratio;
determining two or more candidate regions of interest in the first image;
encoding, by a text encoder, the prompt to obtain an embedding of the prompt;
encoding, by an image encoder, each of the two or more candidate regions of interest to obtain an embedding of the candidate region of interest;
comparing the embedding of each of the two or more candidate regions of interest with the embedding of the prompt to obtain a prompt alignment indicator for each of the two or more candidate regions of interest in the first image, the prompt alignment indicator indicating a level of alignment between each candidate region of interest and the prompt;
determining a target region of interest among the two or more candidate regions of interest in the first image, the target region of interest being determined automatically based on a comparison among two or more prompt alignment indicators for the two or more candidate regions of interest;
cropping the first image at the target region of interest to obtain a second image, the second image having a second aspect ratio that is different from the first aspect ratio; and
causing presentation of the second image on the user device.
2 . The method of claim 1 , wherein the prompt is a text prompt.
3 . The method of claim 1 , wherein the processor-implemented automated image generator comprises a text-to-image machine learning model.
4 . The method of claim 1 , further comprising, prior to the cropping of the first image to obtain the second image, upsampling the first image by applying a uniform scaling factor to a width and a height of the first image.
5 . The method of claim 1 , wherein each candidate region of interest of the two or more candidate regions of interest is generated as a bounding box with respect to the first image.
6 . The method of claim 1 , wherein the prompt alignment indicator is an alignment score.
7 . The method of claim 1 , further comprising: prior to the cropping, automatically adjusting the cropping region such that the cropping region has the second aspect ratio.
8 . The method of claim 1 , wherein the receiving the image generation request comprises:
causing presentation of an input text box in a user interface provided by an interaction client executing on the user device; and
receiving user input comprising the prompt via the input text box in the user interface, wherein the causing presentation of the second image on the user device comprises causing presentation of the second image in the user interface provided by the interaction client.
9 . The method of claim 1 , wherein the first aspect ratio is a 1:1 aspect ratio.
10 . The method of claim 1 , wherein the second aspect ratio has a height that is greater than its width.
11 . The method of claim 1 , wherein the processor-implemented automated image generator comprises a text-to-image machine learning model, the text-to-image machine learning model being trained using a training data set comprising a plurality of training images, each training image having a corresponding text description forming part of the training data set.
12 . The method of claim 1 , wherein the prompt alignment indicator for the candidate region of interest comprises a cosine similarity between the embedding of the candidate region of interest and the embedding of the prompt.
13 . The method of claim 1 , wherein the prompt is randomly generated in response to a user selection of an interactive element in a user interface.
14 . The method of claim 3 , wherein the text-to-image machine learning model is a diffusion model.
15 . The method of claim 10 , wherein the second aspect ratio is a height: width ratio of 3:2.
16 . The method of claim 11 , wherein, for each training image, the corresponding text description comprises a first text description and a second text description, the first text description being different from the second text description, and the second text description being a caption generated using a processor-implemented automated caption generator.
17 . The method of claim 16 , wherein the processor-implemented automated caption generator comprises an image-to-text machine learning model.
18 . A system comprising a memory storing instructions and one or more processors configured by the instructions to perform operations comprising:
receiving, from a user device, an image generation request comprising a prompt;
responsive to receiving the image generation request, generating, by a processor-implemented automated image generator and based on the prompt, a first image having a first aspect ratio;
determining two or more candidate regions of interest in the first image;
encoding, by a text encoder, the prompt to obtain an embedding of the prompt;
encoding, by an image encoder, each of the two or more candidate regions of interest to obtain an embedding of the candidate region of interest;
comparing the embedding of each of the two or more candidate regions of interest with the embedding of the prompt to obtain a prompt alignment indicator for each of the two or more candidate regions of interest in the first image, the prompt alignment indicator indicating a level of alignment between each candidate region of interest and the prompt;
determining a target region of interest among the two or more candidate regions of interest in the first image, the target region of interest being determined automatically based on a comparison among two or more prompt alignment indicators for the two or more candidate regions of interest;
cropping the first image at the target region of interest to obtain a second image, the second image having a second aspect ratio that is different from the first aspect ratio; and
causing presentation of the second image on the user device.
19 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by at least one computer, cause the at least one computer to perform operations comprising:
receiving, from a user device, an image generation request comprising a prompt;
responsive to receiving the image generation request, generating, by a processor-implemented automated image generator and based on the prompt, a first image having a first aspect ratio;
determining two or more candidate regions of interest in the first image;
encoding, by a text encoder, the prompt to obtain an embedding of the prompt;
encoding, by an image encoder, each of the two or more candidate regions of interest to obtain an embedding of the candidate region of interest;
comparing the embedding of each of the two or more candidate regions of interest with the embedding of the prompt to obtain a prompt alignment indicator for each of the two or more candidate regions of interest in the first image, the prompt alignment indicator indicating a level of alignment between each candidate region of interest and the prompt;
determining a target region of interest among the two or more candidate regions of interest in the first image, the target region of interest being determined automatically based on a comparison among two or more prompt alignment indicators for the two or more candidate regions of interest;
cropping the first image at the target region of interest to obtain a second image, the second image having a second aspect ratio that is different from the first aspect ratio; and
causing presentation of the second image on the user device.