IP Library › Granted Patent US 12,750,459
Granted Patent B2
US 12,750,459 · App. 18/115,997 · Granted Sep 29, 2026

Aspect ratio conversion for automated image generation

Inventors: Mykyta Bakunov (Freienbach, CH); Arnab Ghosh (Oxford, GB); Pavel Savchenkov (London, GB); Sergey Smetanin (London, GB); Jian Ren (Marina del Ray, CA)
Assignee: SNAP INC.
H04N7/0122G06T3/40G06T7/70G06T9/00G06V10/25G06T2207/20084G06T2207/20132G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,750,459
App. No.
18/115,997
Granted
Sep 29, 2026
Kind
B2
Abstract

Examples disclosed herein describe aspect ratio conversion techniques for automated image generation. An image generation request comprising a prompt is received from a user device. A processor-implemented automated image generator may generate a first image based on the prompt. The first image has a first aspect ratio. According to some examples, a region of interest is determined in the first image, based on a prompt alignment indicator for the region of interest. The first image is then processed to obtain a second image. The processing includes an automatic cropping operation directed at the region of interest. The second image has a second aspect ratio that is different from the first aspect ratio. The second image is caused to be presented on the user device.

Claims (48)

1 . A method comprising:

receiving, from a user device, an image generation request comprising a prompt;

responsive to receiving the image generation request, generating, by a processor-implemented automated image generator and based on the prompt, a first image having a first aspect ratio;

determining two or more candidate regions of interest in the first image;

encoding, by a text encoder, the prompt to obtain an embedding of the prompt;

encoding, by an image encoder, each of the two or more candidate regions of interest to obtain an embedding of the candidate region of interest;

comparing the embedding of each of the two or more candidate regions of interest with the embedding of the prompt to obtain a prompt alignment indicator for each of the two or more candidate regions of interest in the first image, the prompt alignment indicator indicating a level of alignment between each candidate region of interest and the prompt;

determining a target region of interest among the two or more candidate regions of interest in the first image, the target region of interest being determined automatically based on a comparison among two or more prompt alignment indicators for the two or more candidate regions of interest;

cropping the first image at the target region of interest to obtain a second image, the second image having a second aspect ratio that is different from the first aspect ratio; and

causing presentation of the second image on the user device.

2 . The method of claim 1 , wherein the prompt is a text prompt.

3 . The method of claim 1 , wherein the processor-implemented automated image generator comprises a text-to-image machine learning model.

4 . The method of claim 1 , further comprising, prior to the cropping of the first image to obtain the second image, upsampling the first image by applying a uniform scaling factor to a width and a height of the first image.

5 . The method of claim 1 , wherein each candidate region of interest of the two or more candidate regions of interest is generated as a bounding box with respect to the first image.

6 . The method of claim 1 , wherein the prompt alignment indicator is an alignment score.

7 . The method of claim 1 , further comprising: prior to the cropping, automatically adjusting the cropping region such that the cropping region has the second aspect ratio.

8 . The method of claim 1 , wherein the receiving the image generation request comprises:

causing presentation of an input text box in a user interface provided by an interaction client executing on the user device; and

receiving user input comprising the prompt via the input text box in the user interface, wherein the causing presentation of the second image on the user device comprises causing presentation of the second image in the user interface provided by the interaction client.

9 . The method of claim 1 , wherein the first aspect ratio is a 1:1 aspect ratio.

10 . The method of claim 1 , wherein the second aspect ratio has a height that is greater than its width.

11 . The method of claim 1 , wherein the processor-implemented automated image generator comprises a text-to-image machine learning model, the text-to-image machine learning model being trained using a training data set comprising a plurality of training images, each training image having a corresponding text description forming part of the training data set.

12 . The method of claim 1 , wherein the prompt alignment indicator for the candidate region of interest comprises a cosine similarity between the embedding of the candidate region of interest and the embedding of the prompt.

13 . The method of claim 1 , wherein the prompt is randomly generated in response to a user selection of an interactive element in a user interface.

14 . The method of claim 3 , wherein the text-to-image machine learning model is a diffusion model.

15 . The method of claim 10 , wherein the second aspect ratio is a height: width ratio of 3:2.

16 . The method of claim 11 , wherein, for each training image, the corresponding text description comprises a first text description and a second text description, the first text description being different from the second text description, and the second text description being a caption generated using a processor-implemented automated caption generator.

17 . The method of claim 16 , wherein the processor-implemented automated caption generator comprises an image-to-text machine learning model.

18 . A system comprising a memory storing instructions and one or more processors configured by the instructions to perform operations comprising:

receiving, from a user device, an image generation request comprising a prompt;

responsive to receiving the image generation request, generating, by a processor-implemented automated image generator and based on the prompt, a first image having a first aspect ratio;

determining two or more candidate regions of interest in the first image;

encoding, by a text encoder, the prompt to obtain an embedding of the prompt;

encoding, by an image encoder, each of the two or more candidate regions of interest to obtain an embedding of the candidate region of interest;

comparing the embedding of each of the two or more candidate regions of interest with the embedding of the prompt to obtain a prompt alignment indicator for each of the two or more candidate regions of interest in the first image, the prompt alignment indicator indicating a level of alignment between each candidate region of interest and the prompt;

determining a target region of interest among the two or more candidate regions of interest in the first image, the target region of interest being determined automatically based on a comparison among two or more prompt alignment indicators for the two or more candidate regions of interest;

cropping the first image at the target region of interest to obtain a second image, the second image having a second aspect ratio that is different from the first aspect ratio; and

causing presentation of the second image on the user device.

19 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by at least one computer, cause the at least one computer to perform operations comprising:

receiving, from a user device, an image generation request comprising a prompt;

responsive to receiving the image generation request, generating, by a processor-implemented automated image generator and based on the prompt, a first image having a first aspect ratio;

determining two or more candidate regions of interest in the first image;

encoding, by a text encoder, the prompt to obtain an embedding of the prompt;

encoding, by an image encoder, each of the two or more candidate regions of interest to obtain an embedding of the candidate region of interest;

comparing the embedding of each of the two or more candidate regions of interest with the embedding of the prompt to obtain a prompt alignment indicator for each of the two or more candidate regions of interest in the first image, the prompt alignment indicator indicating a level of alignment between each candidate region of interest and the prompt;

determining a target region of interest among the two or more candidate regions of interest in the first image, the target region of interest being determined automatically based on a comparison among two or more prompt alignment indicators for the two or more candidate regions of interest;

cropping the first image at the target region of interest to obtain a second image, the second image having a second aspect ratio that is different from the first aspect ratio; and

causing presentation of the second image on the user device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2023
From: BAKUNOV, MYKYTA; GHOSH, ARNAB; SAVCHENKOV, PAVEL; SMETANIN, SERGEY; REN, JIAN
To: SNAP INC.
Reel/Frame 063269/0704 →
Continuity (1)
Related Publication 20240297957A1 · Sep 5, 2024
References Cited (99)
US 7214065B2 · Fitzsimmons, Jr. · 2007 [cited by applicant]
US 8867849B1 · Kirkham et al. · 2014 [cited by applicant]
US 10496924B1 · Highnam et al. · 2019 [cited by applicant]
US 11445148B1 · Øhrn · 2022 [cited by applicant]
US 11809688B1 · Parasnis et al. · 2023 [cited by applicant]
US 11947893B1 · Seth · 2024 [cited by applicant]
US 12169626B2 · Zakharov et al. · 2024 [cited by applicant]
US 12205207B2 · Smetanin et al. · 2025 [cited by applicant]
US 20140051402A1 · Qureshi · 2014 [cited by applicant]
US 20140344712A1 · Okazawa et al. · 2014 [cited by applicant]
US 20160125269A1 · Lee et al. · 2016 [cited by applicant]
US 20160188153A1 · Lerner et al. · 2016 [cited by applicant]
US 20190295302A1 · Fu et al. · 2019 [cited by applicant]
US 20200267182A1 · Highnam et al. · 2020 [cited by applicant]
US 20210042796A1 · Khoury et al. · 2021 [cited by applicant]
US 20210209184A1 · Huang et al. · 2021 [cited by applicant]
US 20210352460A1 · Rohde et al. · 2021 [cited by applicant]
US 20220036153A1 · O'Malia et al. · 2022 [cited by applicant]
US 20220101578A1 · Bedi et al. · 2022 [cited by applicant]
US 20220114698A1 · Liu · 2022 [cited by applicant]
US 20220415012A1 · Karmakar · 2022 [cited by examiner]
US 20230025835A1 · Moriya et al. · 2023 [cited by applicant]
US 20230054174A1 · Peled et al. · 2023 [cited by applicant]
US 20230081171A1 · Zhang · 2023 [cited by examiner]
US 20230177878A1 · Sekar et al. · 2023 [cited by applicant]
US 20230215441A1 · Wu · 2023 [cited by applicant]
US 20230222703A1 · Baheti et al. · 2023 [cited by applicant]
US 20230230198A1 · Zhang et al. · 2023 [cited by applicant]
US 20230260164A1 · Yuan et al. · 2023 [cited by applicant]
US 20230262102A1 · Das et al. · 2023 [cited by applicant]
US 20230281789A1 · Sudarsky et al. · 2023 [cited by applicant]
US 20230298224A1 · Aggarwal et al. · 2023 [cited by applicant]
US 20230342284A1 · Easton et al. · 2023 [cited by applicant]
US 20240135610A1 · Kolkin et al. · 2024 [cited by applicant]
US 20240161258A1 · Maschmeyer et al. · 2024 [cited by applicant]
US 20240169622A1 · Xie et al. · 2024 [cited by applicant]
US 20240193821A1 · Denison · 2024 [cited by applicant]
US 20240295953A1 · Zakharov et al. · 2024 [cited by applicant]
US 20240296535A1 · Bakunov et al. · 2024 [cited by applicant]
US 20240296606A1 · Smetanin et al. · 2024 [cited by applicant]
US 20240311960A1 · Feng · 2024 [cited by examiner]
US 20240362830A1 · Zhang et al. · 2024 [cited by applicant]
US 20250148674A1 · Smetanin et al. · 2025 [cited by applicant]
US 20260057503A1 · Bakunov et al. · 2026 [cited by applicant]
CN 113721764A · 2021 [cited by examiner]
CN 115391588 · 2022 [cited by applicant]
EP 3698258 · 2020 [cited by applicant]
WO 2024182144 · 2024 [cited by applicant]
WO 2024182169 · 2024 [cited by applicant]
WO 2024182240 · 2024 [cited by applicant]
WO 2024182438 · 2024 [cited by applicant]
Cheng et al., LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation, Feb. 2023, ARXIV Cornell University Library (Year: 2023). [cited by examiner]
Cheng et al., Layout Diffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation, Feb. 2023, Cornell University Library (Year: 2023). [cited by examiner]
Li, Junnan, “BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation”, arXiv: 2201.12086v2 [cs.CV], (Feb. 15, 2022), 12 pgs. [cited by applicant]
Schuhmann, Christoph, “CLIP+MLP Aesthetic Score Predictor”, [Online] Retrieved from the Internet: <URL: https://github.com/christophschuhmann/improved-aesthetic-predictor>, (Jun. 30, 2022), 2 pgs. [cited by applicant]
Vincent, James, “TikTok Now Offers A Very Basic Text-To-Image AI Generator Directly In The App”, The Verge, (Aug. 15, 2022), 4 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Corrected Notice of Allowability mailed Nov. 6, 2024”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,843, Non Final Office Action mailed Nov. 8, 2024”, 16 pgs. [cited by applicant]
“U.S. Appl. No. 17/844,587, Corrected Notice of Allowability mailed Dec. 16, 2024”, 3 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Notice of Allowance mailed Jul. 26, 2024”, 7 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,971, Notice of Allowance mailed Sep. 17, 2024”, 6 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Non Final Office Action mailed Sep. 26, 2023”, 26 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,971, Non Final Office Action mailed Dec. 19, 2023”, 19 pgs. [cited by applicant]
“Midjourney Banned Words: Understanding the AI image Generator's Restrictions”, [Online]. Retrieved from the Internet: <https://midjourney.co.in/midjourney-banned-words-understanding-the-ai-image-generators-restrictions… [cited by applicant]
“AI Prompt Art Maker Generator”, Emoji World, [Online]. Retrieved from the Internet: <https://apps.apple.com/us/app/ai-prompt-art-maker-generator/id6444807049>, (Dec. 6, 2022), 5 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Response filed Dec. 18, 23 to Non Final Office Action mailed Sep. 26, 2023”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,971, Response filed Mar. 15, 24 to Non Final Office Action mailed Dec. 19, 2023”, 13 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Final Office Action mailed Mar. 20, 2024”, 27 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Response filed May 15, 2024 to Final Office Action mailed Mar. 20, 2024”, 9 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017140, International Search Report mailed May 28, 2024”, 4 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017140, Written Opinion mailed May 28, 2024”, 5 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016539, International Search Report mailed May 31, 2024”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016539, Written Opinion mailed May 31, 2024”, 5 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,971, Notice of Allowance mailed Jun. 3, 2024”, 6 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017545, International Search Report mailed Jun. 7, 2024”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017545, Written Opinion mailed Jun. 7, 2024”, 7 pgs. [cited by applicant]
“Application Serial No. 18 /116,003, Notice of Allowance mailed Jun. 7, 2024”, 7 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016223, International Search Report mailed Jun. 18, 2024”, 4 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016223, Written Opinion mailed Jun. 18, 2024”, 7 pgs. [cited by applicant]
Chen, Shoufa, “DiffusionDet: Diffusion Model for Object Detection”, Arxiv.org, Cornell University Library, 201, Olin Library Cornell University Ithaca, NY 1485, (Nov. 17, 2022), 16 pgs. [cited by applicant]
Cheng, Jiaxin, “LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Feb. 16, 2023), 15 pgs. [cited by applicant]
Dinh, Tan M., “Tise: Bag of Metrics for Text—to—Image Synthesis Evaluation”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Jul. 19, 2022), 34 pgs. [cited by applicant]
Gu, Shuyang, “GIQA: Generated Image Quality Assessment”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Mar. 19, 2020), 26 pgs. [cited by applicant]
Hao, Yaru, “Optimizing Prompts for Text-to-Image Generation”, Arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Dec. 19, 2022), 16 pgs. [cited by applicant]
Li, Yuheng, “GLIGEN: Open-Set Grounded Text-to-Image Generation”, Arxiv.org, Cornell University Library, 201, Olin Library Cornell University Ithaca, NY, 14853, (Jan. 17, 2023), 21 pgs. [cited by applicant]
Oppenlaender, Jonas, “A taxonomy of prompt modifiers for text-to-image generation”, arXiv preprint arXiv:2204.13988v2., (Jul. 31, 2022), 18 pgs. [cited by applicant]
Palli, Praneeth, “Want To Change Wallpaper For A Specific Chat On WhatsApp? Follow These Steps”, [Online]. Retrieved from the Internet: <https://in.mashable.com/tech/31833/want-to-change-wallpaper-for-a-specific-chat-on… [cited by applicant]
Yu, Wenxin, “Blind Image Quality Assessment for a Single Image From Text to Image Synthesis”, IEEE Access, IEEE, USA, vol. 9, (Jul. 1, 2021), 12 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,843, Examiner Interview Summary mailed Jan. 28, 2025”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,843, Notice of Allowance mailed Aug. 1, 2025”, 6 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,843, Response filed Feb. 7, 2025 to Non Final Office Action mailed Nov. 8, 2024”, 14 pgs. [cited by applicant]
“CompVis / stable-diffusion”, [Online]. Retrieved from the Internet: <URL: https://web.archive.org/web/20230228034148/https://github.com/CompVis/stable-diffusion>, (Archived on Feb. 28, 2023), 8 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016223, International Preliminary Report on Patentability mailed Sep. 11, 2025”, 9 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016539, International Preliminary Report on Patentability mailed Sep. 11, 2025”, 7 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017140, International Preliminary Report on Patentability mailed Sep. 11, 2025”, 7 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017545, International Preliminary Report on Patentability mailed Sep. 11, 2025”, 9 pgs. [cited by applicant]
“salesforce / BLIP”, [Online]. Retrieved from the Internet: <URL: https://web.archive.org/web/20230227172725/https://github.com/salesforce/BLIP>, (Archived on Feb. 27, 2023), 6 pgs. [cited by applicant]
“ultralytics / yolov5”, [Online]. Retrieved from the Internet: <URL: https://web.archive.org/web/20230226045622/https://github.com/ultralytics/yolov5>, (Archived on Feb. 26, 2023), 5 pgs. [cited by applicant]
Schuhmann, Christoph, “Laion-Aesthetics”, [Online]. Retrieved from the Internet: <URL: https://laion.ai/blog/laion-aesthetics/>, (Aug. 16, 2022), 4 pgs. [cited by applicant]