IP Library Granted Patent US 12,488,445
Granted Patent B2
US 12,488,445 · App. 18/176,843 · Granted Dec 2, 2025

Automatic image quality evaluation

Inventors: Mykyta Bakunov (Adliswil, CH); Arnab Ghosh (Oxford, GB); Pavel Savchenkov (London, GB); Sergey Smetanin (London, GB); Jian Ren (Marina Del Ray, CA)
Assignee: SNAP INC.
G06T7/0002G06F40/126G06F40/40G06T11/00G06T2207/20081G06T2207/20084G06T2207/30168
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,445
App. No.
18/176,843
Granted
Dec 2, 2025
Kind
B2
Abstract

Examples disclosed herein describe techniques for automatic image quality evaluation. A first set of images generated by a first automated image generator and a second set of images generated by a second automated image generator are accessed. A first machine learning model generates a first quality indicator for each image in the first set of images and the second set of images. A second machine learning model generates a second quality indicator for each image in the first set of images and the second set of images. Based on the generated indicators, a first image from the first set of images and a second image from the second set of images are automatically selected and compared. A first ranking of the first automated image generator and the second automated image generator is generated based on the comparison, and ranking data is caused to be presented on a device.

Claims (55)

1 . A system comprising:

a memory that stores instructions; and

one or more processors configured by the instructions to perform operations comprising:

accessing a first set of images generated by a first automated image generator and a second set of images generated by a second automated image generator, the first set of images and the second set of images having been automatically generated based on a first text prompt;

generating, by a first machine learning model that is trained to evaluate image quality using a first quality metric, a first quality indicator for each image in the first set of images and for each image in the second set of images;

generating, by a second machine learning model that is trained to evaluate image quality using a second quality metric, a second quality indicator for each image in the first set of images and for each image in the second set of images;

automatically selecting, based on the first quality indicators and the second quality indicators, a first image from the first set of images and a second image from the second set of images;

generating, based on an automatic comparison between the first image and the second image, a first ranking of the first automated image generator and the second automated image generator; and

causing presentation, on a device, of ranking data generated at least partially from the first ranking.

2 . The system of claim 1 , wherein the first automated image generator is a first text-to-image machine learning model, and the second automated image generator is a second text-to-image machine learning model.

3 . The system of claim 1 , the operations further comprising, prior to the accessing the first set of images:

generating, by the first automated image generator, the first set of images.

4 . The system of claim 3 , the operations further comprising, prior to the accessing the second set of images:

generating, by the second automated image generator, the second set of images.

5 . The system of claim 1 , wherein the first quality indicator and the second quality indicator are different, and the first quality indicator and the second quality indicator are each one of an aesthetic quality score, an alignment score, or a visual realism score.

6 . The system of claim 5 , the operations further comprising:

automatically determining, based on the respective first quality indicators and second quality indicators, a combined score for each image in the first set of images and a combined score for each image in the second set of images, wherein the first image has a highest combined score from among the first set of images and the second image has a highest combined score from among the second set of images.

7 . The system of claim 5 , wherein the aesthetic quality score for each image is generated by a Multi-Layer Perceptron (MLP) neural network.

8 . The system of claim 5 , wherein generation of the alignment score for each image comprises encoding the image and the first text prompt to obtain respective vectors, and automatically measuring a similarity between the respective vectors using a cosine similarity method.

9 . The system of claim 5 , wherein the visual realism score for each image is automatically generated using a Visual Question Answering (VQA) machine learning model.

10 . The system of claim 9 , wherein generating the visual realism score for each image further comprises automatically generating a prediction of whether the image includes one or more artifacts.

11 . The system of claim 1 , the operations further comprising:

generating, by a third machine learning model that is trained to evaluate image quality using a third quality metric, a third quality indicator for each image in the first set of images and for each image in the second set of images, wherein the first image and the second image are selected based on the first quality indicators, the second quality indicators, and the third quality indicators.

12 . The system of claim 11 , wherein the first quality indicator, the second quality indicator, and the third quality indicator are different, and the first quality indicator, the second quality indicator and the third quality indicator are each one of an aesthetic quality score, an alignment score, or a visual realism score.

13 . The system of claim 12 , the operations further comprising:

automatically determining, based on the respective first quality indicators, second quality indicators and third quality indicators, a combined score for each image in the first set of images and a combined score for each image in the second set of images, wherein the first image has a highest combined score from among the first set of images and the second image has a highest combined score from among the second set of images.

14 . The system of claim 1 , wherein the first ranking is based on an automatic comparison between the first quality indicator generated for the first image and the first quality indicator generated for the second image, and wherein the first quality metric is one of aesthetic quality, alignment, or visual realism.

15 . The system of claim 14 , the operations further comprising:

generating, based on an automatic comparison between the second quality indicator generated for the first image and the second quality indicator generated for the second image, a second ranking of the first automated image generator and the second automated image generator, the ranking data being generated at least partially from the first ranking and the second ranking, wherein the second quality metric differs from the first quality metric and is one of: aesthetic quality, alignment, or visual realism.

16 . The system of claim 3 , wherein the first automated image generator is a text-to-image machine learning model, and wherein the operations further comprise:

training the text-to-image machine learning model on a training data set comprising a plurality of training records, each training record comprising a training image and at least one corresponding text description for the training image.

17 . The system of claim 1 , the operations further comprising:

receiving, from a user device associated with a user, an image generation request comprising the first text prompt;

selecting, based on the automatic comparison between the first image and the second image, one of the first image or the second image as a selected image; and

causing presentation of the selected image on the user device.

18 . The system of claim 1 , the operations further comprising:

accessing a third set of images generated by the first automated image generator and a fourth set of images generated by the second automated image generator, the third set of images and the fourth set of images having been automatically generated based on a second text prompt;

generating, by the first machine learning model, a first quality indicator for each image in the third set of images and for each image in the fourth set of images;

generating, by the second machine learning model, a second quality indicator for each image in the third set of images and for each image in the fourth set of images;

automatically selecting, based on the first quality indicators and the second quality indicators for the third set of images and the fourth set of images, a third image from the third set of images and a fourth image from the fourth set of images; and

generating, based on an automatic comparison between the third image and the fourth image, a third ranking of the first automated image generator and the second automated image generator, wherein the ranking data is generated at least partially from the first ranking and the third ranking.

19 . A method comprising:

accessing, by one or more processors, a first set of images generated by a first automated image generator and a second set of images generated by a second automated image generator, the first set of images and the second set of images having been automatically generated based on a first text prompt;

generating, by the one or more processors and using a first machine learning model that is trained to evaluate image quality using a first quality metric, a first quality indicator for each image in the first set of images and for each image in the second set of images;

generating, by the one or more processors and using a second machine learning model that is trained to evaluate image quality using a second quality metric, a second quality indicator for each image in the first set of images and for each image in the second set of images;

automatically selecting, by the one or more processors, based on the first quality indicators and the second quality indicators, a first image from the first set of images and a second image from the second set of images;

generating, by the one or more processors, based on an automatic comparison between the first image and the second image, a first ranking of the first automated image generator and the second automated image generator; and

causing presentation, on a device, of ranking data generated at least partially from the first ranking.

20 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by at least one processor, cause the at least one processor to perform operations comprising:

accessing a first set of images generated by a first automated image generator and a second set of images generated by a second automated image generator, the first set of images and the second set of images having been automatically generated based on a first text prompt;

generating, by a first machine learning model that is trained to evaluate image quality using a first quality metric, a first quality indicator for each image in the first set of images and for each image in the second set of images;

generating, by a second machine learning model that is trained to evaluate image quality using a second quality metric, a second quality indicator for each image in the first set of images and for each image in the second set of images;

automatically selecting, based on the first quality indicators and the second quality indicators, a first image from the first set of images and a second image from the second set of images;

generating, based on an automatic comparison between the first image and the second image, a first ranking of the first automated image generator and the second automated image generator; and

causing presentation, on a device, of ranking data generated at least partially from the first ranking.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2023
From: BAKUNOV, MYKYTA; GHOSH, ARNAB; SAVCHENKOV, PAVEL; SMETANIN, SERGEY; REN, JIAN
To: SNAP INC.
Reel/Frame 063185/0532 →
Continuity (1)
Related Publication 20240296535A1 · Sep 5, 2024
References Cited (79)
US 7214065B2 · Fitzsimmons, Jr. · 2007 [cited by applicant]
US 8867849B1 · Kirkham et al. · 2014 [cited by applicant]
US 10496924B1 · Highnam et al. · 2019 [cited by applicant]
US 11445148B1 · Øhrn · 2022 [cited by applicant]
US 11809688B1 · Parasnis et al. · 2023 [cited by applicant]
US 11947893B1 · Seth · 2024 [cited by applicant]
US 12169626B2 · Zakharov et al. · 2024 [cited by applicant]
US 12205207B2 · Smetanin et al. · 2025 [cited by applicant]
US 20140051402A1 · Qureshi · 2014 [cited by applicant]
US 20140344712A1 · Okazawa et al. · 2014 [cited by applicant]
US 20160125269A1 · Lee · 2016 [cited by examiner]
US 20160188153A1 · Lerner et al. · 2016 [cited by applicant]
US 20190295302A1 · Fu · 2019 [cited by examiner]
US 20200267182A1 · Highnam et al. · 2020 [cited by applicant]
US 20210042796A1 · Khoury et al. · 2021 [cited by applicant]
US 20210209184A1 · Huang et al. · 2021 [cited by applicant]
US 20210352460A1 · Rohde et al. · 2021 [cited by applicant]
US 20220036153A1 · O'malia et al. · 2022 [cited by applicant]
US 20220101578A1 · Bedi et al. · 2022 [cited by applicant]
US 20220114698A1 · Liu · 2022 [cited by applicant]
US 20230025835A1 · Moriya et al. · 2023 [cited by applicant]
US 20230054174A1 · Peled et al. · 2023 [cited by applicant]
US 20230177878A1 · Sekar et al. · 2023 [cited by applicant]
US 20230215441A1 · Wu · 2023 [cited by applicant]
US 20230222703A1 · Baheti et al. · 2023 [cited by applicant]
US 20230230198A1 · Zhang et al. · 2023 [cited by applicant]
US 20230260164A1 · Yuan et al. · 2023 [cited by applicant]
US 20230262102A1 · Das et al. · 2023 [cited by applicant]
US 20230281789A1 · Sudarsky et al. · 2023 [cited by applicant]
US 20230298224A1 · Aggarwal et al. · 2023 [cited by applicant]
US 20230342284A1 · Easton et al. · 2023 [cited by applicant]
US 20240135610A1 · Kolkin et al. · 2024 [cited by applicant]
US 20240161258A1 · Maschmeyer · 2024 [cited by examiner]
US 20240169622A1 · Xie et al. · 2024 [cited by applicant]
US 20240193821A1 · Denison · 2024 [cited by applicant]
US 20240295953A1 · Zakharov et al. · 2024 [cited by applicant]
US 20240296606A1 · Smetanin et al. · 2024 [cited by applicant]
US 20240297957A1 · Bakunov et al. · 2024 [cited by applicant]
US 20240362830A1 · Zhang · 2024 [cited by examiner]
CN 115391588 · 2022 [cited by applicant]
EP 3698258 · 2020 [cited by applicant]
WO 2024182144 · 2024 [cited by applicant]
WO 2024182169 · 2024 [cited by applicant]
WO 2024182240 · 2024 [cited by applicant]
WO 2024182438 · 2024 [cited by applicant]
Li, Junnan, “BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation”, arXiv: 2201.12086v2 [cs.CV], (Feb. 15, 2022), 12 pgs. [cited by applicant]
Schuhmann, Christoph, “CLIP+MLP Aesthetic Score Predictor”, [Online] Retrieved from the Internet: <URL: https://github.com/christophschuhmann/improved-aesthetic-predictor>, (Jun. 30, 2022), 2 pgs. [cited by applicant]
Vincent, James, “TikTok Now Offers a Very Basic Text-To-Image AI Generator Directly in The App”, The Verge, (Aug. 15, 2022), 4 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Non Final Office Action mailed Sep. 26, 2023”, 26 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,971, Non Final Office Action mailed Dec. 19, 2023”, 19 pgs. [cited by applicant]
“Midjourney Banned Words: Understanding the AI image Generator's Restrictions”, [Online]. Retrieved from the Internet: <https://midjourney.co.in/midjourney-banned-words-understanding-the-ai-image-generators-restrictions… [cited by applicant]
“AI Prompt Art Maker Generator”, Emoji World, [Online]. Retrieved from the Internet: <https://apps.apple.com/us/app/ai-prompt-art-maker-generator/id6444807049>, (Dec. 6, 2022), 5 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Response filed Dec. 18, 2023 to Non Final Office Action mailed Sep. 26, 2023”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,971, Response filed Mar. 15, 2024 to Non Final Office Action mailed Dec. 19, 2023”, 13 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Final Office Action mailed Mar. 20, 2024”, 27 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Response filed May 15, 2024 to Final Office Action mailed Mar. 20, 2024”, 9 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017140, International Search Report mailed May 28, 2024”, 4 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017140, Written Opinion mailed May 28, 2024”, 5 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016539, International Search Report mailed May 31, 2024”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016539, Written Opinion mailed May 31, 2024”, 5 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,971, Notice of Allowance mailed Jun. 3, 2024”, 6 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017545, International Search Report mailed Jun. 7, 2024”, 3 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/017545, Written Opinion mailed Jun. 7, 2024”, 7 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Notice of Allowance mailed Jun. 7, 2024”, 7 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016223, International Search Report mailed Jun. 18, 2024”, 4 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/016223, Written Opinion mailed Jun. 18, 2024”, 7 pgs. [cited by applicant]
Chen, Shoufa, “DiffusionDet: Diffusion Model for Object Detection”, arxiv.org, Cornell University Library, 201, Olin Library Cornell University Ithaca, NY 1485, (Nov. 17, 2022), 16 pgs. [cited by applicant]
Cheng, Jiaxin, “LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Feb. 16, 2023), 15 pgs. [cited by applicant]
Dinh, Tan M., “TISE: Bag of Metrics for Text-to-Image Synthesis Evaluation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Jul. 19, 2022), 34 pgs. [cited by applicant]
Gu, Shuyang, “GIQA: Generated Image Quality Assessment”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Mar. 19, 2020), 26 pgs. [cited by applicant]
Hao, Yaru, “Optimizing Prompts for Text-to-Image Generation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Dec. 19, 2022), 16 pgs. [cited by applicant]
Li, Yuheng, “GLIGEN: Open-Set Grounded Text-to-Image Generation”, arxiv.org, Cornell University Library, 201, Olin Library Cornell University Ithaca, NY, 14853, (Jan. 17, 2023), 21 pgs. [cited by applicant]
Oppenlaender, Jonas, “A taxonomy of prompt modifiers for text-to-image generation”, arXiv preprint arXiv:2204.13988v2., (Jul. 31, 2022), 18 pgs. [cited by applicant]
Palli, Praneeth, “Want To Change Wallpaper For A Specific Chat on WhatsApp? Follow These Steps”, [Online]. Retrieved from the Internet: <https://in.mashable.com/tech/31833/want-to-change-wallpaper-for-a-specific-chat-on… [cited by applicant]
Yu, Wenxin, “Blind Image Quality Assessment for a Single Image From Text to Image Synthesis”, IEEE Access, IEEE, USA, vol. 9, (Jul. 1, 2021), 12 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Notice of Allowance mailed Jul. 26, 2024”, 7 pgs. [cited by applicant]
“U.S. Appl. No. 18/176,971, Notice of Allowance mailed Sep. 17, 2024”, 6 pgs. [cited by applicant]
“U.S. Appl. No. 18/116,003, Corrected Notice of Allowability mailed Nov. 6, 2024”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/844,587, Corrected Notice of Allowability mailed Dec. 16, 2024”, 3 pgs. [cited by applicant]