IP Library › Granted Patent US 12,511,795
Granted Patent B2
US 12,511,795 · App. 18/129,136 · Granted Dec 30, 2025

Method, electronic device, and computer program product for image generation

Inventors: Zijia Wang (Weifang, CN); Zhisong Liu (Shenzhen, CN); Jiacheng Ni (Shanghai, CN); Zhen Jia (Shanghai, CN)
Assignee: Dell Products L.P.
G06T11/00G06T5/77G06V10/761G06V10/764G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,795
App. No.
18/129,136
Granted
Dec 30, 2025
Kind
B2
Abstract

A method in an illustrative embodiment includes: acquiring an image set, where the image set includes a first plurality of images that can be classified into at least two categories; determining a corner case image set in the image set, where the corner case image set includes a second plurality of images that tend to be incorrectly classified; training an image generator with at least some images in the second plurality of images and first guidance associated with the at least some images; and generating an additional image by the trained image generator with a first image in the second plurality of images and additional guidance, where the additional guidance is different from the first guidance. By means of the technical solutions of the present disclosure, an image generation efficiency can be improved, and the quality of generated images can be enhanced, thereby improving user experience.

Claims (63)

1 . A method for image generation, comprising:

acquiring an image set, wherein the image set comprises a first plurality of images that can be classified into at least two categories;

determining a corner case image set in the image set, wherein the corner case image set comprises a second plurality of images that tend to be incorrectly classified;

training an image generator with at least some images in the second plurality of images and first guidance associated with the at least some images; and

generating an additional image by the trained image generator with a first image in the second plurality of images and additional guidance, wherein the additional guidance is different from the first guidance, the additional guidance comprising input information that at least partially mischaracterizes a visual element detected in the first image, and further wherein the additional image is generated by the trained image generator as a modified version of the first image in the second plurality of images, at least in part by: (i) applying a segmentation process to the first image to segment at least the detected visual element into a plurality of segmented portions, and (ii) applying a semantic alignment process to the first image to alter at least a subset of the segmented portions of the first image to provide semantic consistency between the altered segmented portions and the input information, the additional image thereby differing from the first image in the second plurality of images in a manner controlled at least in part by the additional guidance, to provide an expansion of the image set.

2 . The method for image generation according to claim 1 , wherein determining the corner case image set in the image set comprises:

determining the corner case image set utilizing a distance-based surprise adequacy method.

3 . The method for image generation according to claim 2 , wherein determining the corner case image set utilizing the distance-based surprise adequacy method comprises:

determining an image classification space for the at least two categories based on the image set;

determining, for a first image in the first plurality of images, a first Euclidean distance between the first image and other images belonging to the same category in the image classification space and a second Euclidean distance between the first image and images belonging to other categories in the image classification space; and

determining, based on the first Euclidean distance and the second Euclidean distance, whether the first image belongs to the corner case image set.

4 . The method for image generation according to claim 3 , wherein determining, based on the first Euclidean distance and the second Euclidean distance, whether the first image belongs to the corner case image set comprises:

determining, based on a ratio of the first Euclidean distance to the second Euclidean distance, whether the first image belongs to the corner case image set.

5 . The method for image generation according to claim 1 , wherein the first guidance comprises at least one of the following:

image guidance; and

word guidance.

6 . The method for image generation according to claim 1 , further comprising:

acquiring a prompt associated with the first image in the second plurality of images;

generating a processed first image with the first image and the prompt; and

generating the additional image with the first image in the second plurality of images and the additional guidance comprising:

generating the additional image with the processed first image and the additional guidance.

7 . The method for image generation according to claim 6 , wherein the prompt comprises at least one of the following:

an image prompt; and

a word prompt.

8 . The method for image generation according to claim 6 , wherein the first image comprises an object, and the prompt is associated with the object.

9 . The method for image generation according to claim 8 , wherein generating the processed first image with the first image and the prompt comprises:

removing the object from the first image based on the prompt.

10 . An electronic device, comprising:

at least one processing unit; and

at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, wherein the instructions, when executed by the at least one processing unit, cause the electronic device to perform actions comprising:

acquiring an image set, wherein the image set comprises a first plurality of images that can be classified into at least two categories;

determining a corner case image set in the image set, wherein the corner case image set comprises a second plurality of images that tend to be incorrectly classified;

training an image generator with at least some images in the second plurality of images and first guidance associated with the at least some images; and

generating an additional image by the trained image generator with a first image in the second plurality of images and additional guidance, wherein the additional guidance is different from the first guidance, the additional guidance comprising input information that at least partially mischaracterizes a visual element detected in the first image, and further wherein the additional image is generated by the trained image generator as a modified version of the first image in the second plurality of images, at least in part by: (i) applying a segmentation process to the first image to segment at least the detected visual element into a plurality of segmented portions, and (ii) applying a semantic alignment process to the first image to alter at least a subset of the segmented portions of the first image to provide semantic consistency between the altered segmented portions and the input information, the additional image thereby differing from the first image in the second plurality of images in a manner controlled at least in part by the additional guidance, to provide an expansion of the image set.

11 . The electronic device according to claim 10 , wherein determining the corner case image set in the image set comprises:

determining the corner case image set utilizing a distance-based surprise adequacy method.

12 . The electronic device according to claim 11 , wherein determining the corner case image set utilizing the distance-based surprise adequacy method comprises:

determining an image classification space for the at least two categories based on the image set;

determining, for a first image in the first plurality of images, a first Euclidean distance between the first image and other images belonging to the same category in the image classification space and a second Euclidean distance between the first image and images belonging to other categories in the image classification space; and

determining, based on the first Euclidean distance and the second Euclidean distance, whether the first image belongs to the corner case image set.

13 . The electronic device according to claim 12 , wherein determining, based on the first Euclidean distance and the second Euclidean distance, whether the first image belongs to the corner case image set comprises:

determining, based on a ratio of the first Euclidean distance to the second Euclidean distance, whether the first image belongs to the corner case image set.

14 . The electronic device according to claim 10 , wherein the first guidance comprises at least one of the following:

image guidance; and

word guidance.

15 . The electronic device according to claim 10 , further comprising:

acquiring a prompt associated with the first image in the second plurality of images;

generating a processed first image with the first image and the prompt; and

generating the additional image with the first image in the second plurality of images and the additional guidance comprising:

generating the additional image with the processed first image and the additional guidance.

16 . The electronic device according to claim 15 , wherein the prompt comprises at least one of the following:

an image prompt; and

a word prompt.

17 . The electronic device according to claim 15 , wherein the first image comprises an object, and the prompt is associated with the object.

18 . The electronic device according to claim 17 , wherein generating the processed first image with the first image and the prompt comprises:

removing the object from the first image based on the prompt.

19 . A computer program product tangibly stored on a non-transitory computer-readable medium and comprising machine-executable instructions, wherein the machine-executable instructions, when executed by a machine, cause the machine to perform actions comprising:

acquiring an image set, wherein the image set comprises a first plurality of images that can be classified into at least two categories;

determining a corner case image set in the image set, wherein the corner case image set comprises a second plurality of images that tend to be incorrectly classified;

training an image generator with at least some images in the second plurality of images and first guidance associated with the at least some images; and

generating an additional image by the trained image generator with a first image in the second plurality of images and additional guidance, wherein the additional guidance is different from the first guidance, the additional guidance comprising input information that at least partially mischaracterizes a visual element detected in the first image, and further wherein the additional image is generated by the trained image generator as a modified version of the first image in the second plurality of images, at least in part by: (i) applying a segmentation process to the first image to segment at least the detected visual element into a plurality of segmented portions, and (ii) applying a semantic alignment process to the first image to alter at least a subset of the segmented portions of the first image to provide semantic consistency between the altered segmented portions and the input information, the additional image thereby differing from the first image in the second plurality of images in a manner controlled at least in part by the additional guidance, to provide an expansion of the image set.

20 . The computer program product of claim 19 , wherein determining the corner case image set in the image set comprises:

determining the corner case image set utilizing a distance-based surprise adequacy method.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2023
From: WANG, ZIJIA; LIU, ZHISONG; NI, JIACHENG; JIA, ZHEN
To: DELL PRODUCTS L.P.
Reel/Frame 063183/0364 →
Priority Claims (1)
CN 202310183839.2 · Feb 28, 2023 · national
Continuity (1)
Related Publication 20240289998A1 · Aug 29, 2024
References Cited (12)
US 20140369596A1 · Siskind · 2014 [cited by examiner]
US 20180137119A1 · Li · 2018 [cited by examiner]
US 20220036512A1 · Kim · 2022 [cited by examiner]
US 20230047094A1 · Zhao · 2023 [cited by examiner]
A. Radford et al., “Learning Transferable Visual Models From Natural Language Supervision,” International Conference on Machine Learning, arXiv:2103.00020v1, Feb. 26, 2021, 48 pages. [cited by applicant]
J. Kim et al., “Guiding Deep Learning System Testing using Surprise Adequacy,” arXiv:1808.08444v1, Aug. 25, 2018, 12 pages. [cited by applicant]
A. Dosovitskiy et al., “An Image is Worth 16x16 Words Transformers for Image Recognition at Scale,” The International Conference on Learning Representations, arXiv:2010.11929v2, Jun. 3, 2021, 22 pages. [cited by applicant]
D. A. Hudson et al., “Generative Adversarial Transformers,” Proceedings of the 38th International Conference on Machine Learning, Jul. 2021, 13 pages. [cited by applicant]
Q. Chen et al., “Photographic Image Synthesis with Cascaded Refinement Networks,” arXiv:1707.09405v1, Jul. 28, 2017, 10 pages. [cited by applicant]
P. Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, pp. 1125-1134. [cited by applicant]
T.-C. Wang et al., “High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, 10 pages. [cited by applicant]
T. Park et al., “Semantic Image Synthesis with Spatially-Adaptive Normalization,” arXiv:1903.07291v2, Nov. 5, 2019, 19 pages. [cited by applicant]