IP Library Granted Patent US 11,537,277
Granted Patent B2
US 11,537,277 · App. 16/040,220 · Granted Dec 27, 2022

System and method for generating photorealistic synthetic images based on semantic information

Inventors: Raja Bala (Pittsford, NY); Sricharan Kallur Palli Kumar (Mountain View, CA); Matthew A. Shreve (Mountain View, CA)
Assignee: Palo Alto Research Center Incorporated
G06F3/04845G06F3/04847G06N3/0454G06N3/0472G06N3/088G06T5/002G06T7/11G06T7/143G06T7/35G06T11/00G06T2207/20016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,537,277
App. No.
16/040,220
Granted
Dec 27, 2022
Kind
B2
Abstract

Embodiments described herein provide a system for generating semantically accurate synthetic images. During operation, the system generates a first synthetic image using a first artificial intelligence (AI) model and presents the first synthetic image in a user interface. The user interface allows a user to identify image units of the first synthetic image that are semantically irregular. The system then obtains semantic information for the semantically irregular image units from the user via the user interface and generates a second synthetic image using a second AI model based on the semantic information. The second synthetic image can be an improved image compared to the first synthetic image.

Claims (44)

1. A method for generating semantically accurate synthetic images, comprising:

generating a first synthetic image using a first artificial intelligence (AI) model;

presenting the first synthetic image in a user interface, wherein the user interface is configured to receive information indicating image units of the first synthetic image that are semantically irregular;

obtaining semantic information for the semantically irregular image units from a user via the user interface;

incorporating the semantic information into a loss function of a second AI model; and

generating a second synthetic image using the second AI model based on the loss function, wherein the second synthetic image is generated when the loss function is within a threshold and includes one or more improved image units compared to the first synthetic image.

2. The method of claim 1 , further comprising:

obtaining a selection of a highlighting tool for the user interface, wherein a region marked by the highlighting tool is selectable via the user interface, and wherein the highlighting tool corresponds to one of: a grid-based selector, a polygon-based selector, and a freehand selector; and

in response to the highlighting tool being the grid-based selector, setting the highlighting tool at an obtained granularity.

3. The method of claim 2 , further comprising:

obtaining a selection of a region selected by the highlighting tool in the user interface; and

obtaining a weight allocated for the region.

4. The method of claim 1 , wherein the first and second AI models are generative adversarial networks (GANs).

5. The method of claim 1 , wherein the second AI model includes a discriminator that outputs a spatial probability map, wherein a respective element of the spatial probability map indicates the probability of an image unit of the second synthetic image being realistic.

6. The method of claim 5 , wherein the second AI model includes a generator that outputs the second synthetic image in such a way that corresponding image units, which include the one or more improved image units, are different than the semantically irregular image units of the first synthetic image.

7. The method of claim 6 , further comprising generating a spatial weight mask based on the obtained semantic information, wherein the spatial weight mask comprises weights allocated to a respective image unit of the first synthetic image, and wherein the generator and the discriminator determine the semantically irregular image units based on the spatial weight mask.

8. The method of claim 1 , wherein an image unit corresponds to one or more of: a pixel, a pixel block, and a group of pixels of an irregular shape.

9. The method of claim 1 , wherein generating the first and the second synthetic images includes mapping a noise vector into a synthetic image sample.

10. The method of claim 1 , further comprising:

generating a third synthetic image using the first AI model;

automatically obtaining semantic information for semantically irregular image units of the third synthetic image based on an irregularity detection technique; and

generating a fourth synthetic image using the second AI model based on the automatically obtained semantic information.

11. A non-transitory computer-readable storage medium storing instructions that when executed by a computer cause the computer to perform a method for generating semantically accurate synthetic images, the method comprising:

generating a first synthetic image using a first artificial intelligence (AI) model;

presenting the first synthetic image in a user interface, wherein the user interface is configured to receive information indicating image units of the first synthetic image that are semantically irregular;

obtaining semantic information for the semantically irregular image units from a user via the user interface;

incorporating the semantic information into a loss function of a second AI model; and

generating a second synthetic image using the second AI model based on the loss function, wherein the second synthetic image is generated when the loss function is within a threshold and includes one or more improved image units compared to the first synthetic image.

12. The non-transiory computer-readable storage medium of claim 11 , wherein the method further comprises:

obtaining a selection of a highlighting tool for the user interface, wherein a region marked by the highlighting tool is selectable via the user interface, and wherein the highlighting tool corresponds to one of: a grid-based selector, a polygon-based selector, and a freehand selector; and

in response to the highlighting tool being the grid-based selector, setting the highlighting tool at an obtained granularity.

13. The non-transitory computer-readable storage medium of claim 12 , wherein the method further comprises:

obtaining a selection of a region selected by the highlighting tool in the user interface; and

obtaining a weight allocated for the region.

14. The non-transitory computer-readable storage medium of claim 11 , wherein the first and second AI models are generative adversarial networks (GANs).

15. The non-transitory computer-readable storage medium of claim 11 , wherein the second AI model includes a discriminator that outputs a spatial probability map, wherein a respective element of the spatial probability map indicates the probability of an image unit of the second synthetic image being realistic.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the second AI model includes a generator that outputs the second synthetic image in such a way that corresponding image units are different than the semantically irregular image units of the first synthetic image.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the method further comprises generating a spatial weight mask based on the obtained semantic information, wherein the spatial weight mask comprises weights allocated to a respective image unit of the first synthetic image, and wherein the generator and the discriminator determine the semantically irregular image units based on the spatial weight mask.

18. The non-transitory computer-readable storage medium of claim 11 , wherein an image unit corresponds to one or more of: a pixel, a pixel block, and a group of pixels of an irregular shape.

19. The non-transitory computer-readable storage medium of claim 11 , wherein generating the first and the second synthetic images includes mapping a noise vector into a synthetic image sample.

20. The non-transitory computer-readable storage medium of claim 11 , wherein the method further comprises:

generating a third synthetic image using the first AI model;

automatically obtaining semantic information for semantically irregular image units of the third synthetic image based on an irregularity detection technique; and

generating a fourth synthetic image using the second AI model based on the automatically obtained semantic information.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2025
From: XEROX CORPORATION
To: GENESEE VALLEY INNOVATIONS, LLC
Reel/Frame 073562/0677 →
SECOND LIEN NOTES PATENT SECURITY AGREEMENT Recorded Jul 2, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 071785/0550 →
FIRST LIEN NOTES PATENT SECURITY AGREEMENT Recorded Apr 11, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 070824/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT RF 064760/0389 Recorded Feb 13, 2024
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: XEROX CORPORATION
Reel/Frame 068261/0001 →
SECURITY INTEREST Recorded Feb 13, 2024
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 066741/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVAL OF US PATENTS 9356603, 10026651, 10626048 AND INCLUSION OF US PATENT 7167871 PREVIOUSLY RECORDED ON REEL 064038 FRAME 0001. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 28, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064161/0001 →
SECURITY INTEREST Recorded Jun 22, 2023
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 064760/0389 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2023
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: XEROX CORPORATION
Reel/Frame 064038/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2018
From: BALA, RAJA; KALLUR PALLI KUMAR, SRICHARAN; SHREVE, MATTHEW A.
To: PALO ALTO RESEARCH CENTER INCORPORATED
Reel/Frame 046405/0578 →
Continuity (1)
Related Publication 20200026416A1 · Jan 23, 2020
Cited By (1)
US 12,493,798