IP Library › Granted Patent US 12,530,820
Granted Patent B2
US 12,530,820 · App. 16/588,910 · Granted Jan 20, 2026

Image generation using one or more neural networks

Inventor: Ming-Yu Liu (San Jose, CA)
Assignee: NVIDIA Corporation
G06T11/001G06N3/02G06T3/4053G06T7/11G06V10/764G06V10/774G06V10/82G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,820
App. No.
16/588,910
Filed
Sep 30, 2019
Granted
Jan 20, 2026
Kind
B2
Art Unit
2664
USPC
382/173
Abstract

Apparatuses, systems, and techniques are presented to generate or manipulate digital images. In at least one embodiment, a network is trained to generate modified images including user-selected features.

Claims (42)

1 . A processor, comprising:

one or more circuits to use one or more neural networks to replace one or more portions of a first image with one or more portions of a second image based, at least in part, on one or more descriptive labels of the one or more portions of the first image indicated by one or more users.

2 . The processor of claim 1 , wherein the first image comprises an image captured using a camera or an image generated from an initial segmentation mask.

3 . The processor of claim 1 , wherein the one or more circuits are to apply one or more style filters to content to be rendered corresponding to the descriptive labels of the one or more portions of the first image.

4 . The processor of claim 1 , wherein the one or more circuits are further to determine one or more segmentation boundaries of one or more features of the first image, and wherein the segmentation boundaries are enabled to be added, deleted, or modified.

5 . The processor of claim 4 , wherein deletion of a segmentation boundary enables removal of an object represented in the first image.

6 . The processor of claim 1 , wherein the one or more circuits are to use the one or more neural networks to determine a different type of content to be rendered in one or more other portions of the first image.

7 . The processor of claim 1 , wherein the replacement of the one or more portions causes the first image to have a higher resolution.

8 . A system, comprising:

one or more processors to use one or more neural networks to replace one or more portions of a first image with one or more portions of a second image based, at least in part, on one or more descriptive labels of the one or more portions of the first image indicated by one or more users; and memory to store the first image and the second image.

9 . The system of claim 8 , wherein the first image comprises an image captured using a camera or an image generated from an initial segmentation mask.

10 . The system of claim 8 , wherein the one or more processors are to apply one or more style filters to content to be rendered corresponding to the descriptive labels of the one or more portions of the first image.

11 . The system of claim 8 , wherein the one or more processors are further to determine one or more segmentation boundaries of one or more features of the one first image, and wherein the segmentation boundaries are enabled to be added, deleted, or modified.

12 . The system of claim 11 , wherein deletion of a segmentation boundary enables removal of an object represented in the first image.

13 . The system of claim 8 , wherein the one or more processors are to use the one or more neural networks to determine a different type of content to be rendered in one or more other portions of the first image.

14 . The system of claim 8 , wherein the replacement of the one or more portions causes the first image to have a higher resolution.

15 . A method, comprising:

using one or more neural networks to replace one or more portions of a first image with one or more portions of a second image based, at least in part, on descriptive labels of the one or more portions of the first image indicated by one or more users.

16 . The method of claim 15 , wherein the first image comprises an image captured using a camera or an image generated from an initial segmentation mask.

17 . The method of claim 15 , further comprising:

applying one or more style filters to content to be rendered corresponding to the descriptive labels of the one or more portions of the first image.

18 . The method of claim 15 , further comprising:

enabling one or more segmentation boundaries determined of one or more features of the first image to be added, deleted, or modified.

19 . The method of claim 18 , wherein deletion of a segmentation boundary enables removal of an object represented in the first image.

20 . The method of claim 15 , further comprising using the one or more neural networks to determine a different type of content to be rendered in one or more other portions of the first image.

21 . The method of claim 15 , wherein the replacement of the one or more portions causes the first image to have a higher resolution.

22 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

use one or more neural networks to replace one or more portions of a first image with one or more portions of a second image based, at least in part, on descriptive labels of the one or more portions of the first image indicated by one or more users.

23 . The non-transitory machine-readable medium of claim 22 , wherein the first image comprises an image captured using a camera or an image generated from an initial segmentation mask.

24 . The non-transitory machine-readable medium of claim 22 , wherein the one or more processors are to apply one or more style filters to content to be rendered corresponding to the descriptive labels of the one or more portions of the first image.

25 . The non-transitory machine-readable medium of claim 22 , wherein the one or more processors are further to determine one or more segmentation boundaries of one or more features of the first image, and wherein the segmentation boundaries are enabled to be added, deleted, or modified.

26 . The non-transitory machine-readable medium of claim 25 , wherein deletion of a segmentation boundary enables removal of an object represented in the first image.

27 . The non-transitory machine-readable medium of claim 22 , wherein the one or more processors are to use the one or more neural networks to determine a different type of content to be rendered in one or more other portions of the first image.

28 . The non-transitory machine-readable medium of claim 22 , wherein the replacement of the one or more portions causes the first image to have a higher resolution.

29 . A processor comprising:

one or more circuits to train one or more neural networks to replace one or more portions of a first image with one or more portions of a second image based, at least in part, on one or more descriptive labels of the one or more portions of the first image indicated by one or more users.

30 . The processor of claim 29 , wherein the first image comprises an image captured using a camera or an image generated from an initial segmentation mask.

31 . The processor of claim 29 , wherein the one or more circuits are to apply one or more style filters to content to be rendered corresponding to the descriptive labels of the one or more portions of the first image.

32 . The processor of claim 29 , wherein the one or more circuits are further to determine one or more segmentation boundaries of one or more features of the first image, and wherein the segmentation boundaries are enabled to be added, deleted, or modified.

33 . The processor of claim 32 , wherein deletion of a segmentation boundary enables removal of an object represented in the first image.

34 . The processor of claim 29 , wherein the one or more circuits are to use the one or more neural networks to determine a different type of content to be rendered in one or more other portions of the first image.

35 . The processor of claim 29 , wherein the replacement of the one or more portions causes the first image to have a higher resolution.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2019
From: LIU, MING-YU
To: NVIDIA CORPORATION
Reel/Frame 050788/0907 →
Continuity (1)
Related Publication 20210097691A1 · Apr 1, 2021
References Cited (14)
US 20160117800A1 · Korkin · 2016 [cited by examiner]
US 20170294000A1 · Shen · 2017 [cited by examiner]
US 20180189598A1 · Cheung · 2018 [cited by examiner]
US 20190147582A1 · Lee · 2019 [cited by applicant]
US 20200151860A1 · Safdarnejad · 2020 [cited by examiner]
US 20200242774A1 · Park · 2020 [cited by examiner]
CN 108537864A · 2018 [cited by applicant]
Park: “Semantic Image Synthesis With Spatially-Adaptive Normalization”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE. Jun. 15, 2019 (Jun. 15, 2019). pp. 2332-2341. XP033686771. DOI: 1… [cited by examiner]
International Search Report and Written Opinion issued in PCT Application No. PCT/US2020/052665, dated Nov. 27, 2020. [cited by applicant]
Park Taesung et al: “Semantic Image Synthesis With Spatially-Adaptive Normalization”. 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE. Jun. 15, 2019 (Jun. 15, 2019). pp. 2332-2341. XP033… [cited by applicant]
Wang Ting-Chun et al: “High-Resolution Image Synthesis and Semantic Manipulation with Conditional GANs”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Jun. 18, 2018 (Jun. 18, 2018), pp. 8798… [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Office Action for Chinese Application No. 202080063533.8, mailed Feb. 5, 2025, 17 pages. [cited by applicant]
Office Action for Chinese Application No. 202080063533.8, mailed Aug. 2, 2025, 14 pages. [cited by applicant]