IP Library Granted Patent US 12,591,949
Granted Patent B2
US 12,591,949 · App. 17/065,780 · Granted Mar 31, 2026

Image generation using one or more neural networks

Inventor: Ming-Yu Liu (San Jose, CA)
Assignee: NVIDIA Corporation
G06T3/4046G06T7/11G06T2200/24G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,949
App. No.
17/065,780
Granted
Mar 31, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques are presented to generate images. In at least one embodiment, one or more neural networks are used to adjust one or more aspect ratios of one or more objects of one or more images based, at least in part, on input from one or more users.

Claims (42)

1 . One or more processors, comprising:

circuitry to use one or more neural networks to obtain a user selected aspect ratio that specifies an aspect ratio of an image to be generated, and to generate a segmentation mask of the image to be generated, wherein the segmentation mask is generated having an aspect ratio corresponding to the user selected aspect ratio and one or more portions of the segmentation mask that correspond to one or more semantic objects are to be modified based, at least in part, on the user selected aspect ratio.

2 . The one or more processors of claim 1 , wherein the one or more neural networks are further to obtain a user adjusted aspect ratio that is used to adjust the aspect ratio of an image to be generated, and to generate one or more new instances of the segmentation mask using the user adjusted aspect ratio, wherein the one or more portions of the segmentation mask that correspond to the one or more semantic objects are to be modified based, at least in part, on the user adjusted aspect ratios.

3 . The one or more processors of claim 1 , wherein the one or more neural networks are to generate the image using the segmentation mask.

4 . The one or more processors of claim 1 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects are to be modified based, at least in part, on one or more realistic aspect ratios learned by the one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

5 . The one or more processors of claim 1 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects have corresponding one or more aspect ratios modified based, at least in part, on one or more realistic aspect ratios learned by the one or more neural networks that correspond to the one or more semantic objects, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

6 . The one or more processors of claim 1 , wherein the circuitry further is to identify that the user selected aspect ratio is inconsistent with one or more realistic aspect ratios learned by the one or more neural networks and generate a prompt for a user to specify an aspect ratio consistent with the one or more realistic aspect ratios learned by the one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

7 . A system comprising:

one or more processors to cause one or more circuits to use to use one or more neural networks to obtain a user adjusted aspect ratio that is used to adjust the aspect ratio of an image to be generated, and to generate a segmentation mask of the image to be generated, wherein the segmentation mask is generated having an aspect ratio corresponding to the user selected aspect ratio and one or more portions of the segmentation mask that correspond to one or more semantic objects are to be modified based, at least in part, on the user selected aspect ratio.

8 . The system of claim 7 , wherein the one or more neural networks are further to obtain a user adjusted aspect ratio that adjusts the aspect ratio of an image to be generated, and to generate one or more new instances of the segmentation mask using the user adjusted aspect ratios, wherein the one or more portions of the segmentation mask that correspond to the one or more semantic objects are to be modified based, at least in part, on the user adjusted aspect ratios.

9 . The system of claim 7 , wherein the one or more neural networks are to generate the image using the segmentation mask.

10 . The system of claim 7 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects are to be modified based, at least in part, on one or more realistic aspect ratios learned by the one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

11 . The system of claim 7 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects have its aspect ratio modified based, at least in part, on one or more realistic aspect ratios learned by the one or more neural networks that correspond to the one or more semantic objects, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

12 . The one or more processors of claim 1 , wherein input from one or more users explicitly selects an aspect ratio for the image to be generated, and the user selected aspect ratio is to be used by the one or more neural networks to adjust one or more aspect ratios of one or more semantic objects in the image to be generated.

13 . The system of claim 7 , wherein the one or more circuits are further to identify that the user selected aspect ratio is inconsistent with one or more realistic aspect ratios learned by the one or more neural networks and generate a prompt for a user to specify an aspect ratio consistent with the one or more realistic aspect ratios learned by the one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

14 . A method comprising:

obtaining a user selected aspect ratio that specifies an aspect ratio of an image to be generated; and

generating a segmentation mask of the image to be generated, wherein the segmentation mask is generated having an aspect ratio corresponding to the user selected aspect ratio and one or more portions of the segmentation mask that correspond to one or more semantic objects are to be modified based, at least in part, on the user selected aspect ratio.

15 . The method of claim 14 , further comprising:

obtaining a user adjusted aspect ratio that is used to adjust the aspect ratio of an image to be generated; and

generating one or more new instances of the segmentation mask using the user adjusted aspect ratios, wherein the one or more portions of the segmentation mask that correspond to the one or more semantic objects are to be modified based, at least in part, on the user adjusted aspect ratios.

16 . The method of claim 14 , further comprising:

generating the image using the segmentation mask.

17 . The method of claim 14 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects are to be modified based, at least in part, on one or more realistic aspect ratios learned by one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

18 . The method of claim 14 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects have its aspect ratio modified based, at least in part, on one or more realistic aspect ratios learned by one or more neural networks that correspond to the one or more semantic objects, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

19 . The method of claim 14 , further comprising: identifying that the user selected aspect ratio is inconsistent with one or more realistic aspect ratios learned by one or more neural networks and generate a prompt for a user to specify an aspect ratio consistent with the one or more realistic aspect ratios learned by the one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

20 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to cause one or more circuits to at least:

obtain a user adjusted aspect ratio that is used to adjust the aspect ratio of an image to be generated, and to generate a segmentation mask of the image to be generated, wherein the segmentation mask is generated having an aspect ratio corresponding to the user selected aspect ratio and one or more portions of the segmentation mask that correspond to one or more semantic objects are to be modified based, at least in part, on the user selected aspect ratio.

21 . The non-transitory machine-readable medium of claim 20 , wherein the instructions if performed further cause the one or more processors to:

obtain a user adjusted aspect ratio that adjusts the aspect ratio of an image to be generated, and to generate one or more new instances of the segmentation mask using the user adjusted aspect ratios, wherein the one or more portions of the segmentation mask that correspond to the one or more semantic objects are to be modified based, at least in part, on the user adjusted aspect ratios.

22 . The non-transitory machine-readable medium of claim 20 , wherein one or more neural networks generate the image using the segmentation mask.

23 . The non-transitory machine-readable medium of claim 20 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects are to be modified based, at least in part, on one or more realistic aspect ratios learned by one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

24 . The non-transitory machine-readable medium of claim 20 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects have its aspect ratio modified based, at least in part, on one or more realistic aspect ratios learned by one or more neural networks that correspond to the one or more semantic objects, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

25 . The non-transitory machine-readable medium of claim 20 , wherein the one or more circuits are further to identify that the user selected aspect ratio is inconsistent with one or more realistic aspect ratios learned by one or more neural networks and generate a prompt for a user to specify an aspect ratio consistent with the one or more realistic aspect ratios learned by the one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

26 . An image generation system, comprising:

one or more processors to cause one or more circuits to use to use one or more neural networks to obtain a user selected aspect ratio that specifies an aspect ratio of an image to be generated, and to generate a segmentation mask of the image to be generated, wherein the segmentation mask is generated having an aspect ratio corresponding to the user selected aspect ratio and one or more portions of the segmentation mask that correspond to one or more semantic objects are to be modified based, at least in part, on the user selected aspect ratio; and

memory for storing network parameters for the one or more neural networks.

27 . The image generation system of claim 26 , wherein the one or more processors are further to obtain a user adjusted aspect ratio that adjusts the aspect ratio of an image to be generated, and to generate one or more new instances of the segmentation mask using the user adjusted aspect ratios, wherein the one or more portions of the segmentation mask that correspond to the one or more semantic objects are to be modified based, at least in part, on the user adjusted aspect ratios.

28 . The image generation system of claim 26 , wherein the one or more neural networks are to generate the image using the segmentation mask.

29 . The image generation system of claim 26 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects are to be modified based, at least in part, on one or more realistic aspect ratios learned by the one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

30 . The image generation system of claim 26 , wherein the one or more portions of the segmentation mask corresponding to one or more semantic objects have its aspect ratio modified based, at least in part, on one or more realistic aspect ratios learned by the one or more neural networks that correspond to the one or more semantic objects, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

31 . The image generation system of claim 26 , wherein the one or more circuits are further to identify that the user selected aspect ratio is inconsistent with one or more realistic aspect ratios learned by the one or more neural networks and generate a prompt for a user to specify an aspect ratio consistent with the one or more realistic aspect ratios learned by the one or more neural networks, wherein the one or more realistic aspect ratios are based, at least in part, on one or more rules or style guides for one or more types of semantic objects.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2020
From: LIU, MING-YU
To: NVIDIA CORPORATION
Reel/Frame 054185/0122 →
Continuity (1)
Related Publication 20220114698A1 · Apr 14, 2022
References Cited (15)
US 20160057363A1 · Posa · 2016 [cited by examiner]
US 20190147296A1 · Wang · 2019 [cited by applicant]
US 20210042933A1 · Obayashi · 2021 [cited by examiner]
US 20210142479A1 · Phogat · 2021 [cited by examiner]
US 20230036950A1 · Saa-Garriga · 2023 [cited by examiner]
US 20230162321A1 · Seo · 2023 [cited by examiner]
US 20240422284A1 · Vanchinathan · 2024 [cited by examiner]
CN 111724302A · 2020 [cited by applicant]
WO 2020101434A1 · 2020 [cited by applicant]
International Search Report and Written Opinion issued in PCT Application No. PCT/US2021/053807, dated Feb. 2, 2022. [cited by applicant]
Donghoon Lee et al: “Context-Aware Synthesis and Placement of Object Instances”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Dec. 6, 2018 (Dec. 6, 2018), XP080989728, abs… [cited by applicant]
Moab Arar et al: “Image Resizing by Reconstruction from Deep Features”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 17, 2019 (Apr. 17, 2019), XP081170819, abstract; … [cited by applicant]
Yan Bo et al: “Semantic Segmentation Guided Pixel Fusion for Image Retargeting”, IEEE Transactions On Multimedia, IEEE, USA, vol. 22, No. 3, Jul. 31, 2019 (Jul. 31, 2019), pp. 676-687, XP011773757, ISSN: 1520-9210, DOI:… [cited by applicant]
Office Action for Chinese Application No. 202180008601.5, mailed Nov. 27, 2024, 18 pages. [cited by applicant]
Decision of Rejection for Chinese Application No. 20218008601.5, mailed Apr. 24, 2025, 22 pages. [cited by applicant]