IP Library Granted Patent US 12,682,553
Granted Patent B1
US 12,682,553 · App. 17/696,259 · Granted Jul 14, 2026

Controllable image generation using one or more neural networks

Inventors: Umar Iqbal (Fremont, CA); Amit Raj (Atlanta, GA); Pavlo Molchanov (Mountain View, CA); Jan Kautz (Lexington, MA); Koki Nagano (Playa Vista, CA); Sameh Khamis (Alameda, CA)
Assignee: NVIDIA Corporation
G06T15/20G06T7/70G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,553
App. No.
17/696,259
Filed
Mar 16, 2022
Granted
Jul 14, 2026
Kind
B1
Art Unit
2612
USPC
345/419
Abstract

Apparatuses, systems, and techniques are presented to generate one or more images. In at least one embodiment, one or more neural networks are used to generate one or more images of one or more objects including one or more portions of one or more other objects.

Claims (39)

1 . One or more processors, comprising:

circuitry to:

obtain geometric data and texture information for one or more first objects at least partially depicted in one or more first images, and specified poses of one or more second objects at least partially depicted in one or more second images; and

use the geometric data and texture information for the one or more first objects and the specified poses of the one or more second objects as input to one or more neural networks to generate one or more second images to depict the one or more first objects in the same specified poses as the one or more second objects.

2 . The one or more processors of claim 1 , wherein the circuitry is further to use the one or more neural networks to generate the one or more second images including a representation of another object in the specified pose depicting the one or more first objects in the same specified poses as the one or more second objects.

3 . The one or more processors of claim 1 , wherein the circuitry is further to identify the one or more first objects of the one or more first images to be posed across different images.

4 . The one or more processors of claim 1 , wherein the one or more neural networks include a generative model to synthesize the one or more second images, wherein the generative model is trained to synthesize image data for multiple different objects.

5 . The one or more processors of claim 1 , wherein the circuitry is further to train a generative model using at least one video including a representation of different objects.

6 . The one or more processors of claim 1 , wherein the circuitry is further to use the one or more neural networks to generate the one or more second images from one or more viewpoints, wherein the one or more viewpoints include viewpoints not represented in a single video.

7 . A system comprising:

one or more processors to:

obtain geometric data and texture information for one or more first objects at least partially depicted in one or more first images, and specified poses of one or more second objects at least partially depicted in one or more second images; and

use the geometric data and texture information for the one or more first objects and the specified poses of the one or more second objects as input to one or more neural networks to generate one or more second images to depict the one or more first objects in the same specified poses as the one or more second objects.

8 . The system of claim 7 , wherein the one or more processors are further to use the one or more neural networks to generate one or more second images including a representation of another object in the specified pose depicting the one or more first objects in the same specified poses as the one or more second objects.

9 . The system of claim 7 , wherein the one or more processors are further to identify the one or more first objects of the one or more first images and the one or more poses of the one or more second objects.

10 . The system of claim 7 , wherein the one or more neural networks include a generative model to synthesize the one or more second images, wherein the generative model is trained to synthesize image data for multiple different objects.

11 . The system of claim 7 , wherein the one or more processors are further to train a generative model using at least one video including a representation of one or more different objects.

12 . The system of claim 7 , wherein the one or more processors are further to use the one or more neural networks to generate the one or more second images from one or more viewpoints, wherein the one or more viewpoints include viewpoints not represented in a single video.

13 . A method comprising:

obtaining geometric data and texture information for one or more first objects at least partially depicted in one or more first images, and specified poses of one or more second objects at least partially depicted in one or more second images; and

using the geometric data and texture information for the one or more first objects and the specified poses of the one or more second objects as input to one or more neural networks to generate one or more second images to depict the one or more first objects in the same specified poses as the one or more second objects.

14 . The method of claim 13 , further comprising:

using the one or more neural networks to generate one or more second images including a representation of another object in the specified pose depicting the one or more first objects in the same specified poses as the one or more second objects.

15 . The method of claim 13 , further comprising:

identifying the one or more first objects of the one or more first images to be posed across different images.

16 . The method of claim 13 , wherein the one or more neural networks include a generative model to synthesize the one or more second images, wherein the generative model is trained to synthesize image data for multiple different objects.

17 . The method of claim 13 , further comprising:

training a generative model using at least one video including a representation of different objects.

18 . The method of claim 13 , further comprising:

using the one or more neural networks to generate the one or more second images from one or more viewpoints, wherein the one or more viewpoints include different viewpoints not represented in a single video.

19 . An image synthesis system, comprising:

one or more processors to obtain geometric data and texture information for one or more first objects at least partially depicted in one or more first images, and specified poses of one or more second objects at least partially depicted in one or more second images; and

use the geometric data and texture information for the one or more first objects and the specified poses of the one or more second objects as input to one or more neural networks to generate one or more second images to depict the one or more first objects in the same specified poses as the one or more second objects; and

memory for storing network parameters for the one or more neural networks.

20 . The image synthesis system of claim 19 , wherein the one or more processors are further to use the one or more neural networks to generate one or more second images including a representation of another object in the same specified pose depicting the one or more first objects in the specified poses of the one or more second objects.

21 . The image synthesis system of claim 19 , wherein the one or more processors are further to identify the one or more first objects of the one or more first images to be posed across different images.

22 . The image synthesis system of claim 19 , wherein the one or more neural networks include a generative model to synthesize the one or more second images, wherein the generative model is trained to synthesize image data for multiple different objects.

23 . The image synthesis system of claim 19 , wherein the one or more processors are further to train a generative model using at least one video including a representation of different objects.

24 . The image synthesis system of claim 19 , wherein the one or more processors are further to use the one or more neural networks to generate the one or more second images from one or more viewpoints.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2022
From: IQBAL, UMAR; RAJ, AMIT; MOLCHANOV, PAVLO; KAUTZ, JAN; NAGANO, KOKI; KHAMIS, SAMEH
To: NVIDIA CORPORATION
Reel/Frame 059444/0555 →
References Cited (32)
US 11282291B1 · Boardman · 2022 [cited by examiner]
US 11288857B2 · Meshry · 2022 [cited by examiner]
US 12147505B1 · Marsden · 2024 [cited by examiner]
US 12182938B2 · Totty · 2024 [cited by examiner]
US 20180129910A1 · Zia · 2018 [cited by examiner]
US 20190026917A1 · Liao · 2019 [cited by examiner]
US 20200113666A1 · Loomis · 2020 [cited by examiner]
US 20200302251A1 · Shechtman · 2020 [cited by examiner]
US 20210152735A1 · Zhou · 2021 [cited by examiner]
US 20220026920A1 · Ebrahimi Afrouzi · 2022 [cited by examiner]
US 20220067982A1 · Pardeshi · 2022 [cited by examiner]
US 20220122307A1 · Kalarot · 2022 [cited by examiner]
US 20220138840A1 · Sadalgi · 2022 [cited by examiner]
US 20220139037A1 · Li · 2022 [cited by examiner]
US 20220198738A1 · Xu · 2022 [cited by examiner]
US 20230230275A1 · Lin · 2023 [cited by examiner]
US 20230274492A1 · Chen · 2023 [cited by examiner]
US 20230377180A1 · Ambrus · 2023 [cited by examiner]
Chen et al., “Animatable Neural Radiance Fields from Monocular RGB Videos,” Sep. 7, 2021, 12 Pages. [cited by applicant]
He et al., “ARCH++: Animation-Ready Clothed Human Reconstruction Revisited,” Oct. 13, 2021, 11 Pages. [cited by applicant]
Huang et al., “Arbitrary Style Transfer in Realtime with Adaptive Instance Normalization,” IEEE International Conference on Computer Vision, 2017, 10 pages. [cited by applicant]
Huang et al., “ARCH: Animatable Reconstruction of Clothed Humans,” IEEE/CVF Conference on Computer Vision and Pattern Recongnition, Jun. 13, 2020, 10 Pages. [cited by applicant]
IEEE “IEEE Standard for Floating-Point Arithmetric”, Microprocessor Standards Committee of the IEEE Computer Society, IEEE Std 754-2008, dated Jun. 12, 2008. [cited by applicant]
Liu et al., “Neutral Actor: Neural Free-view Synthesis of Human Actors with Pose Control,” Jun. 3, 2021, 15 Pages. [cited by applicant]
Peng et al., “Animatable Neural Radiance Fields for Human Body Modeling,” May 6, 2021, 10 Pages. [cited by applicant]
Peng et al., “Neural Body: Implicit Neural Representations with Structured Latent Codes for Novel View Synthesis of Dynamic Humans,” Mar. 29, 2021, 10 Pages. [cited by applicant]
Raj et al., “ANR: Articulated Neural Rendering for Virtual Avatars,” Dec. 23, 2020, 10 Pages. [cited by applicant]
Saito et al., “PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization,” ICCV, 2019, 11 pages. [cited by applicant]
Saito et al., “PIFuHD: Multi-Level Pixel-Aligned Implicit Function for High-Resolution 3D Human Digitization,” CVPR, 2020, 10 pages. [cited by applicant]
Su et al., “A—NeRF: Surface-free Human 3D Pose Refinement via Neural Rendering,” Feb. 11, 2021, 15 Pages. [cited by applicant]
Yang et al., “Neural Shape, Skeleton, and Skinning Fields for 3D Human Modeling,” Jan. 17, 2021, 16 Pages. [cited by applicant]
Zhi et al., “TexMesh: Reconstructing Detailed Human Texture and Geometry from RGB-D Video,” Sep. 21, 2020, 23 Pages. [cited by applicant]