IP Library › Granted Patent US 12,536,798
Granted Patent B2
US 12,536,798 · App. 17/701,037 · Granted Jan 27, 2026

Image generation using a neural network

Inventor: Chong Yu (Shanghai, CN)
Assignee: NVIDIA Corporation
G06V20/46G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,798
App. No.
17/701,037
Granted
Jan 27, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to generate an image. In at least one embodiment, one or more neural networks are to generate a second image based, at least in part, on a first image and information indicating zero or more differences between the first and second image.

Claims (59)

1 . A processor, comprising:

one or more circuits to:

determine a first image to be a new keyframe based, at least in part, on comparison of a difference between the first image and a current keyframe to a threshold; and

use one or more neural networks to:

generate a first feature map for the first image; and

generate a second feature map of a second image based, at least in part, on the first feature map of the first image and information indicating zero or more differences between the first image and the second image, wherein the first image and the second image are different frames of a video.

2 . The processor of claim 1 , wherein prior to generation of the second feature map, the one or more circuits are further to compress the information indicating zero of more differences using a sparse compression technique based, at least in part, on a number of non-zero values in an array of the information and a number of zero values in the array.

3 . The processor of claim 1 , wherein prior to generation of the second feature map, the one or more circuits are further to determine the second image is a not keyframe based, at least in part, on comparison of a difference between the second image and the first image to the threshold.

4 . The processor of claim 1 , wherein the one or more neural networks are to generate the second feature map of the second image based, at least in part, on a feature map of the differences between the first image and the second image.

5 . The processor of claim 1 , wherein the information indicating the zero or more differences indicates pixel differences between the first image and the second image.

6 . The processor of claim 4 , wherein the one or more neural networks are to generate the second feature map of the second image based, at least in part, on a combination of the first feature map and the feature map of the differences between the first image and the second image.

7 . The processor of claim 1 , wherein the one or more neural networks are to augment the second image with feature map information based, at least in part, on the second feature map.

8 . The processor of claim 1 , wherein the one or more neural networks are to generate a third feature map of a third image based, at least in part, on the first feature map of the first image and information indicating zero or more differences between the first image and the third image, wherein the first image, the second image, and the third image are different frames of a video from a video source.

9 . The processor of claim 1 , wherein the zero or more differences between the first and second images are to be determined using per-pixel subtraction.

10 . The processor of claim 1 , wherein the zero or more differences between the first and second images are to be determined based, at least in part, on applying one or more convolution filters to the first and second images.

11 . A computer-implemented method, comprising:

determining a first image to be a new keyframe based, at least in part, on comparison of a difference between the first image and a current keyframe to a threshold; and

using one or more neural networks to:

generate a first feature map for the first image; and

generate a second feature map of a second image based, at least in part, on the first feature map of the first image and information indicating zero or more differences between the first image and the second image, wherein the first image and the second image are different frames of a video.

12 . The computer-implemented method of claim 11 , further comprising, prior to generation of the second feature map, compressing the information indicating zero of more differences using a sparse compression technique based, at least in part, on a number of non-zero values in an array of the information and a number of zero values in the array.

13 . The computer-implemented method of claim 11 , further comprising, prior to generation of the second feature map, determining the second image is a not keyframe based, at least in part, on comparison of a difference between the second image and the first image to the threshold.

14 . The computer-implemented method of claim 11 , wherein the second feature map of the second image is to be generated based, at least in part, on a feature map of the zero or more differences between the first image and the second image.

15 . The computer-implemented method of claim 11 , wherein the information indicating the zero or more differences indicates pixel differences between the first image and the second image.

16 . The computer-implemented method of claim 11 , further comprising:

compressing the first image prior to using the one or more neural networks to generate the second feature map of the second image.

17 . The computer-implemented method of claim 14 , further comprising:

combining the feature map of the zero or more differences between the first image and the second image with the first feature map to generate a combined feature map as the second feature map.

18 . The computer-implemented method of claim 11 , wherein:

the first feature map is to be used to generate a plurality of feature maps based, at least in part, on the first feature map and information indicating zero or more differences between the first image and respective images of a plurality of other images.

19 . The computer-implemented method of claim 14 , wherein:

a first neural network of the one or more neural networks is to generate the feature map of the zero or more differences between the first image and the second image; and

a second neural network of the one or more neural networks is to generate the second feature map based, at least in part, on the first feature map and the feature map of the zero or more differences between the first image and the second image.

20 . The computer-implemented method of claim 11 , wherein the zero or more differences between the first and second images are to be determined based, at least in part, on applying one or more convolution filters to the first and second images.

21 . A computer system, comprising:

one or more processors and memory storing executable instructions that, if performed by the one or more processors, cause the one or more processors to:

determine a first image to be a new keyframe based, at least in part, on comparison of a difference between the first image and a current keyframe to a threshold; and

use one or more neural networks to:

generate a first feature map for the first image; and

generate a second feature map of a second image based, at least in part, on the first feature map of the first image and information indicating zero or more differences between the first image and the second image, wherein the first image and the second image are different frames of a video.

22 . The computer system of claim 21 , wherein the executable instructions cause the one or more processors to, prior to generation of the second feature map, compress the information indicating zero of more differences using a sparse compression technique based, at least in part, on a number of non-zero values in an array of the information and a number of zero values in the array.

23 . The computer system of claim 21 , wherein the executable instructions cause the one or more processors to, prior to generation of the second feature map, determine the second image is a not keyframe based, at least in part, on comparison of a difference between the second image and the first image to the threshold.

24 . The computer system of claim 21 , wherein the one or more neural networks are to generate the second feature map of the second image based, at least in part, on a feature map of the differences between the first image and the second image.

25 . The computer system of claim 21 , wherein the information indicating the zero or more differences indicates pixel differences between the first image and the second image.

26 . The computer system of claim 25 , wherein the one or more neural networks are to generate the second feature map of the second image based, at least in part, on a combination of the first feature map and the feature map of the differences between the first image and the second image.

27 . The computer system of claim 21 , wherein the one or more neural networks are to augment the second image with feature map information based, at least in part, on the second feature map.

28 . The computer system of claim 21 , wherein the one or more neural networks are to generate a third feature map of a third image based, at least in part, on the first feature map of the first image and information indicating zero or more differences between the first image and the third image, wherein the first image, the second image, and the third image are different frames of a video from a video source.

29 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors,

determine a first image to be a new keyframe based, at least in part, on comparison of a difference between the first image and a current keyframe to a threshold; and

cause one or more neural networks to:

generate a first feature map for the first image; and

generate a second feature map of a second image based, at least in part, on the first feature map of the first image and information indicating zero or more differences between the first image and the second image, wherein the first image and the second image are different frames of a video.

30 . The non-transitory machine-readable medium of claim 29 , wherein the set of instructions, which if performed by the one or more processors, compress, prior to generation of the second feature map, the information indicating zero of more differences using a sparse compression technique based, at least in part, on a number of non-zero values in an array of the information and a number of zero values in the array.

31 . The non-transitory machine-readable medium of claim 29 , wherein the set of instructions, which if performed by the one or more processors, determine, prior to generation of the second feature map, the second image is a not keyframe based, at least in part, on comparison of a difference between the second image and the first image to the threshold.

32 . The non-transitory machine-readable medium of claim 29 , wherein the one or more neural networks are to generate the second feature map of the second image based, at least in part, on a feature map of the differences between the first image and the second image.

33 . The non-transitory machine-readable medium of claim 29 , wherein the information indicating the zero or more differences indicates pixel differences between the first image and the second image.

34 . The non-transitory machine-readable medium of claim 32 , wherein the one or more neural networks are to generate the second feature map of the second image based, at least in part, on a combination of the first feature map and the feature map of the differences between the first image and the second image.

35 . The non-transitory machine-readable medium of claim 29 , wherein the one or more neural networks are to augment the second image with feature map information based, at least in part, on the second feature map.

36 . The non-transitory machine-readable medium of claim 29 , wherein the one or more neural networks are to generate a third feature map of a third image based, at least in part, on the first feature map of the first image and information indicating zero or more differences between the first image and the third image, wherein the first image, the second image, and the third image are different frames of a video from a video source.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2022
From: YU, CHONG
To: NVIDIA CORPORATION
Reel/Frame 059341/0638 →
Priority Claims (1)
WO PCT/CN2022/075485 · Feb 8, 2022 · international
Continuity (1)
Related Publication 20230306739A1 · Sep 28, 2023
References Cited (28)
US 11490128B2 · Zhang · 2022 [cited by examiner]
US 20190197350A1 · Park · 2019 [cited by examiner]
US 20200065632A1 · Guo et al. · 2020 [cited by applicant]
US 20200304802A1 · Habibian · 2020 [cited by examiner]
US 20200304804A1 · Habibian · 2020 [cited by examiner]
US 20210004653A1 · Kim et al. · 2021 [cited by applicant]
US 20210152799A1 · Park et al. · 2021 [cited by applicant]
US 20210281867A1 · Golinski · 2021 [cited by examiner]
US 20210314474A1 · Yang · 2021 [cited by examiner]
US 20210342977A1 · Xia · 2021 [cited by applicant]
US 20220012536A1 · Wang · 2022 [cited by examiner]
US 20220103839A1 · Van Rozendaal · 2022 [cited by examiner]
US 20220295095A1 · Pourreza · 2022 [cited by examiner]
US 20220385907A1 · Zhang · 2022 [cited by examiner]
US 20230057261A1 · Liu · 2023 [cited by examiner]
US 20230074979A1 · Brehmer · 2023 [cited by examiner]
US 20230154169A1 · Habibian · 2023 [cited by examiner]
US 20230269395A1 · Yang · 2023 [cited by examiner]
US 20240223817A1 · Toderici · 2024 [cited by examiner]
US 20250056036A1 · Mohan · 2025 [cited by examiner]
JP 2019004388A · 2019 [cited by applicant]
WO 2021172749A1 · 2021 [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/CN2022/075485, mailed Aug. 30, 2022, filed Feb. 8, 2022, 9 pages. [cited by applicant]
Karpathy et al., “Large-scale Video Classification with Convolutional Neural Networks,” IEEE, Sep. 25, 2014, 8 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Xu et al., “Spatiotemporal CNN for Video Object Segmentation,” IEEE, Apr. 4, 2019, 10 pages. [cited by applicant]