IP Library Granted Patent US 12670600
Granted Patent B2
US 12670600 · App. 17/678,666 · Granted Jun 30, 2026

Disentanglement of image attributes using a neural network

Inventors: Aysegul Dundar (Ankara, TR); Kevin Jonathan Shih (Santa Clara, CA); Animesh Garg (Berkeley, CA); Robert Thomas Pottorff (Santa Clara, CA); Andrew Tao (Los Altos, CA); Bryan Christopher Catanzaro (Los Altos Hills, CA)
Assignee: NVIDIA Corporation
G06T7/194G06N3/084G06N5/046G06N20/00G06T7/70G06V20/20G06T2207/20081G06T2207/20084G06T2207/20212
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670600
App. No.
17/678,666
Granted
Jun 30, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to perform unsupervised keypoint or landmark learning using one or more neural networks. In at least one embodiment, one or more neural networks use pose and appearance information to construct a foreground and a background, which are then used to reconstruct an input image and determine loss values to train the one or more neural networks.

Claims (46)

1 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to train one or more neural networks to identify keypoints on objects by causing the one or more processors to at least:

generate, using Gaussian parameters projected to a heatmap that indicates whether a location in an input image is associated with an object, a reconstructed image that includes:

foreground data generated based, at least in part, on the input image; and

background data generated using a mask of a background of the input image; and

identify one or more keypoints of a foreground object in the input image based, at least in part, on differences between the reconstructed image and the input image, the one or more keypoints being spatial locations on the foreground object.

2 . The non-transitory machine-readable medium of claim 1 , wherein:

one or more annotations are determined based, at least in part, on the foreground data and the background data;

the background data is combined with the foreground data in order to compute training information; and

the training information and the one or more annotations are used to further train the one or more neural networks.

3 . The non-transitory machine-readable medium of claim 2 , wherein the one or more neural networks infer the foreground data based, at least in part, on one or more perturbations applied to the input image.

4 . The non-transitory machine-readable medium of claim 3 , wherein the one or more perturbations modify appearance characteristics of the input image.

5 . The non-transitory machine-readable medium of claim 3 , wherein the one or more perturbations modify pose characteristics of the input image.

6 . The non-transitory machine-readable medium of claim 2 , wherein the training information comprises values used to update the one or more neural networks.

7 . The non-transitory machine-readable medium of claim 2 , wherein the one or more annotations identify a location of the one or more keypoints in the input image.

8 . The non-transitory machine-readable medium of claim 1 , wherein the input image is a frame in video data.

9 . A method, comprising:

training one or more neural networks to identify one or more keypoints of one or more foreground objects, the one or more keypoints corresponding to spatial locations on the one or more foreground objects, by at least:

generating, using Gaussian parameters projected to a heatmap that indicates whether a location in an input image is associated with an object, one or more reconstructed images that include:

foreground data generated based, at least in part, on one or more input images; and

background data generated using one or more background masks of the one or more input images; and

identifying thione or more keypoints of the one or more foreground objects in the one or more input images based, at least in part, on differences between the one or more reconstructed images and the one or more input images.

10 . The method of claim 9 , wherein the method further comprises training the one or more neural networks to:

determine, by the one or more neural networks, the foreground data of the one or more input images;

generate the one or more reconstructed images by combining the background data with the foreground data to facilitate computing training information; and

training the one or more neural networks includes training the one or more neural networks based, at least in part, on the training information.

11 . The method of claim 10 , wherein identifying the one or more keypoints includes determining one or more annotations identifying the one or more keypoints based, at least in part, on the foreground data and the background data.

12 . The method of claim 10 , wherein the foreground data is determined based, at least in part, on one or more perturbations applied to the one or more input images.

13 . The method of claim 12 , wherein the one or more perturbations modify characteristics of one or more objects in the one or more input images.

14 . The method of claim 10 , wherein the training information is computed based on one or more reconstructed images, the one or more reconstructed images based, at least in part, on the foreground data and the background data.

15 . The method of claim 14 , wherein the training information comprises loss values indicating differences between the one or more reconstructed images and the one or more input images.

16 . The method of claim 9 , wherein the one or more input images are one or more frames frame in video data.

17 . One or more processors, comprising circuitry to train one or more neural networks to identify keypoints on objects by causing the one or more processors to at least:

generate, using Gaussian parameters projected to a heatmap that indicates whether a location in an input image is associated with an object, a reconstructed image that includes:

a foreground portion generated based, at least in part, on the input image; and

a background portion generated using a mask of a background of the input image; and

identify one or more keypoints of a foreground object in the input image based, at least in part, on differences between the reconstructed image and the input image, the one or more keypoints corresponding to locations on the foreground object.

18 . The one or more processors of claim 17 , wherein:

the one or more neural networks are trained to infer the foreground portion of the input image;

the foreground portion comprises the one or more keypoints;

the background portion is combined with the foreground portion in order to compute training information; and

the training information is used to further train the one or more neural networks.

19 . The one or more processors of claim 18 , wherein the foreground portion is inferred based, at least in part, on a perturbation applied to the input image.

20 . The one or more processors of claim 19 , wherein the one or more processors are further to use the circuitry to use the one or more neural networks to modify one or more of:

an appearance characteristic associated with the input image, or

a pose information characteristic associated with the input image.

21 . The one or more processors of claim 17 , wherein the heatmap is to enforce unimodal representation of the one or more keypoints.