IP Library › Granted Patent US 12,670,600
Granted Patent B2
US 12,670,600 · App. 17/678,666 · Granted Jun 30, 2026

Disentanglement of image attributes using a neural network

Inventors: Aysegul Dundar (Ankara, TR); Kevin Jonathan Shih (Santa Clara, CA); Animesh Garg (Berkeley, CA); Robert Thomas Pottorff (Santa Clara, CA); Andrew Tao (Los Altos, CA); Bryan Christopher Catanzaro (Los Altos Hills, CA)
Assignee: NVIDIA Corporation
G06T7/194G06N3/084G06N5/046G06N20/00G06T7/70G06V20/20G06T2207/20081G06T2207/20084G06T2207/20212
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,600
App. No.
17/678,666
Filed
Feb 23, 2022
Granted
Jun 30, 2026
Kind
B2
Examiner
BITAR, NANCY
Art Unit
2664
USPC
382/155
Abstract

Apparatuses, systems, and techniques to perform unsupervised keypoint or landmark learning using one or more neural networks. In at least one embodiment, one or more neural networks use pose and appearance information to construct a foreground and a background, which are then used to reconstruct an input image and determine loss values to train the one or more neural networks.

Claims (46)

1 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to train one or more neural networks to identify keypoints on objects by causing the one or more processors to at least:

generate, using Gaussian parameters projected to a heatmap that indicates whether a location in an input image is associated with an object, a reconstructed image that includes:

foreground data generated based, at least in part, on the input image; and

background data generated using a mask of a background of the input image; and

identify one or more keypoints of a foreground object in the input image based, at least in part, on differences between the reconstructed image and the input image, the one or more keypoints being spatial locations on the foreground object.

2 . The non-transitory machine-readable medium of claim 1 , wherein:

one or more annotations are determined based, at least in part, on the foreground data and the background data;

the background data is combined with the foreground data in order to compute training information; and

the training information and the one or more annotations are used to further train the one or more neural networks.

3 . The non-transitory machine-readable medium of claim 2 , wherein the one or more neural networks infer the foreground data based, at least in part, on one or more perturbations applied to the input image.

4 . The non-transitory machine-readable medium of claim 3 , wherein the one or more perturbations modify appearance characteristics of the input image.

5 . The non-transitory machine-readable medium of claim 3 , wherein the one or more perturbations modify pose characteristics of the input image.

6 . The non-transitory machine-readable medium of claim 2 , wherein the training information comprises values used to update the one or more neural networks.

7 . The non-transitory machine-readable medium of claim 2 , wherein the one or more annotations identify a location of the one or more keypoints in the input image.

8 . The non-transitory machine-readable medium of claim 1 , wherein the input image is a frame in video data.

9 . A method, comprising:

training one or more neural networks to identify one or more keypoints of one or more foreground objects, the one or more keypoints corresponding to spatial locations on the one or more foreground objects, by at least:

generating, using Gaussian parameters projected to a heatmap that indicates whether a location in an input image is associated with an object, one or more reconstructed images that include:

foreground data generated based, at least in part, on one or more input images; and

background data generated using one or more background masks of the one or more input images; and

identifying thione or more keypoints of the one or more foreground objects in the one or more input images based, at least in part, on differences between the one or more reconstructed images and the one or more input images.

10 . The method of claim 9 , wherein the method further comprises training the one or more neural networks to:

determine, by the one or more neural networks, the foreground data of the one or more input images;

generate the one or more reconstructed images by combining the background data with the foreground data to facilitate computing training information; and

training the one or more neural networks includes training the one or more neural networks based, at least in part, on the training information.

11 . The method of claim 10 , wherein identifying the one or more keypoints includes determining one or more annotations identifying the one or more keypoints based, at least in part, on the foreground data and the background data.

12 . The method of claim 10 , wherein the foreground data is determined based, at least in part, on one or more perturbations applied to the one or more input images.

13 . The method of claim 12 , wherein the one or more perturbations modify characteristics of one or more objects in the one or more input images.

14 . The method of claim 10 , wherein the training information is computed based on one or more reconstructed images, the one or more reconstructed images based, at least in part, on the foreground data and the background data.

15 . The method of claim 14 , wherein the training information comprises loss values indicating differences between the one or more reconstructed images and the one or more input images.

16 . The method of claim 9 , wherein the one or more input images are one or more frames frame in video data.

17 . One or more processors, comprising circuitry to train one or more neural networks to identify keypoints on objects by causing the one or more processors to at least:

generate, using Gaussian parameters projected to a heatmap that indicates whether a location in an input image is associated with an object, a reconstructed image that includes:

a foreground portion generated based, at least in part, on the input image; and

a background portion generated using a mask of a background of the input image; and

identify one or more keypoints of a foreground object in the input image based, at least in part, on differences between the reconstructed image and the input image, the one or more keypoints corresponding to locations on the foreground object.

18 . The one or more processors of claim 17 , wherein:

the one or more neural networks are trained to infer the foreground portion of the input image;

the foreground portion comprises the one or more keypoints;

the background portion is combined with the foreground portion in order to compute training information; and

the training information is used to further train the one or more neural networks.

19 . The one or more processors of claim 18 , wherein the foreground portion is inferred based, at least in part, on a perturbation applied to the input image.

20 . The one or more processors of claim 19 , wherein the one or more processors are further to use the circuitry to use the one or more neural networks to modify one or more of:

an appearance characteristic associated with the input image, or

a pose information characteristic associated with the input image.

21 . The one or more processors of claim 17 , wherein the heatmap is to enforce unimodal representation of the one or more keypoints.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2022
From: DUNDAR, AYSEGUL; SHIH, KEVIN JONATHAN; GARG, ANIMESH; POTTORFF, ROBERT THOMAS; TAO, ANDREW; CATANZARO, BRYAN CHRISTOPHER
To: NVIDIA CORPORATION
Reel/Frame 059080/0172 →
Continuity (2)
Continuation 16786057 · Feb 10, 2020
Related Publication 20220180528A1 · Jun 9, 2022
References Cited (43)
US 7940960B2 · Okada · 2011 [cited by examiner]
US 8311954B2 · Ning · 2012 [cited by examiner]
US 9280827B2 · Tuzel · 2016 [cited by examiner]
US 10074054B2 · Adams · 2018 [cited by examiner]
US 20190251401A1 · Shechtman · 2019 [cited by examiner]
US 20190361994A1 · Shen et al. · 2019 [cited by applicant]
US 20190377949A1 · Chen · 2019 [cited by examiner]
US 20200085702A1 · Lerebour · 2020 [cited by examiner]
US 20200250812A1 · Ceccaldi · 2020 [cited by examiner]
US 20200357142A1 · Aydin · 2020 [cited by examiner]
US 20210142479A1 · Phogat · 2021 [cited by examiner]
US 20210201077A1 · Lwowski · 2021 [cited by examiner]
US 20210383242A1 · Ostyakov · 2021 [cited by examiner]
Balakrishnan et al., “Synthesizing Images of Humans in Unseen Poses,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 9 pages. [cited by applicant]
Charles et al., “Domain Adaptation for Upper Body Pose Tracking in Signed TV Broadcasts,” British Machine Vision Conference, 2013, 11 pages. [cited by applicant]
Denton et al., “Stochastic Video Generation with a Learned Prior,” In Proceedings of the 35th International Conference on Machine Learning, 2018, 10 pages. [cited by applicant]
Denton et al., “Unsupervised Learning of Disentangled Representations from Video,” Advances in Neural Information Processing Systems, 2017, 10 pages. [cited by applicant]
Dundar et al.,“Unsupervised Disentanglement of Pose, Appearance and Background from Images and Videos,” Jan. 26, 2020, retrieved Dececember 21, 2020 from https://arxiv.org/pdf/2001.09518.pdf, 20 pages. [cited by applicant]
Finn et al., “Unsupervised Learning for Physical Interaction Through Video Prediction,” Advances in Neural Information Processing Systems, 2016, 9 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
International Electrotechnical Commission, “Functional safety of electrical/electronic/programmable electronic safety-related systems,” IEC Standard 61508-1, Apr. 2014, 23 pages. [cited by applicant]
International Organization for Standardization, “Road vehicles—Functional safety,” ISO Standard 26262, https://www.iso.org/obp/ui/#iso:std:iso:26262:-1:ed-1:v1:en, Nov. 11, 2011, 35 pages. [cited by applicant]
Ionescu et al., “Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7): 2013, 15 pages. [cited by applicant]
Jakab et al., “Unsupervised Learning of Object Landmarks through Conditional Image Generation,” Neural Information Processing Systems, 2018, 12 pages. [cited by applicant]
Kanazawa et al., “WarpNet: Weakly Supervised Matching for Single-View Reconstruction,” IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, 9 pages. [cited by applicant]
Lee et al., “Stochastic Adversarial Video Prediction,” Apr. 4, 2018, 26 pages. [cited by applicant]
Liu et al., “Deep Learning Face Attributes in the Wild,” ICCV, 2015, 9 pages. [cited by applicant]
Lorenz et al., “Unsupervised Part-Based Disentangling of Object Shape and Appearance,” CVPR, 2019, 10 pages. [cited by applicant]
Miyato et al., “Spectral Normalization for Generative Adversarial Networks,” ICLR, 2018, 26 pages. [cited by applicant]
Newell et al., “Stacked Hourglass Networks for Human Pose Estimation,” In European Conference on Computer Vision, Jul. 26, 2016, 17 pages. [cited by applicant]
Park et al., “Semantic Image Synthesis with Spatially Adaptive Normalization,” CVPR, 2019, 10 pages. [cited by applicant]
Pfister et al., “Flowing ConvNets for Human Pose Estimation in Videos,” Proceedings of the IEEE International Conference on Computer Vision, 2015, 9 pages. [cited by applicant]
Rhodin et al., “Neural Scene Decomposition for Multi-Person Motion Capture,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 11 pages. [cited by applicant]
Rhodin et al., “Unsupervised Geometry—Aware Representation for 3D Human Pose Estimation,” Proceedings of the European Conference on Computer Vision, 2018, 18 pages. [cited by applicant]
Ronneberger et al., “U-net: Convolutional networks for biomedical image segmentation,” International Conference on Medical Image Computing and Computer-Assisted Intervention, Oct. 5, 2015, 8 pages. [cited by applicant]
Schuldt et al., “Recognizing Human Actions: A Local SVM Approach,” Proceedings of the 17th International Conference on Pattern Recognition, vol. 3, 2004, 5 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles, ” Standard No. J3016-201806, issued Ja… [cited by applicant]
Suwajanakorn et al., “Discovery of Latent 3D Keypoints via End-to-End Geometric Reasoning,” Advances in Neural Information Processing Systems, Jul. 2018, 13 pages. [cited by applicant]
Thewlis et al., “Unsupervised Learning of Object Frames by Dense Equivariant Image Labelling,” Advances in Neural Information Processing Systems, 2017, 12 pages. [cited by applicant]
Thewlis et al., “Unsupervised Learning of Object Landmarks by Factorized Spatial Embeddings,” ICCV, 2017, 10 pages. [cited by applicant]
Zhang et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,” CVPR, 2018, 10 pages. [cited by applicant]
Zhang et al., “Unsupervised Discovery of Object Landmarks as Structural Representations,”CVPR, 2018, 10 pages. [cited by applicant]