IP Library Granted Patent US 12,387,481
Granted Patent B2
US 12,387,481 · App. 16/990,881 · Granted Aug 12, 2025

Enhanced object identification using one or more neural networks

Inventors: Jiwoong Choi (Santa Clara, CA); Jose Manuel Alvarez Lopez (Mountain View, CA)
Assignee: NVIDIA CORPORATION
G06V20/10G06F17/18G06F18/24G06N3/047G06N3/08G06T11/20G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,481
App. No.
16/990,881
Granted
Aug 12, 2025
Kind
B2
Abstract

Apparatuses, systems, and techniques to identify one or more objects in one or more images. In at least one embodiment, one or more objects are identified in one or more images based, at least in part, on a likelihood that one or more objects is different from other objects in one or more images.

Claims (86)

1. One or more processors, comprising: circuitry to compute one or more values indicating a likelihood that one or more objects in one or more images is different from other objects in the one or more images based at least on distributions of a set of properties corresponding to the one or more objects detected in the one or more images, and select a subset of the one or more images to include in a dataset based at least on the one or more values.

2. The one or more processors of claim 1 , wherein the circuitry is to: determine, for the set of properties, probability distributions of a plurality of bounding boxes corresponding to the one or more objects detected in the one or more images;

calculate one or more uncertainty values from the probability distributions; and aggregate the one or more uncertainty values to determine one or more informativeness values for the one or more images.

3. The one or more processors of claim 2 , wherein one or more neural networks are to further:

compare the one or more informativeness values with each other;

select the subset of the one or more images based at least in part on a comparison of the one or more informativeness values with each other; and

train a neural network using the subset of the one or more images.

4. The one or more processors of claim 2 , wherein the circuitry is to: determine a Gaussian mixture model (GMM) from the probability distributions and weights corresponding to the probability distributions; and

use the GMM to calculate an aleatoric uncertainty value and epistemic uncertainty value.

5. The one or more processors of claim 4 , wherein the circuitry is to train the GMM using a mixture density network to regress one or more localization parameters of the GMM using a negative log-likelihood loss function.

6. The one or more processors of claim 4 , wherein the are circuitry is to train the GMM using a mixture density network to regress one or more classification parameters of the GMM using a loss function based at least in part on the probability distributions and the weights.

7. The one or more processors of claim 2 , wherein the probability distributions correspond to localization parameters of the plurality of bounding boxes and classification parameters of objects in the plurality of bounding boxes.

8. The one or more processors of claim 1 , wherein one or more neural networks are used to identify one or more objects in one or more images based, at least in part, on one or more probability distributions for a set of properties of a plurality of bounding boxes detected in the one or more images.

9. A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to compute one or more values indicating a likelihood that one or more objects in one or more images is different from other objects in the one or more images based at least on distributions of a set of properties corresponding to the one or more objects detected in the one or more images, and select a subset of the one or more images to include in a dataset based at least on the one or more values.

10. The non-transitory machine-readable medium of claim 9 , wherein the one or more processors are to use one or more neural networks to:

determine, for the set of properties, a first set of probability distributions and a second set of probability distributions based at least in part on one or more feature maps generated from the one or more images;

determine a first set of uncertainty values based on the first set of probability distributions;

determine a second set of uncertainty values based on the second set of probability distributions; and

aggregate the first set of uncertainty values and the second set of uncertainty values to determine an informativeness score.

11. The non-transitory machine-readable medium of claim 10 , wherein the informativeness score is based at least in part on one or more maximum values of the first set of uncertainty values and the second set of uncertainty values.

12. The non-transitory machine-readable medium of claim 10 , wherein the first set of uncertainty values comprise aleatoric uncertainty values and epistemic uncertainty values.

13. The non-transitory machine-readable medium of claim 9 , wherein:

the first set of probability distributions corresponds to a location of one or more bounding boxes; and

the second set of probability distributions corresponds to one or more different object classifications.

14. The non-transitory machine-readable medium of claim 10 , wherein:

the one or more neural networks are trained using a set of loss functions based at least in part on the first set of probability distributions and the second set of probability distributions; and

the set of loss functions comprise one or more classification loss functions and one or more localization loss functions.

15. The non-transitory machine-readable medium of claim 9 , wherein the one or more processors are to use one or more neural networks to:

determine one or more informativeness scores for the one or more images;

obtain annotations for the subset of the one or more images; and

train a neural network using at least the subset of the one or more images and the annotations.

16. A system comprising:

one or more processors to compute one or more values indicating a likelihood that one or more objects in one or more images is different from other objects in the one or more images based at least on distributions of a set of properties corresponding to the one or more objects detected in the one or more images, and select a subset of the one or more images to include in a dataset based at least on the one or more values; and

one or more memories to store one or more neural networks to compute the one or more values.

17. The system of claim 16 , wherein the one or more processors are to use the one or more neural networks to:

determine, for the set of properties, one or more probability distributions based at least in part on the one or more images;

calculate one or more informativeness scores for the one or more images based at least in part on the one or more probability distributions;

determine the subset of the one or more images based at least in part on the one or more informativeness scores;

obtain annotations for the subset of the one or more images; and

train a neural network using at least the annotations and the subset of the one or more images.

18. The system of claim 17 , wherein:

the one or more neural networks are trained using one or more loss functions based at least in part on the one or more probability distributions; and

the one or more loss functions comprise localization loss functions and classification loss functions.

19. The system of claim 17 , wherein the one or more probability distributions comprise one or more localization probability distributions and one or more classification probability distributions.

20. The system of claim 17 , wherein the one or more probability distributions are defined by a plurality of parameter values that comprise:

a set of mean parameter values;

a set of variance parameter values; and a set of weight parameter values.

21. The system of claim 20 , wherein the one or more processors are to use the one or more neural networks to:

calculate one or more aleatoric uncertainty values at least from the set of variance parameter values and the set of weight parameter values; and

calculate one or more epistemic uncertainty values at least from the set of mean parameter values and the set of weight parameter values.

22. The system of claim 21 , wherein the one or more informativeness scores are calculated based at least in part on the one or more aleatoric uncertainty values and the one or more epistemic uncertainty values.

23. One or more processors, comprising: circuitry to train one or more neural networks to compute one or more values indicating a likelihood that one or more objects in one or more images is different from other objects in the one or more images based at least on distributions of a set of properties corresponding to the one or more objects detected in the one or more images, and select a subset of the one or more images to include in a dataset based at least on the one or more values.

24. The one or more processors of claim 23 , wherein the circuitry is to use the one or more neural networks to:

determine, for the set of properties, a plurality of probability distributions based at least in part on the one or more images, wherein the plurality of probability distributions indicate the likelihood that the one or more objects is different from the other objects in the one or more images;

calculate a plurality of uncertainty values based at least in part on the plurality of probability distributions;

determine a plurality of informativeness scores based at least in part on the plurality of uncertainty values;

compare informativeness scores of the plurality of informativeness scores with each other;

select the subset of the one or more images based at least in part on a comparison of the informativeness scores of the plurality of informativeness scores with each other;

obtain annotations for the subset of the one or more images, wherein the annotations indicate locations and classifications of objects in the subset of the one or more images; and

train a neural network using the subset of the one or more images and the annotations.

25. The one or more processors of claim 24 , wherein the one or more images are obtained from one or more viewpoints of a motor vehicle.

26. The one or more processors of claim 24 , wherein the plurality of probability distributions comprise a plurality of localization probability distributions and a plurality of classification probability distributions.

27. The one or more processors of claim 26 , wherein:

the plurality of localization probability distributions correspond to coordinates of one or more locations of the one or more objects; and

the plurality of classification probability distributions correspond to one or more classifications of the one or more objects.

28. The one or more processors of claim 26 , wherein the circuitry is to use the one or more neural networks to:

calculate a classification loss based at least in part on the plurality of classification probability distributions;

calculate a localization loss based at least in part on the plurality of localization probability distributions; and

update the one or more neural networks based at least in part on the classification loss and the localization loss.

29. The one or more processors of claim 24 , wherein the plurality of informativeness scores are based at least in part on one or more maximum values of the plurality of uncertainty values.

30. A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least: cause circuitry to compute one or more values indicating a likelihood that one or more objects in one or more images is different from other objects in the one or more images based at least on distributions of a set of properties corresponding to the one or more objects detected in the one or more images, and select a subset of the one or more images to include in a dataset based at least on the one or more values.

31. The non-transitory machine-readable medium of claim 30 , wherein the one or more processors are to use one or more neural networks to:

determine, for the set of properties, a first set of probability distributions and a second set of probability distributions based at least in part on one or more feature maps generated from the one or more images;

determine a first set of uncertainty values based on the first set of probability distributions;

determine a second set of uncertainty values based on the second set of probability distributions;

aggregate the first set of uncertainty values and the second set of uncertainty values to determine one or more informativeness scores for the one or more images;

obtain a set of annotations for the subset of the one or more images; and train a neural network using at least the set of annotations and the subset of the one or more images.

32. The non-transitory machine-readable medium of claim 31 , wherein the subset of the one or more images is determined based at least in part on the one or more informativeness scores.

33. The non-transitory machine-readable medium of claim 31 , wherein the first set of uncertainty values and the second set of uncertainty values comprise aleatoric uncertainty values and epistemic uncertainty values.

34. The non-transitory machine-readable medium of claim 31 , wherein each probability distribution of the first set of probability distributions and the second set of probability distributions is represented as a Gaussian Mixture Model (GMM).

35. The non-transitory machine-readable medium of claim 31 , wherein the one or more neural networks are to further:

determine one or more mean parameter values of the first set of probability distributions and the second set of probability distributions;

determine one or more variance parameter values of the first set of probability distributions and the second set of probability distributions; and

determine one or more weight parameter values of the first set of probability distributions and the second set of probability distributions.

36. The non-transitory machine-readable medium of claim 35 , wherein the first set of uncertainty values are determined through a first set of summations based at least in part on one or more products of the one or more variance parameter values and the one or more weight parameter values, and the second set of uncertainty values are determined through a second set of summations based at least in part on the one or more mean parameter values and the one or more weight parameter values.

37. The non-transitory machine-readable medium of claim 36 , wherein the first set of uncertainty values correspond to aleatoric uncertainty and the second set of uncertainty values are correspond to epistemic uncertainty.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2020
From: CHOI, JIWOONG; ALVAREZ LOPEZ, JOSE MANUEL
To: NVIDIA CORPORATION
Reel/Frame 053528/0975 →
Continuity (1)
Related Publication 20220051017A1 · Feb 17, 2022
References Cited (50)
US 6018728A · Spence · 2000 [cited by examiner]
US 6718063B1 · Lennon · 2004 [cited by examiner]
US 7957596B2 · Ofek · 2011 [cited by examiner]
US 8775341B1 · Commons · 2014 [cited by examiner]
US 8990133B1 · Ponulak · 2015 [cited by examiner]
US 9015092B2 · Sinyavskiy · 2015 [cited by examiner]
US 9111226B2 · Richert · 2015 [cited by examiner]
US 9218563B2 · Szatmary · 2015 [cited by examiner]
US 9224090B2 · Piekniewski · 2015 [cited by examiner]
US 9346167B2 · O'Connor · 2016 [cited by examiner]
US 9436909B2 · Piekniewski · 2016 [cited by examiner]
US 11200679B1 · Li · 2021 [cited by examiner]
US 11430225B2 · Evans · 2022 [cited by examiner]
US 20040042663A1 · Yamada · 2004 [cited by examiner]
US 20050152617A1 · Roche · 2005 [cited by examiner]
US 20070076922A1 · Living · 2007 [cited by examiner]
US 20200193607A1 · Sun · 2020 [cited by examiner]
US 20220051017A1 · Choi · 2022 [cited by examiner]
CN 108038853A · 2018 [cited by applicant]
CN 110197286A · 2019 [cited by applicant]
Choi et al., “Uncertainty-Aware Learning from Demonstration Using Mixture Density Networks with Sampling-Free Variance Modeling,” IEEE International Conference on Robotics and Automation, May 21, 2018, 8 pages. [cited by applicant]
He et al., “Deep Mixture Density Network for Probabilistic Object Detection,” Jul. 2, 2020, 8 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2021/045279, mailed Nov. 29, 2021, filed Aug. 9, 2021, 15 pages. [cited by applicant]
Aghdam et al., “Active Learning for Deep Detection Neural Networks,” ICCV, Nov. 20, 2019, 13 pages. [cited by applicant]
Arthur et al., “K-Means++: The Advantages of Careful Seeding,” ACM-SIAM SODA, 2007, 11 pages. [cited by applicant]
Ash et al., “Deep Batch Active Learning by Diverse, Uncertain Gradient Lower Bounds,” Jun. 9, 2019, 21 pages. [cited by applicant]
Atanov et al., “Uncertainty Estimation via Stochastic Batch Normalization,” Mar. 20, 2018, 6 pages. [cited by applicant]
Beluch et al., “The Power of Ensembles for Active Learning in Image Classification,” CVPR, 2018, 10 pages. [cited by applicant]
Brust et al., “Active Learning for Deep Object Detection, ” VISAPP, 2019, 10 pages. [cited by applicant]
Chitta et al., “Large-Scale Visual Active Learning with Deep Probabilistic Ensembles,” Nov. 30, 2018, 10 pages. [cited by applicant]
Chitta et al., “Less is More: An Exploration of Data Redundancy with Active Dataset Subsampling,” May 29, 2019, 10 pages. [cited by applicant]
Desai et al., “An Adaptive Supervision Framework for Active Learning in Object Detection,” BMVC, 2019, 13 pages. [cited by applicant]
Gal “Uncertainty in Deep Learning,” Ph.D. Dissertation, University of Cambridge, 2016, 174 pages. [cited by applicant]
Gal et al., “Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning,” Jun. 6, 2015, 10 pages. [cited by applicant]
Geifman et al., “Bias-Reduced Uncertainty Estimation for Deep Neural Classifiers,” ICLR, 2019, 14 pages. [cited by applicant]
Haussmann et al., “Scalable Active Learning for Object Detection,” NVIDIA, Apr. 9, 2020, 6 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Kao et al., “Localization-Aware Active Learning for Object Detection,” A0CCV, 2018, 35 pages. [cited by applicant]
Lakshminarayanan et al., “Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles,” Advances in Neural Information Processing Systems, 2017, 12 pages. [cited by applicant]
Liu et al., “SSD: Single Shot Multibox Detector,” European Conference on Computer Vision, Springer, Dec. 29, 2016, 17 pages. [cited by applicant]
MacQueen, “Some Methods for Classification and Analysis of Multivariate Observations,” Proceedings of 5-th Berkeley Symposium on Mathematical Statistics and Probability, 1967, 17 pages. [cited by applicant]
Pati et al., “Orthogonal Matching Pursuit: Recursive Function Approximation with Applications to Wavelet Decomposition,” 27th Asilomar Conference Signals Systems and Computers, Nov. 1-3, 1993, 5 pages. [cited by applicant]
Redmon et al., “You Only Look Once: Unified, Realtime Object Detection,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, 10 pages. [cited by applicant]
Roy et al., “Deep Active Learning for Object Detection,” BMVC, 2018, 12 pages. [cited by applicant]
Sener et al., “Active Learning for Convolutional Neural Networks: A Core-Set Approach,” ICLR, 2018, 13 pages. [cited by applicant]
Settles, “Active Learning Literature Survey,” Technical Report, 2010, 67 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Wang et al., “Uncertainty-Based Active Learning via Sparse Modeling for Image Classification,” IEEE Transactions on Image Processing, 28(1): Jan. 2019, 14 pages. [cited by applicant]
Office Action for Chinese Application No. 202180010958.7, mailed Apr. 8, 2025, 21 pages. [cited by applicant]