IP Library Granted Patent US 12,548,310
Granted Patent B1
US 12,548,310 · App. 17/672,228 · Granted Feb 10, 2026

Neural network-based object detection

Inventors: Xinlong Wang (Adelaide, AU); Zhiding Yu (Santa Clara, CA); Shalini De Mello (San Francisco, CA); Anima Anandkumar (Pasadena, CA); Jose Manuel Alvarez Lopez (Mountain View, CA)
Assignee: NVIDIA Corporation
G06V10/82G06N3/045G06N3/088G06N3/0895G06T7/10G06T7/11G06V10/255G06V10/26G06V10/70G06V10/774G06V10/7747G06V10/809G06T2207/20081G06T2207/20084G06T2207/30252
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,310
App. No.
17/672,228
Granted
Feb 10, 2026
Kind
B1
Abstract

Apparatuses, systems, and techniques are presented to detect one or more objects in one or more images. In at least one embodiment, one or more neural networks can be trained to detect one or more objects, in one or more unlabeled images, based at least in part upon one or more predicted segmentations of the one or more objects.

Claims (59)

1 . One or more processors, comprising circuitry to:

cause one or more first neural networks to generate one or more first segmentations of one or more objects in one or more unlabeled images;

cause one or more second neural networks to generate one or more second segmentations of the one or more objects; and

update the one or more second neural networks to detect the one or more objects in the one or more unlabeled images based, at least in part, on one or more projections of the one or more first segmentations and the one or more second segmentations of the one or more objects.

2 . The one or more processors of claim 1 , wherein the circuitry is further to provide the one or more unlabeled images as input to a free mask generator for generating the one or more first segmentations.

3 . The one or more processors of claim 2 , wherein the circuitry is further to cause the free mask generator to generate a plurality of segmentation proposals for the one or more unlabeled images, remove redundant segmentation proposals, and select at least a subset of the plurality of segmentation proposals having highest determined confidence values as the one or more first segmentations.

4 . The one or more processors of claim 2 , wherein the circuitry is further to provide semantic segmentation information for the one or more objects, determined by the free mask generator, as input to the one or more second neural networks for use in training the one or more second neural networks.

5 . The one or more processors of claim 1 , wherein the circuitry is further to cause the one or more second neural networks to utilize the one or more first segmentations as weak labels for the one or more objects to be used in determining a loss value with respect to segmentations inferred for the one or more objects.

6 . The one or more processors of claim 1 , wherein the circuitry is further to perform continued self-training of the one or more second neural networks by providing the one or more first segmentations and semantic embeddings corresponding to the one or more first segmentations for the one or more objects, detected by the one or more first neural networks, as additional training input for the one or more second neural networks.

7 . The one or more processors of claim 1 , wherein the circuitry is further to perform continued self-training of the one or more second neural networks by providing one or more first segmentations for the one or more objects, detected by the one or more first neural networks, as additional training input for the one or more second neural networks.

8 . One or more processors, comprising circuitry to: train one or more second neural networks to detect one or more objects, in one or more unlabeled images based, at least in part, on one or more projections of one or more first segmentations of the one or more objects generated by one or more first neural networks and one or more second segmentations of the one or more objects generated by the one or more second neural networks.

9 . The one or more processors of claim 8 , wherein the circuitry is further to train the one or more first neural networks, at least in part, by providing the one or more unlabeled images as input to a free mask generator for generating the one or more first segmentations.

10 . The one or more processors of claim 9 , wherein the circuitry is further to cause the free mask generator to generate a plurality of segmentation proposals for the one or more unlabeled images, remove redundant segmentation proposals, and select at least a subset of the plurality of segmentation proposals having highest determined confidence values as the one or more first segmentations.

11 . The one or more processors of claim 9 , wherein the circuitry is further to provide semantic segmentation information for the one or more objects, determined by the free mask generator, as input to the one or more second neural networks for use in training the one or more second neural networks.

12 . The one or more processors of claim 9 , wherein the circuitry is further to cause the one or more second neural networks to utilize the one or more first segmentations as weak labels for the one or more objects to be used in determining a loss value with respect to segmentations inferred for the one or more objects.

13 . A system comprising: one or more processors to:

cause one or more first neural networks to generate one or more first segmentations of one or more objects in one or more unlabeled images;

cause one or more second neural networks to generate one or more second segmentations of the one or more objects; and

update the one or more second neural networks to detect the one or more objects in the one or more unlabeled images based, at least in part, on one or more projections of the one or more first segmentations and the one or more second segmentations of the one or more objects.

14 . The system of claim 13 , wherein the one or more processors are further to provide the one or more unlabeled images as input to a free mask generator for generating the one or more first segmentations.

15 . The system of claim 14 , wherein the one or more processors are further to cause the free mask generator to generate a plurality of segmentation proposals for the one or more unlabeled images, remove redundant segmentation proposals, and select at least a subset of the plurality of segmentation proposals having highest determined confidence values as the one or more first segmentations.

16 . The system of claim 14 , wherein the one or more processors are further to provide semantic segmentation information for the one or more objects, determined by the free mask generator, as input to the one or more second neural networks for use in training the one or more second neural networks.

17 . The system of claim 13 , wherein the one or more processors are further to cause the one or more second neural networks to utilize the one or more first segmentations as weak labels for the one or more objects to be used in determining a loss value with respect to segmentations inferred for the one or more objects.

18 . The system of claim 13 , wherein the one or more processors are further to perform continued self-training of the one or more second neural networks by providing one or more first segmentations for the one or more objects, detected by the one or more first neural networks, as additional training input for the one or more second neural networks.

19 . A method comprising:

causing one or more first neural networks to generate one or more first segmentations of one or more objects in one or more unlabeled images;

causing one or more second neural networks to generate one or more second segmentations of the one or more objects; and

updating the one or more second neural networks to detect the one or more objects in the one or more unlabeled images based, at least in part, on one or more projections of the one or more first segmentations and the one or more second segmentations of the one or more objects.

20 . The method of claim 19 , further comprising:

providing the one or more unlabeled images as input to a free mask generator for generating the one or more first segmentations.

21 . The method of claim 20 , further comprising:

causing the free mask generator to generate a plurality of segmentation proposals for the one or more unlabeled images, remove redundant segmentation proposals, and select at least a subset of the plurality of segmentation proposals having highest determined confidence values as the one or more first segmentations.

22 . The method of claim 20 , further comprising:

providing semantic segmentation information for the one or more objects, determined by the free mask generator, as input to the one or more second neural networks for use in training the one or more second neural networks.

23 . The method of claim 19 , further comprising:

causing the one or more second neural networks to utilize the one or more first segmentations as weak labels for the one or more objects to be used in determining a loss value with respect to segmentations inferred for the one or more objects.

24 . The method of claim 19 , further comprising:

performing continued self-training of the one or more second neural networks by providing one or more first segmentations for the one or more objects, detected by the one or more first neural networks, as additional training input for the one or more second neural networks.

25 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

cause one or more first neural networks to generate one or more first segmentations of one or more objects in one or more unlabeled images;

cause one or more second neural networks to generate one or more second segmentations of the one or more objects; and

update the one or more second neural networks to detect the one or more objects in the one or more unlabeled images based, at least in part, on one or more projections of the one or more first segmentations and the one or more second segmentations of the one or more objects.

26 . The non-transitory machine-readable medium of claim 25 , wherein instructions if performed further cause the one or more processors to:

provide the one or more unlabeled images as input to a free mask generator for generating the one or more first segmentations.

27 . The non-transitory machine-readable medium of claim 26 , wherein the one or more processors are further to cause the free mask generator to generate a plurality of segmentation proposals for the one or more unlabeled images, remove redundant segmentation proposals, and select at least a subset of the plurality of segmentation proposals having highest determined confidence values as the one or more first segmentations.

28 . The non-transitory machine-readable medium of claim 26 , wherein the one or more processors are further to provide semantic segmentation information for the one or more objects, determined by the free mask generator, as input to the one or more second neural networks for use in training the one or more second neural networks.

29 . The non-transitory machine-readable medium of claim 25 , wherein the one or more processors are further to cause the one or more second neural networks to utilize the one or more first segmentations as weak labels for the one or more objects to be used in determining a loss value with respect to segmentations inferred for the one or more objects.

30 . The non-transitory machine-readable medium of claim 25 , wherein the one or more processors are further to perform continued self-training of the one or more second neural networks by providing one or more first segmentations for the one or more objects, detected by the one or more first neural networks, as additional training input for the one or more second neural networks.

31 . An object detection system, comprising:

one or more processors to:

cause one or more first neural networks to generate one or more first segmentations of one or more objects in one or more unlabeled images;

cause one or more second neural networks to generate one or more second segmentations of the one or more objects; and

update the one or more second neural networks to detect the one or more objects in the one or more unlabeled images based, at least in part, on one or more projections of the one or more first segmentations and the one or more second segmentations of the one or more objects; and

memory for storing network parameters for the one or more first neural networks and the one or more second neural networks.

32 . The object detection system of claim 31 , wherein the one or more processors are further to provide the one or more unlabeled images as input to a free mask generator for generating the one or more first segmentations.

33 . The object detection system of claim 32 , wherein the one or more processors are further to cause the free mask generator to generate a plurality of segmentation proposals for the one or more unlabeled images, remove redundant segmentation proposals, and select at least a subset of the plurality of segmentation proposals having highest determined confidence values as the one or more first segmentations.

34 . The object detection system of claim 32 , wherein the one or more processors are further to provide semantic segmentation information for the one or more objects, determined by the free mask generator, as input to the one or more second neural networks for use in training the one or more second neural networks.

35 . The object detection system of claim 31 , wherein the one or more processors are further to cause the one or more second neural networks to utilize the one or more first segmentations as weak labels for the one or more objects to be used in determining a loss value with respect to segmentations inferred for the one or more objects.

36 . The object detection system of claim 31 , wherein the one or more processors are further to perform continued self-training of the one or more second neural networks by providing one or more first segmentations for the one or more objects, detected by the one or more first neural networks, as additional training input for the one or more second neural networks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2022
From: WANG, XINLONG; YU, ZHIDING; DE MELLO, SHALINI; ANANDKUMAR, ANIMA; ALVAREZ LOPEZ, JOSE MANUEL
To: NVIDIA CORPORATION
Reel/Frame 059230/0862 →
References Cited (84)
US 12394064B2 · Evans · 2025 [cited by examiner]
US 20200175375A1 · Chen · 2020 [cited by examiner]
US 20200327674A1 · Yang · 2020 [cited by examiner]
US 20200394458A1 · Yu · 2020 [cited by examiner]
US 20210027085A1 · Wu · 2021 [cited by examiner]
US 20210027098A1 · Ge · 2021 [cited by examiner]
US 20210241034A1 · Laradji · 2021 [cited by examiner]
US 20230104262A1 · Lin · 2023 [cited by examiner]
Pathak et al., “Context Encoders: Feature Learning by Inpainting,” IEEE Conference on Computer Vision and Pattern Recognition, Nov. 21, 2016, 12 pages. [cited by applicant]
Pinheiro et al., “Unsupervised Learning of Dense Visual Representations,” Dec. 7, 2020, 14 pages. [cited by applicant]
Russell et al., “Using Multiple Segmentations to Discover Objects and their Extent in Image Collections,” IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Oct. 9, 2006, 8 Pages. [cited by applicant]
Simeoni et al., “Localizing Objects with Self-Supervised Transformers and no Labels,” Sep. 29, 2021, 25 Pages. [cited by applicant]
Sivic et al., “Discovering Object Categories in Image Collections,” Feb. 25, 2005, 14 Pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Standard No. J3016-201806, dated Jun. … [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Tian et al., “BoxInst: High-Performance Instance Segmentation with Box Annotations,” CVPR, 2021, 10 pages. [cited by applicant]
Tian et al., “Conditional Convolutions for Instance Segmentation,” Jul. 26, 2020, 18 Pages. [cited by applicant]
Uijlings et al., “Selective Search for Object Recognition,” IJCV, 2013, 14 pages. [cited by applicant]
Van Gansbeke et al., “Unsupervised Semantic Segmentation by Contrasting Object Mask Proposals,” Aug. 3, 2021, 14 Pages. [cited by applicant]
Vo et al., “Large-Scale Unsupervised Object Discovery,” Nov. 16, 2021, 18 Pages. [cited by applicant]
Vo et al., “Toward Unsupervised, Multi-Object Discovery in Large-Scale Image Collections,” Jul. 6, 2020, 25 Pages. [cited by applicant]
Vo et al., “Unsupervised Image Matching and Object Discovery as Optimization,” Apr. 5, 2019, 10 Pages. [cited by applicant]
Wang et al., “Dense Contrastive Learning for Self-Supervised Visual Pre-Training,” IEEE, Apr. 4, 2021, 10 pages. [cited by applicant]
Wang et al., “SOLO: A Simple Framework for Instance Segmentation,” Jun. 30, 2021, 20 Pages. [cited by applicant]
Wang et al., “SOLO: Segmenting Objects by Locations,” Dec. 15, 2019, 10 pages. [cited by applicant]
Wang et al., “SOLOv2: Dynamic and Fast Instance Segmentation,” Oct. 23, 2020, 17 Pages. [cited by applicant]
Wang et al., “Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation,” Apr. 10, 2021, 11 Pages. [cited by applicant]
Wei et al., “Unsupervised Object Discovery and Co-Localization by Deep Descriptor Transforming,” Jul. 20, 2017, 16 Pages. [cited by applicant]
Wu et al., “Unsupervised Feature Learning via Non-Parametric Instance Discrimination,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Xiao et al., “Region Similarity Representation Learning,” Aug. 18, 2021, 12 pages. [cited by applicant]
Xie et al., “DetCo: Unsupervised Contrastive Learning for Object Detection,” Jul. 23, 2021, 10 Pages. [cited by applicant]
Xie et al., “Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation Learning,” CVPR, 2021, 10 Pages. [cited by applicant]
Zhang et al., “Colorful Image Colorization,” ECCV, Oct. 5, 2016, 29 pages. [cited by applicant]
Zhou et al., “Weakly Supervised Instance Segmentation using Class Peak Response,” CVPR, 2018, 10 pages. [cited by applicant]
Arbelaez et al., “Multiscale Combinatorial Grouping,” CVPR, 2014, 8 Pages. [cited by applicant]
Bar et al., “DETReg: Unsupervised Pretraining with Region Priors for Object Detection,” Dec. 1, 2021, 13 Pages. [cited by applicant]
Bolya et al., “YOLACT: Real-time Instance Segmentation,” ICCV, 2019, 10 pages. [cited by applicant]
Brabandere et al., “Semantic Instance Segmentation with a Discriminative Loss Function,” Aug. 8, 2017, 10 Pages. [cited by applicant]
Caron et al., “Emerging Properties in Self-Supervised Vision Transformers,” May 24, 2021, 21 pages. [cited by applicant]
Caron et al., “Unsupervised Learning of Visual Features by Contrasting Cluster Assignments,” Oct. 15, 2020, 23 Pages. [cited by applicant]
Chaitanya et al., “Contrastive learning of global and local features for medical image segmentation with limited annotations,” Oct. 30, 2020, 18 Pages. [cited by applicant]
Chen et al., “A Simple Framework for Contrastive Learning of Visual Representations,” International Conference on Machine Learning, Proceedings of Machine Learning Research, 2020, 11 pages. [cited by applicant]
Chen et al., “BlendMask: Top-Down Meets Bottom-Up for Instance Segmentation,” CVPR, 2020, 9 pages. [cited by applicant]
Chen et al., “Exploring Simple Siamese Representation Learning,” CVPR, 2021, 9 Pages. [cited by applicant]
Chen et al., “Hybrid Task Cascade for Instance Segmentation,” Apr. 9, 2019, 10 Pages. [cited by applicant]
Chen et al., “Improved Baselines with Momentum Contrastive Learning,” Mar. 9, 2020, 3 pages. [cited by applicant]
Chen et al., “Show, Match and Segment: Joint Weakly Supervised Learning of Semantic Matching and Object Co-segmentation,” Mar. 29, 2020, 16 Pages. [cited by applicant]
Cheng et al., “Pointly-Supervised Instance Segmentation,” Apr. 13, 2021, 13 Pages. [cited by applicant]
Cho et al., “Unsupervised Object Discovery and Localization in the Wild: Part-based Matching with Bottom-up Region Proposals,” May 4, 2015, 10 Pages. [cited by applicant]
Dai et al., “UP-DETR: Unsupervised Pre-training for Object Detection with Transformers,” Apr. 7, 2021, Apr. 7, 2021, 11 Pages. [cited by applicant]
Everingham et al., “The Pascal Visual Object Classes (VOC) Challenge,” IJCV, 2010, 36 pages. [cited by applicant]
Faktor et al., ““Clustering by Composition”—Unsupervised Discovery of Image Categories,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Jun. 2014, 14 Pages. [cited by applicant]
Gao et al, “SSAP: Single-Shot Instance SegmentationWith Affinity Pyramid,” CVF, 2019, 10 Pages. [cited by applicant]
Ghiasi et al., “Simple Copy-Paste is a Strong Data Augmentation Method for Instance Segmentation,” Jun. 23, 2021, 13 Pages. [cited by applicant]
Gidaris et al., “Unsupervised Representation Learning by Predicting Image Rotations,” Mar. 21, 2018, 16 pages. [cited by applicant]
Grill et al., “Bootstrap Your Own Latent A New Approach to Self-Supervised Learning,” Sep. 10, 2020, 35 pages. [cited by applicant]
Gupta et al., “LVIS: A Dataset for Large Vocabulary Instance Segmentation,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 9 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” CVPR, 2016, 9 pages. [cited by applicant]
He et al., “Mask R-CNN,” ICCV, 2017, 9 pages. [cited by applicant]
He et al., “Momentum Contrast for Unsupervised Visual Representation Learning,” CVPR, 2020, 10 pages. [cited by applicant]
Henaff et al., “Efficient Visual Pretraining with Contrastive Detection,” Mar. 19, 2021, 12 Pages. [cited by applicant]
Hsu et al., “Co-attention CNNs for Unsupervised Object Co-segmentation,” Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, 2018, 9 Pages. [cited by applicant]
Hsu et al., “Weakly Supervised Instance Segmentation using the Bounding Box Tightness Prior,” Neural Information Processing Systems, 2019, 12 pages. [cited by applicant]
Hwang et al., “SegSort: Segmentation by Discriminative Sorting of Segments,” Oct. 30, 2019, 17 Pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
Ji et al., “Invariant Information Clustering for Unsupervised Image Classification and Segmentation,” Mar. 26, 2019, 10 Pages. [cited by applicant]
Joulin et al., “Discriminative Clustering for Image Co-Segmentation,” IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Aug. 5, 2010, 8 Pages. [cited by applicant]
Khoreva et al., “Simple Does It: Weakly Supervised Instance and Semantic Segmentation,” IEEE Conference on Computer Vision and Pattern Recognition, Jul. 2017, 10 pages. [cited by applicant]
Kim et al., “Unsupervised detection of regions of interest using iterative link analysis,” MIT Libraries, Dec. 7, 2009, 10 Pages. [cited by applicant]
Lan et al., “DISCOBOX: Weakly Supervised Instance Segmentation and Semantic Correspondence from Box Supervision,” Jun. 5, 2021, 13 Pages. [cited by applicant]
Li et al., “Efficient Self-supervised Vision Transformers for Representation Learning,” Jun. 17, 2021, 24 Pages. [cited by applicant]
Li et al., “Fully Convolutional Instance-Aware Semantic Segmentation,” CVPR, 2017, 9 pages. [cited by applicant]
Lin et al., “Feature Pyramid Networks for Object Detection,” CVPR, 2017, 9 pages. [cited by applicant]
Lin et al., “Focal Loss for Dense Object Detection,” ICCV, 2017, 9 pages. [cited by applicant]
Lin et al., “Microsoft COCO: Common Objects in Context,” European Conference on Computer Vision, Jul. 5, 2014, 14 pages. [cited by applicant]
Liu et al., “Leveraging Instance- , Image- and Dataset-Level Information for Weakly Supervised Instance Segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 13 Pages. [cited by applicant]
Liu et al., “Path Aggregation Network for Instance Segmentation,” CVPR, 2018, 10 pages. [cited by applicant]
Liu et al., “SGN: Sequential Grouping Networks for Instance Segmentation,” CVF, 2017, 9 Pages. [cited by applicant]
Maninis et al., “Convolutional Oriented Boundaries: From Image Segmentation to High-Level Tasks,” Apr. 28, 2017, 14 Pages. [cited by applicant]
Martin et al., “A Database of Human Segmented Natural Images and its Application to Evaluating Segmentation Algorithms and Measuring Ecological Statistics,” Jan. 2001, 11 Pages. [cited by applicant]
Milletari et al., “V-net: Fully convolutional neural networks for volumetric medical image segmentation,” 2016 Fourth International Conference on 3D Vision (3DV), Oct. 25, 2016, 11 pages. [cited by applicant]
Mottaghi et al., “The Role of Context for Object Detection and Semantic Segmentation in theWild,” CVPR, 2014, 8 Pages. [cited by applicant]
Newell et al., “Associative Embedding: End-to-End Learning for Joint Detection and Grouping,” Jun. 9, 2017, 11 Pages. [cited by applicant]
Noroozi et al., “Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles,” European Conference on Computer Vision, Jun. 26, 2016, 17 pages. [cited by applicant]