IP Library › Granted Patent US 12,573,216
Granted Patent B2
US 12,573,216 · App. 16/891,368 · Granted Mar 10, 2026

Machine-learning-based object detection system

Inventors: Can Zhao (Rockville, MD); Daguang Xu (Potomac, MD); Wentao Zhu (Mountain View, CA); Dong Yang (North Bethesda, MD); Ziyue Xu (Reston, VA)
Assignee: NVIDIA Corporation
G06V20/64G06T7/0012G06T11/20G06V10/25G06V10/764G06V10/82G06T2207/10081G06T2207/10088G06T2207/10116G06T2207/20081G06T2207/20084G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,216
App. No.
16/891,368
Granted
Mar 10, 2026
Kind
B2
Abstract

In at least one embodiment, an object detection system uses a neural network to identify and/or locate a set of organs in a medical image. In at least one embodiment, when training to identify and/or locate a particular organ, a subset of incompletely-labeled training images is used that excludes training images for which labels associated with particular organ are unavailable.

Claims (75)

1 . One or more processors, comprising:

circuitry to cause a neural network to be trained by at least:

obtaining a set of incompletely labeled images, at least one image of the set of incompletely labeled images including at least one labeled object and at least one unlabeled object; and

training the neural network to identify a particular object using the set of incompletely labeled images by preventing loss associated with the at least one image from being used to update weights of the neural network as a result of identification, by the neural network, of the at least one unlabeled object in the at least one image.

2 . The one or more processors of claim 1 , wherein the at least one unlabeled object includes an organ within one or more three-dimensional medical images.

3 . The one or more processors of claim 1 , wherein:

the set of incompletely labeled images is a set of incompletely labeled three-dimensional images, where each image in the set of incompletely labeled three-dimensional images includes one or more labels identifying one or more objects and a bounding region for each of one or more objects, the set of incompletely labeled three-dimensional images including the at least one image that lacks labeling for the at least one unlabeled object; and

training the neural network to identify the particular object includes using the set of incompletely labeled three-dimensional images, excluding, for a purpose of training to identify the particular object, an image that lacks a label for the particular object.

4 . The one or more processors of claim 3 , wherein locating an organ of the one or more objects is accomplished by at least:

determining a bounding box around the one or more objects; and

indicating, to a user, the location of the bounding box.

5 . The one or more processors of claim 3 , wherein:

a location of an object is determined using an anchor that indicates where detection begins; and

that anchor describes a three-dimensional region that has a shape based at least in part on a type of the object.

6 . The one or more processors of claim 5 , wherein the three-dimensional region is a rectangular box.

7 . The one or more processors of claim 5 , wherein a size of the anchor is varied over a range of sizes that correspond to size variations of human anatomy.

8 . The one or more processors of claim 1 , wherein the neural network is trained using a loss based at least in part on a prediction of an Intersection over Union.

9 . The one or more processors of claim 8 , wherein the prediction of the Intersection over Union provides a quality assessment of a location of an object.

10 . The one or more processors of claim 1 , wherein the neural network is trained using a loss calculated based, at least in part, on skipping images with unlabeled objects.

11 . The one or more processors of claim 1 , wherein the neural network includes a plurality of region proposal networks associated with a plurality of objects, wherein at least one region proposal network associated with an unlabeled object is skipped in training.

12 . A system, comprising one or more processors to cause a neural network to be trained by at least:

obtaining a set of images, at least one image of the set of images including:

one or more labels that identify one or more objects in the at least one image; and

one or more unlabeled objects; and

training the neural network to identify a particular object using the set of images by preventing loss associated with the at least one image from being used to update weights of the neural network as a result of identification, by the neural network, of the one or more unlabeled objects in the at least one image.

13 . The system of claim 12 , wherein the one or more unlabeled objects comprise an organ within a three-dimensional medical image.

14 . The system of claim 12 , wherein:

the set of images includes one or more bounding regions for the one or more objects and lacks at least one bounding region for the one or more unlabeled objects; and

training the neural network to identify the particular object includes excluding images in the set of images that lack a label for the particular object.

15 . The system of claim 14 , wherein locating an object of the one or more objects is accomplished by at least:

determining a region around the object; and

presenting an indication of the region on an electronic display.

16 . The system of claim 15 , wherein:

the region is defined by a box-shaped region; and

the region is presented on the electronic display as an isometric view of the box-shaped region.

17 . The system of claim 14 , wherein:

the system uses an anchor that indicates, to the system, an estimate of a size and shape of the particular object; and

the estimate is specific to a type of the particular object.

18 . The system of claim 15 , wherein the anchor defines a three-dimensional shape, and a range of sizes for the three-dimensional shape.

19 . The system of claim 18 , wherein the range is determined using a scaling factor associated with a particular image.

20 . The system of claim 13 , wherein the three-dimensional medical image is:

a magnetic resonance imaging (MRI) image,

a computed tomography (CT) scan image, or

an X-ray image.

21 . One or more processors comprising processing circuitry to use a neural network, wherein the neural network is trained by:

obtaining a set of images, at least one image of the set of images including one or more labeled objects and one or more unlabeled objects; and

training the neural network to identify a particular object using the set of images by excluding the at least one image from the training.

22 . The one or more processors of claim 21 , wherein the neural network is trained to identify a plurality of organs within a three-dimensional medical image.

23 . The one or more processors of claim 21 , wherein:

the set of images includes a bounding region for each of the one or more labeled objects and lacks at least one bounding region for the one or more unlabeled objects; and

the neural network is trained to identify the particular object by excluding, for a purpose of training, images of the set of images from the training that lack a label for the particular object.

24 . The one or more processors of claim 21 , wherein the neural network is trained using a loss based at least in part on a prediction of an Intersection over Union.

25 . The one or more processors of claim 21 , wherein the neural network is further trained by at least removing, from a set of three-dimensional images, one or more images that lack labeling for a type of a particular object.

26 . The one or more processors of claim 25 , wherein each image in the set of three-dimensional images represents a portion of a human body with a known composition of organs.

27 . The one or more processors of claim 26 , wherein the set of images represent head, chest, or abdominal scan images.

28 . The one or more processors of claim 27 , wherein a location of individual organs is determined using estimates of a size and shape of each of the individual organs based at least in part on the type of each organ of the individual organs.

29 . The one or more processors of claim 28 , wherein the estimates include a size range around an average object size.

30 . A method comprising, training a neural network, at least in part, by at least:

obtaining a set of incompletely labeled images, at least one image in the set of incompletely labeled images including one or more labeled objects and at least one unlabeled object; and

training the neural network to identify a particular object using the set of incompletely labeled images by excluding the at least one image from the training.

31 . The method of claim 30 , wherein;

the at least one unlabeled object is an organ; and

the set of incompletely labeled images comprise one or more three-dimensional medical images.

32 . The method of claim 30 , wherein:

the set of incompletely labeled images are a set of incompletely labeled three-dimensional images that includes a bounding region for each of the one or more labeled objects and lacks at least one bounding region for the at least one unlabeled object; and

the neural network is trained to identify the particular object by excluding, for a purpose of training to identify the particular object, images of the set of images that lack a label for the particular object.

33 . The method of claim 32 , wherein locating an organ of the one or more labeled objects is accomplished by at least:

determining a bounding box around an object; and

indicating, to a user, the location of the bounding box.

34 . The method of claim 32 , wherein:

a location of an object is determined using an anchor that indicates where detection begins; and

that anchor describes a three-dimensional region that has a shape based at least in part on a type of the object.

35 . The method of claim 34 , wherein the three-dimensional region is a rectangular box.

36 . The method of claim 34 , wherein a size of the anchor is varied over a range of sizes that correspond to size variations of human anatomy.

37 . The method of claim 30 , wherein the neural network is trained using a loss based at least in part on a prediction of an Intersection over Union.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 17, 2020
From: ZHAO, CAN; XU, DAGUANG; ZHU, WENTAO; YANG, DONG; XU, ZIYUE
To: NVIDIA CORPORATION
Reel/Frame 053245/0798 →
Continuity (1)
Related Publication 20210383533A1 · Dec 9, 2021
References Cited (50)
US 10430946B1 · Zhou · 2019 [cited by examiner]
US 20190050982A1 · Song · 2019 [cited by examiner]
US 20200051238A1 · El Harouni · 2020 [cited by examiner]
US 20200058126A1 · Wang · 2020 [cited by examiner]
US 20200285906A1 · Do · 2020 [cited by examiner]
US 20200334810A1 · Accomazzi · 2020 [cited by examiner]
US 20210027098A1 · Ge · 2021 [cited by examiner]
US 20210142485A1 · Koster · 2021 [cited by examiner]
US 20210307841A1 · Buch · 2021 [cited by examiner]
US 20210343014A1 · Haghighi · 2021 [cited by examiner]
US 20210350528A1 · Tang · 2021 [cited by examiner]
US 20220036971A1 · Yoo · 2022 [cited by examiner]
CN 110009599A · 2019 [cited by applicant]
CN 110060263A · 2019 [cited by applicant]
CN 110383292A · 2019 [cited by applicant]
CN 111191784A · 2020 [cited by applicant]
Mlynarski et al., “Anatomically Consistent Segmentation of Organs at Risk in MRI with Convolutional Neural Networks,” Jul. 3, 2019, 40 pages (Year: 2019). [cited by examiner]
Rezatofighi et al.:, “Generalized Intersection over Union: A Metric and A Loss for Bounding Box Regression,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, 9 pages (Year: 2019). [cited by examiner]
Aristovich et al., “Automatic Procedure for Realistic 3D Finite Element Modelling of Human Brain for Bioelectromagnetic Computations,” Journal of Phyics: Conference Series, Institute of Physics Publishing, 238(1): Jul. … [cited by applicant]
Clark et al., “The Cancer Imaging Archive (TCIA): Maintaining and Operating a Public Information Repository,” Journal of digital imaging 26(6): 2013, 14 pages. [cited by applicant]
Durand et al., “Learning a Deep Convnet for Multi-Label Classification with Partial Labels,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, 11 pages. [cited by applicant]
Gibson et al., “Automatic Multi-Organ Segmentation on Abdominal CT with Dense V-Networks,” IEEE Transactions on Medical Imaging 37(8): 2018, 12 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
International Search Report and Written Opinion for Applcation No. PCT/US2021/035506, mailed Sep. 10, 2021, filed Jun. 2, 2021, 15 pages. [cited by applicant]
Landman et al., Miccai Multi-Atlas Labeling Beyond the Cranial Vault—Workshop and Challenge, MICCAI Multi-Atlas Labeling Beyond Cranial Vault—Workshop Challenge, Apr. 15, 2015, 5 pages. [cited by applicant]
Liu et al., “Organ Localization in Pet/CT Images using Hierarchical Conditional Faster R-CNN Method,” Proceedings of the Third International Symposium on Image Computing and Digital Medicine, 2019, 5 pages. [cited by applicant]
Liu et al., “SSD: Single Shot Multibox Detector,” European Conference on Computer Vision, Springer, Dec. 29, 2016, 17 pages. [cited by applicant]
Mamani et al., “Organ Detection in Thorax Abdomen CT using Multi-Label Convolutional Neural Networks,” Medical Imaging Computer-Aided Diagnosis, vol. 10134, International Society for Optics and Photonics, 2017, 6 pages. [cited by applicant]
Mansoor et al, “Region Proposal Networks with Contextual Selective Attention for Real-Time Organ Detection,” IEEE 16th International Symposium on Biomedical Imaging, Dec. 26, 2018, 4 pages. [cited by applicant]
Mlynarski et al., “Anatomically Consistent Segmentation of Organs at Risk in MRI with Convolutional Neural Networks,” Jul. 3, 2019, 40 pages. [cited by applicant]
Nguyen et al., “Semi-Supervised Object Detection with Unlabeled Data,” International Conference on Computer Vision Theory and Applications, 2019, 8 pages. [cited by applicant]
Raudaschl et al.: “Evaluation of Segmentation Methods on Head and Neck CT: Auto-Segmentation Challenge,” Medical physics, 2017, 17 pages. [cited by applicant]
Redmon et al., “YOLOv3: An Incremental Improvement,” Apr. 8, 2018, 6 pages. [cited by applicant]
Redmon et al., “You Only Look Once: Unified, Realtime Object Detection,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, 10 pages. [cited by applicant]
Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” Advances in Neural Information Processing Systems, 2015, 9 pages. [cited by applicant]
Rezatofighi et al., “Generalized Intersection over Union: A Metric and A Loss for Bounding Box Regression,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, 9 pages. [cited by applicant]
Rhee et al., “Active and Semi-Supervised Learning for Object Detection with Imperfect Data,” Cognitive Systems Research 45, 2017, 16 pages. [cited by applicant]
Roth et al., “Data from Pancreas-CT,” The Cancer Imaging Archive, Sep. 16, 2020, 2 pages. [cited by applicant]
Roth et al., “Deeporgan: Multi-level Deep Convolutional Networks for Automated Pancreas Segmentation,” International Conference on Medical Image Computing and Computer-assisted Intervention, Springer, Jun. 23, 2015, 12 … [cited by applicant]
Simpson et al, “A Large Annotated Medical Image Dataset for the Development and Evaluation of Segmentation Algorithms,” Feb. 25, 2019, 15 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Valliéres, et al., “Radiomics Strategies for Risk Assessment of Tumour Failure in Head-and-Neck Cancer,” Scientific Reports, 7(1): 2017, 14 pages. [cited by applicant]
Wang et al., “CT Male Pelvic Organ Segmentation via Hybrid Loss Network with Incomplete Annotation,” IEEE Transactions on Medical Imaging, 39(6): Jun. 2020, 12 pages. [cited by applicant]
Wu et al., “Multi-Label Learning with Missing Labels using Mixed Dependency Graphs,” International Journal of Computer Vision, 126(8): Mar. 31, 2018, 20 pages. [cited by applicant]
Xu et al., “Efficient Multiple Organ Localization in CT Image using 3D Region Proposal Network,” IEEE Transactions on Medical Imaging, 38(8): 2019, 23 pages. [cited by applicant]
Zheng et al., “Distance-IoU Loss: Faster and Better Learning for Bounding Box Regression,” Nov. 19, 2019, 8 pages. [cited by applicant]
Zhu et al., “Anatomynet: Deep Learning for Fast and Fully Automated Whole-Volume Segmentation of Head and Neck Anatomy,” Medical Physics, 46(2): 2019, 13 pages. [cited by applicant]
Office Action for Chinese Application No. 202180005880.X, mailed Aug. 29, 2025, 20 pages. [cited by applicant]
Office Action for Chinese Application No. 202180005880.X, mailed Mar. 26, 2025, 31 pages. [cited by applicant]