IP Library › Granted Patent US 12,423,940
Granted Patent B2
US 12,423,940 · App. 18/102,610 · Granted Sep 23, 2025

Targeted object detection in image processing applications

Inventors: Shekhar Dwivedi (Santa Clara, CA); Gigon Bae (Mountain View, CA)
Assignee: NVIDIA Corporation
G06V10/25G06F18/214G06N3/08G06T7/0014G06T7/11G06T2207/20081G06T2207/20084G06T2207/30008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,940
App. No.
18/102,610
Granted
Sep 23, 2025
Kind
B2
Abstract

Apparatuses, systems, and techniques to train and apply a first machine learning model to identify a plurality of regions of interest within an input image, and to train and apply a plurality of second machine learning models to identify one or more objects within each region of interest identified by the first machine learning model.

Claims (38)

1. A method comprising:

identifying, using a first machine learning model (MLM):

a plurality of reference features within a map of pixel intensities of an input image,

a geometric arrangement of the plurality of reference features, and

one or more regions of interest (ROIs) within the input image, wherein the one or more ROIs are identified based at least on the geometric arrangement; and

providing a first ROI of the one or more ROIs as input to a second MLM to obtain an output of the second MLM, wherein the output of the second MLM indicates one or more objects within the first ROI.

2. The method of claim 1 , wherein the input image is a multi-dimensional image of a first dimensionality, the first MLM identifies the plurality of reference features based at least on processing a plurality of sectional images associated with the multi-dimensional image, at least one sectional image of the plurality of sectional images being of a second dimensionality and representing a section of the input image, and wherein the second dimensionality is lower than the first dimensionality.

3. The method of claim 1 , wherein the first MLM is trained using a plurality of training images having reference features of a type common with a type of the reference features within the input image.

4. The method of claim 1 , wherein the first MLM comprises a neural network with at least one hidden layer.

5. The method of claim 1 , wherein the input image comprises a medical image of a patient and the plurality of reference features are associated with one or more components of an anatomy of the patient depicted by the input image.

6. The method of claim 1 , further comprising:

providing a second ROI of the one or more ROIs as input to a third MLM to obtain an output of the third MLM, wherein the output of the third MLM indicates one or more objects within the second ROI.

7. The method of claim 6 , wherein the first ROI comprises a representation of at least a portion of a first organ of a patient and excludes a representation of any portion of a second organ of the patient.

8. The method of claim 7 , wherein the second ROI comprises the representation of at least a portion of the second organ of the patient.

9. The method of claim 1 , wherein providing the first ROIs to the second MLM comprises providing at least one of a location of the first ROI within the input image or a representation of the first ROI.

10. The method of claim 1 , wherein applying the input image to the first MLM comprises executing one or more computations associated with the first MLM on one or more graphics processing units.

11. The method of claim 1 , wherein the geometric arrangement of the plurality of reference features and the one or more ROIs comprises:

positioning of the one or more ROIs relative to a geometric pattern created by the plurality of reference features.

12. A system comprising one or more processing devices to:

apply an input image to a first machine learning model (MLM) to identify:

a plurality of reference features within a map of pixel intensities of the input image,

a geometric arrangement of the plurality of reference features, and

one or more regions of interest (ROIs) within the input image, wherein the one or more ROIs are identified based at least on the geometric arrangement; and

provide a first ROI of the one or more ROIs as input to a second MLM to obtain an output of the second MLM, wherein the output of the second MLM indicates one or more objects within the first ROI.

13. The system of claim 12 , wherein the input image is a multi-dimensional image of a first dimensionality, the first MLM is to identify the plurality of reference features based at least on processing a plurality of sectional images associated with the multi-dimensional image, at least one sectional image of the plurality of sectional images being of a second dimensionality and representing a section of the input image, and wherein the second dimensionality is lower than the first dimensionality.

14. The system of claim 12 , wherein the first MLM is trained using a plurality of training images having reference features of a type common with a type of the reference features within the input image.

15. The system of claim 12 , wherein the input image is a medical image of a patient and the plurality of reference features are associated with one or more components of an anatomy of the patient depicted by the input image.

16. The system of claim 12 , wherein the first ROI comprises a representation of at least a portion of a first organ of a patient and excludes a representation of any portion of a second organ of the patient.

17. The system of claim 12 , wherein the one or more processing devices comprise one or more graphics processing units.

18. A non-transitory computer-readable storage medium storing instructions that, when executed by a processing device, cause the processing device to:

apply an input image to a first machine learning model (MLM) to identify:

a plurality of reference features within a map of pixel intensities of the input image,

a geometric arrangement of the plurality of reference features, and

one or more regions of interest (ROIs) within the input image, wherein the one or more ROIs are identified based at least on the geometric arrangement; and

provide a first ROI of the one or more ROIs as input to a second MLM to obtain an output of the second MLM, wherein the output of the second MLM indicates one or more objects within the first ROI.

19. The non-transitory computer-readable storage medium of claim 18 , wherein the input image is a multi-dimensional image of a first dimensionality, the first MLM identifies the plurality of reference features based at least on processing a plurality of sectional images associated with the multi-dimensional image, at least one sectional image of the plurality of sectional images being of a second dimensionality and representing a section of the input image, and wherein the second dimensionality is lower than the first dimensionality.

20. The non-transitory computer-readable storage medium of claim 18 , wherein the instructions further cause the processing device to:

provide a second ROI of the one or more ROIs as input to a third MLM to obtain an output of the third MLM, wherein the output of the third MLM indicates one or more objects within the second ROI.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2023
From: DWIVEDI, SHEKHAR; BAE, GIGON
To: NVIDIA CORPORATION
Reel/Frame 064624/0738 →
Continuity (2)
Continuation 17248069 · Jan 7, 2021
Related Publication 20230169746A1 · Jun 1, 2023
References Cited (16)
US 20200268339A1 · Hao et al. · 2020 [cited by applicant]
US 20200294257A1 · Yoo et al. · 2020 [cited by applicant]
US 20200324795A1 · Bojarski et al. · 2020 [cited by applicant]
US 20210383534A1 · Tadross · 2021 [cited by examiner]
US 20220076133A1 · Yang et al. · 2022 [cited by applicant]
CN 111311551A · 2020 [cited by examiner]
DE 112004001468B4 · 2017 [cited by applicant]
WO 2021210723A1 · 2021 [cited by applicant]
Liu, L., Zhang, B., Wang, H.: Organ localization in peUct images using hierarchical conditional faster r-cnn method. In: Proceedings of the Third International Symposium on Image Computing and Digital Medicine. 2019. pp… [cited by applicant]
Mikolajczyk, K., Schmid, C., Zisserman, A.: Human detection based on a probabilistic assembly of robust part detectors. In: European conference on i computer vision. Springer, Berlin, Heidelberg, 2004. pp. 69-82. doi: 1… [cited by applicant]
Xu, X., et al.: Efficient multiple organ localization in CT image using 30 region proposal network. In: IEEE transactions on medical imaging, 2019, 38. Jg., Nr. 8, pp. 1885-1898. doi: 10.1109/TMI.2019.2894854. [cited by applicant]
Zhang, N., et al.: Part-based R-CNNs for fine-grained category detection. In: I European conference on computer vision. Springer, Cham, 2014. pp. 834-849. doi: 10.1007/978-3-319-10590-1_54. [cited by applicant]
Jiang, G., Wang, Z., Liu, H.: Automatic detection of crop rows based on multi-ROIs. In: Expert systems with applications, 2015, 42. Jg., Nr. 5, pp. 2429-2441. doi: 10.1016/j.eswa.2014.10.033. [cited by applicant]
Lei, Y., et al.: Multi-organ segmentation in head and neck MRI using U-Faster-RCNN. In: Medical Imaging 2020: Image Processing. SPIE, 2020. pp. 826-831. doi: 10.1117/12.2549596. [cited by applicant]
Donner, R., et al.: Global localization of 30 anatomical structures by pre-filtered. Hough Forests and discrete optimization. In: Medical image analysis, 2013, 17. Jg., Nr. 8, pp. 1304-1314. doi: 10.1016/j.media.2013.02… [cited by applicant]
Office Action for German Patent Application No. 102021133631.7, dated Jul. 29, 2022, 15 pages. [cited by applicant]