IP Library › Granted Patent US 12,738,014
Granted Patent B2
US 12,738,014 · App. 17/990,508 · Granted Sep 15, 2026

Directed inferencing using input data transformations

Inventor: Shekhar Dwivedi (Santa Clara, CA)
Assignee: NVIDIA Corporation
G06V10/267G06T1/20G06V10/774
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,014
App. No.
17/990,508
Granted
Sep 15, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to segment a region of interest in an input data for a machine learning model by obtaining a plurality of progressively compressed representations of the input data and processing at least two of the compressed representations to obtain matching locations of a region of interest of the input data.

Claims (42)

1 . A method comprising:

generating a plurality of compressed representations of an input image;

processing, in parallel, the plurality of compressed representations of the input image using a plurality of machine learning models (MLMs), wherein a respective MLM of the plurality of MLMs is used to process a respective compressed representation of the plurality of compressed representations of the input image to estimate a location of a region of interest (ROI) in the respective compressed representation of the input image, and wherein a number of neurons in an input layer of the respective MLM corresponds to a number of pixels of the respective compressed representation; and

responsive to at least a predetermined number of the estimated locations of the ROI satisfying a matching condition, generating an image of the ROI based on the input image and one or more of the estimated locations.

2 . The method of claim 1 , wherein generating the plurality of compressed representations of the input image comprises:

obtaining each of a plurality of pixels of a first compressed representation of the plurality of compressed representations of the input image by aggregating two or more pixels of the input image; and

obtaining each of a plurality of pixels of a second compressed representation of the plurality of compressed representations of the input image by aggregating two or more pixels of the first compressed representation.

3 . The method of claim 1 , the matching condition comprises:

the predetermined number of the estimated locations of the ROI being within a predetermined tolerance from each other.

4 . The method of claim 1 , wherein each of the estimated locations of the ROI comprises an identification of a size and position of a bounding box enclosing the ROI.

5 . The method of claim 1 , further comprising:

processing, using a classification MLM, the generated image of the ROI to obtain at least one of:

an identification of an object depicted in the generated image of the ROI, or a medical diagnosis associated with an organ depicted in the generated image of the ROI.

6 . The method of claim 1 , wherein at least one of the predetermined number of the estimated locations of the ROI or a number of compressed representations is determined based on available, for processing of the input image, computing resources and a size of the input image.

7 . The method of claim 1 , wherein at least one of the predetermined number of the estimated locations of the ROI or a number of compressed representations is determined based on parameters of the MLM.

8 . The method of claim 6 , wherein the computing resources available for processing of the input image comprise one or more graphics processing units (GPUs).

9 . The method of claim 6 , wherein at least one of the predetermined number of the estimated locations of the ROI or a number of compressed representations is determined using a mapping table comprising empirical correspondence of the available computing resources and the size of the input image to the at least one of the predetermined number of the estimated locations of the ROI or the number of compressed representations.

10 . The method of claim 6 , wherein at least one of the predetermined number of the estimated locations of the ROI or a number of compressed representations is determined by applying a learned model to a list of the available computing resources and the size of the input image.

11 . A system comprising:

one or more processing devices to:

generate a plurality of compressed representations of an input image;

process, in parallel, the plurality of compressed representations of the input image using a plurality of machine learning models (MLMs), wherein a respective MLM of the plurality of MLMs is used to process a respective compressed representation of the plurality of compressed representations of the input image to estimate a location of a region of interest (ROI) in the respective compressed representation of the input image, and wherein a number of neurons in an input layer of the respective MLM corresponds to a number of pixels of the respective compressed representation; and

responsive to at least a predetermined number of the estimated locations of the ROI satisfying a matching condition, generate an image of the ROI based on the input image and one or more of the estimated locations.

12 . The system of claim 11 , wherein to generate the plurality of compressed representations of the input image, the one or more processing devices are to:

obtain each of a plurality of pixels of a first compressed representation of the plurality of compressed representations of the input image by aggregating two or more pixels of the input image; and

obtain each of a plurality of pixels of a second compressed representation of the plurality of compressed representations of the input image by aggregating two or more pixels of the first compressed representation.

13 . The system of claim 11 , wherein the matching condition comprises:

the predetermined number of the estimated locations of the ROI being within a predetermined tolerance of each other.

14 . The system of claim 11 , wherein the one or more processing devices are further to:

process, using a classification MLM, the generated image of the ROI to obtain at least one of:

an identification of an object depicted in the generated image of the ROI, or

a medical diagnosis associated with an organ depicted in the generated image of the ROI.

15 . The system of claim 11 , wherein at least one of the predetermined number of the estimated locations of the ROI or a number of compressed representations is determined based on available, for processing of the input image, computing resources and a size of the input image.

16 . The system of claim 11 , wherein at least one of the predetermined number of the estimated locations of the ROI or a number of compressed representations is determined based on parameters of the MLM.

17 . The system of claim 15 , wherein the computing resources available for processing of the input image comprise one or more graphics processing units (GPUs).

18 . The system of claim 15 , wherein at least one of the predetermined number of the estimated locations of the ROI or a number of compressed representations is determined using a mapping table comprising empirical correspondence of the available computing resources and the size of the input image to the at least one of the predetermined number of the estimated locations of the ROI or the number of compressed representations.

19 . The system of claim 15 , wherein at least one of the predetermined number of the estimated locations of the ROI or a number of compressed representations is determined by applying a learned model to a list of the available computing resources and the size of the input image.

20 . A processor comprising:

one or more processing units to:

generate a plurality of compressed representations of an input image;

process, in parallel, the plurality of compressed representations of the input image using a plurality of machine learning models (MLMs), wherein a respective MLM of the plurality of MLMs is used to process a respective compressed representation of the plurality of compressed representations of the input image to estimate a location of a region of interest (ROI) in the respective compressed representation of the input image, and wherein a number of neurons in an input layer of the respective MLM corresponds to a number of pixels of the respective compressed representation; and

responsive to at least a predetermined number of the estimated locations of the ROI satisfying a matching condition, generate an image of the ROI based on the input image and one or more of the estimated locations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2023
From: DWIVEDI, SHEKHAR
To: NVIDIA CORPORATION
Reel/Frame 062273/0502 →
Continuity (1)
Related Publication 20240169686A1 · May 23, 2024
References Cited (20)
US 10192327B1 · Toderici · 2019 [cited by examiner]
US 10984560B1 · Appalaraju · 2021 [cited by examiner]
US 11335077B1 · Salmani Rahimi · 2022 [cited by examiner]
US 12260598B2 · Hurwitz · 2025 [cited by examiner]
US 20090109133A1 · Park · 2009 [cited by examiner]
US 20190158813A1 · Rowell · 2019 [cited by examiner]
US 20200320416A1 · Gillian · 2020 [cited by examiner]
US 20210090247A1 · Jeon · 2021 [cited by examiner]
US 20210208236A1 · John Wilson · 2021 [cited by examiner]
US 20210295570A1 · Broyelle · 2021 [cited by examiner]
US 20220021887A1 · Banerjee · 2022 [cited by examiner]
US 20220374156A1 · Naruko · 2022 [cited by examiner]
US 20240077455A1 · Lamarre · 2024 [cited by examiner]
US 20240233414A1 · Schaumberg · 2024 [cited by examiner]
AU 6174199A · 2000 [cited by examiner]
CN 107679250A · 2018 [cited by examiner]
CN 107832835A · 2018 [cited by examiner]
CN 112738533A · 2021 [cited by examiner]
“Morozkin et al., An Image Compression for Embedded Eye-Tracking Applications, 2016 International Symposium on INnovations in Intelligent Systems and Applications (INISTA), pp. 1-5” (Year: 2016). [cited by examiner]
Fahrni G. Three-Dimensional Adaptive Image Compression Concept for Medical Imaging: Application to Computed Tomography Angiography for Peripheral Arteries. J Cardiovasc Dev Dis. Apr. 27, 2022;9(5):137. (Year: 2022). [cited by examiner]