IP Library Granted Patent US 12705877
Granted Patent B2
US 12705877 · App. 18/818,107 · Granted Aug 11, 2026

Parallelized architecture for multi-dimensional data processors

Inventors: Yi Lu (Shanghai, CN); Ching-Yu Hung (Pleasanton, CA); Hanjie Mei (Shanghai, CN); Yen-Te Shih (Zhubei City, TW); Chao Lyu (Shanghai, CN)
Assignee: NVIDIA Corporation
G06V10/94G06T1/20G06V10/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705877
App. No.
18/818,107
Granted
Aug 11, 2026
Kind
B2
Abstract

Aspects of this technical solution can increase speed of processing in low-latency application areas, while maintaining integrity of image feature recognition at those higher speeds. For example, in image-processing environments associated with autonomous navigation (e.g., driving), a large volume of image data is to be rapidly and accurately processed to maintain reliable and up-to-date models of a physical environment. For example, embodiments in accordance with this disclosure can provide high-speed and accurate image feature recognition of input frame data beyond the capability of CPU processing or general GPU processing to achieve.

Claims (74)

1 . A system, comprising:

a plurality of processors to:

determine that first frame data corresponds to a first time and that second frame data corresponds to a second time subsequent to the first time;

provide the first frame data to a first processor of the plurality of processors configured to execute input arranged in two dimensions;

provide, in parallel with the providing the first frame data to the first processor, the second frame data to a second processor of the plurality of processors configured to execute input arranged in one dimension; and

execute the first processor according to the first frame data in parallel with executing the second processor according to the second frame data.

2 . The system of claim 1 , wherein the plurality of processors to:

execute one or more first instructions based on the first frame data.

3 . The system of claim 2 , wherein the one or more first instructions correspond to at least one of a masking operation, a packing operation, or a halo preparation operation.

4 . The system of claim 2 , wherein the plurality of processors to:

execute, subsequent to the executing the one or more first instructions, one or more second instructions based on the first frame data.

5 . The system of claim 4 , wherein the one or more second instructions correspond to a halo preparation operation.

6 . The system of claim 1 , wherein the plurality of processors to:

provide, subsequent to the providing the first frame data to the first processor, the first frame data to the second processor; and

execute, by the second processor subsequent to the providing the first frame data to the first processor, one or more instructions based on the first frame data.

7 . The system of claim 6 , wherein the one or more instructions correspond to an injection operation.

8 . The system of claim 1 , wherein the plurality of processors to:

provide, subsequent to the providing the first frame data to the first processor, third frame data to the first processor, the third frame data corresponding to output of the first processor based on the first frame data; and

execute, by the first processor subsequent to the providing the first frame data to the first processor, one or more instructions based on the first frame data.

9 . The system of claim 8 , wherein the one or more instructions correspond to a non-maximum suppression operation.

10 . The system of claim 8 , wherein the plurality of processors to:

provide, in parallel with the executing the one or more instructions based on the first frame data, the fourth frame data to the second processor, the fourth frame data corresponding to output of the second processor based on the second frame data.

11 . The system of claim 1 , wherein plurality of processors are comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system implemented using a robot;

an aerial system;

a medical system;

a boating system;

a smart area monitoring system;

a system for performing deep learning operations;

a system for performing simulation operations;

a system for generating or presenting virtual reality (VR) content, augmented reality (AR) content, or mixed reality (MR) content;

a system for performing digital twin operations;

a system implemented using an edge device;

a system incorporating one or more virtual machines (VMs);

a system for generating synthetic data;

a system implemented at least partially in a data center;

a system for performing conversational artificial intelligence (AI) operations;

a system for performing generative AI operations;

a system implementing language models;

a system implementing vision language models (VLMs);

a system implementing large language models (LLMs);

a system implementing multi-modal language models;

a system for hosting one or more real-time streaming applications;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets; or

a system implemented at least partially using cloud computing resources.

12 . A system-on-chip (SoC), comprising:

at least one graphics processing unit (GPU); and

a plurality of processors to:

determine that first frame data corresponds to a first time and that second frame data corresponds to a second time subsequent to the first time;

provide the first frame data to a first processor of the plurality of processors configured to execute input arranged in two dimensions;

provide, in parallel with the providing the first frame data to the first processor, the second frame data to a second processor of the plurality of processors configured to execute input arranged in one dimension; and

execute the first processor according to the first frame data in parallel with executing the second processor according to the second frame data.

13 . The SoC of claim 12 , wherein the plurality of processors to:

execute one or more first instructions based on the first frame data.

14 . The SoC of claim 13 , wherein the one or more first instructions correspond to at least one of a masking operation, a packing operation, or a halo preparation operation.

15 . The SoC of claim 13 , wherein the plurality of processors to:

execute, by the first processor subsequent to the executing the one or more first instructions, one or more second instructions based on the first frame data.

16 . The SoC of claim 15 , wherein the one or more second instructions correspond to a halo preparation operation, an injection operation, or a non-maximum suppression operation.

17 . The SoC of claim 12 , wherein the plurality of processors to:

provide, subsequent to the providing the first frame data to the first processor, the first frame data to the second processor; and

execute, by the second processor subsequent to the providing the first frame data to the first processor, one or more instructions based on the first frame data.

18 . The SoC of claim 12 , comprising the at least one GPU to:

provide, subsequent to the providing the first frame data to the first processor, third frame data to the first processor, the third frame data corresponding to output of the first processor based on the first frame data; and

execute, by the first processor subsequent to the providing the first frame data to the first processor, one or more instructions based on the first frame data.

19 . The SoC of claim 12 , wherein the plurality of processors to:

provide, in parallel with the executing the one or more instructions based on the first frame data, the fourth frame data to the second processor, the fourth frame data corresponding to output of the second processor based on the second frame data.

20 . A method performed by a plurality of processors, comprising:

determining that first frame data corresponds to a first time and that second frame data corresponds to a second time subsequent to the first time;

providing the first frame data to a first processor of the plurality of processors configured to execute input arranged in two dimensions;

providing, in parallel with the providing the first frame data to the first processor, the second frame data to a second processor of the plurality of processors configured to execute input arranged in one dimension; and

executing the first processor according to the first frame data in parallel with executing the second processor according to the second frame data.