IP Library Patent Application 18984852
Patent Application
App. No. 18/984,852

DATA ACCESS IN A HETEROGENEOUS PROCESSING SYSTEM WITH MULTIPLE PROCESSORS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/984,852
Abstract

A data access method and apparatus for a heterogeneous processing system includes a host processor, a first processor coupled to a first memory, a second processor coupled to a second memory, and switch and bus circuitry that communicatively couples the host processor, the first processor, and the second processor. The host processor maps virtual addresses of the second memory to physical addresses of the switch and bus circuitry. The first processor is configured to directly access the second memory using the mapped physical addresses, and may be configured to directly access the second memory for reading and writing data while executing an application.

Claims (47)

1 . A method for accessing data in a heterogeneous processing system using a dataflow graph having a plurality of nodes connected by edges, wherein the heterogeneous processing system includes a host processor, a first processor coupled to a first memory, a second processor coupled to a second memory, and switch and bus circuitry that communicatively couples the host processor, the first processor, and the second processor, the method comprising:

executing at least a portion of a first node of the plurality of nodes of the dataflow graph using the first processor;

executing at least a portion of a second node of the plurality of nodes of the dataflow graph using the second processor;

mapping virtual addresses of the second memory to physical addresses of the switch and bus circuitry; and

configuring the first processor to directly access the second memory using the mapped physical addresses; and

directly accessing, by the first processor, the second memory through the switch and bus circuitry.

2 . The method of claim 1 , wherein the method is for implementing a machine learning system using dataflow graphs.

3 . The method of claim 1 , wherein configuring the first processor comprises configuring a reconfigurable dataflow unit.

4 . The method of claim 1 , wherein configuring the first processor comprises configuring a compute engine.

5 . The method of claim 1 , wherein the first processor comprises a reconfigurable processor, comprising:

an array of coarse-grained reconfigurable units, each including an address generation unit, a plurality of memory units, and a plurality of compute units interconnected by an array-level network;

a top-level network coupled to the address generation unit of the array of coarse-grained reconfigurable units; and

an interface coupled between the top-level network and the switch and bus circuitry;

wherein configuring the first processor comprises configuring the address generation unit of the array of coarse-grained reconfigurable unit in the reconfigurable processor to map virtual addresses of the second memory to physical addresses of the switch and bus circuitry.

6 . The method of claim 1 , further comprising:

generating and storing, by programming the second processor, while executing the portion of the second node, to execute a first part of an application to generate and store first data into the second memory; and

configuring the first processor to directly access, by the first processor, the first data from the second memory using mapped physical addresses while executing a second part of the application the portion of the first node.

7 . The method of claim 1 , further comprising configuring the first processor to write second data generated by the second part of the application into the first memory portion of the first node.

8 . The method of claim 1 , wherein the heterogeneous system includes a host memory coupled to the host processor; and

wherein the heterogeneous system is configured to provide the first processor with direct access to the host memory and to write second data output from executing the portion of the second node part of the application directly into the host memory.

9 . The method of claim 1 , wherein the heterogeneous system is configured to provide the first processor to execute a first part of an application to generating, by the first processor while executing the portion of the first node, the first data and directly writing the first data into the second memory using mapped physical addresses.

10 . The method of claim 1 , the heterogeneous system including a host memory coupled to the host processor,

executing the portion of the first node using the first processor to directly read first data from the host memory

generating second data using the first data; and

while executing the application, directly writing the second data into the second memory using mapped physical addresses.

11 . A method of claim 1 , wherein the heterogeneous system includes mapping the virtual addresses of the second memory to the physical addresses of the switch and bus circuitry; and

wherein the first processor is configured to directly access the second memory using the mapped physical addresses.

12 . The method of claim 1 , wherein the second processor is programmed to execute a first part of an application to generate and write first data into the second memory; and

wherein the first processor is configured to directly access the first data from the second memory using mapped physical addresses while executing a second part of the application.

13 . The method of claim 12 , wherein the first processor is configured to write second data generated by the second part of the application into the first memory.

14 . The method of claim 13 , wherein the first processor is configured to directly access the host memory and to write second data output from executing the second part of the application directly into the host memory.

15 . The method of claim 1 , wherein the first processor is configured to execute a first part of an application to generate first data and to directly write the first data into the second memory using mapped physical addresses.

16 . The method of claim 1 , wherein the first processor is configured to directly read first data from the host memory while executing an application using the first data to generate second data; and

directly writing the second data into the second memory using mapped physical addresses while executing the application.

17 . A heterogeneous processing system, comprising:

a host processor;

a first processor coupled to a first memory, wherein the first processor comprises a reconfigurable processor comprising:

an array of coarse-grained reconfigurable units comprising, an address generation unit, a plurality of memory units, and a plurality of compute units interconnected by an array-level network;

a top-level network coupled to the address generation unit of the array of coarse-grained reconfigurable units; and

an interface coupled between the top-level network and an external port of the first processor;

a second processor coupled to a second memory; and

switch and bus circuitry that communicatively couples the host processor, the extremal port of the first processor, and the second processor;

wherein the host processor is configured to map virtual addresses of the second memory to physical addresses of the switch and bus circuitry; and

wherein the first processor can directly access the second memory using the mapped physical addresses.

18 . The heterogeneous processing system of claim 17 , wherein the second processor is programmed to execute a first part of an application to generate and store first data into the second memory, and wherein the first processor is configured to directly access the first data from the second memory using the mapped physical addresses while executing a second part of the application using the first data.

19 . The heterogeneous processing system of claim 17 , wherein the first processor is further configured to store second data output from executing the second part of the application into the first memory.

20 . The heterogeneous processing system of claim 17 , wherein the first processor may be a reconfigurable processor, a reconfigurable dataflow unit, or a compute engine.

Assignments (2)
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Apr 18, 2025
From: SAMBANOVA SYSTEMS, INC.
To: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY, AS AGENT
Reel/Frame 070892/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 19, 2025
From: GOEL, ARNAV; SANGHVI, NEAL; BAI, JIAYU; ZHENG, QI; KUMAR, RAVINDER
To: SAMBANOVA SYSTEMS, INC.
Reel/Frame 070548/0063 →