IP Library Granted Patent US 12,248,430
Granted Patent B2
US 12,248,430 · App. 17/945,045 · Granted Mar 11, 2025

Overlay layer for network of processor cores

Inventors: Davor Capalija (Toronto, CA); Ivan Matosevic (Toronto, CA); Jasmina Vasiljevic (Toronto, CA); Utku Aydonat (Toronto, CA); Andrew Lewycky (Toronto, CA); S. Alexander Chin (Toronto, CA); Ljubisa Bajic (Toronto, CA); Alex Cejkov (Toronto, CA); Milos Trajkovic (Toronto, CA)
Assignee: Tenstorrent Inc.
G06F15/7807G06F5/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,248,430
App. No.
17/945,045
Granted
Mar 11, 2025
Kind
B2
Abstract

Methods and systems related to the efficient execution of complex computations by a multicore processor and the movement of data among the various processing cores in the multicore processor are disclosed. A multicore processor includes a set of processing cores and associated sets of processing pipelines, core controllers, routers, and network interface units. The multicore processor also includes a computation layer, for conducting computations using the set of processing cores, with executable instructions for the set of processing pipelines which are executed by the set of core controllers. The multicore processor also includes a network-on-chip layer, for connecting the set of processing cores in the multicore processor, with executable instructions for the set of routers and the set of network interface units. The multicore processor also includes a set of programmable controllers, with executable instructions for reformatting computational data from the computation layer for transmission through the network-on-chip layer.

Claims (66)

1. A multicore processor stack, stored on non-transitory computer readable media in a multicore processor, comprising:

a computation layer, for conducting computations using a set of processing cores in the multicore processor, with executable instructions for a set of processing pipelines in the set of processing cores;

a network-on-chip layer, for connecting the set of processing cores in the multicore processor, with executable instructions for a set of routers and a set of network interface units in the multicore processor;

a set of programmable controllers, with executable instructions for reformatting computational data from the computation layer for transmission through the network-on-chip layer, wherein each processing core in the set of processing cores has a programmable controller from the set of programmable controllers; and

a NoC overlay layer that logically isolates the computation layer and the network-on-chip layer.

2. The multicore processor stack of claim 1 , wherein:

reformatting computational data includes at least one of changing a data type and changing a number of data elements used to represent the computational data.

3. The multicore processor stack of claim 1 , further comprising:

a set of data reformatting blocks, wherein the reformatting blocks in the set of data reformatting blocks are individually and operatively coupled to corresponding processing pipelines in the set of processing pipelines; and

wherein the set of data reformatting blocks reformat the computational data in response to an execution of the executable instructions for reformatting the computational data by the set of programmable controllers.

4. The multicore processor stack of claim 1 , further comprising:

a set of main memories, wherein each processing core in the set of processing cores has a main memory from the set of main memories; and

a set of data reformatting blocks, wherein the reformatting blocks in the set of data reformatting blocks are individually and operatively coupled to corresponding memories in the set of main memories; and

wherein the set of data reformatting blocks reformat the computational data in response to an execution of the executable instructions for reformatting the computational data by the set of programmable controllers.

5. The multicore processor stack of claim 1 , wherein:

the programmable controllers are firmware controllers.

6. The multicore processor stack of claim 1 , wherein the NoC overlay layer is hardware-instantiated via the set of programmable controllers and logically isolates the computation layer from the network-on-chip layer in that:

the programmable controllers execute the executable instructions for reformatting computational data from the computation layer without using the computation layer and asynchronously to an execution of the executable instructions for the set of processing pipelines.

7. The multicore processor stack of claim 1 , wherein:

the NoC overlay layer includes executable instructions for an application data flow graph; and

a single compiler generates the executable instructions for the set of processing pipelines, the executable instructions for the set of routers and the set of network interface units, the executable instructions for the application data flow graph, and the executable instructions for reformatting the computational data from the computation layer.

8. The multicore processor stack of claim 1 , wherein:

the computation layer and the NoC overlay layer are communicatively connected via an interface; and

the interface is an application programming interface.

9. The multicore processor stack of claim 8 , wherein:

the interface includes a pull message.

10. The multicore processor stack of claim 1 , wherein:

the NoC overlay layer distributively instantiates a network-on-chip overlay graph across the set of processing cores; and

a path of the network-on-chip overlay graph begins on a first processing core and terminates with a pull computation data API call response on a second processing core.

11. The multicore processor stack of claim 10 , wherein the NoC overlay layer logically isolates the computation layer from the network-on-chip layer in that:

the NoC overlay layer implements the path of the network-on-chip overlay graph asynchronously to a pull computation data API call corresponding to the pull computation data API call response.

12. The multicore processor stack of claim 10 , wherein:

the network-on-chip overlay graph is configurable via an application programming interface.

13. The multicore processor stack of claim 1 , wherein:

the NoC overlay layer distributively instantiates a network-on-chip overlay graph across the set of processing cores; and

the network-on-chip overlay graph is configured via a compiler offline.

14. The multicore processor stack of claim 1 , wherein:

the NoC overlay layer distributively instantiates a network-on-chip overlay graph across the set of processing cores; and

the network-on-chip overlay graph is configured by a runtime system during an application execution.

15. A multicore processor comprising:

a set of processing cores;

a set of processing pipelines, wherein each processing core in the set of processing cores has a processing pipeline from the set of processing pipelines;

a set of core controllers, wherein each processing core in the set of processing cores has a core controller from the set of core controllers;

a set of routers, wherein each processing core in the set of processing cores has a router from the set of routers;

a set of network interface units, wherein each processing core in the set of processing cores has a network interface unit from the set of network interface units;

a computation layer, for conducting computations using the set of processing cores in the multicore processor, with executable instructions for the set of processing pipelines in the set of processing cores, wherein the executable instructions for the set of processing pipelines in the set of processing cores are executed by the set of core controllers;

a network-on-chip layer, for connecting the set of processing cores in the multicore processor, with executable instructions for the set of routers and the set of network interface units; and

a set of programmable controllers, with executable instructions for reformatting computational data from the computation layer for transmission through the network-on-chip layer, wherein each processing core in the set of processing cores has a programmable controller from the set of programmable controllers.

16. The multicore processor of claim 15 , further comprising:

a NoC overlay layer that: (i) is in communication with both the computation layer and the network-on-chip layer; (ii) administrates a network-on-chip overlay graph for administrating an exchange of data among the processing cores of the multicore processor; and (iii) includes the executable instructions for reformatting computational data from the computation layer for transmission through the network-on-chip layer; and

wherein the network-on-chip overlay graph executes asynchronously from the execution of the executable instructions by the computation layer.

17. A multicore processor stack, stored on non-transitory computer readable media in a multicore processor, comprising:

a computation layer, for conducting computations using a set of processing cores in the multicore processor, with executable instructions for a set of processing pipelines in the set of processing cores;

a network-on-chip layer, for connecting the set of processing cores in the multicore processor, with executable instructions for a set of routers and a set of network interface units in the multicore processor; and

a NoC overlay layer that logically isolates the computation layer and the network-on-chip layer;

wherein the NoC overlay layer is hardware-instantiated via a set of programmable controllers in the multicore processor, wherein each processing core in the set of processing cores has a programmable controller from the set of programmable controllers; and

wherein the set of programmable controllers reformat computational data from the computation layer for transmission through the network-on-chip layer.

18. A multicore processor stack, stored on non-transitory computer readable media in a multicore processor, comprising:

a computation layer, for conducting computations using a set of processing cores in the multicore processor, with executable instructions for a set of processing pipelines in the set of processing cores;

a network-on-chip layer, for connecting the set of processing cores in the multicore processor, with executable instructions for a set of routers and a set of network interface units in the multicore processor;

a NoC overlay layer that logically isolates the computation layer and the network-on-chip layer;

a set of programmable controllers, with executable instructions for reformatting computational data from the computation layer for transmission from a first memory on a first processing core in the multicore processor to a second processing core in the multicore processor, wherein the executable instructions convert the computational data from a first format to a second format;

a first set of kernels on the first processing core for conducting computations using data in the first format; and

a second set of kernels on the second processing core for conducting computations using data in the second format;

wherein the first processing core does not have kernels for conducting computations using data in the second format; and

wherein the second processing core does not have kernels for conducting computations using data in the first format.

Assignments (3)
CHANGE OF NAME Recorded Feb 23, 2025
From: TENSTORRENT INC.
To: TENSTORRENT AI INC.
Reel/Frame 070298/0922 →
CHANGE OF NAME Recorded Feb 23, 2025
From: TENSTORRENT AI INC.
To: TENSTORRENT AI ULC
Reel/Frame 070298/0944 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2022
From: CAPALIJA, DAVOR; MATOSEVIC, IVAN; VASILJEVIC, JASMINA; AYDONAT, UTKU; LEWYCKY, ANDREW; CHIN, S. ALEXANDER; BAJIC, LJUBISA; CEJKOV, ALEX; TRAJKOVIC, MILOS
To: TENSTORRENT INC.
Reel/Frame 061111/0722 →
Continuity (3)
Continuation In Part 16942492 · Jul 29, 2020
Provisional Application 62882065 · Aug 2, 2019
Related Publication 20230041130A1 · Feb 9, 2023
References Cited (27)
US 9658675B1 · Witek · 2017 [cited by examiner]
US 20080181115A1 · Soulie et al. · 2008 [cited by applicant]
US 20090260013A1 · Heil et al. · 2009 [cited by applicant]
US 20100161938A1 · Heddes et al. · 2010 [cited by applicant]
US 20110085550A1 · Lecler et al. · 2011 [cited by applicant]
US 20150103822A1 · Gianchandani et al. · 2015 [cited by applicant]
US 20170153993A1 · Palmer et al. · 2017 [cited by applicant]
US 20170220499A1 · Gray · 2017 [cited by applicant]
US 20190258796A1 · Paczkowski et al. · 2019 [cited by applicant]
US 20200067816A1 · Elsherbini et al. · 2020 [cited by applicant]
CN 101171573A · 2008 [cited by applicant]
CN 106155814B · 2019 [cited by applicant]
EP 2328076A1 · 2011 [cited by applicant]
Summons to Attend Oral Proceedings dated Jul. 8, 2024 from European Application No. 20188876.5, 11 pages. [cited by applicant]
First Office Action from Chinese Application No. 202010765332.4 dated Jan. 18, 2023, 5 pages. [cited by applicant]
First Examination Report from EP Application No. 20188876.5 dated Jan. 25, 2024, 13 pages. [cited by applicant]
Carara et al., “Differentiated Communication Services for NoC-Based MPSoCs”, IEEE Transactions on Computers, IEEE, vol. 63, No. 3, Jun. 5, 2012, pp. 595-608. [cited by applicant]
Extended European Search Report dated Mar. 4, 2021 from European Application No. 20188876.5, 12 pages. [cited by applicant]
Final Office Action dated Nov. 19, 2021 from U.S. Appl. No. 16/942,492, 37 pages. [cited by applicant]
Grzegorz Chmaj et al., “Overlay-NoC and H-Phy based computing using Modern Chip MultiProcessors”, 978-1-4673-0818-2/12 IEEE, 2012, 6 pages. [cited by applicant]
Meloni et al., “System Adaptivity and Fault-Tolerance in NoC-based MPSoCs: The MADNESS Project Approach”, Digital System Design, 2012 15th Euromicro Conference, IEEE, Sep. 5, 2012, pp. 517-524. [cited by applicant]
Nachiket Kapre, “Marathon: Statically-Scheduled Conflict-Free Routing on FPGA Overlay NoCs”, IEEE, 2016, 8 pages. [cited by applicant]
Nonfinal Office Action dated May 26, 2021 from U.S. Appl. No. 16/942,492, 50 pages. [cited by applicant]
Notice of Allowance dated May 18, 2022 from U.S. Appl. No. 16/942,492, 33 pages. [cited by applicant]
Tobias Bjerregaard et al., “A Survey of Research and Practices of Network-on-Chip”, ACM Computing Surveys, ACM, vol. 38, No. 1, Jun. 29, 2006, 51 pages. [cited by applicant]
Venkata Yaswanth Raparti at al., “Memory-Aware Circuit Overlay NoCs for Latency Optimized GPGPU Architectures”, 17th Int'l. Symposium on Quality Electronic Design, 978-1-5090-1213-8/16, IEEE, 2016, 6 pages. [cited by applicant]
Venkata Yaswanth Raparti at al., “DAPPER: Data Aware Approximate NoC for GPGPU Architectures”, IEEE, 2018, 8 pages. [cited by applicant]