IP Library Granted Patent US 12,475,064
Granted Patent B2
US 12,475,064 · App. 18/259,235 · Granted Nov 18, 2025

Network on Chip processing system

Inventor: Stefan Blixt (Bålsta, SE)
Assignee: Telesis Innovation AB
G06F15/7825G06F15/17337
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,064
App. No.
18/259,235
Granted
Nov 18, 2025
Kind
B2
Abstract

A Network on Chip, NoC, processing system configured to perform data processing. The NoC processing system is configured for interconnection with a control processor connectable to said NoC processing system. The NoC processing system comprises a plurality of microcode-programmable Processing Elements, PEs, organized in multiple clusters, each cluster comprising a multitude of said programmable PEs, the functionality of each microcode-programmable PE being defined by internal microcode in a microprogram memory associated with the PE. The clusters of programmable PE are arranged on a Network on Chip, NoC, the NoC having a root and a plurality of peripheral nodes, wherein the clusters of PE are arranged at peripheral nodes of the NoC, and the NoC is connectable to the control processor. Each cluster further includes a Cluster Controller, CC, and an associated Cluster Memory, CM, shared by the multitude of programmable PE within the cluster.

Claims (32)

1 . A Network on Chip, NoC, processing system configured to perform data processing, wherein said NoC processing system is configured for interconnection with a control processor connectable to said NoC processing system, said NoC processing system comprising:

a plurality of microcode-programmable Processing Elements, PEs, organized in multiple clusters, each cluster comprising a multitude of said programmable PEs, the functionality of each microcode-programmable Processing Element being defined by internal microcode in a microprogram memory associated with the Processing Element,

wherein said multiple clusters of programmable Processing Elements are arranged on a Network on Chip, NoC, said Network on Chip having a root and a plurality of peripheral nodes, wherein said multiple clusters of Processing Elements are arranged at peripheral nodes of the Network on Chip, and wherein said Network on Chip is connectable to said control processor at said root;

wherein each cluster further includes a Cluster Controller, CC, and an associated Cluster Memory, CM, shared by the multitude of programmable Processing Elements within the cluster;

wherein the Cluster Memory of each cluster is configured to store at least microcode received from an on-chip or off-chip data source under the control of said connectable control processor, and each Processing Element within a cluster is configured to request microcode from the associated Cluster Memory;

wherein the Cluster Controller of each cluster is configured to enable transfer of microcode from the associated Cluster Memory to at least a subset of the Processing Elements within the cluster in response to requests from the corresponding Processing Elements.

2 . The NoC processing system of claim 1 , wherein each Processing Element, instructed or initiated by said control processor, via said Network on Chip, is configured to retrieve microcode from the associated Cluster Memory by a request to the Custer Controller of the corresponding cluster.

3 . The NoC processing system of claim 1 , wherein said processing system enables distribution of microcode in two independent phases, an initial distribution phase related to the Cluster Memories storing microcode received from an on-chip or off-chip data source under the control of the connectable control processor, and a further request-based distribution phase controlled or at least initiated by individual Processing Elements in each cluster.

4 . The NoC processing system of claim 1 , wherein each Processing Element has a data-from-memory register configured to receive microcode from the Cluster Memory, the microcode being transferrable to the microprogram memory of the Processing Element via said data-from-memory register.

5 . The NoC processing system of claim 4 , wherein each Processing Element is configured to transfer the microcode from the data-from-memory register to the microprogram memory of the Processing Element.

6 . The NoC processing system of claim 1 , wherein said NoC processing system is configured for enabling at least part of the microcode to be broadcast simultaneously to the Cluster Memories of at least a subset of the clusters.

7 . The NoC processing system of claim 1 , wherein said NoC processing system is configured to be controlled by said connectable control processor, which is allowed to control transfer of: microcode, input data, and/or application parameters for a data flow application, to selected locations in said Cluster Memories, and transfer of intermediate data from and to, and/or results from, the Cluster Memories; and

wherein Processing Elements of at least selected clusters are configured to start, as instructed or initiated by said control processor, processing of the input data, using the microcode and application parameters that they have received.

8 . The NoC processing system of claim 7 , wherein said NoC processing system is configured for enabling said control processor to control at least part of the data transfers between the root and the clusters to be performed simultaneously with the execution of arithmetic work by the Processing Elements.

9 . The NoC processing system of claim 1 , wherein each Processing Element includes one or more Multiply-Accumulate (MAC) units for performing MAC operations, and the execution of the MAC operations is controlled by local microcode and/or corresponding control bits originating from the microprogram memory associated with the Processing Element or produced by the execution of microcode in the microprogram memory.

10 . The NoC processing system of claim 1 , wherein said NoC processing system includes a switch block for connection to data channels of the Network on Chip leading to and from the clusters of programmable Processing Elements arranged at peripheral nodes of the Network on Chip, and

wherein said NoC processing system includes a NoC register for temporarily holding data to be distributed to the clusters of Processing Elements via said switch block, and/or for holding data from the clusters of Processing Elements.

11 . The NoC processing system of claim 10 , wherein said switch block includes a multitude of switches and control logic for controlling the multitude of switches to support a set of different transfer modes over said data channels of the Network on Chip.

12 . The NoC processing system of claim 10 , wherein the Network on Chip is a star-shaped network connecting the switch block with the clusters, with a point-to-point channel to each cluster.

13 . The NoC processing system of claim 1 , wherein said NoC processing system includes a control network for transfer of commands and/or status information related to data transfers to be performed over the Network on Chip and/or for controlling the Cluster Controllers, wherein commands or instructions for initiating retrieval and/or execution of microcode by the Processing Elements are transferred via said control network.

14 . The NoC processing system of claim 1 , wherein said Cluster Controller, as controlled by commands from the control processor over the Network on Chip, is configured to keep the Processing Elements of the cluster in reset or hold/wait mode until microcode for the Processing Elements has been stored in the associated Cluster Memory, and then activate the Processing Elements,

wherein each Processing Element, when starting after reset or hold/wait mode, is configured to output a request for microcode to the Cluster Controller.

15 . The NoC processing system of claim 1 , wherein said Cluster Controller includes multiple interfaces including a NoC interface with the Network on Chip, a PE interface with the Processing Elements within the cluster, as well as a CM interface with the Cluster Memory of the cluster.

16 . The NoC processing system of claim 1 , wherein said connectable control processor is an on-chip processor integrated on the same chip as the NoC.

17 . The NoC processing system of claim 1 , wherein said connectable control processor is an external processor connectable for controlling the NoC.

18 . The NoC processing system of claim 17 , wherein said external processor is a computer or server processor, in which at least part of functionality of the control processor is performed by a computer program, when executed by the computer or server processor.

19 . The NoC processing system of claim 1 , wherein said connectable control processor is configured to execute object code representing a data flow model, created by a compiler, for enabling said control processor to control transfer of microcode, input data and application parameters of a data flow application in predefined order to the Processing Elements of the clusters for execution and to receive and deliver end results.

20 . The NoC processing system of claim 19 , wherein the data flow model is a Convolutional Neural Network (CNN) data flow model and/or a Digital Signal Processing (DSP) data flow model, comprising a number of data processing layers, which are executed layer by layer.

21 . The NoC processing system of claim 1 , wherein said Cluster Controller is configured to receive requests, referred to as cluster broadcast requests, from at least a subset of said multiple Processing Elements within the cluster for broadcasting of microcode and/or data from said Cluster Memory to said at least a subset of said multiple Processing Elements; and

wherein said Cluster Controller is configured to initiate said broadcasting of said microcode and/or data in response to the received cluster broadcast requests only after broadcast requests have been received from all Processing Elements of said at least a subset of said multiple Processing Elements that are participating in the broadcast.

22 . The NoC processing system of claim 1 , wherein said NoC processing system is an accelerator system for an overall data processing system.

23 . The NoC processing system of claim 1 , wherein said Network on Chip is organized as a star network or fat tree network, with said root as a central hub and branches, or a hierarchy of branches, reaching out to the plurality of peripheral nodes, also referred to as leaf nodes, where said multiple clusters of Processing Elements are arranged.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 8, 2025
From: IMSYS AB (PUBL)
To: TELESIS INNOVATION AB
Reel/Frame 072504/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2023
From: BLIXT, STEFAN
To: TELESIS INNOVATION AB
Reel/Frame 064434/0410 →
Continuity (2)
Provisional Application 63130089 · Dec 23, 2020
Related Publication 20250307204A1 · Oct 2, 2025
References Cited (48)
US 5287470A · Simpson · 1994 [cited by applicant]
US 5345563A · Uihlein et al. · 1994 [cited by applicant]
US 5890007A · Zinguuzi · 1999 [cited by applicant]
US 6018782A · Hartmann · 2000 [cited by applicant]
US 6145072A · Shams et al. · 2000 [cited by applicant]
US 8060727B2 · Blixt · 2011 [cited by applicant]
US 9553590B1 · Manohararajah · 2017 [cited by examiner]
US 11221929B1 · Katz · 2022 [cited by examiner]
US 20030033490A1 · Gappisch et al. · 2003 [cited by applicant]
US 20060253660A1 · Hall · 2006 [cited by examiner]
US 20070159488A1 · Danskin et al. · 2007 [cited by applicant]
US 20070283037A1 · Burns et al. · 2007 [cited by applicant]
US 20080037650A1 · Stojancic et al. · 2008 [cited by applicant]
US 20080189514A1 · McConnell · 2008 [cited by examiner]
US 20100091787A1 · Muff et al. · 2010 [cited by applicant]
US 20100191814A1 · Heddes et al. · 2010 [cited by applicant]
US 20110307459A1 · Jacob (Yaakov) · 2011 [cited by applicant]
US 20120124324A1 · Park et al. · 2012 [cited by applicant]
US 20120290815A1 · Takahashi · 2012 [cited by applicant]
US 20140156907A1 · Palmer · 2014 [cited by applicant]
US 20140310467A1 · Shalf et al. · 2014 [cited by applicant]
US 20170078385A1 · Dress · 2017 [cited by applicant]
US 20170116153A1 · Takada · 2017 [cited by applicant]
US 20170147513A1 · Hilton et al. · 2017 [cited by applicant]
US 20170153993A1 · Palmer · 2017 [cited by examiner]
US 20170220499A1 · Gray · 2017 [cited by applicant]
US 20170230447A1 · Harsha et al. · 2017 [cited by applicant]
US 20170286329A1 · Fernando · 2017 [cited by applicant]
US 20180232148A1 · Saeed · 2018 [cited by applicant]
US 20180254942A1 · Tocker · 2018 [cited by examiner]
US 20190042245A1 · Toll et al. · 2019 [cited by applicant]
US 20190138237A1 · Palmer · 2019 [cited by applicant]
US 20190158427A1 · Harsha et al. · 2019 [cited by applicant]
US 20190191814A1 · Stuempfig et al. · 2019 [cited by applicant]
US 20190303328A1 · Balski et al. · 2019 [cited by applicant]
US 20200201690A1 · Sankaralingam et al. · 2020 [cited by applicant]
US 20200301865A1 · Davies · 2020 [cited by applicant]
US 20210181957A1 · Diril · 2021 [cited by examiner]
US 20220100601A1 · Baum · 2022 [cited by examiner]
US 20220101043A1 · Katz · 2022 [cited by examiner]
US 20220103186A1 · Kaminitz · 2022 [cited by examiner]
US 20240054081A1 · Blixt · 2024 [cited by examiner]
US 20250252433A1 · Ramilo · 2025 [cited by examiner]
JP 2009104521A · 2009 [cited by applicant]
WO 2017120270A1 · 2017 [cited by applicant]
Hassan, et al. “An Enhanced Network-on-chip Simulation for Cluster-based Routing”, Procedia Computer Science vol. 94, 2016, pp. 410-417. [cited by applicant]
International Search Report and Written Opinion for corresponding Application No. PCT/SE2021/051294, issued on Feb. 18, 2022. [cited by applicant]
Caputa “Efficient High-Speed On-Chip Global Interconnects”, Linköping Studies in Science and Technology, Dissertation No. 992, Linköping University 2006. [cited by applicant]