IP Library Granted Patent US 12,730,769
Granted Patent B2
US 12,730,769 · App. 18/192,631 · Granted Sep 8, 2026

Reconfigurable, streaming-based clusters of processing elements, and multi-modal use thereof

Inventors: Michele Rossi (Bareggio, IT); Giuseppe Desoli (San Fermo Della Battaglia, IT); Thomas Boesch (Rovio, CH)
Assignee: STMicroelectronics International N.V.
G06F13/4022G06F13/1668
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,769
App. No.
18/192,631
Granted
Sep 8, 2026
Kind
B2
Abstract

A hardware accelerator includes processing elements of a neural network, each processing element having a memory; a stream switch; stream engines coupled to functional circuits via the stream switch, wherein the stream engines, in operation, generate data streaming requests to stream data to and from functional circuits of the plurality of functional circuits; a first system bus interface coupled to the stream engines; a second system bus interface coupled to the processing elements; and mode control circuitry, which, in operation, sets respective modes of operation for the plurality of processing elements. The modes of operation include: a compute mode of operation in which the processing element performs computing operations using the memory associated with the processing element; and a memory mode of operation in which the memory associated with the processing element performs memory operations, bypassing the stream switch, via the second system bus interface.

Claims (60)

1 . A hardware accelerator, comprising:

a processing cluster including a crossbar switch interconnecting a plurality of processing elements within the processing cluster, each processing element of the interconnected processing elements having a respective internal memory within the processing element, the processing cluster coupled to a stream switch;

the stream switch;

a plurality of stream engines coupled to a plurality of functional circuits via the stream switch, wherein the plurality of stream engines, in operation, generate data streaming requests to stream data to and from functional circuits of the plurality of functional circuits;

a first system bus interface coupled to the plurality of stream engines;

a second system bus interface coupled to each processing element of the interconnected processing elements, independent from the stream switch and the crossbar switch; and

mode control circuitry, which, in operation, sets respective modes of operation for individual processing elements of the interconnected processing elements of the processing cluster, wherein the modes of operation for an individual processing element include:

a compute mode of operation in which the processing element performs computing operations using the respective internal memory within the processing element; and

a memory mode of operation in which the respective internal memory within the processing element performs memory operations, bypassing the stream switch, via the second system bus interface.

2 . The hardware accelerator of claim 1 , wherein the plurality of processing elements includes a plurality of processing circuits and a memory associated with the plurality of processing circuits.

3 . The hardware accelerator of claim 1 , wherein at least one processing element of the plurality of processing elements comprises one or more In-Memory Computing (IMC) elements.

4 . The hardware accelerator of claim 1 , wherein the plurality of processing elements form one or more clusters of processing elements, a cluster including a reconfigurable crossbar switch, wherein the reconfigurable crossbar switch is coupled to the stream switch, and the reconfigurable crossbar switch, in operation, streams data to, from and between processing elements of the cluster.

5 . The hardware accelerator of claim 4 , wherein the at least one of the one or more clusters comprises a reconfigurable memory network, wherein the memory network is coupled to memories of the plurality of processing elements, and the memory network, in operation, transfers data to, from, and between processing elements of the processing cluster.

6 . The hardware accelerator of claim 5 , comprising:

configuration registers, which, in operation, store configuration information for configuring the memory network, the configuration information being based on modes of operation associated with individual processing elements of the plurality of processing elements.

7 . The hardware accelerator of claim 1 , comprising:

a broadcast network, wherein the broadcast network is coupled to the stream switch, and the broadcast network, in operation, streams data from the stream switch to the processing elements.

8 . The hardware accelerator of claim 1 , wherein the mode control circuitry comprises one or more configuration registers.

9 . The hardware accelerator of claim 1 , wherein the mode control circuitry comprises respective configuration registers embedded in the processing elements.

10 . The hardware accelerator of claim 1 , wherein, in the memory mode of operation, the respective internal memory within the processing element stores at least one of: feature data, kernel data, or partial sum data associated with a convolutional operation.

11 . A system, comprising:

a host device; and

a hardware accelerator, the hardware accelerator including:

a processing cluster including a crossbar switch interconnecting a plurality of processing elements within the processing cluster, each processing element of the interconnected processing elements having a respective internal memory within the processing element, the processing cluster coupled to a stream switch;

the stream switch;

a plurality of stream engines coupled to a plurality of functional circuits via the stream switch, wherein the plurality of stream engines, in operation, generate data streaming requests to stream data to and from functional circuits of the plurality of functional circuits;

a first system bus interface coupled to the plurality of stream engines;

a second system bus interface coupled to each processing element of the interconnected processing elements, independent from the stream switch and the crossbar switch; and

mode control circuitry, which, in operation, sets respective modes of operation for individual processing elements of the interconnected processing elements, wherein the modes of operation for an individual processing element include:

a compute mode of operation in which the processing element performs computing operations using the respective internal memory within the processing element; and

a memory mode of operation in which the respective internal memory within the processing element performs memory operations, bypassing the stream switch, via the second system bus interface.

12 . The system of claim 11 , wherein the plurality of processing elements includes a plurality of processing circuits and a memory associated with the plurality of processing circuits.

13 . The system of claim 11 , wherein at least one processing element of the plurality of processing elements comprises one or more In-Memory Computing (IMC) elements.

14 . The system of claim 11 , wherein the plurality of processing elements form one or more clusters of processing elements, a cluster including a reconfigurable crossbar switch, wherein the reconfigurable crossbar switch is coupled to the stream switch, and the reconfigurable crossbar switch, in operation, streams data to, from and between processing elements of the cluster.

15 . The system of claim 14 , wherein the at least one of the one or more clusters comprises a reconfigurable memory network, wherein the memory network is coupled to memories of the plurality of processing elements, and the memory network, in operation, transfers data to, from, and between processing elements of the processing cluster.

16 . The system of claim 15 , comprising:

configuration registers, which, in operation, store configuration information for configuring the memory network, the configuration information being based on modes of operation associated with individual processing elements of the plurality of processing elements.

17 . The system of claim 11 , comprising:

a broadcast network, wherein the broadcast network is coupled to the stream switch, and the broadcast network, in operation, streams data from the stream switch to the processing elements.

18 . The system of claim 11 , wherein the mode control circuitry comprises one or more configuration registers.

19 . The system of claim 11 , wherein the mode control circuitry comprises respective configuration registers embedded in the processing elements.

20 . The system of claim 11 , wherein, in the memory mode of operation, the respective internal memory within the processing element stores at least one of: feature data, kernel data, or partial sum data associated with a convolutional operation.

21 . A method, comprising:

streaming data between stream engines of a plurality of stream engines of a hardware accelerator and functional circuits of a plurality of functional circuits of the hardware accelerator via a stream switch, wherein the plurality of functional circuits includes at least one cluster including a crossbar switch interconnecting a plurality of processing elements within the at least one cluster, wherein the at least one cluster is coupled to the stream switch, and wherein a dedicated system bus interface is coupled to each processing element of the plurality of processing elements, independent from the stream switch and the crossbar switch; and

setting respective modes of operation for individual processing elements of the interconnected processing elements, wherein the modes of operation for an individual processing element include:

a compute mode of operation in which the processing element performs computing operations using a respective internal memory within the processing element; and

a memory mode of operation in which the respective internal memory within the processing element performs memory operations, via the dedicated system bus interface that bypasses the stream switch.

22 . The method of claim 21 , comprising streaming data to, from and between processing elements of the cluster via a reconfigurable crossbar switch, wherein the reconfigurable crossbar switch is coupled to the stream switch.

23 . The method of claim 21 , comprising transferring data to, from, and between processing elements of the cluster via a reconfigurable memory network, wherein the reconfigurable memory network is coupled to memories of the plurality of processing elements.

24 . The method of claim 23 , comprising storing configuration information in configuration registers, wherein the configuration information operates to configure the memory network based on modes of operation associated with individual processing elements of the plurality of processing elements.

25 . The method of claim 21 , comprising storing at least one of: feature data, kernel data, or partial sum data associated with a convolutional operation in a respective internal memory within a processing unit operating in the memory mode of operation.

26 . A non-transitory computer-readable medium comprising instructions which, when executed by a computer, cause the computer to:

configure a stream switch to stream data between stream engines of a plurality of stream engines of a hardware accelerator and functional circuits of a plurality of functional circuits of the hardware accelerator, wherein the plurality of functional circuits includes at least one cluster including a crossbar switch interconnecting a plurality of processing elements within the at least one cluster, wherein the at least one cluster is coupled to the stream switch, and wherein a dedicated system bus interface is coupled to each processing element of the plurality of processing elements, independent from the stream switch and the crossbar switch; and

set respective modes of operation for individual processing elements of the interconnected processing elements, wherein the modes of operation for an individual processing element include:

a compute mode of operation in which the processing element performs computing operations using a respective internal memory within the processing element; and

a memory mode of operation in which the respective internal memory within the processing element performs memory operations, via the dedicated system bus interface that the stream switch.

27 . The non-transitory computer-readable medium of claim 26 , wherein the contents configure a crossbar switch to stream data to, from and between processing elements of the cluster, wherein the crossbar switch is coupled to the stream switch.

28 . The non-transitory computer-readable medium of claim 26 , wherein the contents configure a memory network to transfer data to, from, and between processing elements of the cluster, wherein the memory network is coupled to memories of the plurality of processing elements.

29 . The non-transitory computer-readable medium of claim 28 , wherein the contents include configuration information the operates to configure the memory network based on modes of operation associated with individual processing elements of the plurality of processing elements.

30 . The non-transitory computer-readable medium of claim 26 , wherein the contents configure a processing unit, in the memory mode of operation, to store at least one of: feature data, kernel data, or partial sum data associated with a convolutional operation.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2024
From: STMICROELECTRONICS S.R.L.
To: STMICROELECTRONICS INTERNATIONAL N.V.
Reel/Frame 069180/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2023
From: BOESCH, THOMAS
To: STMICROELECTRONICS INTERNATIONAL N.V.
Reel/Frame 064380/0971 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2023
From: ROSSI, MICHELE; DESOLI, GIUSEPPE
To: STMICROELECTRONICS S.R.L.
Reel/Frame 064380/0974 →
Continuity (2)
Provisional Application 63485669 · Feb 17, 2023
Related Publication 20240281397A1 · Aug 22, 2024
References Cited (25)
US 11100193B2 · Gu et al. · 2021 [cited by applicant]
US 11263077B1 · Seznayov et al. · 2022 [cited by applicant]
US 11323391B1 · McColgan et al. · 2022 [cited by applicant]
US 11562115B2 · Boesch et al. · 2023 [cited by applicant]
US 20080007928A1 · Salama · 2008 [cited by examiner]
US 20180189229A1 · Desoli et al. · 2018 [cited by applicant]
US 20180189641A1 · Boesch et al. · 2018 [cited by applicant]
US 20180189642A1 · Boesch et al. · 2018 [cited by applicant]
US 20180218275A1 · Arrigoni et al. · 2018 [cited by applicant]
US 20180341495A1 · Culurciello · 2018 [cited by examiner]
US 20190266479A1 · Singh · 2019 [cited by examiner]
US 20200272779A1 · Boesch et al. · 2020 [cited by applicant]
US 20200294575A1 · O et al. · 2020 [cited by applicant]
US 20210073450A1 · Boesch et al. · 2021 [cited by applicant]
US 20210150318A1 · Kwak · 2021 [cited by applicant]
US 20210402898A1 · Alvarez · 2021 [cited by examiner]
US 20220101086A1 · Cappetta et al. · 2022 [cited by applicant]
US 20220107911A1 · Fleming et al. · 2022 [cited by applicant]
US 20220156566A1 · Jeon et al. · 2022 [cited by applicant]
US 20220318610A1 · Seo et al. · 2022 [cited by applicant]
US 20220343144A1 · Liang et al. · 2022 [cited by applicant]
US 20240281646A1 · Rossi et al. · 2024 [cited by applicant]
Chi et al., “PRIME: A Novel Processing-in-memory Architecture for Neural Network Computation in ReRAM-based Main Memory,” 2016 ACM/IEEE 43rd Annual International Symposium on Computer Architecture, pp. 27-39. [cited by applicant]
Huawei Technologies Co., Ltd., “CANN 3.3.0—Software Installation Guide (ascend-deployer),” Issue 01, Aug. 4, 2021, 166 pages. [cited by applicant]
U.S. Appl. No. 18/192,629, filed Mar. 29, 2023. [cited by applicant]