IP Library › Granted Patent US 12,455,851
Granted Patent B1
US 12,455,851 · App. 18/906,697 · Granted Oct 28, 2025

Adaptive buffer sharing in multi-core reconfigurable streaming-based architectures

Inventors: Michele Rossi (Bareggio, IT); Thomas Boesch (Rovio, CH); Giuseppe Desoli (San Fermo Della Battaglia, IT)
Assignee: STMicroelectronics International N.V.
G06F15/82G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,455,851
App. No.
18/906,697
Granted
Oct 28, 2025
Kind
B1
Abstract

A hardware accelerator includes a plurality of functional circuits, a stream switch, and a plurality of neural network processing cores coupled to the plurality of functional circuits via the stream switch to stream data to and from functional circuits of the plurality of functional circuits, wherein the neural network processing cores include at least one sender core having a buffer whose buffer content is sharable with at least one receiver core of the neural network processing cores via at least one of a dedicated output stream link or dedicated virtual channel on a pre-existing output stream link, and wherein the at least one of the dedicated output stream link or dedicated virtual channel is dedicated to sharing the buffer content via the stream switch.

Claims (30)

1 . A hardware accelerator, comprising:

a plurality of functional circuits;

a stream switch; and

a plurality of neural network processing cores coupled to the plurality of functional circuits via the stream switch to stream data to and from functional circuits of the plurality of functional circuits, wherein the neural network processing cores include at least one sender core having a buffer whose buffer content is sharable with at least one receiver core of the neural network processing cores via at least one of a dedicated output stream link or dedicated virtual channel on a pre-existing output stream link, and wherein the at least one of the dedicated output stream link or dedicated virtual channel is dedicated to sharing the buffer content via the stream switch.

2 . The hardware accelerator of claim 1 , wherein buffer access patterns of the at least one sender core are compatible with buffer access patterns of the at least one receiver core.

3 . The hardware accelerator of claim 1 , wherein the at least one sender core and the at least one receiver core are programmed with dedicated configuration registers to configure the sharing of the buffer content.

4 . The hardware accelerator of claim 1 , wherein the at least one sender core and the at least one receiver core are implemented with In-Memory Computing (IMC) devices.

5 . The hardware accelerator of claim 4 , wherein the at least one sender core has at least one IMC device as the buffer locked in memory mode and has other IMC devices operating in compute mode.

6 . The hardware accelerator of claim 4 , wherein the at least one receiver core has all IMC devices operating in compute mode.

7 . The hardware accelerator of claim 1 , wherein the at least one sender core uses a first subset of the buffer content for computation and the at least one receiver core uses a second subset of the buffer content for computation.

8 . The hardware accelerator of claim 1 , wherein the hardware accelerator is a neural processing unit (NPU).

9 . A system, comprising:

a host device; and

a hardware accelerator, the hardware accelerator including:

a plurality of functional circuits;

a stream switch; and

a plurality of neural network processing cores coupled to the plurality of functional circuits via the stream switch to stream data to and from functional circuits of the plurality of functional circuits, wherein the neural network processing cores include at least one sender core having a buffer whose buffer content is sharable with at least one receiver core of the neural network processing cores via at least one of a dedicated output stream link or dedicated virtual channel on a pre-existing output stream link, and wherein the at least one of the dedicated output stream link or dedicated virtual channel is dedicated to sharing the buffer content via the stream switch.

10 . The system of claim 9 , wherein buffer access patterns of the at least one sender core are compatible with buffer access patterns of the at least one receiver core.

11 . The system of claim 9 , wherein the at least one sender core and the at least one receiver core are programmed with dedicated configuration registers to configure the sharing of the buffer content.

12 . The system of claim 9 , wherein the at least one sender core and the at least one receiver core are implemented with In-Memory Computing (IMC) devices.

13 . The system of claim 12 , wherein the at least one sender core has at least one IMC device as the buffer locked in memory mode and has other IMC devices operating in compute mode.

14 . The system of claim 12 , wherein the at least one receiver core has all IMC devices operating in compute mode.

15 . The system of claim 9 , wherein the at least one sender core uses a first subset of the buffer content for computation and the at least one receiver core uses a second subset of the buffer content for computation.

16 . The system of claim 9 , wherein the hardware accelerator is a neural processing unit (NPU).

17 . A method, comprising:

streaming data between a plurality of neural network processing cores of a hardware accelerator and a plurality of functional circuits of the hardware accelerator via a stream switch,

wherein the neural network processing cores include at least one sender core having a buffer whose buffer content is sharable with at least one receiver core of the neural network processing cores via at least one of a dedicated output stream link or dedicated virtual channel on a pre-existing output stream link, and wherein the at least one of the dedicated output stream link or dedicated virtual channel is dedicated to sharing the buffer content via the stream switch.

18 . The method of claim 17 , wherein buffer access patterns of the at least one sender core are compatible with buffer access patterns of the at least one receiver core.

19 . The method of claim 17 , wherein the at least one sender core and the at least one receiver core are programmed with dedicated configuration registers to configure the sharing of the buffer content.

20 . The method of claim 17 , wherein the at least one sender core uses a first subset of the buffer content for computation and the at least one receiver core uses a second subset of the buffer content for computation.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2024
From: ROSSI, MICHELE; DESOLI, GIUSEPPE
To: STMICROELECTRONICS S.R.L.
Reel/Frame 069193/0345 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2024
From: BOESCH, THOMAS
To: STMICROELECTRONICS INTERNATIONAL N.V.
Reel/Frame 068953/0877 →
References Cited (13)
US 8327187B1 · Metcalf · 2012 [cited by examiner]
US 11119765B2 · Ahmed · 2021 [cited by examiner]
US 20110314255A1 · Krishna · 2011 [cited by examiner]
US 20190303295A1 · Steinmacher-Burow · 2019 [cited by examiner]
US 20200004690A1 · Mathew · 2020 [cited by examiner]
US 20210097047A1 · Billa et al. · 2021 [cited by applicant]
US 20210241806A1 · Chawla · 2021 [cited by examiner]
US 20220101086A1 · Cappetta et al. · 2022 [cited by applicant]
US 20220101108A1 · Akopyan · 2022 [cited by examiner]
US 20220318609A1 · Fick · 2022 [cited by examiner]
US 20230058989A1 · Hammarlund et al. · 2023 [cited by applicant]
US 20230131698A1 · Noguera · 2023 [cited by examiner]
US 20240220777A1 · Girardi et al. · 2024 [cited by applicant]