IP Library › Granted Patent US 12,566,570
Granted Patent B2
US 12,566,570 · App. 17/946,311 · Granted Mar 3, 2026

Write combine buffer (WCB) for deep neural network (DNN) accelerator

Inventors: Martin-Thomas Grymel (Leixlip, IE); David Thomas Bernard (Kilcullen, IE); Martin Power (Dublin, IE); Niall Hanrahan (Galway, IE); Kevin Brady (Newry, GB)
Assignee: Intel Corporation
G06F3/0656G06F3/0604G06F3/0673G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,570
App. No.
17/946,311
Granted
Mar 3, 2026
Kind
B2
Abstract

A compute tile includes a WCB that receives a workload of writing an output tensor of a convolution into a local memory of the compute tile. The local memory may be a SRAM. The WCB receives write transactions. A write transaction includes a data block, which is a part of the output tensor, and metadata describing one or more attributes of the data block. The WCB may store write transactions in its internal buffers. The WCB may determine whether to combine two write transactions, e.g., based on an operation mode or metadata in the write transactions. In embodiments where the WCB determines to combine the two write transactions, the WCB may combine the two write transactions into a new write transaction and write the new write transaction into the local memory or an internal memory of the WCB. The total number of write transactions for the workload can be reduced.

Claims (100)

1 . A method of deep learning, the method comprising:

storing a first write transaction in an internal memory, wherein the first write transaction comprising a first data block, and the first data block is a result of one or more multiply-accumulate (MAC) operations performed by a compute tile for a convolutional layer in a deep neural network (DNN);

storing a second write transaction in a buffer, wherein the second write transaction comprising a second data block, and the second data block is a result of one or more other MAC operations performed by the compute tile for the convolutional layer in the DNN;

determining whether to combine the first write transaction and the second write transaction;

in response to determining to combine the first data block and the second data block, generating a combined write transaction by combining the first write transaction with the second write transaction; and

writing the combined write transaction into a memory at an address in the memory, wherein the memory is inside the compute tile.

2 . The method of claim 1 , wherein the compute tile produces an output tensor of the convolutional layer by performing a convolution with one or more filters, and the output tensor includes the first data block and the second data block.

3 . The method of claim 1 , wherein determining whether to combine the first write transaction and the second write transaction comprises:

receiving an instruction that specifies an operation mode of a write combine buffer;

determining that the operation mode is a bypass mode; and

determining not to combine the first write transaction and the second write transaction.

4 . The method of claim 1 , further comprising:

generating the first write transaction or the second write transaction by combining a third write transaction and a fourth write transaction,

wherein each of the third write transaction and the fourth write transaction comprises a data block that is a result of one or more additional MAC operations performed by the compute tile.

5 . The method of claim 1 , wherein:

the first write transaction further comprises first metadata that specifies one or more attributes of the first data block,

the second write transaction further comprises second metadata that specifies one or more attributes of the second data block, and

determining whether to combine the first write transaction and the second write transaction comprises determining whether to combine the first write transaction and the second write transaction based on the first metadata and the second metadata.

6 . The method of claim 5 , wherein determining whether to combine the first write transaction and the second write transaction comprises:

determining whether the first metadata or the second metadata indicates that all bytes in the first write transaction or the second write transaction are enabled, wherein an enabled byte is to be written into the memory.

7 . The method of claim 5 , wherein:

the first metadata specifies a first memory address for the first data block,

the second metadata specifies a second memory address for the second data block, and

determining whether to combine the first data block and the second data block comprises determining whether the first memory address matches the second memory address.

8 . The method of claim 1 , wherein:

an output tensor of the convolutional layer includes the first data block and the second data block,

the output tensor comprises one or more halo regions,

data in each of the one or more halo regions is to be provided to another array of MAC units for performing further MAC operations, and

determining whether to combine the first data block and the second data block comprises determining whether the first data block and the second data block are in a same halo region of the one or more halo regions.

9 . The method of claim 1 , wherein:

the first write transaction further comprises metadata that specifies one or more attributes of the first data block, and

storing the first write transaction in the internal memory comprises:

determining a memory location for the first write transaction based on the metadata, and

storing the first write transaction at the memory location in the internal memory.

10 . The method of claim 1 , further comprising:

determining that there is no write activity at the memory at a first time; and

writing one or more write transactions stored in the internal memory into the memory at a second time,

wherein there is a predetermined delay between the first time and the second time.

11 . One or more non-transitory computer-readable media storing instructions executable to perform operations for deep learning, the operations comprising:

storing a first write transaction in an internal memory, wherein the first write transaction comprising a first data block, and the first data block is a result of one or more multiply-accumulate (MAC) operations performed by a compute tile for a convolutional layer in a deep neural network (DNN);

storing a second write transaction in a buffer, wherein the second write transaction comprising a second data block, and the second data block is a result of one or more other MAC operations performed by the compute tile for the convolutional layer in the DNN;

determining whether to combine the first write transaction and the second write transaction;

in response to determining to combine the first data block and the second data block, generating a combined write transaction by combining the first write transaction with the second write transaction; and

writing the combined write transaction into a memory at an address in the memory, wherein the memory is inside the compute tile.

12 . The one or more non-transitory computer-readable media of claim 11 , wherein the compute tile produces an output tensor of the convolutional layer by performing a convolution with one or more filters, and the output tensor includes the first data block and the second data block.

13 . The one or more non-transitory computer-readable media of claim 11 , wherein determining whether to combine the first write transaction and the second write transaction comprises:

receiving an instruction that specifies an operation mode of a write combine buffer;

determining that the operation mode is a bypass mode; and

determining not to combine the first write transaction and the second write transaction.

14 . The one or more non-transitory computer-readable media of claim 11 , wherein the operations further comprise:

generating the first write transaction or the second write transaction by combining a third write transaction and a fourth write transaction,

wherein each of the third write transaction and the fourth write transaction comprises a data block that is a result of one or more additional MAC operations performed by the compute tile.

15 . The one or more non-transitory computer-readable media of claim 11 , wherein:

the first write transaction further comprises first metadata that specifies one or more attributes of the first data block,

the second write transaction further comprises second metadata that specifies one or more attributes of the second data block, and

determining whether to combine the first write transaction and the second write transaction comprises determining whether to combine the first write transaction and the second write transaction based on the first metadata and the second metadata.

16 . The one or more non-transitory computer-readable media of claim 15 , wherein determining whether to combine the first write transaction and the second write transaction comprises:

determining whether the first metadata or the second metadata indicates that all bytes in the first write transaction or the second write transaction are enabled, wherein an enabled byte is to be written into the memory.

17 . The one or more non-transitory computer-readable media of claim 15 , wherein:

the first metadata specifies a first memory address for the first data block,

the second metadata specifies a second memory address for the second data block, and

determining whether to combine the first data block and the second data block comprises determining whether the first memory address matches the second memory address.

18 . The one or more non-transitory computer-readable media of claim 11 , wherein:

an output tensor of the convolutional layer includes the first data block and the second data block,

the output tensor comprises one or more halo regions,

data in each of the one or more halo regions is to be provided to another array of MAC units for performing further MAC operations, and

determining whether to combine the first data block and the second data block comprises determining whether the first data block and the second data block are in a same halo region of the one or more halo regions.

19 . The one or more non-transitory computer-readable media of claim 11 , wherein:

the first write transaction further comprises metadata that specifies one or more attributes of the first data block, and

storing the first write transaction in the internal memory comprises:

determining a memory location for the first write transaction based on the metadata, and

storing the first write transaction at the memory location in the internal memory.

20 . The one or more non-transitory computer-readable media of claim 11 , wherein the operations further comprise:

determining that there is no write activity at the memory at a first time; and

writing one or more write transactions stored in the internal memory into the memory at a second time,

wherein there is a predetermined delay between the first time and the second time.

21 . A deep neural network (DNN) accelerator, the DNN accelerator comprising:

an array of multiple-accumulate (MAC) units configured to execute a convolution on an input tensor with a number of filters to produce an output tensor;

a memory; and

a write combine buffer (WCB) that is configured to:

store a first write transaction in an internal memory of the WCB, wherein the first write transaction comprising a first data block, and the first data block is a result of one or more MAC operations performed by the array of MAC units,

store a second write transaction in a buffer of the WCB, wherein the second write transaction comprising a second data block, and the second data block is a result of one or more other MAC operations performed by the array of MAC units,

determine whether to combine the first write transaction and the second write transaction,

in response to determining to combine the first data block and the second data block, generate a combined write transaction by combining the first write transaction with the second write transaction, and

write the combined write transaction into the memory at an address in the memory.

22 . The DNN accelerator of claim 21 , wherein the WCB is configured to determine whether to combine the first write transaction and the second write transaction by:

receiving an instruction that specifies an operation mode of the WCB;

determining that the operation mode is a bypass mode; and

determining not to combine the first write transaction and the second write transaction.

23 . The DNN accelerator of claim 21 , wherein the WCB is further configured to:

generate the first write transaction or the second write transaction by combining a third write transaction and a fourth write transaction,

wherein each of the third write transaction and the fourth write transaction comprises a data block that is a result of one or more additional MAC operations performed by the array of MAC units.

24 . The DNN accelerator of claim 21 , wherein:

the first write transaction further comprises first metadata that specifies one or more attributes of the first data block,

the second write transaction further comprises second metadata that specifies one or more attributes of the second data block, and

the WCB is configured to determine whether to combine the first write transaction and the second write transaction comprises determining whether to combine the first write transaction and the second write transaction based on the first metadata and the second metadata.

25 . The DNN accelerator of claim 21 , wherein the WCB is further configured to:

determine that there is no write activity at the memory at a first time; and

write one or more write transactions stored in the internal memory into the memory at a second time,

wherein there is a predetermined delay between the first time and the second time.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2022
From: GRYMEL, MARTIN-THOMAS; BERNARD, DAVID THOMAS; POWER, MARTIN; HANRAHAN, NIALL; BRADY, KEVIN
To: INTEL CORPORATION
Reel/Frame 061720/0647 →
Continuity (1)
Related Publication 20230020929A1 · Jan 19, 2023
References Cited (14)
US 9710265B1 · Temam · 2017 [cited by examiner]
US 11514291B2 · Baum · 2022 [cited by examiner]
US 11531873B2 · Boesch · 2022 [cited by examiner]
US 11934824B2 · Yudanov · 2024 [cited by examiner]
US 20200110983A1 · Chen · 2020 [cited by examiner]
US 20220180187A1 · Kim · 2022 [cited by examiner]
US 20220188611A1 · Hazanchuk · 2022 [cited by examiner]
Sze, Vivienne et al Efficient processing of deep neural networks. Synthesis Lectures on Computer Architecture. Jun. 24, 2020;15(2):1-341, retrieved from https://people.csail.mit.edu on Sep. 16, 2022. [cited by applicant]
Williams, Samuel et al “Roofline: An insightful visual performance model for floating-point programs and multicore architectures”, Commun. ACM 52 4 (Apr. 2009, 65-76 (https://doi.org/10.1145/1498765.1498785). [cited by applicant]
Chakrabarti et al., Scaling Intel(R) Software Guard Extensions Applications with Intel(R) SGX Card, (c) 2019 978-1-4503-7226—Aug. 19, 2006, 9 pages. [cited by applicant]
Shawahna et al., FPGA-Based Accelerators of Deep Learning Networks for Learning and Classification: A Review, (c) 2018 IEEE Access, 37 pages. [cited by applicant]
Shen et al., Escher: A CNN Accelerator with Flexible Buffering to Minimize Off-Chip Transfer, 2017 IEEE 25th Annual International Symposium on Field-Programmable Custom Computing Machines, 8 pages. [cited by applicant]
Zheng et al., Hardware Architecture Exploration for Deep Neural Networks, Arabian Journal for Science and Engineering (2021) 46:9703-9712, 10 pages. [cited by applicant]
Qiu et al., Going Deeper with Embedded FPGA Platform for Convolutional Neural Network, (c) 2016 ACM, 978-1-4503-3856—Jan. 16, 2002, 10 pages. [cited by applicant]