IP Library › Granted Patent US 12,645,924
Granted Patent B2
US 12,645,924 · App. 16/973,087 · Granted Jun 2, 2026

Hardware circuit for accelerating neural network computations

Inventors: Ravi Narayanaswami (San Jose, CA); Dong Hyuk Woo (San Jose, CA); Suyog Gupta (Sunnyvale, CA); Uday Kumar Dasari (Union City, CA)
Assignee: Google LLC
G06N3/063G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,924
App. No.
16/973,087
Filed
Dec 8, 2020
Granted
Jun 2, 2026
Kind
B2
Art Unit
2128
USPC
706/41
Abstract

Methods, systems, and apparatus, including computer-readable media, are described for a hardware circuit configured to implement a neural network. The circuit includes multiple super tiles. Each super tile includes a unified memory for storing inputs to a neural network layer and weights for the layer. Each super tile includes multiple compute tiles. Each compute tile executes a compute thread that is used to perform the computations to generate an output for the neural network layer. Each super tile includes arbitration logic coupled to the unified memory and each compute tile. The arbitration logic is configured to: pass inputs stored in the unified memory to the compute tiles; pass weights stored in the unified memory to the compute tiles; and pass, to the unified memory, the output generated for the layer based on computations performed at the compute tiles using the inputs and the weights for the layer.

Claims (97)

1 . A circuit for a hardware accelerator configured to implement a neural network comprising a plurality of neural network layers and to perform computations to generate an output for a neural network layer, the circuit comprising:

a plurality of super tiles for executing computations related to the plurality of neural network layers, each super tile of the plurality of super tiles comprising:

a plurality of compute tiles, wherein each compute tile is configured to execute a compute thread used to perform the computations to generate the output;

a unified memory coupled to each of the plurality of compute tiles and configured to store inputs to the neural network layer and a plurality of weights for the neural network layer; and

an arbitration logic unit coupled to the unified memory and each compute tile of the plurality of compute tiles, wherein the arbitration logic unit is configured to:

pass one or more of the inputs stored in the unified memory to each of the compute tiles;

pass a respective set of weights stored in the unified memory to each of the compute tiles; and

pass, to the unified memory, the output generated for the neural network layer based on computations performed at each of the compute tiles using one or more of the inputs and the respective set of weights,

wherein each super tile of the plurality of super tiles is configured to:

process sub-partitions of an input tensor on two or more compute threads to generate at least a portion of the output, and

provide from the unified memory to the two or more compute threads at least one portion of the input tensor that is shared for processing between the two or more compute threads.

2 . The circuit of claim 1 , comprising a respective controller for each super tile, the respective controller being configured to generate one or more control signals that are used to:

store each of the inputs to the neural network layer in a corresponding location of the unified memory, each of the corresponding locations being identified by a respective address;

store each weight of the plurality of weights for the neural network layer in a corresponding location of the unified memory, each of the corresponding locations being identified by a respective address; and

cause the arbitration logic to pass one or more inputs to a compute cell of a particular compute tile and pass a respective set of weights to the particular compute tile.

3 . The circuit of claim 1 , wherein the controller is configured to:

determine a partitioning of addresses in the unified memory for storing respective batches of inputs to be passed to a corresponding compute tile of a super tile, wherein each partition of addresses is assigned to a respective compute tile of the super tile.

4 . The circuit of claim 3 , wherein:

a respective address in a partition of addresses corresponds to an input in a batch of inputs that form a sample of input features;

the sample of input features comprises multiple sets of input features; and

the sets of input features correspond to images or streams of audio data.

5 . The circuit of claim 3 , wherein the arbitration logic unit is configured to:

obtain, for a first partition of addresses, a first batch of inputs from memory locations identified by addresses in the partition of addresses; and

pass the first batch of inputs to cells of a first compute tile, wherein the first compute tile is assigned to receive each input in the first batch of inputs based on the determined partitioning of addresses in the unified memory.

6 . The circuit of claim 1 , wherein for each respective super tile:

each compute tile of the plurality of compute tiles is configured to execute two or more compute threads in parallel at the compute tile; and

each compute tile executes a compute thread to perform multiplications between one or more inputs to the neural network layer and a weight for the neural network layer to generate a partial output for the neural network layer.

7 . The circuit of claim 6 , wherein for each respective super tile:

each compute tile of the plurality of compute tiles is configured to perform a portion of the computations to generate the output for the neural network layer in response to executing two or more compute threads in parallel at the compute tile; and

in response to performing the portion of the computations, generate one or more partial outputs that are used to generate the output for the neural network layer.

8 . The circuit of claim 1 , wherein:

a first portion of operations that are performed using the compute thread corresponds to a first set of tensor operations for traversing one or more dimensions of a first multi-dimensional tensor; and

the first multi-dimensional tensor is an input tensor comprising data elements corresponding to the inputs stored in the unified memory.

9 . The circuit of claim 8 , wherein:

a second portion of operations that are performed using the compute thread corresponds to a second set of tensor operations for traversing one or more dimensions of a second multi-dimensional tensor that is different than the first multi-dimensional tensor; and

the second multi-dimensional tensor is a weight tensor comprising data elements corresponding to the plurality of weights stored in the unified memory.

10 . The circuit if claim 1 , wherein each super tile of the plurality of super tiles is configured to:

partition the input tensor into sub-partitions, wherein partitioning is carried out along X-Y dimensions of the input tensor.

11 . The circuit of claim 1 , wherein the input tensor corresponds to an image, and wherein the at least one portion of the input tensor that is shared for processing between two or more compute threads correspond to at least one portion of the image shared for processing on the two or more compute threads.

12 . The circuit of claim 1 , wherein each super tile of the plurality of super tiles is configured to determine the at least one portion of the input tensor based on a size of the input tensor being greater than that which can be fully stored in a memory of the compute tile of the plurality of compute tiles.

13 . A method for performing computations to generate an output for a neural network layer of a neural network comprising a plurality of neural network layers using a circuit for a hardware accelerator configured to implement the neural network, the method comprising:

receiving, at a super tile of a plurality of super tiles configured to execute computations related to the plurality of neural networks, inputs to the neural network layer and a plurality of weights for the neural network layer;

storing, in a unified memory of the super tile, the inputs to the neural network layer and the plurality of weights for the neural network layer, wherein the unified memory is coupled to each of a plurality of compute tiles included in the super tile;

passing, using an arbitration logic unit of the super tile, one or more of the inputs stored in the unified memory to each compute tile of the plurality of compute tiles in the super tile, wherein the arbitration logic unit is coupled to the unified memory and each compute tile of the plurality of compute tiles;

passing, using the arbitration logic unit of the super tile, a respective set of weights stored in the unified memory to each of the compute tiles;

executing a compute thread at each of the compute tiles in the super tile to perform the computations to generate the output for the neural network layer;

generating the output for the neural network layer based on computations performed using one or more of the inputs and the respective set of weights at each of the compute tiles;

generating at least a portion of the output based on processing of sub-partitions of an input tensor on two or more compute threads; and

providing, from the unified memory to the two or more compute threads, at least one portion of the input tensor that is shared for processing between the two or more compute threads.

14 . The method of claim 13 , comprising:

passing, using the arbitration logic unit and to the unified memory, the output generated for the neural network layer; and

passing, using a respective controller of the super tile, the output generated for the neural network layer to another super tile at the circuit.

15 . The method of claim 14 , comprising:

generating control signals by the respective controller of the super tile;

storing, based on the control signals, each of the inputs to the neural network layer in a corresponding location of the unified memory, each of the corresponding locations being identified by a respective address;

storing, based on the control signals, each weight of the plurality of weights for the neural network layer in a corresponding location of the unified memory, each of the corresponding locations being identified by a respective address; and

causing, based on the control signals, the arbitration logic to pass one or more inputs to a compute cell of a particular compute tile and pass a respective set of weights to the particular compute tile.

16 . The method of claim 15 , comprising:

determining, by the controller, a partitioning of addresses in the unified memory for storing respective batches of inputs to be passed to a corresponding compute tile of a super tile, wherein each partition of addresses is assigned to a respective compute tile of the super tile.

17 . The method of claim 16 , wherein:

a respective address in a partition of addresses corresponds to an input in a batch of inputs that form a sample of input features;

the sample of input features comprises multiple sets of input features; and

the sets of input features correspond to images or streams of audio data.

18 . The method of claim 16 , comprising, for a first partition of addresses:

obtaining, by the arbitration logic unit, a first batch of inputs from memory locations identified by addresses in the partition of addresses; and

passing the first batch of inputs to cells of a first compute tile, wherein the first compute tile is assigned to receive each input in the first batch of inputs based on the determined partitioning of addresses in the unified memory.

19 . The method of claim 13 , comprising, for each respective super tile:

executing two or more compute threads in parallel at each compute tile of the plurality of compute tiles; and

wherein each compute tile executes a compute thread to perform multiplications between one or more inputs to the neural network layer and a weight for the neural network layer to generate a partial output for the neural network layer.

20 . The method of claim 19 , comprising:

for each respective super tile:

performing, at each compute tile of the plurality of compute tiles, a portion of the computations to generate the output for the neural network layer in response to executing the two or more compute threads in parallel at the compute tile; and

in response to performing the portion of the computations, generating one or more partial outputs that are used to generate the output for the neural network layer.

21 . The method of claim 13 , wherein:

a first portion of operations that are performed using the compute thread corresponds to a first set of tensor operations for traversing one or more dimensions of a first multi-dimensional tensor; and

the first multi-dimensional tensor is an input tensor comprising data elements corresponding to the inputs stored in the unified memory.

22 . The method of claim 21 , wherein:

a second portion of operations that are performed using the compute thread corresponds to a second set of tensor operations for traversing one or more dimensions of a second multi-dimensional tensor that is different than the first multi-dimensional tensor; and

the second multi-dimensional tensor is a weight tensor comprising data elements corresponding to the plurality of weights stored in the unified memory.

23 . The method of claim 13 , comprising:

partitioning the input tensor into sub-partitions of the input tensor, wherein the partitioning is carried out along X-Y dimensions of the input tensor.

24 . The method of claim 13 , wherein the input tensor corresponds to an image, and wherein the at least one portion of the input tensor that is shared for processing between two or more compute threads correspond to at least one portion of the image shared for processing on the two or more compute threads.

25 . The method of claim 13 , wherein the input tensor corresponds to an image, and wherein the at least one portion of the input tensor that is shared for processing between two or more compute threads correspond to at least one portion of the image shared for processing on the two or more compute threads.

26 . A system-on-chip comprising:

a circuit for a hardware accelerator configured to implement a neural network comprising a plurality of neural network layers and to perform computations to generate an output for a neural network layer;

a host controller configured to access memory that is external to the circuit for the hardware accelerator, wherein the memory is configured to store data for processing at the neural network layer;

a host interface configured to exchange data communications between the circuit for the hardware accelerator and the host controller; and

a plurality of super tiles, for executing computations related to the plurality of neural network layers, disposed in the circuit, each super tile of the plurality of super tiles comprising:

a plurality of compute tiles, where in each compute tile is configured to execute a compute thread used to perform the computations to generate the output;

a unified memory coupled to each of the plurality of compute tiles and configured to store inputs to the neural network layer and a plurality of weights for the neural network layer, wherein the inputs and the plurality of weights correspond to the data stored in the memory accessible by the host controller; and

an arbitration logic unit coupled to the unified memory and each compute tile of the plurality of compute tiles, wherein the arbitration logic unit is configured to:

pass one or more of the inputs stored in the unified memory to each of the compute tiles;

pass a respective set of weights stored in the unified memory to each of the compute tiles;

pass, to the unified memory, the output generated for the neural network layer based on computations performed at each of the compute tiles using one or more of the inputs and the respective set of weights,

wherein each super tile of the plurality of super tiles is configured to:

process sub-partitions of an input tensor on two or more compute threads to generate at least a portion of the output, and

provide from the unified memory to the two or more compute threads at least one portion of the input tensor that is shared for processing between the two or more compute threads.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 3, 2021
From: NARAYANASWAMI, RAVI; WOO, DONG HYUK; GUPTA, SUYOG; DASARI, UDAY KUMAR
To: GOOGLE LLC
Reel/Frame 055132/0465 →
Continuity (1)
Related Publication 20210326683A1 · Oct 21, 2021
References Cited (36)
US 9269123B2 · Schneider · 2016 [cited by applicant]
US 20180046900A1 · Dally · 2018 [cited by examiner]
US 20180285254A1 · Baum et al. · 2018 [cited by applicant]
US 20190205746A1 · Nurvitadhi et al. · 2019 [cited by applicant]
US 20190220742A1 · Kuo · 2019 [cited by examiner]
US 20190370631A1 · Fais · 2019 [cited by examiner]
US 20200192720A1 · Liu · 2020 [cited by examiner]
US 20200409664A1 · Li · 2020 [cited by examiner]
US 20210125042A1 · Han · 2021 [cited by examiner]
CN 109190758A · 2019 [cited by applicant]
KR 1020190066058 · 2019 [cited by applicant]
TW 201911140A · 2019 [cited by applicant]
WO WO2019212688 · 2019 [cited by applicant]
Cadambi et al., “A Programmable Parallel Accelerator for Learning and Classification,” PACT '10, Sep. 11-15, 2010, pp. 273-283. (Year: 2010). [cited by examiner]
Song et al., “HyPar: Towards Hybrid Parallelism for Deep Learning Accelerator Array,” 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA), 2019, pp. 56-68. (Year: 2019). [cited by examiner]
Shao et al., “Simba: Scaling Deep-Learning Inference with Multi-Chip-Module-Based Architecture,” Micro-52, Oct. 12-16, 2019, pp. 14-27. (Year: 2019). [cited by examiner]
Shao et al. (Year: 2019). [cited by examiner]
Cadambi et al. (Year: 2010). [cited by examiner]
Song et al. (Year: 2019). [cited by examiner]
Lasalle et al., “A Scalable Tile-Based Framework for Region-Merging Segmentation,” May 4, 2015, IEEE Transactions on Geoscience and Remote Sensing (vol. 53, Issue: 10, Oct. 2015), pp. 5473-5485. (Year: 2015). [cited by examiner]
Wang et al. “A Unified Tensor Level Set for Image Segmentation,” Nov. 3, 2009, IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics) (vol. 40, Issue: 3, Jun. 2010), pp. 857-867. (Year: 2009). [cited by examiner]
Office Action in Japanese Appln. No. 2022-517255, mailed on Sep. 26, 2023, 7 pages (with English translation). [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2019/067648, dated Jun. 30, 2022, 9 pages. [cited by applicant]
Duran et al., “A many core implementation of the direct n body problem,” High Performance Parallelism Pearls, Nov. 2017, 16 pages. [cited by applicant]
Lan et al., “High performance implementation of 3d convolutional neural networks on a gpu,” Computational intelligence and neuroscience, Nov. 2017, 9 pages. [cited by applicant]
PCT International Search Report and Written Opinion in International Appln. No. PCT/US2019/067648, dated Oct. 6, 2020, 14 pages. [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2022-517255, dated Jun. 18, 2024, 5 pages (with English translation). [cited by applicant]
Office Action in Korean Appln. No. 10-2022-7008209, dated Jun. 17, 2024, 17 pages (with English translation). [cited by applicant]
Office Action in European Appln. No. 19842493.9, mailed on Jul. 8, 2024, 7 pages. [cited by applicant]
Office Action in Taiwanese Appln. No. 11321015770, mailed on Oct. 7, 2024, 26 pages (with machine translation). [cited by applicant]
Office Action in Chinese Appln. No. 201980100413.8, mailed on Dec. 20, 2024, 16 pages (with English translation). [cited by applicant]
Office Action in Korean Appln. No. 10-2022-7008209, dated Feb. 4, 2025, 7 pages (with English translation). [cited by applicant]
Office Action in European Appln. No. 19842493.9, mailed on May 30, 2025, 9 pages. [cited by applicant]
Office Action in Japanese Appln. No. 2024-114737, mailed on Jul. 22, 2025, 5 pages (with English translation). [cited by applicant]
Notice of Allowance in Korean Appln. No. 10-2022-7008209, mailed on Sep. 29, 2025, 4 pages (with English translation). [cited by applicant]
Office Action in Korean Appln. No. 10-2022-7008209, mailed on Feb. 4, 2025, 6 pages (with English translation). [cited by applicant]