IP Library › Granted Patent US 12,561,395
Granted Patent B2
US 12,561,395 · App. 18/444,249 · Granted Feb 24, 2026

Low latency matrix multiply unit

Inventors: Andrew Everett Phelps (Middleton, WI); Norman Paul Jouppi (Palo Alto, CA)
Assignee: Google LLC
G06F17/16G06F5/015G06F9/30036G06F9/30101G06F15/8046G06F9/30032G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,395
App. No.
18/444,249
Granted
Feb 24, 2026
Kind
B2
Abstract

Methods, systems, and apparatus for a matrix multiply unit implemented as a systolic array of cells are disclosed. Each cell of the matrix multiply includes: a weight matrix register configured to receive a weight input from either a transposed or a non-transposed weight shift register; a transposed weight shift register configured to receive a weight input from a horizontal direction to be stored in the weight matrix register; a non-transposed weight shift register configured to receive a weight input from a vertical direction to be stored in the weight matrix register; and a multiply unit that is coupled to the weight matrix register and configured to multiply the weight input of the weight matrix register with a vector data input in order to obtain a multiplication result.

Claims (29)

1 . A processor comprising:

a processing core having an array of cells;

a first plurality of shift chains corresponding to rows of the array of cells;

a second plurality of shift chains corresponding to columns of the array of cells,

wherein the processor is configured to execute instructions that cause the processor to perform operations comprising:

selectively loading weight values using the first plurality of shift chains corresponding to the rows of the array of cells or using the second plurality of shift chains corresponding to the columns of the array of cells, and

performing one or more matrix multiplication operations using the loaded weight values and a vector input.

2 . The processor of claim 1 , further comprising alternately loading weight values using the first plurality of shift chains and the second plurality of shift chains.

3 . The processor of claim 2 , wherein alternately loading the weight values corresponds to different steps of a neural network training process.

4 . The processor of claim 1 , wherein the first plurality of shift chains corresponds to transposed weights and the second plurality of shift chains corresponds to non-transposed weights or wherein the first plurality of shift chains corresponds to non-transposed weights and the second plurality of shift chains corresponds to transposed weights.

5 . The processor of claim 1 , wherein the operations further comprises loading the vector input using the first plurality of shift chains or the second plurality of shift chains.

6 . The processor of claim 1 , further comprising a plurality of weight registers coupled to the first plurality of shift chains and the second plurality of shift chains.

7 . The processor of claim 6 , further comprising a plurality of multiplexors configured to select between loading the weight registers with data from the first plurality of shift chains or the second plurality of shift chains.

8 . The processor of claim 1 , wherein the processing core further includes one or more of a scalar memory, a vector memory, a scalar processing unit, one or more vector registers, a transposed unit, or a reduction and permutation unit.

9 . A method performed by a processor, wherein the processor comprises a processing core having an array of cells, a first plurality of shift chains, and a second plurality of shift chains, the method comprising:

selectively loading weight values using the first plurality of shift chains corresponding to the rows of the array of cells or using the second plurality of shift chains corresponding to the columns of the array of cells; and

performing one or more matrix multiplication operations using the loaded weight values and a vector input.

10 . The method of claim 9 , further comprising alternately loading weight values using the first plurality of shift chains and the second plurality of shift chains.

11 . The method of claim 10 , wherein alternately loading the weight values corresponds to different steps of a neural network training process.

12 . The method of claim 9 , wherein the first plurality of shift chains corresponds to transposed weights and the second plurality of shift chains corresponds to non-transposed weights or wherein the first plurality of shift chains corresponds to non-transposed weights and the second plurality of shift chains corresponds to transposed weights.

13 . The method of claim 9 , wherein the operations further comprises loading the vector input using the first plurality of shift chains or the second plurality of shift chains.

14 . The method of claim 9 , further comprising a plurality of weight registers coupled to the first plurality of shift chains and the second plurality of shift chains.

15 . The method of claim 14 , further comprising a plurality of multiplexors configured to select between loading the weight registers with data from the first plurality of shift chains or the second plurality of shift chains.

16 . The method of claim 9 , wherein the processing core further includes one or more of a scalar memory, a vector memory, a scalar processing unit, one or more vector registers, a transposed unit, or a reduction and permutation unit.

17 . A non-transitory computer program product storing instructions that, when executed by a programmable processor comprising a processing core having an array of cells, a first plurality of shift chains, and a second plurality of shift chains, cause the processor to perform operations comprising:

selectively loading weight values using the first plurality of shift chains corresponding to the rows of the array of cells or using the second plurality of shift chains corresponding to the columns of the array of cells; and

performing one or more matrix multiplication operations using the loaded weight values and a vector input.

18 . The non-transitory computer program of claim 17 , further comprising alternately loading weight values using the first plurality of shift chains and the second plurality of shift chains.

19 . The non-transitory computer program of claim 18 , wherein alternately loading the weight values corresponds to different steps of a neural network training process.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2024
From: PHELPS, ANDREW EVERETT; JOUPPI, NORMAN PAUL
To: GOOGLE INC.
Reel/Frame 067298/0589 →
ENTITY CONVERSION Recorded May 2, 2024
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 067304/0589 →
Continuity (7)
Continuation 18111468 · Feb 17, 2023
Continuation 17210293 · Mar 23, 2021
Continuation 16915286 · Jun 29, 2020
Continuation 16529662 · Aug 1, 2019
Continuation 15983037 · May 17, 2018
Provisional Application 62507766 · May 17, 2017
Related Publication 20240303297A1 · Sep 12, 2024
References Cited (60)
US 4720780A · Dolecek · 1988 [cited by applicant]
US 5471627A · Means et al. · 1995 [cited by applicant]
US 6023742A · Ebeling et al. · 2000 [cited by applicant]
US 7161995B1 · Mazahreh et al. · 2007 [cited by applicant]
US 8543634B1 · Xu et al. · 2013 [cited by applicant]
US 8620984B2 · Mazahreh et al. · 2013 [cited by applicant]
US 9747546B2 · Ross et al. · 2017 [cited by applicant]
US 10635740B2 · Phelps et al. · 2020 [cited by applicant]
US 11989259B2 · Phelps et al. · 2024 [cited by applicant]
US 20040122887A1 · Macy · 2004 [cited by applicant]
US 20110040821A1 · Eichenberger et al. · 2011 [cited by applicant]
US 20110125819A1 · Mazahreh et al. · 2011 [cited by applicant]
US 20140289445A1 · Savich · 2014 [cited by applicant]
US 20160342892A1 · Ross · 2016 [cited by applicant]
US 20160342893A1 · Ross et al. · 2016 [cited by applicant]
US 20170103314A1 · Ross et al. · 2017 [cited by applicant]
US 20170103316A1 · Ross et al. · 2017 [cited by applicant]
US 20180336163A1 · Phelps et al. · 2018 [cited by applicant]
US 20180336164A1 · Phelps et al. · 2018 [cited by applicant]
US 20190311243A1 · Whatmough et al. · 2019 [cited by applicant]
US 20200226202A1 · Phelps et al. · 2020 [cited by applicant]
US 20240362298A1 · Phelps et al. · 2024 [cited by applicant]
US 20250232003A1 · Phelps et al. · 2025 [cited by applicant]
CN 1774709A · 2006 [cited by applicant]
CN 103975302A · 2014 [cited by applicant]
CN 105117372A · 2015 [cited by applicant]
CN 103246773 · 2016 [cited by applicant]
JP H03131965 · 1991 [cited by applicant]
JP 2013511891 · 2013 [cited by applicant]
JP 2017021483A · 2017 [cited by applicant]
JP 2018521374 · 2018 [cited by applicant]
KR 1020170007151 · 2017 [cited by applicant]
TW 201525856 · 2015 [cited by applicant]
TW 201706871 · 2017 [cited by applicant]
TW 201706873 · 2017 [cited by applicant]
WO WO2016186810 · 2016 [cited by applicant]
Chen et al., “Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks,” Paper, Presented at Proceedings of the ACM/IEEE 43rd Annual International Symposium on Computer Architectur… [cited by applicant]
Office Action in Chinese Appln. No. 201880004328.7, mailed on Nov. 11, 2022, 13 pages (with English translation). [cited by applicant]
Office Action in Taiwanese Appln. No. 113135506, mailed on Oct. 29, 2024, 16 pages (with machine translation). [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2024-024596, mailed on Apr. 1, 2025, 5 pages (with English translation). [cited by applicant]
Office Action in Brazilian Appln. No. 1120190229167, mailed on May 12, 2025, 9 pages (with machine translation). [cited by applicant]
Extended European Search Report in European Appln. No. 20198533.0, mailed on Mar. 5, 2021, 7 pages. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2018/033261, mailed on Nov. 19, 2019, 8 pages. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2018/033270, mailed on Nov. 19, 2019, 7 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2018/033261, mailed on Oct. 15, 2018, 15 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2018/033270, mailed on Oct. 15, 2018, 14 pages. [cited by applicant]
Kato et al., “Highly Parallel Neurocomputer” Fujitsu Limited, May 1991, pp. 213-221. [cited by applicant]
Khan et al., “Two-dimensional multirate systolic array design for artificial neural networks,” First Great Lakes Symposium on Kalamazoo, Mar. 1, 1991, 8 pages. [cited by applicant]
Office Action in Chinese Appln. No. 201880004576.1, mailed on Nov. 11, 2022, 12 pages (with English translation). [cited by applicant]
Office Action in Indian Appln. No. 201947035951, mailed on Jun. 12, 2021, 6 pages (with English translation). [cited by applicant]
Office Action in Japanese Appln. No. 2019-553237, mailed on Dec. 1, 2020, 7 pages (with English translation). [cited by applicant]
Office Action in Japanese Appln. No. 2022-138332, mailed on Jul. 25, 2023, 6 pages (with English translation). [cited by applicant]
Office Action in Korean Appln. No. 10-2019-7026961, mailed on Jan. 21, 2021, 9 pages (with English translation). [cited by applicant]
Office Action in Korean Appln. No. 10-2019-7026961, mailed on Jul. 21, 2020, 8 pages (with English translation). [cited by applicant]
Office Action in Taiwanese Appln. No. 109105071, mailed on Nov. 30, 2020, 7 pages (with English translation). [cited by applicant]
Office Action in Taiwanese Appln. No. 110130201, mailed on Sep. 30, 2021, 13 pages (with English translation). [cited by applicant]
Office Action in Taiwanese Appln. No. 107116872, mailed on Apr. 23, 2019, 6 pages (with English translation). [cited by applicant]
Office Action in Taiwanese Appln. No. 111127382, mailed on Aug. 30, 2022, 11 pages (with English translation). [cited by applicant]
Ramacher, “Synapse-A Neurocomputer that Synthesized Neural Algorithms on a Parallel Systolic Engine,” Journal of Parallel and Distributed Computing, Mar. 1, 1992, 13 pages. [cited by applicant]
Search Report in European Appln. No. 20188875.7, mailed on Nov. 27, 2020, 8 pages. [cited by applicant]