IP Library › Granted Patent US 12,524,658
Granted Patent B2
US 12,524,658 · App. 17/694,598 · Granted Jan 13, 2026

Special purpose neural network training chip

Inventors: Thomas Norrie (Mountain View, CA); Olivier Temam (Antony, FR); Andrew Everett Phelps (Middleton, WI); Norman Paul Jouppi (Palo Alto, CA)
Assignee: Google LLC
G06N3/063G06F9/3001G06F9/30032G06F9/30036G06F9/30141G06F9/3885G06F9/3887G06F17/16G06N3/08G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,658
App. No.
17/694,598
Filed
Mar 14, 2022
Granted
Jan 13, 2026
Kind
B2
Art Unit
2123
USPC
706/25
Abstract

Methods, systems, and apparatus including a special purpose hardware chip for training neural networks are described. The special-purpose hardware chip may include a scalar processor configured to control computational operation of the special-purpose hardware chip. The chip may also include a vector processor configured to have a 2-dimensional array of vector processing units which all execute the same instruction in a single instruction, multiple-data manner and communicate with each other through load and store instructions of the vector processor. The chip may additionally include a matrix multiply unit that is coupled to the vector processor configured to multiply at least one two-dimensional matrix with a second one-dimensional vector or two-dimensional matrix in order to obtain a multiplication result.

Claims (29)

1 . A special-purpose hardware chip for training neural networks, the special-purpose hardware chip comprising:

a scalar processor configured to control computational operation of the special-purpose hardware chip;

a vector processor having a 2-dimensional array of vector processing units; and

a matrix multiply unit that is coupled to the vector processor and configured to multiply at least a first two-dimensional matrix with a first one-dimensional vector or a second two-dimensional matrix in order to obtain a multiplication result,

wherein the vector processor includes a plurality of lanes, wherein each vector processing unit of the 2-dimensional array of vector processing units in the vector processor is located in a respective lane of the plurality of lanes, wherein one or more vector processing units of the 2-dimensional array of vector processing units that are located in the same lane are configured to communicate with one another through respective load and store instructions.

2 . The special-purpose hardware chip of claim 1 , further comprising:

a vector memory configured to provide private memory to the vector processor.

3 . The special-purpose hardware chip of claim 1 , further comprising:

a scalar memory configured to provide private memory to the scalar processor.

4 . The special-purpose hardware chip of claim 1 , further comprising:

a transpose unit configured to perform a transposition operation of a matrix.

5 . The special-purpose hardware chip of claim 1 , further comprising:

a reduction and permutation unit configured to perform a reduction on numbers and permute the numbers among different lanes of a vector array.

6 . The special-purpose hardware chip of claim 1 , further comprising:

a first memory configured to store data of the special-purpose hardware chip.

7 . The special-purpose hardware chip of claim 1 , further comprising a sparse computation core.

8 . The special-purpose hardware chip of claim 1 , further comprising:

an interface; and

an inter-chip interconnect, which connects the interface or resources on the special purpose hardware chip to other special-purpose hardware chips or resources.

9 . The special-purpose hardware chip of claim 8 , further comprising:

a first memory, wherein the inter-chip interconnect connects the interface and the first memory to other special-purpose hardware chips.

10 . The special-purpose hardware chip of claim 8 , wherein the interface is a host interface to a host computer.

11 . The special-purpose hardware chip of claim 8 , wherein the interface is a standard network interface to a network of host computers.

12 . The special-purpose hardware chip of claim 8 , further comprising a scalar memory and a vector memory.

13 . The special-purpose hardware chip of claim 8 , wherein a scalar instruction set of the instructions includes arithmetic operations used in address calculations, load/store instructions, and branch instructions, and wherein the remaining instructions of the instructions encode instructions for the vector processor and the matrix multiply unit.

14 . The special-purpose hardware chip of claim 1 , wherein each vector processing unit in the 2-dimensional array of vector processing units includes 32 registers.

15 . The special-purpose hardware chip of claim 1 , wherein each vector processing unit in the 2-dimensional array of vector processing units is configured to perform at least one of a floating point operation or integer operation.

16 . The special-purpose hardware chip of claim 1 , wherein each vector processing unit in the 2-dimensional array of vector processing units is configured to execute two respective arithmetic logic unit (ALU) instructions, a respective load instruction, and a respective store instruction in each clock cycle.

17 . The special-purpose hardware chip of claim 16 , wherein each vector processing unit in the 2-dimensional array of vector processing units is configured to compute respective offset memory addresses for executing the respective load and store instructions in each clock cycle.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2022
From: NORRIE, THOMAS; TEMAM, OLIVIER; PHELPS, ANDREW EVERETT; JOUPPI, NORMAN PAUL
To: GOOGLE LLC
Reel/Frame 059391/0964 →
Continuity (3)
Continuation 15983056 · May 17, 2018
Provisional Application 62507771 · May 17, 2017
Related Publication 20220261622A1 · Aug 18, 2022
References Cited (57)
US 5423051A · Fuller et al. · 1995 [cited by applicant]
US 5544336A · Kato et al. · 1996 [cited by applicant]
US 5872988A · Duranton · 1999 [cited by applicant]
US 7305540B1 · Trivedi et al. · 2007 [cited by applicant]
US 9898441B2 · Narayanaswami et al. · 2018 [cited by applicant]
US 10055692B1 · Mclaren et al. · 2018 [cited by applicant]
US 20030182518A1 · Nakanishi · 2003 [cited by applicant]
US 20130073495A1 · Izhikevich et al. · 2013 [cited by applicant]
US 20150378734A1 · Hansen · 2015 [cited by examiner]
US 20160342889A1 · Thorson et al. · 2016 [cited by applicant]
US 20160342891A1 · Ross et al. · 2016 [cited by applicant]
US 20160378465A1 · Venkatesh et al. · 2016 [cited by applicant]
US 20170097884A1 · Werner et al. · 2017 [cited by applicant]
US 20170132496A1 · Shoaib et al. · 2017 [cited by applicant]
US 20190347125A1 · Sankaran · 2019 [cited by examiner]
US 20190354862A1 · Ross et al. · 2019 [cited by applicant]
US 20200167530A1 · Buchanan · 2020 [cited by examiner]
CN 102521179A · 2012 [cited by applicant]
CN 104011672A · 2014 [cited by applicant]
CN 106503797A · 2017 [cited by applicant]
EP 0581828 · 1997 [cited by applicant]
EP 1569128 · 2015 [cited by applicant]
JP H04290155A · 1992 [cited by applicant]
JP 5841071 · 2016 [cited by applicant]
JP 2017138966A · 2017 [cited by applicant]
JP 2018513475 · 2018 [cited by applicant]
JP 2018531467 · 2018 [cited by applicant]
KR 20160046623A · 2016 [cited by applicant]
WO WO2002084451 · 2002 [cited by applicant]
WO WO2013101210 · 2013 [cited by applicant]
WO WO2016171909 · 2016 [cited by applicant]
WO WO2016171928 · 2016 [cited by applicant]
WO WO2016171928A1 · 2016 [cited by applicant]
WO WO2016186801 · 2016 [cited by applicant]
WO WO2017068318 · 2017 [cited by applicant]
EP Office Action in European Application No. 18735433.7, dated Dec. 17, 2020, 7 pages. [cited by applicant]
IN Office Action in Indian Application No. 201947035705, dated Jun. 14, 2021, 7 pages (with English translation). [cited by applicant]
International Search Report and Written Opinion issued in International Application No. PCT/US2018/033215, mailed on Aug. 22, 2018, 15 pages. [cited by applicant]
JP Office Action in Japanese Application No. 2019-549507, dated Apr. 6, 2021, 5 pages (with English translation). [cited by applicant]
KR Office Action in Korean Application No. 10-2019-7026557, dated Oct. 20, 2020, 8 pages (with English translation). [cited by applicant]
KR Office Action in Korean Application No. 10-2019-7026557, dated Apr. 28, 2021, 8 pages (with English translation). [cited by applicant]
PCT International Preliminary Report on Patentability in International Application No. PCT/US2018/033215, dated Nov. 19, 2020, 9 pages. [cited by applicant]
TW Office Action in Taiwan Application No. 107116869, dated May 15, 2019, 10 pages (with English translation). [cited by applicant]
TW Office Action in Taiwan Application No. 107116869m dated Mar. 17, 2020, 5 pages (with English translation). [cited by applicant]
EP Extended Search Report in European Appln. No. 22171943.8, dated Oct. 4, 2022, 6 pages. [cited by applicant]
Extended Search Report in European Appln. No. 24163748.7, mailed on Jul. 10, 2024, 6 pages. [cited by applicant]
Notice of Allowance in Japanese Appln. No. 2023-114361, mailed on Aug. 27, 2024, 7 pages (with English translation). [cited by applicant]
Notice of Allowance in Korean Application No. 10-2022-7045015, mailed on Jan. 24, 2024, 3 pages (with English translation). [cited by applicant]
Notice of Allowance in Taiwan Application No. 112130161, mailed on Jun. 14, 2024, 21 pages (with English translation). [cited by applicant]
Shimokawa et al., “Digital Neurocomputer, Multineuro,” Toshiba Review, Dec. 1991, 46(12):931-934 (English abstract). [cited by applicant]
Saito Yasutake, “Deep Learning—First Edition” O'Reilly Japan, Sep. 2016, 50 pages. [cited by applicant]
CN Office Action in Chinese Appln. No. 201880018006.8, dated Dec. 1, 2022, 26 pages (with English Translation). [cited by applicant]
JP Office Action in Japanese Patent Appln. No. 2021-142529, dated Nov. 22, 2022, 7 pages (with English Translation). [cited by applicant]
Office Action in Taiwanese Appln. No. 111120440, dated Feb. 3, 2023, 4 pages (with English Translation). [cited by applicant]
Jiyu et al., “A Basic-Block Reordering Algorithm Based on Neural Networks,” Acta Scientiarum Naturalium Universitatis Pekinensis, Jan. 2011, 47:9-16 (with English abstract). [cited by applicant]
Office Action in Chinese Appln. No. 202310655432.5, mailed on Aug. 22, 2025, 29 pages (with English translation). [cited by applicant]
Office Action in Korean Appln. No. 10-2022-7045015, mailed on Nov. 27, 2025, 4 pages (with English translation). [cited by applicant]