IP Library › Granted Patent US 12,189,472
Granted Patent B2
US 12,189,472 · App. 18/386,641 · Granted Jan 7, 2025

Error checking for systolic array computation

Inventors: Doe Hyun Yoon (Foster City, CA); Norman Paul Jouppi (Palo Alto, CA)
Assignee: Google LLC
G06F11/1004G06F15/8046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,472
App. No.
18/386,641
Granted
Jan 7, 2025
Kind
B2
Abstract

Aspects of the disclosure are directed to a computation unit implementing a systolic array and configured for detecting errors while processing data on the systolic array. Checksum circuit in communication with a systolic array is configured to compute checksums and perform error detection while the systolic array processes input data. Instead of pre-generating checksums in input matrices, input matrices can be directly fed into the systolic array through the checksum circuit. On the output side, the checksum circuit can generate and compare checksums with checksums in an output matrix generated by the systolic array. Error checking the operations to generate the output matrix can be performed without delaying the operations of the systolic array, and without preprocessing the input matrices.

Claims (52)

1. A computation unit comprising:

one or more processing elements configured to:

receive first input elements from a first input matrix and second input elements from a second input matrix; and

generate an output matrix from the first input elements and the second input elements; and

an output checksum circuit configured to receive the output matrix, and determine, from the output matrix, an occurrence of one or more errors in the generation of the output matrix.

2. The computation unit of claim 1 , wherein, to determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuit is further configured to:

generate, from at least one row of the output matrix, a row checksum;

compare the row checksum with a checksum in an output checksum column; and

determine, from the comparison, whether the row checksum matches with the checksum in the output checksum column within a predetermined threshold.

3. The computation unit of claim 2 , wherein, to determine whether the row checksum matches with the checksum in the output checksum column within a predetermined threshold, the output checksum circuit is further configured to determine whether an absolute difference between the row checksum and the checksum in the output checksum column are within a predetermined threshold.

4. The computation unit of claim 1 , wherein to determine the occurrence of one or more errors in the generation of the output matrix, the output checksum circuit is further configured to:

generate, from at least one column of the output matrix, a column checksum;

compare the column checksum with a checksum in an output checksum row; and

determine, from the comparison, whether the column checksum matches with the checksum in the output checksum row within a predetermined threshold.

5. The computation unit of claim 1 , wherein, in response to the determination of the occurrence of one or more errors in the generation of the output matrix, the output checksum circuit is further configured to send an indication of the occurrence of one or more errors to one or more devices.

6. The computation unit of claim 5 , wherein the one or more processing elements are further configured to, in response to sending the indication of the occurrence of one or more errors to the one or more devices, receive an adjusted voltage higher than a critical voltage for the computation unit.

7. The computation unit of claim 6 , wherein the one or more processing elements are further configured to receive a first voltage lower than the critical voltage for the computation unit until receiving the adjusted voltage.

8. The computation unit of claim 1 , wherein the first input elements and the second input elements are floating point values.

9. A method for data processing by a computation unit, the method comprising:

receiving, by one or more processing elements, first input elements from a first input matrix;

receiving, by the one or more processing elements, second input elements from a second input matrix;

generating, by the one or more processing elements, an output matrix from the first input elements and the second input elements;

receiving, by an output checksum circuit, the output matrix; and

determining, by the output checksum circuit, from the output matrix, an occurrence of one or more errors in the generation of the output matrix.

10. The method of claim 9 , wherein determining the occurrence of one or more errors in the generation of the output matrix further comprises:

generating, from at least one row of the output matrix, a row checksum;

comparing the row checksum with a checksum in an output checksum column; and

determining, from the comparison, whether the row checksum matches with the checksum in the output checksum column within a predetermined threshold.

11. The method of claim 10 , wherein determining whether the row checksum matches with the checksum in the output checksum column within a predetermined threshold further comprises determining whether an absolute difference between the row checksum and the checksum in the output checksum column are within a predetermined threshold.

12. The method of claim 9 , wherein determining the occurrence of one or more errors in the generation of the output matrix further comprises:

generating, from at least one column of the output matrix, a column checksum;

comparing the column checksum with a checksum in an output checksum row; and

determining, from the comparison, whether the column checksum matches with the checksum in the output checksum row within a predetermined threshold.

13. The method of claim 9 , further comprising sending, by the output checksum circuit, an indication of the occurrence of one or more errors to one or more devices.

14. The method of claim 13 , further comprising receiving, by the one or more processing elements, an adjusted voltage higher than a critical voltage for the computation unit.

15. The method of claim 14 , further comprising receiving, by the one or more processing elements, a first voltage lower than the critical voltage for the computation unit until receiving the adjusted voltage.

16. The method of claim 9 , wherein the first input elements and the second input elements are floating point values.

17. A non-transitory computer readable medium for storing instructions that, when executed by one or more processors, cause the one or more processors to performing operations for data processing by a computation unit, the operations comprising:

receiving, by one or more processing elements, first input elements from a first input matrix;

receiving, by the one or more processing elements, second input elements from a second input matrix;

generating, by the one or more processing elements, an output matrix from the first input elements and the second input elements;

receiving, by an output checksum circuit, the output matrix; and

determining, by the output checksum circuit, from the output matrix, an occurrence of one or more errors in the generation of the output matrix.

18. The non-transitory computer readable medium of claim 17 , wherein determining the occurrence of one or more errors in the generation of the output matrix further comprises:

generating, from at least one row of the output matrix, a row checksum;

comparing the row checksum with a checksum in an output checksum column; and

determining, from the comparison, whether the row checksum matches with the checksum in the output checksum column within a predetermined threshold.

19. The non-transitory computer readable medium of claim 18 , wherein determining whether the row checksum matches with the checksum in the output checksum column within a predetermined threshold further comprises determining whether an absolute difference between the row checksum and the checksum in the output checksum column are within a predetermined threshold.

20. The non-transitory computer readable medium of claim 17 , wherein determining the occurrence of one or more errors in the generation of the output matrix further comprises:

generating, from at least one column of the output matrix, a column checksum;

comparing the column checksum with a checksum in an output checksum row; and

determining, from the comparison, whether the column checksum matches with the checksum in the output checksum row within a predetermined threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2023
From: YOON, DOE HYUN; JOUPPI, NORMAN PAUL
To: GOOGLE LLC
Reel/Frame 065446/0878 →
Continuity (4)
Continuation 17961623 · Oct 7, 2022
Continuation 17410558 · Aug 24, 2021
Provisional Application 63222549 · Jul 16, 2021
Related Publication 20240061742A1 · Feb 22, 2024
References Cited (31)
US 5323335A · Mitchell · 1994 [cited by examiner]
US 8406334B1 · Rao · 2013 [cited by examiner]
US 11232062B1 · Volpe · 2022 [cited by examiner]
US 11308026B1 · Volpe · 2022 [cited by examiner]
US 11308027B1 · Volpe · 2022 [cited by examiner]
US 20030018675A1 · Asai · 2003 [cited by examiner]
US 20150188572A1 · Silberman · 2015 [cited by examiner]
US 20190236049A1 · Vantrease · 2019 [cited by examiner]
US 20200334322A1 · Liu et al. · 2020 [cited by applicant]
US 20210157548A1 · Elmer · 2021 [cited by examiner]
CN 110933728A · 2020 [cited by applicant]
JP 2021508125A · 2021 [cited by applicant]
J. Chen et al., “Fault Tolerant One-sided Matrix Decompositions on Heterogeneous Systems with GPUs,” SC18: International Conference for High Performance Computing, Networking, Storage and Analysis, Dallas, TX, USA, 2018… [cited by examiner]
Huang and Abraham. Algorithm-Based Fault Tolerance for Matnx Operations. Jun. 1984. IEEE Transactions on Computers, vol. C-33, No. 6, pp. 518-528. Retrieved from the Internet: <https://graal.ens-lyon.fr/˜abenoit/CR02/pa… [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/036946 dated Nov. 14, 2022. 16 pages. [cited by applicant]
J.-. Han and D. C. Krishnan, “Linear arithmetic code and its application in fault-tolerant systolic array,” Proceedings. IEEE Energy and Information Technologies in the Southeast, 1989, pp. 1015-1020 vol. 3, doi: 10.110… [cited by applicant]
Jacobs et al. Overhead and reliability analysis of algorithm-based fault tolerance in FPGA systems. Field Frogrammable Logic and Ai?flications (FFL), 2012 22nd International Conference on, IEEE, [ Online] Aug. 29, 2012 … [cited by applicant]
Kung and Lam. Fault-Tolerant VLSI Systolic Arrays and Two-Level Pipelining. Nov. 28, 1983. Proceedings of SPIE 0431, Real-Time Signal Processing VI. 17 pages. Retrieved from the Internet: <https://www.spiedigitallibrary… [cited by applicant]
Li et al. Efficient Soft-Error Detection for Low-precision Deep Learning Recommendation Models. Feb. 27, 2021. 8 pages. Retrieved from the Internet: <https://arxiv.org/pdf/2103.00130.pdf>. [cited by applicant]
Nair and Abraham. Real-Number Codes for Fault-Tolerant Matrix Operations on Processor Arrays. Apr. 1990. IEEE Transactions on Computers, vol. 39, No. 4, pp. 426-435. Retrieved from the Internet: <http://pduwork.pbworks.… [cited by applicant]
Pandey et al. GreenTPU: Improving Timing Error Resilience of a Near-Threshold Tensor Processing Unit. Jun. 2019. DAC '19: Proceedings of the 56th Annual Design Automation Conference 2019. Article No. 173. 6 pages. Retri… [cited by applicant]
Paul et al. Voltage Scaling for Partitioned Systolic Array in a Reconfigurable Platform. Feb. 13, 2021. 6 pages. Retrieved from the Internet: <https://arxiv.org/pdf/2102.06888.pdf>. [cited by applicant]
Safarpour and Silvén. Algorithm Level Error Detection in Low Voltage Systolic Array. Jun. 2021. IEEE Transactions on Circuits and Systems—II. pp. 1-5. Retrieved from the Internet: <https://www.techrxiv.org/articles/prep… [cited by applicant]
Wang and Jen. Redundancy design for a fault tolerant systolic array. May 1990. IEE Proceedings, vol. 137, Pt. E, No. 3. pp. 218-226. Retrieved from the Internet: <https://ir.nctu.edu.tw/bitstream/11536/4117/1/A1990DC360… [cited by applicant]
Wu et al. Fault Tolerant Matrix-Matrix Multiplication: Correcting Soft Errors On-Line. Nov. 13, 2011. Colorado School of Mines. 18 pages. Retrieved from the Internet: <https://www.csm.ornl.gov/srt/conferences/Scala/2011… [cited by applicant]
Zhang et al. Analyzing and Mitigating the Impact of Permanent Faults on a Systolic Array Based Neural Network Accelerator. Feb. 17, 2018. 6 pages. Retrieved from the Internet: <https://arxiv.org/pdf/1802.04657.pdf>. [cited by applicant]
Zhang et al. ThUnderVolt: Enabling Aggressive Voltage Underscaling and Timing Error Resilience for Energy Efficient Deep Learning Accelerators. Mar. 13, 2018. 7 pages. Retrieved from the Internet: <https://arxiv.org/pdf… [cited by applicant]
Zhao et al. FT-CNN: Algorithm-Based Fault Tolerance for Convolutional Neural Networks. Sep. 7, 2020. 13 pages. Retrieved from the Internet: <https://arxiv.org/pdf/2003.12203.pdf>. [cited by applicant]
Ernst et al. Razor: A Low-Power Pipeline Based on Circuit-Level Timing Speculation. Dec. 2003. 36th Annual International Symposium on Microarchitecture (Micro-36), 12 pages. Retrieved from the Internet <https://web.eecs… [cited by applicant]
Gundi et al. Effort: Enhancing Energy Efficiency and Error Resilience of a Near-Threshold Tensor Processing Unit. 2020. 25th Asia and South Pacific Design Automation Conference (ASP-DAC), pp. 241-246. [cited by applicant]
Hari et al. Making Convolutions Resilient via Algorithm-Based Error Detection Techniques. Jun. 8, 2020. 12 pages. Retrieved from the Internet <https://arxiv.org/pdf/2006.04984.pdf>. [cited by applicant]