IP Library Granted Patent US 12,287,756
Granted Patent B2
US 12,287,756 · App. 18/376,494 · Granted Apr 29, 2025

General-purpose systolic array

Inventors: Reginald Clifford Young (Palo Alto, CA); Trevor Gale (Boston, MA); Sushma Honnavara-Prasad (Los Gatos, CA); Paolo Mantovani (New York, NY)
Assignee: GOOGLE LLC
G06F15/8046G06F15/8069G06F15/8084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,287,756
App. No.
18/376,494
Granted
Apr 29, 2025
Kind
B2
Abstract

A systolic array cell is described, the cell including two general-purpose arithmetic logic units (ALUs) and register-file. A plurality of the cells may be configured in a matrix or array, such that the output of the first ALU in a first cell is provided to a second cell to the right of the first cell, and the output of the second ALU in the first cell is provided to a third cell below the first cell. The two ALUs in each cell of the array allow for processing of a different instruction in each cycle.

Claims (54)

1. A processor cell, comprising:

a crossbar switch;

a first floating point logic unit (FPLU) configured to:

receive, as input to the first FPLU, output from the crossbar switch; and

provide, as output from the first FPLU, input to one or more additional processing cells;

a second FPLU configured to:

receive, as input to the second FPLU, output from the crossbar switch; and

provide, as output from the second FPLU, input to the one or more additional processing cells; and

a register file configured to:

receive, as input to the register file, the outputs from the first FPLU and the second FPLU; and

provide, as output from the register file, one or more inputs to the crossbar switch.

2. The processor cell of claim 1 , wherein at least one of the first or second FPLUs comprises an 8-bit, 16-bit, 32-bit, or 64-bit FPLU.

3. The processor cell of claim 1 , wherein the crossbar switch is configured to receive, as input to the crossbar switch, output from the one or more additional processing cells.

4. The processor cell of claim 1 , wherein at least one of the first or second FPLUs comprises a multiplier.

5. The processor cell of claim 1 , further comprising a multiplexer configured to:

receive, as input to the multiplexer, the outputs from the first FPLU and the second FPLU; and

provide, as output from the multiplexer, an input to the register file.

6. The processor cell of claim 1 , wherein the first FPLU is coupled to a first output of the crossbar switch and the second FPLU is coupled to a second output of the crossbar switch.

7. A computation unit, comprising:

a plurality of processor cells having a first output of a first processor cell provided as input to a second processor cell and a second output of the first processor cell provided as input to a third processor cell, wherein each processor cell comprises:

a crossbar switch;

a plurality of floating point logic units (FPLUs), each FPLU configured to:

receive, as input to the FPLU, output from the crossbar switch; and

provide, as output from the FPLU, input to the second or third processor cell; and

a register file configured to:

receive, as input to the register file, the outputs from the plurality of FPLUs; and

provide, as output from the register file, one or more inputs to the crossbar switch.

8. The computation unit of claim 7 , wherein at least one of the plurality of FPLUS comprises an 8-bit, 16-bit, 32-bit, or 64-bit FPLU.

9. The computation unit of claim 7 , wherein at least one of the plurality of processor cells further comprises a multiplexer configured to:

receive, as input to the multiplexer, the outputs from the plurality of FPLUs; and

provide, as output from the multiplexer, an input to the register file.

10. The computation unit of claim 7 , wherein at least one of the plurality of FPLUs comprises a multiplier.

11. The computation unit of claim 7 , wherein the plurality of FPLUs comprises a first FPLU coupled to a first output of the crossbar switch and a second FPLU coupled to a second output of the crossbar switch.

12. The computation unit of claim 7 , wherein a crossbar switch of the first processor cell is configured to receive, as input to the crossbar switch, output from a fourth processor cell.

13. The computation unit of claim 7 , wherein the computation unit is configured to receive two source vectors and produce at least one result vector per cycle.

14. A computing system, comprising:

one or more memories;

one or more processors in communication with the one or more memories;

a plurality of cells in communication with the one or more processors, the plurality of cells having a first output of a first cell provided as input to a second cell and a second output of the first cell provided as input to a third cell, wherein each cell comprises:

a crossbar switch;

a plurality of floating point logic units (FPLUs), each FPLU configured to:

receive, as input to the FPLU, output from the crossbar switch; and

provide, as output from the FPLU, input to the second or third cell;

and a register file configured to:

receive, as input to the register file, the outputs from the plurality of FPLUs; and

provide, as output from the register file, one or more inputs to the crossbar switch.

15. The computing system of claim 14 , wherein the one or more processors comprise at least one of a scalar core or a vector processing unit.

16. The computing system of claim 14 , wherein the one or more memories comprise a vector data cache or a level one cache.

17. The computing system of claim 14 , further comprising a sequencer configured to control instructions sent to the one or more processors and the plurality of cells.

18. The computing system of claim 14 , wherein the plurality of cells is configured to receive two source vectors and produce at least one result vector per cycle.

19. The computing system of claim 14 , wherein at least one of the plurality of cells further comprises a multiplexer configured to:

receive, as input to the multiplexer, the outputs from the plurality of FPLUs; and

provide, as output from the multiplexer, an input to the register file.

20. The computing system of claim 14 , wherein a crossbar switch of the first cell is configured to receive, as input to the crossbar switch, output from a fourth cell.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2025
From: GOOGLE LLC
To: GDM HOLDING LLC
Reel/Frame 071465/0754 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2023
From: YOUNG, REGINALD CLIFFORD; GALE, TREVOR; HONNAVARA-PRASAD, SUSHMA; MANTOVANI, PAOLO
To: GOOGLE LLC
Reel/Frame 065117/0757 →
Continuity (2)
Continuation 17703479 · Mar 24, 2022
Related Publication 20240078212A1 · Mar 7, 2024
References Cited (21)
US 9158575B2 · Smith · 2015 [cited by examiner]
US 9170812B2 · Vorbach et al. · 2015 [cited by applicant]
US 10528321B2 · Bittner et al. · 2020 [cited by applicant]
US 10698976B2 · Phelps et al. · 2020 [cited by applicant]
US 10824938B2 · Barik et al. · 2020 [cited by applicant]
US 11494627B1 · Li et al. · 2022 [cited by applicant]
US 20130311532A1 · Olsen · 2013 [cited by applicant]
US 20140337853A1 · Kim · 2014 [cited by examiner]
US 20150268963A1 · Etsion · 2015 [cited by examiner]
US 20180307495A1 · Ould-Ahmed-Vall · 2018 [cited by examiner]
US 20200279169A1 · Hoskins et al. · 2020 [cited by applicant]
US 20200341942A1 · Ray · 2020 [cited by examiner]
US 20210201466A1 · Chen · 2021 [cited by examiner]
US 20210406010A1 · Ahmed · 2021 [cited by applicant]
Annaratone et al. The Warp Computer: Architecture, Implementation, and Performance. IEEE Transactions on Computers, IEEE, USA, vol. C-19, No. 12, Dec. 31, 1987 (Dec. 31, 1987), pp. 1523-1538. [cited by applicant]
Annaratore, E. et al., Warp Architecture and Implementation, 1986, IEEE, pp. 346-356. (Year: 1986). [cited by applicant]
Borkar et al. iWarp: An Integrated Solution to High-Speed Parallel Computing. Supercomputing '88. Yvol. I, Proceedings. Orlando, FL, USA Nov. 14-18, 1988, Washington, DC, USA, IEEE Comput. Soc. Pr, US, Nov. 14, 1988 (No… [cited by applicant]
Fisher et al. Architecture of the PSC: A Programmable Systolic Chip. Proceedings of the Annual Symposium on Computer Architecture. Stockholm, 1983; [Proceedings of the Annual International Symposium on Computer Architec… [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2023/013016 dated May 26, 2023. 18 pages. [cited by applicant]
Lee et al. Super Semi-systolic Array-Based Application-Specific PLO Architecture. Mar. 31, 2006 (Mar. 31, 2006), SAT 2015 18th International Conference, Austin, TX, USA, Sep. 24-27, 2015; [Lecture Notes in Computer Scie… [cited by applicant]
Xu, J. et al., Accelerating Matrix Processing for MIMO Systems, 2021, IEEE, 6 pages. (Year: 2021). [cited by applicant]