IP Library Granted Patent US 12,399,714
Granted Patent B2
US 12,399,714 · App. 18/074,990 · Granted Aug 26, 2025

Vector processing unit

Inventors: William Lacy (Madison, WI); Gregory Michael Thorson (Waunakee, WI); Christopher Aaron Clark (Madison, WI); Norman Paul Jouppi (Palo Alto, CA); Thomas Norrie (Mountain View, CA); Andrew Everett Phelps (Middleton, WI)
Assignee: Google LLC
G06F9/3001G06F7/588G06F9/30032G06F9/30036G06F9/30043G06F9/30098G06F9/3887G06F9/3891G06F13/36G06F13/4068G06F13/4282G06F15/8053G06F15/8092G06F17/16G06F15/8046G06N3/063G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,399,714
App. No.
18/074,990
Granted
Aug 26, 2025
Kind
B2
Abstract

A vector processing unit is described, and includes processor units that each include multiple processing resources. The processor units are each configured to perform arithmetic operations associated with vectorized computations. The vector processing unit includes a vector memory in data communication with each of the processor units and their respective processing resources. The vector memory includes memory banks configured to store data used by each of the processor units to perform the arithmetic operations. The processor units and the vector memory are tightly coupled within an area of the vector processing unit such that data communications are exchanged at a high bandwidth based on the placement of respective processor units relative to one another, and based on the placement of the vector memory relative to each processor unit.

Claims (40)

1. A system comprising:

a first vector processing unit (VPU) lane;

a vector memory co-located with the first VPU lane, the vector memory having a plurality of memory banks;

a second VPU lane that is within a distance to the vector memory co-located with the first VPU lane such that data traverses the distance in a single clock cycle; and

a matrix unit configured to perform matrix multiplication on data that is received from the first VPU lane and the second VPU lane.

2. The system of claim 1 , wherein each of the first VPU lane and the second VPU lane includes a respective vector memory having a plurality of memory banks.

3. The system of claim 1 , wherein each of the first VPU lane and the second VPU lane is a VPU sub-lane.

4. The system of claim 1 , wherein each of the first VPU lane and the second VPU lane is a respective computing resource of an integrated circuit die section of the system.

5. The system of claim 4 , wherein each of the first VPU lane and the second VPU lane comprises a plurality of VPU sub-lanes.

6. The system of claim 4 , wherein a first resource within a VPU sub-lane of the first VPU lane is within a distance to a second resource within the VPU sub-lane of the first VPU lane such that data traverses the distance in a single clock cycle.

7. The system of claim 6 , wherein a first resource within a VPU sub-lane of the second VPU lane is within a distance to a second resource within the VPU sub-lane of the second VPU lane such that data traverses the distance in a single clock cycle.

8. The system of claim 4 , further comprising:

an external memory coupled to each of the first VPU lane and the second VPU lane; and

an inter-chip interconnect that interconnects each of the external memory, the first VPU lane, and the second VPU lane.

9. The system of claim 8 , wherein the external memory is external to the integrated circuit die section.

10. The system of claim 8 , wherein each of the external memory and the inter-chip interconnect is configured to exchange data with the vector memory and the first VPU lane.

11. The system of claim 4 , wherein the vector memory is included in the first VPU lane.

12. The system of claim 11 , further comprising:

a plurality of second VPU lanes; and

a respective vector memory in each second VPU lane of the plurality of second VPU lanes.

13. The system of claim 4 , wherein:

the matrix unit is external to the integrated circuit die section; and

the data traverses a distance between the matrix unit and at least the first VPU lane in a single clock cycle.

14. The system of claim 4 , wherein the data comprises at least 1024 vector operands.

15. The system of claim 1 , wherein the data is represented as a multi-dimensional vector and the system further comprises:

a permute unit configured to reshape or rearrange the data with reference to the multi-dimensional vector.

16. The system of claim 1 , further comprising:

a cross-lane unit configured to move data between the first VPU lane and the second VPU lane.

17. A system comprising:

an external memory;

an inter-chip interconnect;

at least one vector processing unit (VPU) lane;

wherein each VPU lane of the at least one VPU lane comprises corresponding vector memory having a plurality of memory banks,

wherein each of the external memory and the inter-chip interconnect is configured to exchange data with the vector memory of the at least one VPU lane,

wherein each VPU lane of the at least one VPU lane comprises a plurality of VPU sub-lanes, and

wherein each VPU lane of the at least one VPU lane is within a distance to the corresponding vector memory such that data traverses the distance in a single clock cycle; and

a matrix unit configured to perform matrix multiplication on vector operands corresponding to the data, wherein the vector operands are received from the at least one VPU lane.

18. The system of claim 17 , wherein the data is represented as a multi-dimensional vector and the system further comprises:

a permute unit configured to reshape or rearrange the data with reference to the multi-dimensional vector, and

a cross-lane unit configured to move data between two or more VPU lanes of the system.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2023
From: LACY, WILLIAM; THORSON, GREGORY MICHAEL; CLARK, CHRISTOPHER AARON; JOUPPI, NORMAN PAUL; NORRIE, THOMAS; PHELPS, ANDREW EVERETT
To: GOOGLE INC.
Reel/Frame 063657/0695 →
CHANGE OF NAME Recorded May 16, 2023
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 063664/0832 →
Continuity (5)
Continuation 17327957 · May 24, 2021
Continuation 16843015 · Apr 8, 2020
Continuation 16291176 · Mar 4, 2019
Continuation 15454214 · Mar 9, 2017
Related Publication 20230297372A1 · Sep 21, 2023
References Cited (79)
US 4128880A · Cray, Jr. · 1978 [cited by examiner]
US 4150434A · Shibayama et al. · 1979 [cited by applicant]
US 4636942A · Chen et al. · 1987 [cited by applicant]
US 4755931A · Abe · 1988 [cited by examiner]
US 5067095A · Peterson et al. · 1991 [cited by applicant]
US 5327365A · Fujisaki et al. · 1994 [cited by applicant]
US 5758176A · Agarwal · 1998 [cited by examiner]
US 5790821A · Pflum · 1998 [cited by applicant]
US 5805875A · Asanovic · 1998 [cited by applicant]
US 5825677A · Agarwal et al. · 1998 [cited by applicant]
US 6539368B1 · Chernikov et al. · 2003 [cited by applicant]
US 7681013B1 · Trivedi · 2010 [cited by examiner]
US 8782376B2 · Knowles · 2014 [cited by examiner]
US 9529571B2 · Van Kampen · 2016 [cited by examiner]
US 10261786B2 · Lacy · 2019 [cited by examiner]
US 10915318B2 · Lacy · 2021 [cited by examiner]
US 11016764B2 · Lacy · 2021 [cited by examiner]
US 11520581B2 · Lacy · 2022 [cited by examiner]
US 20050251644A1 · Maher · 2005 [cited by examiner]
US 20070150697A1 · Sachs · 2007 [cited by applicant]
US 20080091924A1 · Jouppi · 2008 [cited by examiner]
US 20080294870A1 · Rajopadhye et al. · 2008 [cited by applicant]
US 20090150647A1 · Mejdrich et al. · 2009 [cited by applicant]
US 20100122064A1 · Vorbach · 2010 [cited by examiner]
US 20100257329A1 · Khailany et al. · 2010 [cited by applicant]
US 20110219207A1 · Suh et al. · 2011 [cited by applicant]
US 20120089792A1 · Fahs et al. · 2012 [cited by applicant]
US 20140365548A1 · Mortensen · 2014 [cited by applicant]
US 20150120631A1 · Serrano et al. · 2015 [cited by applicant]
US 20160163016A1 · Gould et al. · 2016 [cited by applicant]
US 20160283240A1 · Mishra et al. · 2016 [cited by applicant]
US 20160342889A1 · Thorson et al. · 2016 [cited by applicant]
US 20170161064A1 · Vasilyev et al. · 2017 [cited by applicant]
US 20170371654A1 · Bajic et al. · 2017 [cited by applicant]
US 20180004530A1 · Vorbach · 2018 [cited by examiner]
US 20190087716A1 · Du et al. · 2019 [cited by applicant]
CN 104603766A · 2015 [cited by applicant]
CN 105278917A · 2016 [cited by applicant]
CN 105930902 · 2016 [cited by applicant]
CN 106293640A · 2017 [cited by applicant]
CN 208061184U · 2018 [cited by applicant]
EP 0232827 · 1987 [cited by applicant]
GB 2484906 · 2012 [cited by applicant]
TW 270192 · 1996 [cited by applicant]
TW 201617977 · 2016 [cited by applicant]
TW 201638788A · 2016 [cited by applicant]
WO WO199120027 · 1991 [cited by applicant]
Arora.,“The Architecture and Evolution of CPU-GPU Systems for General Purpose Computing,” Jan. 1, 2012, 12 pages. [cited by applicant]
ausairpower.net [online], “Vector Processing Futures,” last updated on Jan. 27, 2014, retrieved on Nov. 3, 2016, retrieved from URL<http://www.ausairpower.net/OSR-0600.html>, 10 pages. [cited by applicant]
Calhoun et al., “Stream Vector Processing Unit: Stream Processing Using SIMD on a General Purpose Processor,” Elec525, Spring 2004, retrieved on Feb. 16, 2017, retrieved from URL<http://www.owlnet.rice.edu/˜elec525/proj… [cited by applicant]
CN Office Action issued in Chinese Appln. No. 201721706109.2, dated May 11, 2018, 4 pages. [cited by applicant]
EP Office Action in European Appln. No. 17199241.5, dated Jun. 21, 2022, 8 pages. [cited by applicant]
EP Office Action in European Appln. No. 17199241.5, dated Oct. 16, 2020, 7 pages. [cited by applicant]
Extended European Search Report issued in European Appln. No. 17199241.5, dated Jun. 7, 2018, 14 pages. [cited by applicant]
GB Office Action in Great Britain Application No. GB2003781.8, dated Jan. 25, 2021, 4 pages (with English translation). [cited by applicant]
GB Office Action in Great Britain Appln. No. GB1717851.8, dated Nov. 20, 2019, 4 pages. [cited by applicant]
GB Office Action issued in British Application No. GB1717851.8, dated Apr. 13, 2018, 8 pages. [cited by applicant]
International Preliminary Report on Patentability issued in International Appln. No. PCT/US2017/058561, mailed on Sep. 10, 2019, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2017058561, mailed on Feb. 13, 2018, 21 pages. [cited by applicant]
Kah-Hyong et al., “Efficient Hardware Accelerators for the Computation of Tchebichef Moments,” IEEE Transactions on Circuits and Systems for Video Technology, Mar. 2012, 22(3):414-425. [cited by applicant]
Manadhata et al., “Vector Processors,” retrieved on Feb. 16, 2017, retrieved from URL<http://www.cs.cmu.edu/afs/cs.cmu.edu/academic/class/15740-f03/www/lectures/vector.pdf>, 4 pages. [cited by applicant]
Patterson, “Lecture 6: Vector Processing,” Powerpoint, Spring 1998, retrieved on Feb. 16, 2017, retrieved from URL<https://people.eecs.berkeley.edu/˜pattrsn/252S98/Lec06-vector.pdf>, 60 pages. [cited by applicant]
Soliman et al. “A shared matrix unit for a chip multi-core processor,” Journal of Parallel and Distributed Computing, Mar. 21, 2013, 11 pages. [cited by applicant]
TW Office Action in Taiwan Appln. No. 108110038, dated Sep. 24, 2020, 7 pages (with English translation). [cited by applicant]
TW Office Action in Taiwanese Appln. No. 108110038, dated Jul. 2, 2019, 4 pages (with English translation). [cited by applicant]
wikipedia.org [online] “SerDes,” Jun. 9, 2016, retrieved on Feb. 5, 2018, retrieved from URL<https://en.wikipedia.org/w/index.php?title=SerDes&oldid=724463696>, 3 pages. [cited by applicant]
wikipedia.org [online], “PowerPC,” last edited on Feb. 29, 2020, retrieved on May 6, 2020, retrieved from URL<https://en.wikipedia.org/wiki/PowerPC>. [cited by applicant]
Office Action in Chinese Appln. No. 201711296156.9, dated Apr. 28, 2023, 12 pages (with English Translation). [cited by applicant]
Fowers et al., “A High Memory Bandwidth FPGA Accelerator for Sparse Matrix-Vector Multiplication,” Presented at 22nd Annual International Symposium on Field-Programmable Custom Computing Machines, Boston, MA, USA, Jul. … [cited by applicant]
Office Action in Chinese Appln. No. 201711296156.9, dated Jan. 20, 2023, 13 pages (with English Translation). [cited by applicant]
Yan, “Design and Optimization of Parallel Vector Memory Access Units,” China's Master's Theses, Mar. 15, 2016, 95 pages (with English Abstract). [cited by applicant]
Office Action in DE Appln. No. 10 2017 125 348.3, dated Mar. 21, 2023, 27 pages (with English Translation). [cited by applicant]
Patel et al., “Accelerator Architectures,” IEEE Micro, Jul.-Aug. 2008, pp. 4-12. [cited by applicant]
wikipedia.com [online], “PCI Express,” Nov. 2, 2002, retrieved on Apr. 20, 2023, retrieved from URL<https://en.wikipedia.org/w/index.php?title=PCI Express&oldid=693750119>, 23 pages. [cited by applicant]
wikipedia.com [online], “Hardware random number generator,” Dec. 22, 2002, retrieved on Apr. 20, 2023, retrieved from URL<https://en.wikipedia.org/w/index.php?%20%E2%80%94%20title=Hardware_%20random_%20number_%20generat… [cited by applicant]
Office Action in Taiwanese Appln. No. 112103841, dated Sep. 1, 2023, 7 pages (with English translation). [cited by applicant]
Notice of Allowance in Taiwanese Appln. No. 113113006, mailed on Sep. 3, 2024, 21 pages (with English translation). [cited by applicant]
Shaaban, “Introduction to Vector Processing,” Presentation, Presented as Part of Class EECC 722, Rochester Institute of Technology (RIT), Oct. 10, 2012, 14 pages. [cited by applicant]
Office Action in European Appln. No. 25152213.2, mailed on Mar. 31, 2025, 10 pages [cited by applicant]