IP Library Granted Patent US 11,520,581
Granted Patent B2
US 11,520,581 · App. 17/327,957 · Granted Dec 6, 2022

Vector processing unit

Inventors: William Lacy (Madison, WI); Gregory Michael Thorson (Waunakee, WI); Christopher Aaron Clark (Madison, WI); Norman Paul Jouppi (Palo Alto, CA); Thomas Norrie (Mountain View, CA); Andrew Everett Phelps (Middleton, WI)
Assignee: Google LLC
G06F9/3001G06F7/588G06F9/30032G06F9/30036G06F9/30043G06F9/30098G06F9/3891G06F13/36G06F13/4068G06F13/4282G06F15/8053G06F15/8092G06F17/16G06F15/8046G06N3/063G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,520,581
App. No.
17/327,957
Granted
Dec 6, 2022
Kind
B2
Abstract

A vector processing unit is described, and includes processor units that each include multiple processing resources. The processor units are each configured to perform arithmetic operations associated with vectorized computations. The vector processing unit includes a vector memory in data communication with each of the processor units and their respective processing resources. The vector memory includes memory banks configured to store data used by each of the processor units to perform the arithmetic operations. The processor units and the vector memory are tightly coupled within an area of the vector processing unit such that data communications are exchanged at a high bandwidth based on the placement of respective processor units relative to one another, and based on the placement of the vector memory relative to each processor unit.

Claims (24)

1. A system comprising:

a plurality of vector processing units; and

a plurality of matrix units coupled to the plurality of vector processing units such that data communications can be exchanged, each matrix unit being configured to perform multiplications between weights of a neural network and activation inputs to generate accumulated values,

wherein each vector processing unit is arranged in a corresponding vector processing unit (VPU) lane, and wherein each vector processing unit comprises:

a plurality of processor units arranged across multiple sub-lanes of the VPU lane, wherein each processor units comprises an arithmetic logic unit (ALU) configured to perform arithmetic operations associated with vectorized computations for a multi-dimensional data array; and

a corresponding vector memory in data communication with the plurality of processor units, wherein the vector memory includes memory banks configured to store data used by the plurality of processor units to perform the arithmetic operations,

wherein the plurality of processor units and the corresponding vector memory are tightly coupled within an area of the vector processing unit such that data communications can be exchanged at a high bandwidth based on the placement of respective processor units relative to one another and based on the placement of the vector memory relative to each processor unit.

2. The system of claim 1 , wherein each processor unit of the plurality of processor units comprises at least one ALU.

3. The system of claim 1 , wherein each vector processor unit comprises 16 ALUs.

4. The system of claim 1 , wherein the vector memory comprises static random access memory (SRAM).

5. The system of claim 1 , wherein the system is configured to allow transfer of 32 bytes between the vector memory and the plurality of processor units during a single clock cycle.

6. The system of claim 1 , wherein the vector processing unit is configured to perform vector computations based on concurrent use of two or more of the ALUs.

7. The system of claim 1 , wherein each ALU is configured to perform a 32-bit arithmetic operation between streams of vector data that represent operands for the arithmetic operation.

8. The system of claim 1 , wherein the plurality of matrix units and the plurality of vector processing units represent a processor core of an integrated circuit chip; and

the processor core is confugured to processs a single instruction stream at least across the multiple sub-lanes.

9. The system of claim 1 , wherein:

units of the system are configured to operate on streams of data;

a first stream of data progresses in a first direction toward the plurality of matrix units; and

a second, different stream of data progresses in a second direction away from the plurality of matrix units.

10. The system of claim 1 , wherein at least one processor unit comprises a plurality of ALUs, and wherein multiple ALUs within a single processor unit are configured to execute arithmetic operations simultaneously during a single processor clock cycle.

11. The system of claim 1 , wherein multiple ALUs within a single processor unit are configured to execute arithmetic operations simultaneously during a single processor clock cycle.

12. The system of claim 1 , wherein the system is configured to perform at least 2048 operations in a single clock cycle.

13. The system of claim 1 , wherein each operation includes a 32-bit word.

14. The system of claim 1 , wherein a VPU lane is configured to move 8 vectors from a corresponding memory unit of the VPU lane to 8 sub-lanes of the VPU lane within a single clock cycle.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2021
From: LACY, WILLIAM; THORSON, GREGORY MICHAEL; CLARK, CHRISTOPHER AARON; JOUPPI, NORMAN PAUL; NORRIE, THOMAS; PHELPS, ANDREW EVERETT
To: GOOGLE INC.
Reel/Frame 056569/0482 →
CHANGE OF NAME Recorded Jun 16, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 057226/0966 →
Continuity (4)
Continuation 16843015 · Apr 8, 2020
Continuation 16291176 · Mar 4, 2019
Continuation 15454214 · Mar 9, 2017
Related Publication 20210357212A1 · Nov 18, 2021
Cited By (1)
US 12,399,714