IP Library Granted Patent US 10,592,583
Granted Patent B2
US 10,592,583 · App. 16/283,913 · Granted Mar 17, 2020

Permuting in a matrix-vector processor

Inventors: Dong Hyuk Woo (San Jose, CA); Gregory Michael Thorson (Waunakee, WI); Andrew Everett Phelps (Middleton, WI); Olivier Temam (Antony, FR); Jonathan Ross (Menlo Park, CA); Christopher Aaron Clark (Madison, WI)
Assignee: Google LLC
G06F17/16G06F7/76G06F9/30032G06F9/30036G06N3/063G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,592,583
App. No.
16/283,913
Granted
Mar 17, 2020
Kind
B2
Abstract

A circuit comprises an input register configured to receive an input vector of elements, a control register configured to receive a control vector of elements, wherein each element of the control vector corresponds to a respective element of the input vector, and wherein each element specifies a permutation of a corresponding element of the input vector, and a permute execution circuit configured to generate an output vector of elements corresponding to a permutation of the input vector. Generating each element of the output vector comprises accessing, at the input register, a particular element of the input vector, accessing, at the control register, a particular element of the control vector corresponding to the particular element of the input vector, and outputting the particular element of the input vector as an element at a particular position of the output vector that is selected based on the particular element of the control vector.

Claims (40)

1. A matrix-vector permuting system, comprising:

a staggered memory read flattener configured to:

perform a staggered read of data corresponding to an input matrix or vector;

flatten the data; and

output a flattened vector;

an input register configured to receive an input vector of input elements corresponding to the flattened vector;

a permute execution circuit configured to generate an output vector from the input vector, wherein generating each output element of the output vector comprises:

accessing, at a particular position of a control register, a particular control element of a control vector corresponding to a particular position of the output vector;

selecting a particular position of the input register based on the particular control element of the control vector corresponding to the particular position of the output vector;

accessing, at the particular position of the input register, a particular input element of the input vector; and

outputting the particular input element of the input vector as an output element at the particular position of the output vector; and

a staggered memory writer configured to receive the output elements and write the output elements into memory.

2. The system of claim 1 , wherein the permute execution circuit is configured to output the output vector as a staggered output such that a single output element of the output vector is output at each cycle in an order beginning with a lowest-order bit position of the output vector.

3. The system of claim 1 , wherein the input vector is received as a staggered input such that each input element of the input vector is received at each cycle in an order beginning with a lowest-order bit position of the input vector.

4. The system of claim 3 , wherein receiving the input vector as a staggered input comprises:

receiving, at a flattening register, each input element of the input vector in a separate cycle in an order beginning with a lowest-order bit position of the flattening register; and

popping all input elements of the input vector simultaneously from the flattening register to the input register.

5. The system of claim 4 , wherein popping all input elements of the input vector simultaneously from the flattening register to the input register comprises:

determining that a highest-order bit of the flattening register has received valid data; and

popping all input elements of the input vector simultaneously from the flattening register to the input register in response to determining that the highest-order bit of the flattening register has received valid data.

6. The system of claim 4 , wherein popping all input elements of the input vector simultaneously from the flattening register to the input register comprises:

determining that the flattening register has received a number of input elements of the input vector equal to a dimension of the input vector; and

popping all input elements of the input vector simultaneously from the flattening register to the input register in response to determining that the flattening register has received the number of input elements of the input vector equal to the dimension of the input vector.

7. The system of claim 1 , wherein each control element of the control vector specifies a number of positions to rotate the input element in the corresponding position of the input vector.

8. The system of claim 1 , wherein the control vector is received from an off-chip processor that is separate from the circuit.

9. The system of claim 1 , wherein the permute execution circuit comprises a memory crossbar.

10. The circuit of claim 1 , wherein the permute execution circuit comprises multiple one-to-many multiplexors, and wherein each control element of the control vector is a control signal for controlling the output of a corresponding multiplexor of the permute execution circuit.

11. The system of claim 1 , wherein the input vector of input elements corresponds to a row of an input matrix or a column of an input matrix.

12. A circuit for permuting an input matrix, the circuit comprising:

a staggered memory read flattener configured to:

perform a staggered read of data corresponding to an input matrix or vector;

flatten the data; and

output a flattened vector;

an input register configured to receive an input vector of input elements corresponding to the flattened vector;

a permute execution circuit configured to generate an output vector from the input vector, wherein generating each output element of the output vector comprises:

accessing, at a particular position of a control register, a particular control element of a control vector corresponding to a particular position of the output vector;

selecting a particular position of the input register based on the particular control element of the control vector corresponding to the particular position of the output vector;

accessing, at the particular position of the input register, a particular input element of the input vector; and

outputting the particular input element of the input vector as an output element at the particular position of the output vector; and

a staggered memory writer configured to receive the output elements and write the output elements into memory.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2019
From: WOO, DONG HYUK; THORSON, GREGORY MICHAEL; PHELPS, ANDREW EVERETT; TEMAM, OLIVIER; CLARK, CHRISTOPHER AARON
To: GOOGLE INC.
Reel/Frame 048428/0086 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2019
From: ROSS, JONATHAN
To: GOOGLE LLC
Reel/Frame 048428/0167 →
ENTITY CONVERSION Recorded Feb 25, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 048430/0001 →
Continuity (4)
Continuation 15966275 · Apr 30, 2018
Continuation 15496418 · Apr 25, 2017
Provisional Application 62460394 · Feb 17, 2017
Related Publication 20190258694A1 · Aug 22, 2019