IP Library Granted Patent US 6,898,691
Granted Patent B2
US 6,898,691 · App. 10/164,040 · Granted May 24, 2005

Rearranging data between vector and matrix forms in a SIMD matrix processor

Assignee: Intrinsity, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 6,898,691
App. No.
10/164,040
Granted
May 24, 2005
Kind
B2
Abstract

This invention discloses a group of instructions, block 4 and block 4 v, in a matrix processor 16 that rearranges data between vector and matrix forms of an A×B matrix of data 120 where the data matrix includes one or more 4×4 sub-matrices of data 160-166. The instructions of this invention simultaneously swaps row or columns between the first 140, second 142, third 144, and fourth 146 matrix registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between the different individual matrix registers, or swapping columns between the different individual matrix registers. Additionally, successive iterations or combinations of the block 4 and or block 4 v instructions perform standard tensor matrix operations from the following group of matrix operations: transpose, shuffle, and deal.

Claims (58)

1. A group of instructions in a Matrix Processor that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

16 processing elements where an individual processing element (PE) comprises one or more PE register entries in a PE register file;

a mesh row column interconnect that couples said processing elements into a 4×4 matrix processing array;

a first, second, third, and fourth matrix register wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register;

wherein said first, second, third, and fourth matrix registers simultaneously swaps row or columns between said first, second, third, and fourth matrix registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix registers, or swapping columns between said first, second, third, and fourth matrix registers.

2. A Matrix Processor that includes instructions that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

16 processing elements where an individual processing element (PE) comprises one or more PE register entries in a PE register file;

a mesh row column interconnect that couples said processing elements into a 4×4 matrix processing array;

a first, second, third, and fourth matrix register wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register;

wherein said first, second, third, and fourth matrix registers simultaneously swaps row or columns between said first, second, third, and fourth matrix registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix registers, or swapping columns between said first, second, third, and fourth matrix registers.

3. A system that includes a Matrix Processor with instructions that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

16 processing elements where an individual processing element (PE) comprises one or more PE register entries in a PE register file;

a mesh row column interconnect that couples said processing elements into a 4×4 matrix processing array;

a first, second, third, and fourth matrix register wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register;

wherein said first, second, third, and fourth matrix registers simultaneously swaps row or columns between different said first, second, third, and fourth registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix registers, or swapping columns between said first, second, third, and fourth matrix registers.

4. A method to make a Matrix Processor that includes instructions that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

providing 16 processing elements where an individual processing element (PE) comprises one or more PE register entries in a PE register file;

coupling said processing elements into a 4×4 matrix processing array with a mesh row column interconnect;

providing a first, second, third, and fourth matrix register wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register;

wherein said first, second, third, and fourth matrix registers simultaneously swaps row or columns between said first, second, third, and fourth matrix registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix registers, or swapping columns between said first, second, third, and fourth matrix registers.

5. A method to use instructions in a Matrix Processor that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

providing 16 processing elements where an individual processing element (PE) comprises one or more PE register entries in a PE register file;

providing a mesh row column interconnect that couples said processing elements into a 4×4 matrix processing array;

providing a first, second, third, and fourth matrix register wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register; and

simultaneously swapping row or columns between said first, second, third, and fourth matrix registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix registers, or swapping columns between said first, second, third, and fourth matrix registers.

6. A dependent claim according to claim 1 , 2 , 3 , 4 , or 5 wherein successive iterations or combinations of the instructions perform standard tensor matrix operations from the following group of matrix operations: transpose, shuffle, and deal.

7. A dependent claim according to claim 1 , 2 , 3 , 4 , or 5 wherein the swapping of rows or columns converts the data in the data matrix into one of the following matrix data orders: 4 vectors of the larger data matrix to a 4×4 data sub-matrix in row major order, and 4 vectors of the larger data matrix to a 4×4 data sub-matrix in column major order.

8. A group of instructions in a Matrix Processor that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

16 processing elements where an individual processing element (PE) comprises 16 PE register entries in a PE register file;

a mesh row column interconnect that couples said processing elements into a 4×4 matrix processing array;

16 matrix registers wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register;

wherein a group of said 16 matrix registers comprises a first, second, third, and fourth matrix register of said 16 matrix registers that simultaneously swaps row or columns between said first, second, third, and fourth matrix registers of said group of matrix registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix register of said group of matrix registers, or swapping columns between said first, second, third, and fourth matrix register of said group of matrix registers; and

wherein the swapping of rows or columns converts the data in the data matrix into one of the following matrix data orders: 4 vectors of the larger data matrix to a 4×4 data sub-matrix in row major order, and 4 vectors of the larger data matrix to a 4×4 data sub-matrix in column major order.

9. A Matrix Processor that includes instructions that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

16 processing elements where an individual processing element (PE) comprises one or more PE register entries in a PE register file;

a mesh row column interconnect that couples said processing elements into a 4×4 matrix processing array;

16 matrix registers wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register;

wherein a group of said 16 matrix registers comprises a first, second, third, and fourth matrix register of said 16 matrix registers that simultaneously swaps row or columns between said first, second, third, and fourth matrix register of said group of matrix registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix register of said group of matrix registers, or swapping columns between said first, second, third, and fourth matrix register of said group of matrix registers; and

wherein the swapping of rows or columns converts the data in the data matrix into one of the following matrix data orders: 4 vectors of the larger data matrix to a 4×4 data sub-matrix in row major order, and 4 vectors of the larger data matrix to a 4×4 data sub-matrix in column major order.

10. A system that includes a Matrix Processor with instructions that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

16 processing elements where an individual processing element (PE) comprises one or more PE register entries in a PE register file;

a mesh row column interconnect that couples said processing elements into a 4×4 matrix processing array;

16 matrix registers wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register;

wherein a group of said 16 matrix registers comprises a first, second, third, and fourth matrix register that simultaneously swaps row or columns between said first, second, third, and fourth matrix register of said group of matrix registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix register of said group of matrix registers, or swapping columns between said first, second, third, and fourth matrix register of said group of matrix registers; and

wherein the swapping of rows or columns converts the data in the data matrix into one of the following matrix data orders: 4 vectors of the larger data matrix to a 4×4 data sub-matrix in row major order, and 4 vectors of the larger data matrix to a 4×4 data sub-matrix in column major order.

11. A method to make a Matrix Processor that includes instructions that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

providing 16 processing elements where an individual processing element (PE) comprises one or more PE register entries in a PE register file;

coupling said processing elements into a 4×4 matrix processing array with a mesh row column interconnect;

providing 16 matrix registers wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register;

wherein a group of said 16 matrix registers comprises a first, second, third, and fourth matrix register that simultaneously swaps row or columns between said first, second, third, and fourth matrix register of said group of matrix registers according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix register of said group of matrix registers, or swapping columns between said first, second, third, and fourth matrix register of said group of matrix registers; and

wherein the swapping of rows or columns converts the data in the data matrix into one of the following matrix data orders: 4 vectors of the larger data matrix to a 4×4 data sub-matrix in row major order, and 4 vectors of the larger data matrix to a 4×4 data sub-matrix in column major order.

12. A method to use instructions in a Matrix Processor that rearranges data between vector and matrix forms of an A×B matrix of data where the data matrix includes one or more 4×4 sub-matrices of data, comprising:

providing 16 processing elements where an individual processing element (PE) comprises one or more PE register entries in a PE register file;

providing a mesh row column interconnect that couples said processing elements into a 4×4 matrix processing array;

providing 16 matrix registers wherein an individual matrix register comprises an individual PE register entry from each said PE register file from each said individual processing element that are then combined together to from said individual matrix register; and

simultaneously swapping row or columns between a group of said 16 matrix registers that comprise a first, second, third, and fourth matrix register according to the instructions that perform predefined matrix tensor operations on the data matrix that includes one of the following group of operations: swapping rows between said first, second, third, and fourth matrix register of said group of matrix registers, or swapping columns between said first, second, third, and fourth matrix register of said group of matrix registers;

wherein the swapping of rows or columns converts the data in the data matrix into one of the following matrix data orders: 4 vectors of the larger data matrix to a 4×4 data sub-matrix in row major order, and 4 vectors of the larger data matrix to a 4×4 data sub-matrix in column major order.

13. A dependent claim according to claim 8 , 9 , 10 , 11 , or 12 wherein successive iterations or combinations of the instructions perform standard tensor matrix operations from the following group of matrix operations: transpose, shuffle, and deal.

Assignments (7)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2010
From: INTRINSITY, INC.
To: APPLE INC.
Reel/Frame 024380/0329 →
RELEASE OF SECURITY INTEREST Recorded Jan 11, 2008
From: SILICON VALLEY BANK
To: INTRINSITY, INC
Reel/Frame 020525/0485 →
GRANT OF SECURITY INTEREST Recorded Dec 13, 2007
From: INTRINSITY INC.
To: PATENT SKY LLC
Reel/Frame 020234/0365 →
RELEASE OF SECURITY INTEREST Recorded Dec 7, 2007
From: ADAMS CAPITAL MANAGEMENT III, L.P.
To: INTRINSITY, INC.
Reel/Frame 020206/0340 →
SECURITY AGREEMENT Recorded Apr 17, 2007
From: INTRINSITY, INC.
To: ADAMS CAPITAL MANAGEMENT III, L.P.
Reel/Frame 019161/0661 →
SECURITY AGREEMENT Recorded Feb 12, 2007
From: INTRINSITY, INC.
To: SILICON VALLEY BANK
Reel/Frame 018923/0404 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2003
From: OLSON, TIMOTHY A.; BLOMGREN, JAMES S.; HARLE, CHRISTOPHE
To: INTRINSITY, INC.
Reel/Frame 014056/0642 →
Continuity (3)
Provisional Application 6037417400 · Apr 19, 2002
Provisional Application 6029641000 · Jun 6, 2001
Related Publication 20020198911A1 · Dec 26, 2002