IP Library Granted Patent US 10,372,787
Granted Patent B2
US 10,372,787 · App. 15/839,229 · Granted Aug 6, 2019

Hardware accelerator pre-configured with coefficients for matrix-transform operations

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,372,787
App. No.
15/839,229
Granted
Aug 6, 2019
Kind
B2
Abstract

A special-purpose hardware accelerator may include a cache configured to store an input matrix related to performing a convolution operation and a matrix-multiplication subsystem pre-configured with matrix-transform coefficients for performing matrix-transform operations. The matrix-multiplication subsystem may perform the convolution operation by (1) reading the input matrix from the cache, (2) transforming the input matrix via matrix multiplication, (3) transforming, via matrix multiplication, a parameter matrix that includes convolution parameters for performing the convolution operation, (4) applying the transformed parameter matrix to the transformed input matrix via an element-wise multiplication operation, and then (5) performing an inverse-transformation operation on the results of the element-wise multiplication operation to create an output matrix for the convolution operation. Various other systems and methods are also disclosed.

Claims (77)

1. A special-purpose hardware accelerator comprising:

a cache configured to store an input matrix related to performing a convolution operation; and

a matrix-multiplication subsystem pre-configured with matrix-transform coefficients for performing matrix-transform operations, wherein the matrix-multiplication subsystem is configured to perform the convolution operation by:

reading the input matrix from the cache;

transforming the input matrix via matrix multiplication;

transforming, via matrix multiplication, a parameter matrix that comprises convolution parameters for performing the convolution operation;

applying the transformed parameter matrix to the transformed input matrix via an element-wise multiplication operation; and

performing an inverse-transformation operation on the results of the element-wise multiplication operation to create an output matrix for the convolution operation.

2. The special-purpose hardware accelerator of claim 1 , wherein the matrix-multiplication subsystem uses the pre-configured matrix-transform coefficients to at least one of:

transform the input matrix;

transform the parameter matrix; or

perform the inverse-transformation operation on the results of the element-wise multiplication operation.

3. The special-purpose hardware accelerator of claim 1 , wherein the pre-configured matrix-transform coefficients comprise at least one of:

pre-configured coefficients for transforming the input matrix;

pre-configured coefficients for transforming the parameter matrix; or

pre-configured coefficients for performing the inverse-transformation operation on the results of the element-wise multiplication operation.

4. The special-purpose hardware accelerator of claim 3 , wherein the pre-configured matrix-transform coefficients comprise at least one of:

a transposed version of the pre-configured coefficients for transforming the input matrix;

a transposed version of the pre-configured coefficients for transforming the parameter matrix; or

a transposed version of the pre-configured coefficients for performing the inverse-transformation operation on the results of the element-wise multiplication operation.

5. The special-purpose hardware accelerator of claim 1 , wherein the matrix-multiplication subsystem comprises at least one of:

a dot-product engine configured to perform the matrix-transform operations; or

an element-wise multiplier configured to perform the element-wise multiplication operation.

6. The special-purpose hardware accelerator of claim 5 , wherein the matrix-multiplication subsystem transforms the input matrix on-the-fly when reading the input matrix from the cache to the element-wise multiplier.

7. The special-purpose hardware accelerator of claim 5 , wherein the parameter matrix is stored in the cache and the matrix-multiplication subsystem transforms the parameter matrix on-the-fly when reading the parameter matrix from the cache to the element-wise multiplier.

8. The special-purpose hardware accelerator of claim 1 , wherein the matrix-multiplication subsystem performs the inverse-transformation operation on-the-fly when storing the output matrix to the cache.

9. The special-purpose hardware accelerator of claim 1 , wherein:

the input matrix comprises the entirety of an input volume for the convolution operation;

transforming the input matrix via matrix multiplication comprises transforming the entire input volume via matrix multiplication; and

the output matrix comprises the entirety of an output volume for the convolution operation.

10. The special-purpose hardware accelerator of claim 1 , wherein:

the input matrix comprises an initial portion of an input volume for the convolution operation;

transforming the input matrix via matrix multiplication comprises transforming the initial portion of the input volume via matrix multiplication; and

creating the output matrix comprises creating an initial portion of an output volume for the convolution operation.

11. The special-purpose hardware accelerator of claim 10 , wherein performing the convolution operation further comprises:

receiving at least one additional portion of the input volume from the cache;

transforming the additional portion of the input volume via matrix multiplication;

applying, via an additional element-wise multiplication operation, the transformed parameter matrix to the additional portion of the input volume that was transformed; and

performing an additional inverse-transformation operation on the results of the additional element-wise multiplication operation to create an additional portion of the output volume for the convolution operation.

12. The special-purpose hardware accelerator of claim 1 , wherein the special-purpose hardware accelerator is configured to:

pose each of a plurality of element-wise multiplication operations as a plurality of dot-product operations; and

batch the plurality of dot-product operations into a single matrix-multiplication operation for processing by the matrix-multiplication subsystem.

13. A computing system comprising:

a memory device configured to store an input matrix related to performing a convolution operation; and

a special-purpose hardware accelerator comprising a matrix-multiplication subsystem that is pre-configured with matrix-transform coefficients for performing matrix-transform operations, wherein the matrix-multiplication subsystem is configured to perform the convolution operation by:

reading the input matrix from the memory device;

transforming the input matrix via matrix multiplication;

transforming, via matrix multiplication, a parameter matrix that comprises convolution parameters for performing the convolution operation;

applying the transformed parameter matrix to the transformed input matrix via an element-wise multiplication operation; and

performing an inverse-transformation operation on the results of the element-wise multiplication operation to create an output matrix for the convolution operation.

14. The computing system of claim 13 , wherein:

the matrix-multiplication subsystem comprises a cache; and

reading the input matrix from the memory device comprises:

storing the input matrix from the memory device into the cache; and

reading the input matrix from the cache.

15. The computing system of claim 14 , wherein the matrix-multiplication subsystem comprises at least one of:

a dot-product engine configured to perform the matrix-transform operations; or

an element-wise multiplier configured to perform the element-wise multiplication operation.

16. The computing system of claim 15 , wherein:

the matrix-multiplication subsystem transforms the input matrix on-the-fly when reading the input matrix from the cache to the element-wise multiplier; and

the parameter matrix is stored in the cache and the matrix-multiplication subsystem transforms the parameter matrix on-the-fly when reading the parameter matrix from the cache to the element-wise multiplier.

17. The computing system of claim 14 , wherein the matrix-multiplication subsystem performs the inverse-transformation operation on-the-fly when storing the output matrix to the cache.

18. The computing system of claim 13 , wherein the matrix-multiplication subsystem uses the pre-configured matrix-transform coefficients to at least one of:

transform the input matrix;

transform the parameter matrix; or

perform the inverse-transformation operation on the results of the element-wise multiplication operation.

19. The computing system of claim 13 , wherein the pre-configured matrix-transform coefficients comprise at least one of:

pre-configured coefficients for transforming the input matrix;

pre-configured coefficients for transforming the parameter matrix; or

pre-configured coefficients for performing the inverse-transformation operation on the results of the element-wise multiplication operation.

20. A computer-implemented method comprising:

reading, from a cache of a special-purpose hardware accelerator, an input matrix related to performing a convolution operation, wherein the special-purpose hardware accelerator is pre-configured with matrix-transform coefficients for performing matrix-transform operations; and

performing, using the special-purpose hardware accelerator, the convolution operation by:

transforming the input matrix via matrix multiplication;

transforming, via matrix multiplication, a parameter matrix that comprises convolution parameters for performing the convolution operation;

applying the transformed parameter matrix to the transformed input matrix via an element-wise multiplication operation; and

performing an inverse-transformation operation on the results of the element-wise multiplication operation to create an output matrix for the convolution operation.

Assignments (2)
CHANGE OF NAME Recorded Jan 27, 2022
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058871/0336 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2018
From: PARK, JONG SOO; ROTEM, NADAV; SMELYANSKIY, MIKHAIL; DIRIL, ABDULKADIR UTKU
To: FACEBOOK, INC.
Reel/Frame 044540/0577 →