IP Library Granted Patent US 12,688,043
Granted Patent B2
US 12,688,043 · App. 18/125,416 · Granted Jul 21, 2026

Matrix multiplication in a dynamically spatially and dynamically temporally dividable architecture

Inventors: Jesse Garrett Beu (Duluth, GA); Thomas Christopher Grocutt (Cambridge, GB)
Assignee: Arm Limited
G06F9/30145G06F9/3001G06F9/30036G06F9/30098
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,043
App. No.
18/125,416
Filed
Mar 23, 2023
Granted
Jul 21, 2026
Kind
B2
Art Unit
2184
USPC
712/220
Abstract

A data processing apparatus includes first vector registers and second vector registers, both dynamically spatially and dynamically temporally dividable. Decode circuitry receives one or more matrix multiplication instructions that indicate a set of first elements in the first vector registers and a set of second elements in the second vector registers, and in response to receiving the matrix multiplication instructions they generate a matrix multiplication operation. The matrix multiplication operation causes one or more execution units to perform a matrix multiplication of the set of first elements by the set of second elements and an average bit width of the first elements is different to an average bit width of the second elements.

Claims (70)

1 . A data processing apparatus comprising:

first vector registers and second vector registers, both configured to be dynamically spatially and dynamically temporally divided; and

decode circuitry configured to receive one or more matrix multiplication instructions comprising an indication of a set of first elements in the first vector registers and a set of second elements in the second vector registers, and in response to receiving the matrix multiplication instructions to generate a matrix multiplication operation, wherein

the matrix multiplication operation is configured to cause one or more execution units to perform a matrix multiplication of the set of first elements by the set of second elements; and

an average bit width of the first elements is different from an average bit width of the second elements,

wherein:

the matrix multiplication instructions comprise a compressed matrix multiplication instruction that comprises an indication of compression data;

the first elements comprise a single row of n activations, where n is an integer;

the second elements comprise m groups of n compressed weights, where m is an integer greater than 1;

the compression data indicates how the n compressed weights are decompressed to form mn uncompressed weights.

2 . The data processing apparatus according to claim 1 , wherein

the matrix multiplication instructions comprise an indication of a further result register;

the result register is configured to store the first set of bits of the result of the matrix multiplication; and

the further result register is configured to store the second set of bits of the result of the matrix multiplication.

3 . The data processing apparatus according to claim 1 , wherein

the matrix multiplication multiplies fewer rows of the first set of elements than a number of columns of the second set of elements.

4 . The data processing apparatus according to claim 1 , wherein

the matrix multiplication is of one row of the first set of elements and two columns of the second set of elements.

5 . The data processing apparatus according to claim 1 , wherein the bit width of the second elements is four bits or less.

6 . The data processing apparatus according to claim 1 , wherein the bit width of the second elements is one bit.

7 . The data processing apparatus according to claim 6 , wherein

the one or more matrix multiplication instructions comprise an indicator value, or the data processing apparatus comprises a selection register configured to store the indicator value; and

the indicator value is configured to indicate a subset of the weights that are used in the matrix multiplication during a particular beat of the data processing apparatus.

8 . The data processing apparatus according to claim 6 , wherein

bits of at least one of the indication of the set of first elements in the first vector registers and the set of second elements in the second vector registers are used to indicate a subset of the weights that are used in the matrix multiplication during a particular beat of the data processing apparatus.

9 . The data processing apparatus according to claim 1 , wherein the second elements are signed.

10 . The data processing apparatus according to claim 1 , wherein the weights are extended prior to the matrix multiplication.

11 . The data processing apparatus according to claim 1 , wherein

the compression data comprises a plurality of portions, each applicable to one of a plurality of matrix multiplication instructions including the compressed matrix multiplication instruction; and

the compressed matrix multiplication instruction comprises a compression data selector configured to select one of the portions of the compression data to be applied to form the n uncompressed weights.

12 . The data processing apparatus according to claim 11 , wherein

the compression data is applicable to a plurality of the matrix multiplication instructions;

at least some of the matrix multiplication instructions indicate different second elements from each other; and

the compression data comprises a number of items.

13 . The data processing apparatus according to claim 12 , wherein

the compression data is applicable to more than m groups of n weights.

14 . The data processing apparatus according to claim 12 , wherein

the items are ordered in the compression data according to a beat in which they are used within the plurality of matrix multiplication operations.

15 . The data processing apparatus according to claim 12 , wherein

the items are ordered in the compression data such that items used in a same beat of a same single matrix multiplication operation are adjacent.

16 . The data processing apparatus according to claim 11 , wherein

the compression data selector is least significant bits of another parameter of the compressed matrix multiplication instruction.

17 . The data processing apparatus according to claim 11 , wherein

the compression data selector is least significant bits of an address of the second elements.

18 . The data processing apparatus according to claim 17 , wherein

the compression data selector is combinable with a stub address of the first elements to form an address of the first elements; and

the compression data selector is combinable with a stub address of a result register to form an address of result register into which at least a part of the result of the matrix multiplication is stored.

19 . The data processing apparatus according to claim 11 , comprising:

multiplexer circuitry configured to select from between the activations to match with the uncompressed weights that are non-zero to provide as an input to the matrix multiplication.

20 . The data processing apparatus according to claim 19 , wherein

the multiplexer circuitry is configured to select from between a subset of the activations to match with the uncompressed weights that are non-zero.

21 . A data processing apparatus comprising:

first vector registers and second vector registers, both configured to be dynamically spatially and dynamically temporally divided; and

decode circuitry configured to receive one or more matrix multiplication instructions comprising an indication of a set of first elements in the first vector registers and a set of second elements in the second vector registers, and in response to receiving the matrix multiplication instructions to generate a matrix multiplication operation, wherein

the matrix multiplication operation is configured to cause one or more execution units to perform a matrix multiplication of the set of first elements by the set of second elements; and

an average bit width of the first elements is different from an average bit width of the second elements, wherein

the matrix multiplication instructions comprise an uncompressed matrix multiplication instruction;

the first elements comprise a single group of n activations, where n is an integer;

the second elements comprise m groups of n weights, where m is an integer greater than 1; and

a bit width of the second elements is 1/m times a bit width of the first elements.

22 . A non-transitory, computer readable medium containing a computer program for controlling a host data processing apparatus to provide an instruction execution environment comprising:

first data structures and second data structures, both configured to be dynamically spatially and dynamically temporally divided; and

decode logic configured to receive one or more matrix multiplication instructions comprising an indication of a set of first elements in the first data structures and a set of second elements in the second data structures, and in response to receiving the matrix multiplication instructions to generate a matrix multiplication operation, wherein

the matrix multiplication operation is configured to cause execution logic to perform a matrix multiplication of the set of first elements by the set of second elements; and

an average bit width of the first elements is different from an average bit width of the second elements,

wherein:

the matrix multiplication instructions comprise a compressed matrix multiplication instruction that comprises an indication of compression data;

the first elements comprise a single row of n activations, where n is an integer;

the second elements comprise m groups of n compressed weights, where m is an integer greater than 1;

the compression data indicates how the n compressed weights are decompressed to form mn uncompressed weights.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2023
From: BEU, JESSE GARRETT; GROCUTT, THOMAS CHRISTOPHER
To: ARM LIMITED
Reel/Frame 063078/0691 →
Continuity (1)
Related Publication 20240320005A1 · Sep 26, 2024
References Cited (41)
US 7376812B1 · Sanghavi · 2008 [cited by examiner]
US 8555034B2 · Damron · 2013 [cited by examiner]
US 10599428B2 · Grocutt · 2020 [cited by applicant]
US 11379556B2 · Mattina et al. · 2022 [cited by applicant]
US 12175246B2 · Baum · 2024 [cited by examiner]
US 20140195783A1 · Karthikeyan et al. · 2014 [cited by applicant]
US 20180189239A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20190042242A1 · Das et al. · 2019 [cited by applicant]
US 20190121837A1 · Azizi · 2019 [cited by examiner]
US 20200050918A1 · Chen · 2020 [cited by examiner]
US 20200210517A1 · Baum · 2020 [cited by examiner]
US 20200310793A1 · Rubanovich et al. · 2020 [cited by applicant]
US 20200349216A1 · Das Sarma · 2020 [cited by examiner]
US 20210089925A1 · Partovi Nia · 2021 [cited by examiner]
US 20210210840A1 · Xu · 2021 [cited by examiner]
US 20210389948A1 · Beu et al. · 2021 [cited by applicant]
US 20220365783A1 · Shin · 2022 [cited by examiner]
US 20220366008A1 · Shin · 2022 [cited by examiner]
US 20240320292A1 · Beu · 2024 [cited by examiner]
EP 3602277 · 2020 [cited by applicant]
GB 2409059A · 2005 [cited by examiner]
GB 2409066A · 2005 [cited by examiner]
GB 2548601 · 2019 [cited by applicant]
WO WO2005026974A1 · 2005 [cited by examiner]
WO 2018174925 · 2018 [cited by applicant]
WO WO2021041447A1 · 2021 [cited by examiner]
WO WO2024063874A1 · 2024 [cited by examiner]
Machine Translation of Chinese Patent Application CN 107851015 A, 2018. (Year: 2018). [cited by examiner]
Machine Translation of Chinese Patent Application CN 113094096 A, 2021. (Year: 2021). [cited by examiner]
Machine Translation of Japanese Patent Application JP 4426099 B2, 2010. (Year: 2010). [cited by examiner]
‘Grokking RISC-V Vector Processing’ by Erik Engheim, Jan. 5, 2023. (Year: 2023). [cited by examiner]
‘Regularization of Neural Networks using DropConnect’ by Wan et al., 2013. (Year: 2013). [cited by examiner]
‘A Mixed-Precision RISC-V Processor for Extreme-Edge DNN Inference’ by Ottavi et al., 2020. (Year: 2020). [cited by examiner]
“Risc-V “V” Vector—Extension Version 1.0” May 2, 2023. (Year: 2023). [cited by examiner]
‘Vitruvius+: An Area-Efficient RISC-V Decoupled Vector Coprocessor for High Performance Computing Applications’ by Minervini et al., 2023. (Year: 2023). [cited by examiner]
‘CK-12 Algebra II with Trigonometry’ by Jordan and Dirga, Feb. 28, 2015. (Year: 2015). [cited by examiner]
International Search Report and Written Opinion of the International Searching Authority for PCT/GB2024/050262 mailed Jun. 21, 2024, 24 pages. [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority for PCT/GB2024/050277 mailed Jul. 2, 2024, 24 pages. [cited by applicant]
U.S. Appl. No. 18/125,432, filed Mar. 23, 2023, Beu et al. [cited by applicant]
Partial International Search Report for PCT/GB2024/050262 mailed Apr. 30, 2024, 19 pages. [cited by applicant]
Partial International Search Report for PCT/GB2024/050277 mailed May 8, 2024, 19 pages. [cited by applicant]