IP Library › Granted Patent US 12,461,745
Granted Patent B2
US 12,461,745 · App. 18/521,000 · Granted Nov 4, 2025

Systems for performing instructions to quickly convert and use tiles as 1D vectors

Inventors: Bret Toll (Hillsboro, OR); Christopher J. Hughes (Santa Clara, CA); Dan Baum (Haifa, IL); Elmoustapha Ould-Ahmed-Vall (Gilbert, AZ); Raanan Sade (Portland, OR); Robert Valentine (Kiryat Tivon, IL); Mark J. Charney (Lexington, MA); Alexander F. Heinecke (San Jose, CA)
Assignee: Intel Corporation
G06F9/30145G06F9/30032G06F9/30036G06F9/30038G06F9/30109
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,461,745
App. No.
18/521,000
Granted
Nov 4, 2025
Kind
B2
Abstract

Disclosed embodiments relate to systems for performing instructions to quickly convert and use matrices (tiles) as one-dimensional vectors. In one example, a processor includes fetch circuitry to fetch an instruction having fields to specify an opcode, locations of a two-dimensional (2D) matrix and a one-dimensional (1D) vector, and a group of elements comprising one of a row, part of a row, multiple rows, a column, part of a column, multiple columns, and a rectangular sub-tile of the specified 2D matrix, and wherein the opcode is to indicate a move of the specified group between the 2D matrix and the 1D vector, decode circuitry to decode the fetched instruction; and execution circuitry, responsive to the decoded instruction, when the opcode specifies a move from 1D, to move contents of the specified 1D vector to the specified group of elements.

Claims (41)

1 . A processor comprising:

a plurality of vector registers, each vector register of the plurality of vector registers to store a plurality of matrix data elements;

execution circuitry to execute an instruction to load a first plurality of source data elements from a first column and a second column of a first source tile to a first vector register of the plurality of vector registers, wherein a first subset of the first plurality of source data elements from the first column are to be interleaved with a second subset of the first plurality of source data elements from the second column within the first vector register, the first source tile comprising group of rows and columns of a first source matrix stored in a memory;

the execution circuitry further comprising:

a set of multipliers to perform a parallel multiplication of each data element of the first plurality of source data elements stored in the first vector register with a corresponding source data element of a second plurality of source data elements stored in a second vector register of the plurality of vector registers to generate a corresponding plurality of products, the second plurality of source data elements from a second source tile of a second source matrix to be multiplied with the first source matrix; and

accumulator circuitry to add groups of the corresponding plurality of products to corresponding accumulated data elements of an accumulation matrix to generate corresponding result data elements of a result matrix.

2 . The processor of claim 1 , wherein the accumulated data elements and result data elements have a width which is at least two times a width of the data elements of the first and second plurality of source data elements.

3 . The processor of claim 2 , wherein the accumulated data elements and result data elements comprise 32-bit floating-point data elements and the data elements of the first and second plurality of source data elements comprise 16-bit floating point data elements.

4 . The processor of claim 2 , wherein the accumulated data elements and result data elements comprise 32-bit integer data elements and the data elements of the first and second plurality of source data elements comprise 8-bit or 4-bit integer data elements.

5 . The processor of claim 1 , wherein source data elements from the first and second columns are interleaved in the first vector register in a sequential order in which source data elements from a same row are adjacent and in which at least one data element from a row is adjacent to another data element from a next sequential row when loaded in the first vector register.

6 . The processor of claim 1 , wherein the instruction includes a first operand to indicate an address of the first source tile in the memory and a second operand to indicate the first vector register.

7 . The processor of claim 1 , wherein the plurality of rows and columns of the first source matrix comprises 16 rows and 16 columns.

8 . A system comprising:

a memory; and

matrix operations circuitry coupled to the memory, the matrix operations circuitry comprising:

a plurality of vector registers, each vector register of the plurality of vector registers to store a plurality of matrix data elements;

execution circuitry to execute an instruction to load a first plurality of source data elements from a first column and a second column of a first source tile to a first vector register of the plurality of vector registers, wherein a first subset of the first plurality of source data elements from the first column are to be interleaved with a second subset of the first plurality of source data elements from the second column within the first vector register, the first source tile comprising group of rows and columns of a first source matrix stored in the memory;

the execution circuitry further comprising:

a set of multipliers to perform a parallel multiplication of each data element of the first plurality of source data elements stored in the first vector register with a corresponding source data element of a second plurality of source data elements stored in a second vector register of the plurality of vector registers to generate a corresponding plurality of products, the second plurality of source data elements from a second source tile of a second source matrix to be multiplied with the first source matrix; and

accumulator circuitry to add groups of the corresponding plurality of products to corresponding accumulated data elements of an accumulation matrix to generate corresponding result data elements of a result matrix.

9 . The system of claim 8 , wherein the accumulated data elements and result data elements have a width which is at least two times a width of the data elements of the first and second plurality of source data elements.

10 . The system of claim 9 , wherein the accumulated data elements and result data elements comprise 32-bit floating-point data elements and the data elements of the first and second plurality of source data elements comprise 16-bit floating point data elements.

11 . The system of claim 9 , wherein the accumulated data elements and result data elements comprise 32-bit integer data elements and the data elements of the first and second plurality of source data elements comprise 8-bit or 4-bit integer data elements.

12 . The system of claim 8 , wherein source data elements from the first and second columns are interleaved in the first vector register in a sequential order in which source data elements from a same row are adjacent and in which at least one data element from a row is adjacent to another data element from a next sequential row when loaded in the first vector register.

13 . The system of claim 8 , wherein the instruction includes a first operand to indicate an address of the first source tile in the memory and a second operand to indicate the first vector register.

14 . The system of claim 8 , wherein the plurality of rows and columns of the first source matrix comprises 16 rows and 16 columns.

15 . A system comprising:

a memory;

a processor; and

matrix operations circuitry coupled to the memory and the processor, the matrix operations circuitry comprising:

a plurality of vector registers, each vector register of the plurality of vector registers to store a plurality of matrix data elements;

execution circuitry to execute an instruction to load a first plurality of source data elements from a first column and a second column of a first source tile to a first vector register of the plurality of vector registers, wherein a first subset of the first plurality of source data elements from the first column are to be interleaved with a second subset of the first plurality of source data elements from the second column within the first vector register, the first source tile comprising group of rows and columns of a first source matrix stored in the memory;

the execution circuitry further comprising:

a set of multipliers to perform a parallel multiplication of each data element of the first plurality of source data elements stored in the first vector register with a corresponding source data element of a second plurality of source data elements stored in a second vector register of the plurality of vector registers to generate a corresponding plurality of products, the second plurality of source data elements from a second source tile of a second source matrix to be multiplied with the first source matrix; and

accumulator circuitry to add groups of the corresponding plurality of products to corresponding accumulated data elements of an accumulation matrix to generate corresponding result data elements of a result matrix.

16 . The system of claim 15 , wherein the accumulated data elements and result data elements have a width which is at least two times a width of the data elements of the first and second plurality of source data elements.

17 . The system of claim 16 , wherein the accumulated data elements and result data elements comprise 32-bit floating-point data elements and the data elements of the first and second plurality of source data elements comprise 16-bit floating point data elements.

18 . The system of claim 16 , wherein the accumulated data elements and result data elements comprise 32-bit integer data elements and the data elements of the first and second plurality of source data elements comprise 8-bit or 4-bit integer data elements.

19 . The system of claim 15 , wherein source data elements from the first and second columns are interleaved in the first vector register in a sequential order in which source data elements from a same row are adjacent and in which at least one data element from a row is adjacent to another data element from a next sequential row when loaded in the first vector register.

20 . The system of claim 15 , wherein the instruction includes a first operand to indicate an address of the first source tile in the memory and a second operand to indicate the first vector register.

21 . The system of claim 15 , wherein the plurality of rows and columns of the first source matrix comprises 16 rows and 16 columns.

Continuity (5)
Continuation 17549363 · Dec 13, 2021
Continuation 17549221 · Dec 13, 2021
Continuation 17240882 · Apr 26, 2021
Continuation 16145066 · Sep 27, 2018
Related Publication 20240103867A1 · Mar 28, 2024
References Cited (129)
US 5247632A · Newman · 1993 [cited by applicant]
US 5475822A · Sibigtroth et al. · 1995 [cited by applicant]
US 5892962A · Cloutier · 1999 [cited by applicant]
US 6161219A · Ramkumar et al. · 2000 [cited by applicant]
US 6175892B1 · Sazzad et al. · 2001 [cited by applicant]
US 6212112B1 · Naura et al. · 2001 [cited by applicant]
US 6332186B1 · Elwood et al. · 2001 [cited by applicant]
US 6877020B1 · Bratt et al. · 2005 [cited by applicant]
US 7003542B2 · Devir · 2006 [cited by applicant]
US 7209939B2 · Castrapel et al. · 2007 [cited by applicant]
US 7590300B2 · Tadas et al. · 2009 [cited by applicant]
US 7725521B2 · Chen et al. · 2010 [cited by applicant]
US 7792895B1 · Juffa et al. · 2010 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 7912889B1 · Juffa et al. · 2011 [cited by applicant]
US 7932910B2 · Hansen et al. · 2011 [cited by applicant]
US 8392487B1 · Mesh et al. · 2013 [cited by applicant]
US 8984043B2 · Ginzburg et al. · 2015 [cited by applicant]
US 9442723B2 · Yang et al. · 2016 [cited by applicant]
US 9906359B2 · Gueron · 2018 [cited by applicant]
US 9960907B2 · Gueron · 2018 [cited by applicant]
US 10535114B2 · Bolz · 2020 [cited by applicant]
US 10990396B2 · Toll · 2021 [cited by examiner]
US 11579880B2 · Toll · 2023 [cited by examiner]
US 11714648B2 · Toll · 2023 [cited by examiner]
US 11954489B2 · Toll · 2024 [cited by examiner]
US 20030126176A1 · Devir · 2003 [cited by applicant]
US 20040111587A1 · Nair et al. · 2004 [cited by applicant]
US 20050055535A1 · Moyer et al. · 2005 [cited by applicant]
US 20050063589A1 · Angrilli et al. · 2005 [cited by applicant]
US 20050193050A1 · Sazegari · 2005 [cited by applicant]
US 20060101245A1 · Nair et al. · 2006 [cited by applicant]
US 20060190517A1 · Guerrero · 2006 [cited by applicant]
US 20070186210A1 · Hussain et al. · 2007 [cited by applicant]
US 20080046681A1 · Sandon et al. · 2008 [cited by applicant]
US 20080071851A1 · Zohar et al. · 2008 [cited by applicant]
US 20080140994A1 · Khailany et al. · 2008 [cited by applicant]
US 20080208942A1 · Won et al. · 2008 [cited by applicant]
US 20090043836A1 · Dupaquis et al. · 2009 [cited by applicant]
US 20090292758A1 · Brokenshire et al. · 2009 [cited by applicant]
US 20090300091A1 · Brokenshire et al. · 2009 [cited by applicant]
US 20090300249A1 · Moyer et al. · 2009 [cited by applicant]
US 20100070742A1 · Dowling · 2010 [cited by applicant]
US 20100180100A1 · Lu et al. · 2010 [cited by applicant]
US 20100325187A1 · Juffa et al. · 2010 [cited by applicant]
US 20110055517A1 · Eichenberger et al. · 2011 [cited by applicant]
US 20110107060A1 · Mcallister et al. · 2011 [cited by applicant]
US 20110153707A1 · Ginzburg · 2011 [cited by applicant]
US 20110262052A1 · Kimura · 2011 [cited by applicant]
US 20120079252A1 · Sprangle · 2012 [cited by applicant]
US 20120113133A1 · Shpigelblat · 2012 [cited by applicant]
US 20120137074A1 · Kim et al. · 2012 [cited by applicant]
US 20120254588A1 · Adrian et al. · 2012 [cited by applicant]
US 20120314774A1 · Yang et al. · 2012 [cited by applicant]
US 20130159665A1 · Kashyap · 2013 [cited by applicant]
US 20130198487A1 · Pusdesris et al. · 2013 [cited by applicant]
US 20130305020A1 · Valentine et al. · 2013 [cited by applicant]
US 20140149480A1 · Catanzaro et al. · 2014 [cited by applicant]
US 20140201450A1 · Haugen · 2014 [cited by applicant]
US 20140208065A1 · Ould-Ahmed-Vall · 2014 [cited by applicant]
US 20150067302A1 · Gueron · 2015 [cited by applicant]
US 20150199266A1 · Franchetti et al. · 2015 [cited by applicant]
US 20160011870A1 · Plotnikov et al. · 2016 [cited by applicant]
US 20170337156A1 · Yadavalli · 2017 [cited by applicant]
US 20180074824A1 · Sazegari et al. · 2018 [cited by applicant]
US 20180113708A1 · Corbal et al. · 2018 [cited by applicant]
US 20180336163A1 · Phelps et al. · 2018 [cited by applicant]
CN 102541774A · 2012 [cited by applicant]
CN 103959233A · 2014 [cited by applicant]
EP 3001306A1 · 2016 [cited by applicant]
GB 0893354A · 1962 [cited by applicant]
KR 1020110079495A · 2011 [cited by applicant]
WO 2004053841A2 · 2004 [cited by applicant]
WO 2016003740A1 · 2016 [cited by applicant]
WO 2016105727A1 · 2016 [cited by applicant]
WO 2018007782A1 · 2018 [cited by applicant]
WO 2018125250A1 · 2018 [cited by applicant]
WO 2018154273A1 · 2018 [cited by applicant]
Communication pursuant to Article 94(3) EPC, EP App. No. 19182737.7, Sep. 29, 2021, 4 pages. [cited by applicant]
Corrected Notice of Allowability, U.S. Appl. No. 15/201,442, Jan. 22, 2019, 5 pages. [cited by applicant]
Corrected Notice of Allowability, U.S. Appl. No. 15/201,442, Mar. 11, 2019, 2 pages. [cited by applicant]
Corrected Notice of Allowability, U.S. Appl. No. 16/145,066, Jul. 28, 2020, 7 pages. [cited by applicant]
European Search Report and Search Opinion, EP App. No. 22217001.1, Mar. 30, 2023, 7 pages. [cited by applicant]
Extended European Search Report and search Opinion for Application No. 22200756.9, Jan. 30, 2023, 07 pages. [cited by applicant]
Extended European Search Report and Written Opinion, EP App. No. 19182737.7, Apr. 7, 2020, 7 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 17/549,363, Oct. 5, 2023, 8 pages. [cited by applicant]
Intention to grant, EP App. No. 19182737.7, Oct. 23, 2023, 8 pages. [cited by applicant]
International Preliminary Report on Patentability, PCT App. No. PCT/US2017/036038, Jan. 17, 2019, 14 pages. [cited by applicant]
International Preliminary Report on Patentability, PCT App. No. PCT/US2017/040546, Oct. 3, 2019, 10 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/036038, Sep. 5, 2017 , 15 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040534, Jan. 3, 2018, 11 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040536, Dec. 20, 2017, 11 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040537, Dec. 20, 2017, 11 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040540, Jan. 3, 2018, 14 pages. [cited by applicant]
International Search Report and Written Opinion, PCT App. No. PCT/US2017/040546, Jan. 24, 2018, 15 pages. [cited by applicant]
Lahr, David L., “Distractions: Timing Matrix Multiplication in SciDB and Setting the Number of Worker Instances in SciDB and Running Matrix Multiplication Piecemeal”, Available Online at <http://dllahr.blogspot.com/2012… [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 15/201,442, May 4, 2018, 11 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 16/398,200, Jul. 28, 2020, 16 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/240,882, Jul. 6, 2022, 7 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/549,221, Sep. 15, 2022, 9 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/549,363, Jan. 4, 2023, 9 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/145,066, Dec. 28, 2020, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 15/201,442, Dec. 14, 2018, 5 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/145,066, May 27, 2020, 9 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/240,882, Jan. 19, 2023, 3 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/240,882, Oct. 11, 2022, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/549,221, Apr. 19, 2023, 3 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/549,221, Jun. 22, 2023, 2 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/549,221, Mar. 13, 2023, 8 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/549,363, Dec. 20, 2023, 7 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/549,363, Feb. 12, 2024, 3 pages. [cited by applicant]
Office Action, EP App. No. 19182737.7, Mar. 27, 2023, 5 pages. [cited by applicant]
Office Action, EP App. No. 19182737.7, Sep. 12, 2022, 4 pages. [cited by applicant]
Office Action, EP App. No. 22200756.9, Dec. 22, 2023, 05 pages. [cited by applicant]
European search report and Search Opinion, EP App. No. 24153306.6, Jun. 19, 2024, 08 pages. [cited by applicant]
European search report and Search Opinion, EP App. No. 24157718.8, Jul. 1, 2024, 07 pages. [cited by applicant]
Intention to Grant, EP App. No. 22200756.9, May 8, 2024, 6 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 18/399,014, Aug. 1, 2024, 7 pages. [cited by applicant]
Office Action , EP App. No. 24157718.8, Feb. 27, 2025, 05 pages. [cited by applicant]
Office Action, EP App. No. 22217001.1, Mar. 10, 2025, 06 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 18/399,014, Dec. 10, 2024, 7 pages. [cited by applicant]
First Office Action, CN App. No. 202111564181.7, Mar. 28, 2025, 8 pages of Original Document only. [cited by applicant]
First Office Action, CN App. No. 202210287299.8, Apr. 30, 2025, 12 pages of Original Document only. [cited by applicant]
Kayaaslan, Enver et al., “Semi-two-dimensional Partitioning for Parallel Sparse Matrix-vector Multiplication”, 29th IEEE International Parallel and Distributed Processing Symposium (IPDPS), May 2015, 12 pages. [cited by applicant]
Office Action, EP App. No. 24153306.6, May 15, 2025, 5 pages. [cited by applicant]
Wei, Hongchang et al., “Asymptotic Fitting Optimization Technology for Source-to-Source Compile System on CPU-GPU Architecture”, Computer Engineering and Applications, vol. 52, No. 21, Mar. 2016, 6 pages. [cited by applicant]
Xia, Youshen et al., “Super-resolution Image Reconstruction using a Two-dimensional Subgradient Algorithm”, IEEE International Conference on Audio, Language and Image Processing (ICALIP), Jul. 2016, pp. 229-234. [cited by applicant]
Zhang, Shengliang et al., “A Two-dimensional PCA Based on Between-class Scatter Matrix”, Computer Engineering, vol. 32. No. 11, Jun. 2006, 3 pages. [cited by applicant]
Second Office Action, CN App. No. 202111564181.7, Aug. 23, 2025, 06 pages of Original Document only. [cited by applicant]