IP Library › Granted Patent US 12,632,714
Granted Patent B1
US 12,632,714 · App. 17/453,293 · Granted May 19, 2026

Interconnect mode for computational arrays

Inventors: Sundeep Amirineni (Cedar Park, TX); Paul Gilbert Meyer (Jericho, VT); Ron Diamant (San Jose, CA); Qingrui Liu (Burlingame, CA)
Assignee: Amazon Technologies, Inc.
G06N3/063G06F7/78G06F12/0802G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,714
App. No.
17/453,293
Granted
May 19, 2026
Kind
B1
Abstract

A processing engine array is provided with an interconnect mode of operation to use the array as an interconnect to move data elements to different locations in memory such as to perform a matrix transpose operation. In this interconnect mode of operation, although computations are still being performed in the array, the computations are not carried out to modify or change the values of the data elements, but are instead carried out to rearrange the data elements in memory. As such, the computations carried out in the interconnect mode of operation can deviate from the expected behavior of floating-point calculations. A mode selection signal can be used to provide the proper outputs of the processing elements of the array depending on the mode of operation.

Claims (48)

1 . A neural network processor comprising:

a state buffer memory;

a results buffer memory; and

a processing engine array having a plurality of processing elements,

wherein the neural network processor is operable to execute a set of instructions to perform a transpose operation to transpose a tensor, wherein the transpose operation includes:

loading data elements of the tensor from the state buffer memory into the processing engine array;

performing a matrix multiplication on the data elements loaded in the processing engine array with an identity matrix to generate a set of outputs; and

storing the set of outputs in the results buffer memory, and

wherein the processing engine array is configured to set a multiplication result to a zero value for the transpose operation in response to multiplying a data element representing infinity stored in the processing engine array with an element of zero from the identity matrix, and

wherein each processing element of the processing engine array includes a multiplication result selection circuit operable to select the multiplication result from a plurality of values including the zero value, a NaN value, and an infinity value.

2 . The neural network processor of claim 1 , wherein the data elements of the tensor are loaded from row partitions of the state buffer memory into corresponding rows of processing elements of the processing engine array, and the set of outputs are stored in column partitions of the results buffer memory, wherein the column partitions in the results buffer memory are mapped to the row partitions in the state buffer memory.

3 . The neural network processor of claim 1 , wherein the multiplication result selection circuit is operable to select the NaN value as a result of multiplying infinity with zero when the processing engine array is not performing a transpose operation.

4 . The neural network processor of claim 1 , wherein the multiplication result selection circuit is operable to select an infinity value as a result of multiplying infinity with a non-zero number.

5 . An integrated circuit device comprising:

a processing element including a multiplication circuit and an adder circuit coupled to the multiplication circuit,

wherein the processing element is operable to generate a multiplication result based on multiplying two multiplication operands using the multiplication circuit, and to generate a partial sum output based on a partial sum input using the adder circuit,

wherein in an interconnect mode of operation, the processing element is operable to generate the partial sum output having a same value as the partial sum input in response to the two multiplication operands being zero and infinity,

wherein the processing element further comprises a multiplication result selection circuit that is operable, in the interconnect mode of operation, to set the multiplication result to a zero value in response to the two multiplication operands being zero and infinity, and

wherein the multiplication result selection circuit is operable in a compute mode of operation, to set the multiplication result to a not-a-number (NaN) value in response to the two multiplication operands being zero and infinity.

6 . The integrated circuit device of claim 5 , wherein the multiplication result selection circuit is operable, in the interconnect mode of operation, to set the multiplication result to the zero value in response to the two multiplication operands being zero and NaN.

7 . The integrated circuit device of claim 5 , wherein the processing element further comprises an output selection circuit that is operable, in the interconnect mode of operation, to select the partial sum input as the partial sum output in response to the two multiplication operands being zero and infinity.

8 . The integrated circuit device of claim 7 , wherein the output selection circuit is operable, in a compute mode of operation, to select a not-a-number (NaN) value as the partial sum output in response to the two multiplication operands being zero and infinity.

9 . The integrated circuit device of claim 5 , wherein the processing element is part of an array of processing elements, and wherein the interconnect mode of operation is used to perform a matrix transpose operation using the array of processing elements.

10 . The integrated circuit device of claim 9 , wherein the matrix transpose operation includes multiplying an input matrix with an identity matrix.

11 . The integrated circuit device of claim 10 , wherein the matrix transpose operation includes preloading the input matrix into weight cache registers of the array, and subsequently shifting the identity matrix into feature map registers of the array.

12 . The integrated circuit device of claim 5 , wherein the processing element is operable to dynamically switch between the interconnect mode of operation and a compute mode of operation based on an instruction being executed by the integrated circuit device.

13 . A method comprising:

receiving a request to perform a transpose operation on a tensor;

in response to receiving the request to perform the transpose operation, asserting a mode signal to cause a computational array to use a zero value as a result of multiplying zero with infinity;

loading data elements of the tensor from a first memory into the computational array;

performing an identity multiplication on the data elements loaded into the computational array, wherein performing the identity multiplication includes shifting an identity matrix into the computational array on feature map input buses;

storing a result of the identity multiplication in a second memory; and

loading the result from the second memory into the first memory.

14 . The method of claim 13 , further comprising:

receiving a request to perform a matrix multiplication operation; and

deasserting the mode signal to cause the computational array to use a not-a-number value as a result of multiplying zero with infinity.

15 . The method of claim 14 , wherein the mode signal controls multiplication result selection circuits in the computational array.

16 . The method of claim 14 , wherein the mode signal controls partial sum output selection circuits in the computational array.

17 . The method of claim 14 , wherein the data elements of the tensor loaded into the computational array are stored in weight cache registers of the computational array.

18 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to execute a compiler, the compiler performing operations including:

receiving a set of operations to be performed by an execution engine having a computational array;

determining that the set of operations includes a matrix transpose operation to be performed on a tensor; and

generating a set of compiled instructions operable to transpose the tensor by:

asserting a mode signal to cause the computational array to use a zero value as a result of multiplying zero with infinity;

loading data elements of the tensor from a first memory into the computational array;

performing an identity multiplication on the data elements loaded into the computational array, wherein performing the identity multiplication includes shifting an identity matrix into the computational array on feature map input buses;

storing a result of the identity multiplication in a second memory; and

loading the result from the second memory into the first memory.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2021
From: AMIRINENI, SUNDEEP; MEYER, PAUL GILBERT; DIAMANT, RON; LIU, QINGRUI
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 057998/0989 →
References Cited (33)
US 4760543A · Ligtenberg et al. · 1988 [cited by applicant]
US 5321639A · Krishnamoorthy et al. · 1994 [cited by applicant]
US 5842035A · Nishikawa · 1998 [cited by applicant]
US 6167502A · Pechanek · 2000 [cited by examiner]
US 6438747B1 · Schreiber et al. · 2002 [cited by applicant]
US 6675187B1 · Greenberger · 2004 [cited by applicant]
US 6754687B1 · Kurak, Jr. et al. · 2004 [cited by applicant]
US 7596678B2 · Beaumont · 2009 [cited by applicant]
US 8327114B1 · Cismas · 2012 [cited by examiner]
US 8473540B1 · Rao et al. · 2013 [cited by applicant]
US 10489479B1 · Shalev · 2019 [cited by examiner]
US 10884707B1 · Li et al. · 2021 [cited by applicant]
US 11036827B1 · Zejda et al. · 2021 [cited by applicant]
US 20090063529A1 · Gustavson et al. · 2009 [cited by applicant]
US 20100211337A1 · Clark · 2010 [cited by examiner]
US 20110119467A1 · Cadambi et al. · 2011 [cited by applicant]
US 20110161625A1 · Pechanek · 2011 [cited by applicant]
US 20130103925A1 · Meeker · 2013 [cited by applicant]
US 20150169369A1 · Baskaran · 2015 [cited by examiner]
US 20160020824A1 · Ulrich et al. · 2016 [cited by applicant]
US 20170168991A1 · Baskaran · 2017 [cited by examiner]
US 20170200094A1 · Bruestle · 2017 [cited by examiner]
US 20180260690A1 · Young · 2018 [cited by examiner]
US 20180314671A1 · Zhang et al. · 2018 [cited by applicant]
US 20190205787A1 · Duriseti · 2019 [cited by examiner]
US 20210096823A1 · Li et al. · 2021 [cited by applicant]
Torbjorn Viem Ness; Low Power Floating Point Unit for RISC-V; Jul. 2018; Norwegian University of Science & Technology; Whole document. (Year: 2018). [cited by examiner]
Baker, K., “Singular Value Decomposition Tutorial”, [email protected], Mar. 19, 2005, pp. 1-24. [cited by applicant]
Johnson, J. R., et al., “A Methodology for Designing, Modifying, and Implementing Fourier Transform Algorithms on Various Architectures”, [cited by applicant]
Lin, Jui-Chieh, et al., “Bit Matrix Transpose with Tensor Product and Perfect Shuffling”, [cited by applicant]
O'Leary, Dianne P., “Systolic Arrays for Matrix Transpose and Other Reorderings”, [cited by applicant]
Srivastava, Nitish, et al., “T2S-Tensor: Productively Generating High-Performance Spatial Hardware for Dense Tensor Computations”, [cited by applicant]
Zekri, Ahmed S., “Restructuring and Implementations of 2D Matrix Transpose Algorithm Using SSE4 Vector Instructions”, [cited by applicant]