IP Library › Granted Patent US 12,596,928
Granted Patent B2
US 12,596,928 · App. 17/792,117 · Granted Apr 7, 2026

Movement of tensor data during reshape operation

Inventors: Arun Chauhan (Belmont, CA); Fatih Mehmet Bakir (Santa Barbara, CA); Phitchaya Mangpo Phothilimthana (Mountain View, CA); Dong Hyuk Woo (San Jose, CA)
Assignee: Google LLC
G06N3/082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,928
App. No.
17/792,117
Granted
Apr 7, 2026
Kind
B2
Abstract

A method of performing a reshape operation specified in a reshape layer of a neural network model is described. The reshape operation reshapes an input tensor with an input tensor shape to an output tensor with an output tensor shape. The tensor data that has to be reshaped is directly routed between tile memories of the hardware accelerator in an efficient manner. This advantageously optimizes usage of memory space and allows any number and type of neural network models to be run on the hardware accelerator.

Claims (45)

1 . A method of performing a reshape operation specified in a reshape layer of a neural network model, the method comprising:

transmitting, to a reshape solver, an input tensor with an input tensor shape and an output tensor with an output tensor shape;

receiving, from the reshape solver, data identifying chunks of tensor data to be moved within memories of a hardware accelerator, a source computing unit on the hardware accelerator from where each corresponding chunk is to be moved, and a target computing unit on the hardware accelerator to where each corresponding chunk is to be moved;

transmitting, to a constraint based solver, the received data and a maximum number of time steps over which the reshape operation is to be performed;

receiving, from the constraint based solver, a schedule based on the received data, a number of computing units within the hardware accelerator, and a maximum number of time steps, the schedule indicating routes for transferring the chunks of tensor data between the memories of the hardware accelerator;

updating the schedule by removing cyclical routes or merging chunks moving over a single route or both to generate an updated schedule;

compiling the updated schedule to generate compiled data; and

transmitting the compiled data to the hardware accelerator.

2 . The method of claim 1 , further comprising:

creating, by the hardware accelerator, buffers in response to the compiled data to temporarily store corresponding chunks being moved within memories of the hardware accelerator.

3 . The method of claim 1 , wherein the reshape solver is programmed with a first set of one or more constraints.

4 . The method of claim 1 , wherein:

the transmitting to the reshape solver, the receiving from the reshape solver, the transmitting to the constraint based solver, and the receiving from the constraint based solver are performed by a central processing unit (CPU) that implements a compiler; and

the compiling of the schedule and the transmitting to the hardware accelerator are performed by the compiler.

5 . The method of claim 1 , further comprising:

updating the received data to remove one or more chunks for which the source computing unit and the target computing unit are adjacently arranged within the hardware accelerator, the updating being performed subsequent to the receiving from the reshape solver of the received data and prior to the transmitting to the constraint based solver of the received data.

6 . The method of claim 2 , wherein a storage capacity of each buffer depends on the reshape operation.

7 . The method of claim 3 , wherein the constraint based solver is programmed with a second set of one or more constraints.

8 . A system that performs a reshape operation specified in a reshape layer of a neural network model, the reshape operation configured to reshape an input tensor with an input tensor shape to an output tensor with an output tensor shape, the system comprising:

at least one programmable processor; and

a machine-readable medium storing instructions that, when executed by the at least one programmable processor, cause the at least one programmable processor to:

transmit, to a reshape solver, the input tensor and the output tensor;

receive, from the reshape solver, data identifying chunks of tensor data to be moved within memories of a hardware accelerator, a source computing unit on the hardware accelerator from where each corresponding chunk is to be moved, and a target computing unit on the hardware accelerator to where each corresponding chunk is to be moved;

transmit, to a constraint based solver, the received data, a number of computing units within the hardware accelerator, and a maximum number of time steps over which the reshape operation is to be performed;

receive, from the constraint based solver, a schedule based on the received data, the number of computing units within the hardware accelerator, and the maximum number of time steps, the schedule indicating routes for transferring the chunks of tensor data between the memories of the hardware accelerator;

updating the schedule by removing cyclical routes or merging chunks moving over a single route or both to generate an updated schedule;

compile the updated schedule to generate compiled data; and

transmit the compiled data to the hardware accelerator.

9 . The system of claim 8 , wherein the hardware accelerator is configured to create buffers in response to the compiled data to temporarily store corresponding chunks being moved within memories of the hardware accelerator.

10 . The system of claim 8 , wherein the reshape solver is programmed with a first set of one or more constraints.

11 . The system of claim 8 , wherein:

the at least one programmable processor is a central processing unit (CPU) that implements a compiler; and

the compiling of the schedule and the transmitting of the compiled data to the hardware accelerator are performed by the compiler.

12 . The system of claim 8 , wherein the at least one programmable processor is configured to update the received data to remove one or more chunks for which the source computing unit and the target computing unit are adjacently arranged within the hardware accelerator, the updating being performed subsequent to the receiving from the reshape solver of the received data and prior to the transmitting to the constraint based solver of the received data.

13 . The system of claim 9 , wherein a storage capacity of each buffer depends on the reshape operation.

14 . The system of claim 10 , wherein the constraint based solver is programmed with a second set of one or more constraints.

15 . A non-transitory computer program product storing instructions that, when executed by at least one programmable processor, cause the at least one programmable processor to:

transmit, to a reshape solver, an input tensor with an input tensor shape and an output tensor with an output tensor shape;

receive, from the reshape solver, data identifying chunks of tensor data to be moved within memories of a hardware accelerator, a source computing unit on the hardware accelerator from where each corresponding chunk is to be moved, and a target computing unit on the hardware accelerator to where each corresponding chunk is to be moved;

transmit, to a constraint based solver, the received data, a number of computing units within the hardware accelerator, and a maximum number of time steps over which a reshape operation is to be performed;

receive, from the constraint based solver, a schedule based on the received data, the number of computing units within the hardware accelerator, and the maximum number of time steps, the schedule indicating routes for transferring the chunks of tensor data between the memories of the hardware accelerator;

updating the schedule by removing cyclical routes or merging chunks moving over a single route or both to generate an updated schedule;

compile the updated schedule to generate compiled data; and

transmit the compiled data to the hardware accelerator.

16 . The non-transitory computer program product of claim 15 , wherein the at least one programmable processor is configured to update the received data to remove one or more chunks for which the source computing unit and the target computing unit are adjacently arranged within the hardware accelerator, the updating being performed subsequent to the receiving from the reshape solver of the received data and prior to the transmitting to the constraint based solver of the received data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 27, 2022
From: CHAUHAN, ARUN; BAKIR, FATIH MEHMET; PHOTHILIMTHANA, PHITCHAYA MANGPO; WOO, DONG HYUK
To: GOOGLE LLC
Reel/Frame 060636/0152 →
Continuity (1)
Related Publication 20230052942A1 · Feb 16, 2023
References Cited (12)
US 20170053676A1 · Marchese · 2017 [cited by examiner]
US 20170116037A1 · Li · 2017 [cited by examiner]
US 20180285727A1 · Baum · 2018 [cited by examiner]
US 20180322382A1 · Mellempudi et al. · 2018 [cited by applicant]
US 20180324393A1 · Ryan · 2018 [cited by examiner]
US 20190042221A1 · Krishnaiyer et al. · 2019 [cited by applicant]
US 20190042925A1 · Choe · 2019 [cited by examiner]
US 20200042856A1 · Datta · 2020 [cited by examiner]
US 20210074270A1 · Ahn · 2021 [cited by examiner]
Dai et al., “On solving multi-commodity flow problems: An experimental evaluation.” Chinese Journal of Aeronautics 30.4, Jun. 2017, 1481-1492. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2020/025676, mailed on Oct. 13, 2022, 12 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2020/025676, mailed on Dec. 23, 2020, 18 pages. [cited by applicant]