IP Library Granted Patent US 11,915,138
Granted Patent B2
US 11,915,138 · App. 16/793,993 · Granted Feb 27, 2024

Method and device for reducing a size of a neural network model

Inventors: Weifeng Zhang (San Mateo, CA); Guoyang Chen (San Mateo, CA); Yu Pu (San Mateo, CA); Yongzhi Zhang (San Mateo, CA); Yuan Xie (San Mateo, CA)
Assignee: Alibaba Group Holding Limited
G06N3/082G06N3/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,915,138
App. No.
16/793,993
Granted
Feb 27, 2024
Kind
B2
Abstract

Methods and apparatus for reducing a size of a neural network model, the method including: compressing data of the neural network model; identifying structure information of a vector register, wherein the structure information includes a number of registers included in the vector register; comparing a number of elements in the compressed data with a first condition, wherein the first condition is determined based on the number of registers in the vector register; and in response to the number of elements satisfying the first condition, associating the compressed data with the vector register to enable loading the compressed data to the vector register.

Claims (46)

1. A method for reducing a size of a neural network model, comprising:

compressing data of the neural network model;

identifying structure information of a vector register, wherein the structure information includes a number of registers included in the vector register;

comparing a number of elements in the compressed data of the neural network model with a first condition, wherein the first condition is determined based on the number of registers in the vector register;

in response to the number of elements satisfying the first condition, associating the compressed data of the neural network model with the vector register to enable loading the compressed data of the neural network model to the vector register;

in response to the number of elements satisfying the first condition, sending an indication to end compression of the data;

comparing the number of elements in the compressed data of the neural network model with a second condition, wherein the second condition is determined based on the number of registers in the vector register, and is different from the first condition; and

adjusting a structure of the vector register in response to the number of elements satisfying the second condition,

wherein satisfying the first condition comprises being equal to or smaller than the number of registers in the vector register, and wherein satisfying the second condition comprises being equal to or smaller than one half of the number of registers in the vector register.

2. The method of claim 1 , wherein compressing the data of the neural network model comprises pruning of the data.

3. The method of claim 1 , wherein the data is a weight matrix of the neural network model.

4. The method of claim 1 , wherein adjusting the structure of the vector register comprises:

reducing the number of registers in the vector register by one half.

5. The method of claim 1 , wherein compressing the data of the neural network model comprises a first operation of compression, the method further comprising:

performing a second operation of compression of the data in response to the number of elements not satisfying the first condition.

6. The method of claim 1 , further comprising:

generating an instruction set based on the association between the compressed data and the vector register to load the compressed data to the vector register.

7. The method of claim 1 , wherein the vector register is part of a group of vector registers that include elements that are executed simultaneously.

8. An apparatus for reducing a size of a neural network model, comprising:

a memory storing a set of instructions; and

one or more processors configured to execute the set of instruction to cause the apparatus to perform:

compressing data of the neural network model,

identifying structure information of a vector register, wherein the structure information includes a number of registers included in the vector register,

comparing a number of elements in the compressed data of the neural network model with a first condition, wherein the first condition is determined based on the number of registers in the vector register,

in response to the number of elements satisfying the first condition, associating the compressed data of the neural network model with the vector register to enable loading the compressed data of the neural network model to the vector register,

in response to the number of elements satisfying the first condition, sending an indication to end compression of the data,

comparing the number of elements in the compressed data of the neural network model with a second condition, wherein the second condition is determined based on the number of registers in the vector register, and is different from the first condition; and

adjusting a structure of the vector register in response to the number of elements satisfying the second condition,

wherein satisfying the first condition comprises being equal to or smaller than the number of registers in the vector register, and wherein satisfying the second condition comprises being equal to or smaller than one half of the number of registers in the vector register.

9. A non-transitory computer readable medium that stores a set of instructions that is executable by at least one processor of a computer to cause the computer to perform a method for reducing a size of a neural network model, the method comprising:

compressing data of the neural network model;

identifying structure information of a vector register, wherein the structure information includes a number of registers included in the vector register;

comparing a number of elements in the compressed data of the neural network model with a first condition, wherein the first condition is determined based on the number of registers in the vector register;

in response to the number of elements satisfying the first condition, associating the compressed data of the neural network model with the vector register to enable loading the compressed data of the neural network model to the vector register;

in response to the number of elements satisfying the first condition, sending an indication to end compression of the data;

comparing the number of elements in the compressed data of the neural network model with a second condition, wherein the second condition is determined based on the number of registers in the vector register, and is different from the first condition; and

adjusting a structure of the vector register in response to the number of elements satisfying the second condition,

wherein satisfying the first condition comprises being equal to or smaller than the number of registers in the vector register, and wherein satisfying the second condition comprises being equal to or smaller than one half of the number of registers in the vector register.

10. The non-transitory computer readable medium of claim 9 , wherein compressing the data of the neural network model comprises pruning of the data.

11. The non-transitory computer readable medium of claim 9 , wherein the data is the weight matrix of a neural network model.

12. The non-transitory computer readable medium of claim 9 , wherein adjusting the structure of the vector register comprises:

reducing the number of registers in the vector register by one half.

13. The non-transitory computer readable medium of claim 9 , wherein compressing the data of the neural network model comprises a first operation of compression, the method further comprising:

performing a second operation of compression of the data in response to the number of elements not satisfying the first condition.

14. The non-transitory computer readable medium of claim 9 , wherein the set of instructions that are executable by the at least one processor of a computer to cause the computer to further perform:

generating an instruction set based on the association between the compressed data and the vector register to load the compressed data to the vector register.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2026
From: ALIBABA GROUP HOLDING LIMITED
To: CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PRIVATE LIMITED
Reel/Frame 075478/0225 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2020
From: ZHANG, WEIFENG; CHEN, GUOYANG; PU, YU; ZHANG, YONGZHI; XIE, YUAN
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 052470/0447 →
Continuity (1)
Related Publication 20210256380A1 · Aug 19, 2021