IP Library Patent Application 17678886
Patent Application
App. No. 17/678,886

QUANTIZATION METHOD, QUANTIZATION DEVICE, AND RECORDING MEDIUM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
17/678,886
Abstract

A quantization method executed by a computer includes: searching for quantization step sizes of parameters of a target layer by using a second inference contribution degree and quantization errors before and after quantization of the parameters of the target layer, the second inference contribution degree indicating a degree of influence of a layer next to the target layer and being obtained using a first inference contribution degree calculated in advance, the layer next to the target layer including second neurons as elements, and the first inference contribution degree indicating a degree of influence of each of layers that constitute a model composed of a neural network and each include first neurons as elements on an inference result obtained by using the model; and quantizing the parameters by using the quantization step sizes obtained as a result of the searching.

Claims (23)

1 . A quantization method executed by a computer, the quantization method comprising:

searching for quantization step sizes of a plurality of parameters of a target layer by using a second inference contribution degree and quantization errors before and after quantization of the plurality of parameters of the target layer, the second inference contribution degree indicating a degree of influence of a layer next to the target layer and being obtained using a first inference contribution degree calculated in advance, the layer next to the target layer including a plurality of second neurons as elements, the first inference contribution degree indicating a degree of influence of each of a plurality of layers that constitute a model composed of a neural network and each include a plurality of first neurons as elements on an inference result obtained by using the model, and the target layer and the layer next to the target layer being included in the plurality of layers; and

quantizing the plurality of parameters by using the quantization step sizes obtained as a result of the searching.

2 . The quantization method according to claim 1 ,

wherein the searching for the quantization step sizes of the plurality of parameters is performed by using an evaluation equation including a product value of the quantization errors and the second inference contribution degree such that the evaluation equation is minimized.

3 . The quantization method according to claim 1 , further comprising:

calculating first neuron values of the first neurons by performing inference by inputting, to the model, each item of data that constitutes an inference contribution degree calculation dataset that is at least a portion of a training dataset used to train the model;

calculating, for each of the first neurons, an accumulated value by accumulating the first neuron values calculated for all items of the data that constitutes the inference contribution degree calculation dataset; and

calculating, as the first inference contribution degree, a value obtained by normalizing the accumulated value of each of the first neurons for each of the plurality of layers.

4 . The quantization method according to claim 1 ,

wherein the plurality of parameters are at least either a plurality of intermediate values of the target layer or a plurality of weights assigned to the second neurons.

5 . The quantization method according to claim 4 ,

wherein the model is a convolutional neural network, and

the intermediate values are feature maps of the target layer.

6 . A quantization device comprising:

a processor; and

a memory,

wherein the processor performs the following by using the memory:

searching for quantization step sizes of a plurality of parameters of a target layer by using a second inference contribution degree and quantization errors before and after quantization of the plurality of parameters of the target layer, the second inference contribution degree indicating a degree of influence of a layer next to the target layer and being obtained using a first inference contribution degree calculated in advance, the layer next to the target layer including a plurality of second neurons as elements, the first inference contribution degree indicating a degree of influence of each of a plurality of layers that constitute a model composed of a neural network and each include a plurality of first neurons as elements on an inference result obtained by using the model, and the target layer and the layer next to the target layer being included in the plurality of layers; and

quantizing the plurality of parameters by using the quantization step sizes obtained as a result of the searching.

7 . A non-transitory computer-readable recording medium for use in a computer, the recording medium having a computer program recorded thereon for causing the computer to execute:

searching for quantization step sizes of a plurality of parameters of a target layer by using a second inference contribution degree and quantization errors before and after quantization of the plurality of parameters of the target layer, the second inference contribution degree indicating a degree of influence of a layer next to the target layer and being obtained using a first inference contribution degree calculated in advance, the layer next to the target layer including a plurality of second neurons as elements, the first inference contribution degree indicating a degree of influence of each of a plurality of layers that constitute a model composed of a neural network and each include a plurality of first neurons as elements on an inference result obtained by using the model, and the target layer and the layer next to the target layer being included in the plurality of layers; and

quantizing the plurality of parameters by using the quantization step sizes obtained as a result of the searching.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2024
From: PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO., LTD.
To: PANASONIC AUTOMOTIVE SYSTEMS CO., LTD.
Reel/Frame 066709/0702 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2022
From: MURATA, NORIFUMI
To: PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO., LTD.
Reel/Frame 061173/0921 →