IP Library › Granted Patent US 12,430,533
Granted Patent B2
US 12,430,533 · App. 17/283,922 · Granted Sep 30, 2025

Neural network processing apparatus, neural network processing method, and neural network processing program

Inventors: Takato Yamada (Tokyo, JP); Antonio Tomas Nevado Vilchez (Tokyo, JP)
Assignee: MAXWELL, INC.
G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,533
App. No.
17/283,922
Granted
Sep 30, 2025
Kind
B2
Abstract

A CNN processing apparatus ( 1 ) includes an input buffer ( 10 ) configured to store an input signal A given to a CNN, a weight buffer ( 11 ) configured to store weights U, a convolutional operation unit ( 12 ) configured to perform a convolutional operation including a product-sum operation of the input signal A and the weights U, a storage unit 16 configured to store a table ( 160 ) which is configured to associate an input and an output of conversion-quantization processing with each other, wherein the input is an operation result of the convolutional operation, and the output is a result of the conversion-quantization processing of converting the input value based on a predetermined condition and quantizing the converted value by reducing a bit accuracy of the converted data, and a processing unit ( 14 ) configured to acquire the output of the conversion-quantization processing corresponding to the operation result by the operation unit by referring to the table ( 160 ).

Claims (50)

1. A neural network processing apparatus comprising:

a first memory configured to store an image data given to a neural network;

a second memory configured to store weights of the neural network;

an operation unit configured to perform a convolutional operation of the neural network including a product-sum operation of the image data and the weights;

a third memory configured to store a table which is configured to associate an input and an output of a conversion-quantization process with each other; and

a processing unit configured to acquire the output of the conversion-quantization process corresponding to an operation result of the convolutional operation by the operation unit;

wherein the conversion-quantization process includes:

a conversion process that converts an input value based on a predetermined condition to obtain a converted data having a bit precision, and

a quantization process that quantizes the converted data by reducing the bit precision of the converted data, and

the processing unit is configured to perform the conversion process and the quantization process together and the output of the conversion and quantization process is acquired by referring to the table.

2. The neural network processing apparatus according to claim 1 , wherein

the table associates a plurality of input sections obtained by dividing the input of the conversion-quantization process into a plurality of continuous sections and the output of the conversion-quantization process with each other, and

the processing unit comprises:

an input determination unit configured to compare the operation result of the convolutional operation by the operation unit with the plurality of input sections and determine an input section including the operation result; and

an output acquisition unit configured to acquire the output of the conversion-quantization process corresponding to the input section according to a determination result of the input determination unit by referring to the table.

3. The neural network processing apparatus according to claim 1 , wherein

the table associates a plurality of thresholds set in advance for the input of the conversion-quantization process and the output of the conversion-quantization process with each other, and

the processing unit comprises:

a first threshold processing unit configured to compare the operation result of the convolutional operation by the operation unit with the plurality of thresholds and output a threshold corresponding to the operation result; and

an output acquisition unit configured to acquire the output of the conversion-quantization process corresponding to the threshold output by the first threshold processing unit by referring to the table.

4. The neural network processing apparatus according to claim 2 , wherein

the table associates division information for identifying an input section in which the output of the conversion-quantization process monotonously increases and an input section in which the output of the conversion-quantization process monotonously decreases, a plurality of thresholds set in advance for the input of the conversion-quantization process, and the output of the conversion-quantization process corresponding to each of the plurality of thresholds with each other,

the input determination unit determines, based on the division information, the input section of the conversion-quantization process to which the operation result of the convolutional operation by the operation unit belongs, and

the processing unit comprises:

a second threshold processing unit configured to compare the operation result of the convolutional operation by the operation unit with the plurality of thresholds in the input section determined by the input determination unit and output a threshold corresponding to the operation result; and

an output acquisition unit configured to acquire the output of the conversion-quantization process corresponding to the threshold output by the second threshold processing unit by referring to the table.

5. The neural network processing apparatus according to claim 1 , wherein

the neural network is a multilayer neural network including at least one intermediate layer.

6. The neural network processing apparatus according to claim 1 , wherein

processing of converting the operation result of the convolutional operation by the operation unit based on the predetermined condition, which is included in the conversion-quantization process, includes at least one of decision of the operation result by an activation function and normalization of the operation result.

7. A neural network processing method comprising:

a first step of storing an image data given to a neural network in a first memory;

a second step of storing weights of the neural network in a second memory;

a third step of performing a convolutional operation of the neural network including a product-sum operation of the image data and the weights;

a fourth step of storing, in a third memory, a table which is configured to associate an input and an output of a conversion-quantization process with each other; and

a fifth step of acquiring the output of the conversion-quantization process corresponding to an operation result of the convolutional operation in the third step by referring to the table,

wherein the conversion-quantization process includes:

a conversion process that converts an input value based on a predetermined condition, and

a quantization process that quantizes the converted data by reducing a bit precision of the converted data, and

the fifth step of executing the conversion process and the quantization process together and the output of the conversion and quantization process is acquired by referring to the table.

8. A non-transitory computer-readable storage medium storing a neural network processing program configured to cause a computer to execute:

a first step of storing an image data given to a neural network in a first memory;

a second step of storing weights of the neural network in a second memory;

a third step of performing a convolutional operation of the neural network including a product-sum operation of the image data and the weights;

a fourth step of storing, in a third memory, a table which is configured to associate an input and an output of conversion-quantization process with each other; and

a fifth step of acquiring the output of the conversion-quantization process corresponding to an operation result of the convolutional operation in the third step by referring to the table,

wherein the conversion-quantization process comprising:

a conversion process that converts an input value based on a predetermined condition, and

a quantization process that quantizes the converted data by reducing a bit precision of the converted data, and

the fifth step of executing the conversion process and the quantization process together and the output of the conversion and quantization process is acquired by referring to the table.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING PARTY NAME PREVIOUSLY RECORDED AT REEL: 69824 FRAME: 203. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 14, 2025
From: LEAPMIND INC.
To: MAXELL, LTD.
Reel/Frame 070521/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2025
From: LEAPMIND INC.
To: MAXELL, INC.
Reel/Frame 069824/0203 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2021
From: YAMADA, TAKATO; NEVADO VILCHEZ, ANTONIO TOMAS
To: LEAPMIND INC.
Reel/Frame 058370/0794 →
Priority Claims (1)
JP 2018-192021 · Oct 10, 2018 · national
Continuity (1)
Related Publication 20210232894A1 · Jul 29, 2021
References Cited (55)
US 5255347A · Matsuba · 1993 [cited by examiner]
US 5847952A · Samad · 1998 [cited by examiner]
US 6389404B1 · Carson · 2002 [cited by examiner]
US 10740432B1 · Diamant · 2020 [cited by examiner]
US 20020059152A1 · Carson · 2002 [cited by examiner]
US 20020181799A1 · Matsugu · 2002 [cited by examiner]
US 20060242257A1 · Kuribayashi · 2006 [cited by applicant]
US 20070183326A1 · Igarashi et al. · 2007 [cited by applicant]
US 20110239032A1 · Kato et al. · 2011 [cited by applicant]
US 20160179434A1 · Herrero Abellanas et al. · 2016 [cited by applicant]
US 20160328646A1 · Lin et al. · 2016 [cited by applicant]
US 20160328647A1 · Lin et al. · 2016 [cited by applicant]
US 20170277977A1 · Kitamura · 2017 [cited by applicant]
US 20180032866A1 · Son et al. · 2018 [cited by applicant]
US 20180046903A1 · Yao et al. · 2018 [cited by applicant]
US 20180211152A1 · Migacz et al. · 2018 [cited by applicant]
US 20180247182A1 · Motoya et al. · 2018 [cited by applicant]
US 20180253641A1 · Yachide et al. · 2018 [cited by applicant]
US 20190251424A1 · Zhou et al. · 2019 [cited by applicant]
US 20220292337A1 · Tian · 2022 [cited by examiner]
US 20230004790A1 · Kang · 2023 [cited by examiner]
US 20230162013A1 · Terashima · 2023 [cited by examiner]
CN 101018196A · 2007 [cited by applicant]
CN 107636697A · 2018 [cited by applicant]
CN 108230277A · 2018 [cited by applicant]
CN 108345939A · 2018 [cited by applicant]
CN 108364061A · 2018 [cited by applicant]
EP 1819105A1 · 2007 [cited by applicant]
JP 06259585A · 1994 [cited by applicant]
JP 07248841A · 1995 [cited by applicant]
JP 11039274A · 1999 [cited by applicant]
JP 2002342308A · 2002 [cited by applicant]
JP 2006154992A · 2006 [cited by applicant]
JP 2006301894A · 2006 [cited by applicant]
JP 2007214795A · 2007 [cited by applicant]
JP 2010041457A · 2010 [cited by applicant]
JP 2010134697A · 2010 [cited by applicant]
JP 2017174039A · 2017 [cited by applicant]
JP 2018135069A · 2018 [cited by applicant]
JP 2018142049A · 2018 [cited by applicant]
JP 2018147182A · 2018 [cited by applicant]
WO 2018101275A1 · 2018 [cited by applicant]
WO 2018140294A1 · 2018 [cited by applicant]
Office Action received for Chinese Patent Application No. 201980066531.1, mailed on Oct. 27, 2023, 17 pages (8 pages of English Translation and 9 pages of Original Document). [cited by applicant]
Hideki Aso et al., “Deep Learning”, Kindaikagaku-sha, Nov. 2015. [cited by applicant]
Hirose et al., “Reduction of memory capacity of deep neural networks by deatomisation”, Incorporated Electronic Information Communication Engineers, vol. 117, No. 45, May 22-24, 2017, pp. 39-44 (English Abstract Submitt… [cited by applicant]
Office Action received for Japanese Patent Application No. 2021-079516, mailed on Mar. 29, 2022, 8 pages (4 pages of English Translation and 4 pages of Office Action). [cited by applicant]
Notice of Reasons for Refusal received for Japanese Patent Application No. 2021-079516, mailed on Nov. 1, 2022, 8 pages (5 pages of English Translation and 3 pages of Original Document). [cited by applicant]
Yuri et al., “The concept for Yasuhiko Nakajima and the distributed neural network by edge computing”, IEICE Technical Report, Japan, general incorporated foundation Institute of Electronics, Information and Communicati… [cited by applicant]
He et al., “Deep residual learning for image recognition”, In Proc. of CVPR, Dec. 2015, pp. 1-12. [cited by applicant]
International Preliminary Report on Patentability received for PCT Patent Application No. PCT/JP2019/035492, mailed on Apr. 22, 2021, 16 pages (10 pages of English Translation and 6 pages of Original Document). [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/JP2019/035492, mailed on Nov. 26, 2019, 18 pages (10 pages of English Translation and 8 pages of Original Document). [cited by applicant]
Office Action received for Japanese Patent Application No. 2020-504046, mailed on Oct. 22, 2020, 13 pages (6 pages of English Translation and 7 pages of Office Action). [cited by applicant]
Notice of Reasons for Refusal received for Japanese Patent Application No. 2021-079516, mailed on Apr. 9, 2024, 34 pages (21 pages of English Translation and 13 pages of Original Document). [cited by applicant]
Office Action received for Chinese Patent Application No. 201980066531.1, mailed on Apr. 7, 2024, 16 pages (9 pages of English Translation and 7 pages of Original Document). [cited by applicant]