IP Library Granted Patent US 12,561,562
Granted Patent B2
US 12,561,562 · App. 17/672,627 · Granted Feb 24, 2026

Accuracy of low-bitwidth neural networks by regularizing the higher-order moments of weights and hidden states

Inventors: Nicholas Dronen (Newton, MA); Tyler J. Kenney (Boston, MA); Tomo Lazovich (Cambridge, MA); Ayon Basumallik (Framingham, MA); Darius Bunandar (Boston, MA)
Assignee: Lightmatter, Inc.
G06N3/08G06F17/16G06N3/065
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,562
App. No.
17/672,627
Granted
Feb 24, 2026
Kind
B2
Abstract

Methods and systems for training neural networks using low-bitwidth accelerators are described. The methods described herein use moment-penalization functions. For example, a method comprises producing a modified data set by training a neural network using a moment-penalization function and the data set. The moment-penalization function is configured to penalize a moment associated with the neural network. Training the neural network in turn comprises quantizing the data set to obtain a fixed-point data set so that the fixed-point data set represents the data set in a fixed-point representation, and passing the fixed-point data set through an analog accelerator. The inventors have recognized that training a neural network using a modified objective function augments the accuracy and robustness of the neural network notwithstanding the use of low-bitwidth accelerators.

Claims (50)

1 . A method comprising:

receiving a data set;

producing a modified data set by training a neural network using a modified objective function and the data set, wherein training the neural network comprises:

modifying an objective function associated with the neural network by adding a moment-penalization function to an objective function to produce the modified objective function, wherein the moment-penalization function is configured to penalize a moment associated with the neural network;

quantizing the data set to obtain a fixed-point data set so that the fixed-point data set represents the data set in a fixed-point representation; and

passing the fixed-point data set through an analog accelerator to produce the modified data set.

2 . The method of claim 1 , wherein the modified data set exhibits a distribution selected from the group consisting of:

a normal distribution,

a uniform distribution, and

a bimodal distribution.

3 . The method of claim 1 , wherein the data set exhibits a leptokurtic distribution.

4 . The method of claim 1 , wherein the moment-penalization function is configured to penalize a kurtosis associated with the data set.

5 . The method of claim 4 , wherein adding the moment-penalization function to the objective function comprises adding a a kurtosis-penalization function to the objective function.

6 . The method of claim 5 , wherein the kurtosis-penalization function is configured to penalize the kurtosis associated with a set of weights of the neural network and/or a set of hidden layers of the neural network.

7 . The method of claim 1 , wherein the moment-penalization function is configured to penalize a skewness associated with the data set.

8 . The method of claim 7 , wherein adding the moment-penalization function to the objective function comprises adding a skewness-penalization function to the objective function.

9 . The method of claim 8 , wherein the skewness-penalization function is configured to penalize the skewness associated with a set of weights of the neural network and/or a set of hidden layers of the neural network.

10 . The method of claim 1 , wherein quantizing the data set to obtain the fixed-point data set comprises quantizing the data set using an b-bit quantizer with b≤12.

11 . The method of claim 1 , wherein the analog accelerator comprises a photonic accelerator, and wherein passing the fixed-point data set through the analog accelerator comprises performing matrix-matrix multiplication in an optical domain.

12 . The method of claim 11 , wherein performing matrix-matrix multiplication in the optical domain comprises encoding light with a plurality of weights representing the neural network.

13 . A system comprising:

an analog accelerator; and

at least one computer hardware processor to perform:

receiving a data set;

producing a modified data set by training a neural network using a modified objective function and the data set, wherein training the neural network comprises:

modifying an objective function associated with the neural network by adding a moment-penalization function to the objective function to produce the modified objective function, wherein the moment-penalization function is configured to penalize a moment associated with the neural network;

quantizing the data set to obtain a fixed-point data set so that the fixed-point data set represents the data set in a fixed-point representation; and

passing the fixed-point data set through the analog accelerator to produce the modified data set.

14 . The system of claim 13 , wherein the modified data set exhibits a distribution selected from the group consisting of:

a normal distribution,

a uniform distribution, and

a bimodal distribution.

15 . The system of claim 13 , wherein the data set exhibits a leptokurtic distribution.

16 . The system of claim 13 , wherein the moment-penalization function is configured to penalize a kurtosis associated with the data set.

17 . The system of claim 16 , wherein adding the moment-penalization function to the objective function comprises adding a kurtosis-penalization function to the objective function.

18 . The system of claim 13 , wherein the moment-penalization function is configured to penalize a skewness associated with the data set.

19 . The system of claim 18 , wherein adding the moment-penalization function to the objective function comprises adding a skewness-penalization function to the objective function.

20 . The system of claim 13 , wherein quantizing the data set to obtain the fixed-point data set comprises quantizing the data set using an b-bit quantizer with b≤12.

21 . The system of claim 13 , wherein the analog accelerator comprises a photonic accelerator, and wherein passing the fixed-point data set through the analog accelerator comprises performing matrix-matrix multiplication in an optical domain.

22 . At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one processor, cause the at least one processor to perform a method comprising:

receiving a data set;

producing a modified data set by training a neural network using a modified objective function and the data set, wherein training the neural network comprises:

modifying an objective function associated with the neural network by adding a moment-penalization function to the objective function to produce the modified objective function, wherein the moment-penalization function is configured to penalize a moment associated with the neural network;

quantizing the data set to obtain a fixed-point data set so that the fixed-point data set represents the data set in a fixed-point representation; and

passing the fixed-point data set through an analog accelerator to produce the modified data set.

23 . The at least one non-transitory computer-readable storage medium of claim 22 , wherein the modified data set exhibits a distribution selected from the group consisting of:

a normal distribution,

a uniform distribution, and

a bimodal distribution.

24 . The at least one non-transitory computer-readable storage medium of claim 22 , wherein the data set exhibits a leptokurtic distribution.

Assignments (4)
TERMINATION OF IP SECURITY AGREEMENT Recorded Nov 5, 2024
From: EASTWARD FUND MANAGEMENT, LLC
To: LIGHTMATTER, INC.
Reel/Frame 069304/0700 →
RELEASE OF SECURITY INTEREST Recorded Mar 31, 2023
From: EASTWARD FUND MANAGEMENT, LLC
To: LIGHTMATTER, INC.
Reel/Frame 063209/0966 →
SECURITY INTEREST Recorded Dec 27, 2022
From: LIGHTMATTER, INC.
To: EASTWARD FUND MANAGEMENT, LLC
Reel/Frame 062230/0361 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2022
From: DRONEN, NICHOLAS; KENNEY, TYLER J.; LAZOVICH, TOMO; BASUMALLIK, AYON; BUNANDAR, DARIUS
To: LIGHTMATTER, INC.
Reel/Frame 059079/0253 →
Continuity (2)
Provisional Application 63150032 · Feb 16, 2021
Related Publication 20220261645A1 · Aug 18, 2022
References Cited (43)
US 10534189B2 · Miller · 2020 [cited by applicant]
US 10740693B2 · Lazovich et al. · 2020 [cited by applicant]
US 10763974B2 · Bunandar et al. · 2020 [cited by applicant]
US 10803259B2 · Kenney et al. · 2020 [cited by applicant]
US 11315019B2 · Nachum · 2022 [cited by examiner]
US 20170351293A1 · Carolan et al. · 2017 [cited by applicant]
US 20180046913A1 · Yu · 2018 [cited by examiner]
US 20180274900A1 · Mower et al. · 2018 [cited by applicant]
US 20190370652A1 · Shen et al. · 2019 [cited by applicant]
US 20200134324A1 · Zeng · 2020 [cited by examiner]
US 20200250518A1 · Khoury et al. · 2020 [cited by applicant]
US 20200387798A1 · Hewage · 2020 [cited by examiner]
US 20200394523A1 · Liu · 2020 [cited by examiner]
US 20210089906A1 · Lazovich · 2021 [cited by applicant]
US 20210125066A1 · Lazovich · 2021 [cited by applicant]
US 20210241156A1 · Lai · 2021 [cited by examiner]
US 20210405682A1 · Gould et al. · 2021 [cited by applicant]
US 20220092424A1 · Almazán Manzanares · 2022 [cited by examiner]
Ulseth et al. (“Training of mixed-signal optical convolutional neural network with reduced quantization level”, 2020 arxiv) (Year: 2020). [cited by examiner]
Ververidis et al. (“Gaussian Mixture Modeling by Exploiting the Mahalanobis Distance”, IEEE vol. 56, No. 7, Jul. 2008) (Year: 2008). [cited by examiner]
Shkolnik et al. (“Robust Quantization: One Model to Rule Them All”, NeurIPS 2020) (Year: 2020). [cited by examiner]
Cao et al. (“Deep Priority Hashing”, MM'18, Oct. 22-26, 2018, pp. 1653-1661) (Year: 2018). [cited by examiner]
Duchi et al. (“Variance-based Regularization with Convex Objectives”, Journal of Machine Learning Research 19 (2018) 1-55) (Year: 2018). [cited by examiner]
Alfadly et al. (“Analytical Moment Regularizer for Gaussian Robust Networks”, 2019) (Year: 2019). [cited by examiner]
Abu-Mostafa et al., Optical neural computers. Scientific American 256.3 (1987):88-95. [cited by applicant]
Amit et al., Spin-glass models of neural networks. Physical Review A. 1985;32(2):1007-1018. [cited by applicant]
Atabaki et al., Integrating photonics with silicon nanoelectronics for the next generation of systems on a chip. Nature. 2018;556(7701):349-354. 10 pages. DOI: 10.1038/s41586-018-0028-z. [cited by applicant]
Bruck et al., On the power of neural networks for solving hard problems. American Institute of Physics. 1988. pp. 137-143. 7 pages. [cited by applicant]
Canziani et al., Evaluation of neural network architectures for embedded systems. Circuits and Systems (ISCAS). 2017 IEEE International Symposium. 4 pages. [cited by applicant]
Chen et al., DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning. ACM Sigplan Notices. 2014;49:269-283. [cited by applicant]
Graves et al., Hybrid computing using a neural network with dynamic external memory. Nature. 2016;538. 21 pages. DOI:10.1038/nature20101. [cited by applicant]
Lecun et al., Deep learning. Nature. 2015;521:436-444. DOI:10.1038/nature14539. [cited by applicant]
Misra et al., Artificial neural networks in hardware: A survey of two decades of progress. Neurocomputing. 2010;74:239-255. [cited by applicant]
Schmidhuber, Deep learning in neural networks: An overview. Neural Networks. 2015;61:85-117. [cited by applicant]
Shafiee et al., Isaac: A convolutional neural network accelerator with in-situ analog arithmetic in crossbars. ACM/IEEE 43rd Annual International Symposium on Computer Architecture. Oct. 2016. 13 pages. [cited by applicant]
Shen et al., Deep learning with coherent nanophotonic circuits. Nature Photonics. 2017;11:441-6. DOI: 10.1038/NPHOTON.2017.93. [cited by applicant]
Shkolnik et al., Robust Quantization: One Model to Rule Them All. 34 [cited by applicant]
Solli et al., Analog optical computing. Nature Photonics. 2015;9:704-6. [cited by applicant]
Sun et al., Single-chip microprocessor that communicates directly using light. Nature. 2015;528:534-8. DOI: 10.1038/nature16454. [cited by applicant]
Tait et al., Broadcast and weight: An integrated network for scalable photonic spike processing. Journal of Lightwave Technology. 2014;32(21):3427-39. DOI: 10.1109/JLT.2014.2345652. [cited by applicant]
Tait et al., Chapter 8 Photonic Neuromorphic Signal Processing and Computing. Springer, Berlin, Heidelberg. 2014. pp. 183-222. [cited by applicant]
Tait et al., Neuromorphic photonic networks using silicon photonic weight banks. Science Reports. 2017;7:7430. 10 pages. [cited by applicant]
Wu et al., Integer Quantization for Deep Learning Inference: Principles and Empirical Evaluation. NVIDIA. Apr. 20, 2020. 20 pages. arXiv:2004.09602v1. [cited by applicant]