IP Library › Granted Patent US 12,417,389
Granted Patent B2
US 12,417,389 · App. 17/100,651 · Granted Sep 16, 2025

Efficient mixed-precision search for quantizers in artificial neural networks

Inventors: Zichuan Liu (San Jose, CA); Kun Wan (San Jose, CA); Xin Lu (San Jose, CA)
Assignee: Adobe Inc.
G06N3/084G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,389
App. No.
17/100,651
Granted
Sep 16, 2025
Kind
B2
Abstract

A method for performing efficient mixed-precision search for an artificial neural network (ANN) includes training the ANN by sampling selected candidate quantizers of a bank of candidate quantizer and updating network parameters for a next iteration based on outputs of layers of the ANN. The outputs are computed by processing quantized data with operators (e.g., convolution). The quantizers converge to optimal bit-widths that reduce classification losses bounded by complexity constrains.

Claims (75)

1. A computer-implemented method for performing a mixed-precision search to design an optimal artificial neural network (ANN) that comprises different bit-widths for different layers of the ANN, the method comprising:

receiving data at a data node of the ANN, wherein the ANN has a network architecture comprising the data node and an operator node with both a quantization edge and a parallel full-precision edge, the quantization edge contains a folding quantizer that includes a set of candidate quantizers with pre-defined bit-widths that are alternatively selectable, and the set of candidate quantizers are not concurrently executed; and

training the ANN based on:

independently sampling a candidate quantizer from the set of candidate quantizers of the folding quantizer that are alternatively selectable for a first training iteration of training the ANN, wherein the candidate quantizer represents weights and activations of the ANN;

computing quantized data based on the candidate quantizer sampled from the set of candidate quantizers, the quantized data comprising a low-precision format with a low bit-width relative to a full-precision format of the received data at the data node;

computing a weighted sum of the quantized data sampled by activating the candidate quantizer sampled from the set of candidate quantizers and unquantized data from the parallel full-precision edge, wherein the unquantized data comprises unsampled candidate quantizers that were replaced with skip connections;

computing an output of the operator node based on the weighted sum; and

updating network parameters of the ANN based on the output of the operator node computed from the weighted sum by updating unsampled candidate quantizers from the set of candidate quantizers to perform a subsequent training iteration of the ANN;

converging the folding quantizer to a fixed optimal quantizer representing a low-precision bit-width relative to a full-precision bit-width of the data; and

based on converging the folding quantizer to the fixed optimal quantizer, generating a trained ANN that comprises different bit-widths for different layers of the ANN.

2. The computer-implemented method of claim 1 further comprising:

converging an additional folding quantizer to an additional fixed optimal quantizer representing an additional low-precision bit-width that is different from the low-precision bit-width relative to the full-precision bit-width of the data, wherein the folding quantizer is for a first layer of the ANN and the additional folding quantizer is for a second layer of the ANN.

3. The computer-implemented method of claim 1 , comprising:

for a first layer of the ANN, converging the folding quantizer to the fixed optimal quantizer representing the low-precision bit-width relative to the full-precision bit-width of the data;

receiving an activation input for a second layer of the ANN that includes an additional folding quantizer for a weight; and

for the second layer of the ANN, converging the additional folding quantizer to an additional fixed optimal quantizer representing an additional low-precision bit-width relative to a full-precision bit-width of the weight.

4. The computer-implemented method of claim 1 , wherein training the ANN comprises:

reducing a number of the set of candidate quantizers available for sampling for the subsequent training iteration.

5. The computer-implemented method of claim 1 , wherein training the ANN comprises:

computing a classification loss of the ANN;

computing a gradient loss of the network parameters; and

updating the network parameters to reduce the classification loss.

6. The computer-implemented method of claim 1 , wherein training the ANN comprises:

selecting the candidate quantizer from the set of candidate quantizers for sampling is based on a learned probability distribution for the set of candidate quantizers, wherein the learned probability distribution is tied to network parameters of the ANN.

7. The computer-implemented method of claim 1 , wherein training the ANN comprises:

selecting the candidate quantizer from the set of candidate quantizers for sampling to reduce a classification loss of the ANN bounded by a complexity constraint.

8. The computer-implemented method of claim 1 , wherein training the ANN comprises:

for a first layer of the ANN and based on converging the folding quantizer to the fixed optimal quantizer, updating a network parameter of the unsampled candidate quantizers from the set of candidate quantizers; and

for a second layer of the ANN and based on converging an additional folding quantizer to an additional fixed optimal quantizer, updating a network parameter of additional unsampled candidate quantizers from an additional set of candidate quantizers.

9. The computer-implemented method of claim 1 , wherein training the ANN comprises:

using a common weight for each candidate quantizer of the set of candidate quantizers.

10. The computer-implemented method of claim 1 , wherein training the ANN comprises:

using an independent weight for each candidate quantizer of the set of candidate quantizers.

11. The computer-implemented method of claim 1 , wherein training the ANN comprises:

updating the network parameters of the unsampled candidate quantizers from the set of candidate quantizers and the sampled candidate quantizer from the set of candidate quantizers for the subsequent training iteration.

12. A non-transitory computer readable medium storing instructions thereon that, when executed by at least one processor, cause a computing device to perform operations to design an optimal artificial neural network (ANN) that comprises different bit-widths for different layers of the ANN comprising:

receiving data at a data node of the ANN, wherein the ANN has a network architecture comprising the data node and an operator node with both a quantization edge and a parallel full-precision edge, the quantization edge contains a folding quantizer that includes a set of candidate quantizers with pre-defined bit-widths that are alternatively selectable, and the set of candidate quantizers are not concurrently executed; and

training the ANN based on:

independently sampling a candidate quantizer from the set of candidate quantizers of the folding quantizer that are alternatively selectable for a first training iteration of training the ANN, wherein the candidate quantizer represents weights and activations of the ANN;

computing quantized data based on the candidate quantizer sampled from the set of candidate quantizers, the quantized data comprising a low-precision format with a low bit-width relative to a full-precision format of the received data at the data node;

computing a weighted sum of the quantized data sampled by activating the candidate quantizer sampled from the set of candidate quantizers and unquantized data from the parallel full-precision edge, wherein the unquantized data comprises unsampled candidate quantizers that were replaced with skip connections;

computing an output of the operator node based on the weighted sum; and

updating network parameters of the ANN based on the output of the operator node computed from the weighted sum by updating unsampled candidate quantizers from the set of candidate quantizers to perform a subsequent training iteration of the ANN;

converging the folding quantizer to a fixed optimal quantizer representing a low-precision bit-width relative to a full-precision bit-width of the data; and

based on converging the folding quantizer to the fixed optimal quantizer, generating a trained ANN that comprises different bit-widths for different layers of the ANN.

13. The non-transitory computer readable medium of claim 12 , further comprising:

converging an additional folding quantizer to an additional fixed optimal quantizer representing an additional low-precision bit-width that is different from the low-precision bit-width, wherein the folding quantizer is for a first layer of the ANN and the additional folding quantizer is for a second layer of the ANN.

14. The non-transitory computer readable medium of claim 12 , comprising:

for a first layer of the ANN, converging the folding quantizer to the fixed optimal quantizer representing the low-precision bit-width relative to the full-precision bit-width of the data;

receiving an activation input for a second layer of the ANN that includes an additional folding quantizer for a weight; and

for the second layer of the ANN, converging the additional folding quantizer to an additional fixed optimal quantizer representing an additional low-precision bit-width relative to a full-precision bit-width of the weight.

15. The non-transitory computer readable medium of claim 12 wherein training the ANN comprises:

selecting the candidate quantizer from the set of candidate quantizers for sampling is based on a learned probability distribution for the set of candidate quantizers, wherein the learned probability distribution is tied to network parameters of the ANN.

16. A system comprising:

at least one memory device comprising an artificial neural network (ANN) that comprises different bit-widths for different layers of the ANN, a data node, and an operator node; and

at least one processor configured to cause the system to:

receive data at the data node of the ANN, wherein the ANN has a network architecture comprising the data node and the operator node with both a quantization edge and a parallel full-precision edge-, the quantization edge contains a folding quantizer that includes a set of candidate quantizers with pre-defined bit-widths that are alternatively selectable and the set of candidate quantizers are not concurrently executed; and

train the ANN by:

independently sample a candidate quantizer from the set of candidate quantizers of the folding quantizer that are alternatively selectable for a first training iteration of training the ANN, wherein the candidate quantizer represents weights and activations of the ANN;

computing quantized data based on the candidate quantizer sampled from the set of candidate quantizers, the quantized data comprising a low-precision format with a low bit-width relative to a full-precision format of the received data at the data node;

computing a weighted sum of the quantized data sampled by activating the candidate quantizer sampled from the set of candidate quantizers and unquantized data from the parallel full-precision edge, wherein the unquantized data comprises unsampled candidate quantizers that were replaced with skip connections;

computing an output of the operator node based on the weighted sum; and

updating network parameters of the ANN based on the output of the operator node computed from the weighted sum by updating unsampled candidate quantizers from the set of candidate quantizers to perform a subsequent training iteration of the ANN;

converge the folding quantizer to a fixed optimal quantizer representing a low-precision bit-width relative to a full-precision bit-width of the data; and

based on converging the folding quantizer to the fixed optimal quantizer, generate a trained ANN that comprises different bit-widths for different layers of the ANN.

17. The system of claim 16 , wherein the at least one processor is further configured to cause the system to converge an additional folding quantizer to an additional fixed optimal quantizer representing an additional low-precision bit-width that is different from the low-precision bit-width, wherein the folding quantizer is for a first layer of the ANN and the additional folding quantizer is for a second layer of the ANN.

18. The system of claim 16 , wherein the at least one processor is further configured to cause the system to:

for a first layer of the ANN, converge the folding quantizer to the fixed optimal quantizer representing the low-precision bit-width relative to the full-precision bit-width of the data;

receive an activation input for a second layer of the ANN that includes an additional folding quantizer for a weight; and

for the second layer of the ANN, converge the additional folding quantizer to an additional fixed optimal quantizer representing an additional low-precision bit-width relative to a full-precision bit-width of the weight.

19. The system of claim 16 , wherein the at least one processor is further configured to cause the system to reduce a number of the set of candidate quantizers available for sampling for the subsequent training iteration.

20. The system of claim 16 , wherein the at least one processor is further configured to cause the system to:

compute a classification loss of the ANN;

compute a gradient loss of the network parameters; and

update the network parameters to reduce the classification loss.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 23, 2020
From: LIU, ZICHUAN; LU, XIN; WAN, KUN
To: ADOBE INC.
Reel/Frame 054442/0475 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2020
From: LIU, ZICHUAN; LU, XIN; WAN, KUN
To: ADOBE INC.
Reel/Frame 054437/0212 →
Continuity (1)
Related Publication 20220164666A1 · May 26, 2022
References Cited (46)
US 20200097831A1 · Wang · 2020 [cited by examiner]
US 20200293893A1 · Georgiadis · 2020 [cited by examiner]
US 20210005208A1 · Beack · 2021 [cited by examiner]
US 20210218414A1 · Malhotra · 2021 [cited by examiner]
US 20210232890A1 · Li · 2021 [cited by examiner]
Bender, G. “Understanding and Simplifying One-Shot Architecture Search,” 2019, 10 pages. [cited by applicant]
Bengio, et al., “Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation,” arXiv preprint arXiv:1308.3432, 2013, 12 pages. [cited by applicant]
Bi, et al., “Stabilizing DARTS with Amended Gradient Estimation on Architectural Parameters,” arXiv preprint arXiv:1910.11831, 2019, 22 pages. [cited by applicant]
Brock, et al., “Smash: One-Shot Model Architecture Search Through Hypernetworks,” arXiv preprint arXiv:1708.05344, 2017, 21 pages. [cited by applicant]
Cai, et al., “Deep Learning with Low Precision by Half-wave Gaussian Quantization,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 5918-5926. [cited by applicant]
Cai, et al., “ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware,” arXiv preprint arXiv:1812.00332, 2018, 13 pages. [cited by applicant]
Cai, et al., “Rethinking Differentiable Search for Mixed-Precision Neural Networks,” arXiv preprint arXiv:2004.05795, 2020, 10 pages. [cited by applicant]
Chen, et al., “Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2019, pp. 3434-3443. [cited by applicant]
Choi, et al, “Pact: Parameterized Clipping Activation for Quantized Neural Networks,” arXiv preprint arXiv:1805.06085, 2018, 15 pages. [cited by applicant]
Chu, et al., “Fairnas: Rethinking Evaluation Fairness of Weight Sharing Neural Architecture Search,” arXiv preprint arXiv:1907.01845, 2019, 23 pages. [cited by applicant]
Courbariaux, et al., “Binaryconnect: Training Deep Neural Networks with binary weights during propagations,” in Advances in neural information processing systems, 2015, pp. 3123-3131. [cited by applicant]
Deng, et al., “Imagenet: A Large-Scale Hierarchical Image Database,” in 2009 IEEE conference on computer vision and pattern recognition. IEEE, 2009, pp. 248-255. [cited by applicant]
Guo, et al., “Single Path One-Shot Neural Architecture Search with Uniform Sampling,” arXiv preprint arXiv:1904.00420, 2019, 16 pages. [cited by applicant]
Han, et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” arXiv preprint arXiv:1510.00149, 2015, 14 pages. [cited by applicant]
Han, et al., “Learning both Weights and Connections for Efficient Neural Network,” in Advances in neural information processing systems, 2015, pp. 1135-1143. [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770-778. [cited by applicant]
Howard, et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” arXiv preprint arXiv:1704.04861, 2017, 9 pages. [cited by applicant]
Howard, et al., “Searching for MobileNetv3,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2019, pp. 1314-1324. [cited by applicant]
Hubara, et al., “Binarized Neural Networks: Training Neural Networks with Weights and Activations Constrained to +1 or −1,” in Advances in neural information processing systems, 2016, pp. 4107-4115. [cited by applicant]
Jung, et al., “Learning to Quantize Deep Networks by Optimizing Quantization Intervals with Task Loss,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4350-4359. [cited by applicant]
Krizhevsky, A., “Cifar-10 and cifar-100 datasets,” 2009. [Online]. Available: https://www.cs.toronto.edu/˜kriz/cifar.html, 3 pages. [cited by applicant]
Lillicrap, et al., “Continuous Control with Deep Reinforcement Learning,” arXiv preprint arXiv:1509.02971, 2015, 14 pages. [cited by applicant]
Liu, et al., “Darts: Differentiable Architecture Search,” arXiv preprint arXiv:1806.09055, 2018, 13 pages. [cited by applicant]
Pham, et al., “Efficient Neural Architecture Search via Parameter Sharing,” arXiv preprint arXiv:1802.03268, 2018. 11 pages. [cited by applicant]
Rastegari, et al, “XNOR-NET: ImageNet Classification Using Binary Convolutional Neural Networks,” in European conference on Computer Vision. Springer, 2016, pp. 525-542. [cited by applicant]
Sandler, et al., “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 4510-4520. [cited by applicant]
Schulman, et al., “Proximal Policy Optimization Algorithms,” arXiv preprint arXiv:1707.06347, 2017, 12 pages. [cited by applicant]
Singh, et al., “HetConv: Heterogeneous Kernel-Based Convolutions for Deep CNNs,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 4830-4839. [cited by applicant]
Stamoulis, et al., “Single-Path NAS: Designing Hardware-Efficient ConvNets in less than 4 Hours,” arXiv preprint arXiv:1904.02877, 2019, 16 pages. [cited by applicant]
Tan, et al., “MnasNet: Platform-Aware Neural Architecture Search for Mobile,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 2820-2828. [cited by applicant]
Wang, et al., “HAQ: Hardware-Aware Automated Quantization with Mixed Precision,” in Proceedings of the IEEE conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 8612-8620. [cited by applicant]
Wu, et al. “Mixed Precision Quantization Of ConvNets via Differentiable Neural Architecture Search,” arXiv preprint arXiv:1812.00090, 2018, 11 pages. [cited by applicant]
Wu, et al., “FBNet: Hardware-Aware Efficient ConvNet Design via Differentiable Neural Architecture Search,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 10 734-10 74… [cited by applicant]
Xie, et al., “SNAS: Stochastic Neural Architecture Search,” arXiv preprint arXiv:1812.09926, 2018, 17 pages. [cited by applicant]
Xu, et al., “PC-DARTS: Partial Channel Connections for Memory-Efficient Differentiable Architecture Search,” arXiv preprint arXiv:1907.05737, 2019, 13 pages. [cited by applicant]
You, et al., “GreedyNAS: Towards Fast One-Shot NAS with Greedy Supernet,” arXiv preprint arXiv:2003.11236, 2020. 15 [ages. [cited by applicant]
Zhang, et al., “LQ-Nets: Learned Quantization for Highly Accurate and Compact Deep Neural Networks,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 365-382. [cited by applicant]
Zhou, et al., “DoReFa-NET: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients,” arXiv preprint arXiv:1606.06160, 2016, 13 pages. [cited by applicant]
Zhu, et al., “Trained Ternary Quantization,” arXiv preprint arXiv:1612.01064, 2016, 10 pages. [cited by applicant]
Zhuang, et al., “Towards Effective Low-bitwidth Convolutional Neural Networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 7920-7928. [cited by applicant]
Zoph, et al., “Neural Architecture Search with Reinforcement Learning,” arXiv preprint arXiv:1611.01578, 2016, 16 pages. [cited by applicant]