IP Library Granted Patent US 12,499,365
Granted Patent B2
US 12,499,365 · App. 17/574,433 · Granted Dec 16, 2025

Artificial neural networks generated by low discrepancy sequences

Inventors: Alexander Keller (Berlin, DE); Matthijs Jules Van Keirsbilck (Berlin, DE)
Assignee: NVIDIA CORPORATION
G06N3/082G06F18/214G06N3/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,365
App. No.
17/574,433
Granted
Dec 16, 2025
Kind
B2
Abstract

Artificial neural networks (ANNs) are computing systems that imitate a human brain by learning to perform tasks by considering examples. These ANNs are typically created by connecting several layers of neural units using connections, where each neural unit is connected to every other neural unit either directly or indirectly to create fully connected layers within the ANN. However, by representing an artificial neural network utilizing paths from an input of the ANN to an output of the ANN, a complexity of the ANN may be reduced, and the ANN may be trained and implemented in a much faster manner when compared to fully connected layers within the ANN. More specifically, the ANN may be trained sparse from scratch in order to avoid a more expensive procedure of training the ANN and compressing it afterwards.

Claims (45)

1 . A method to accelerate training of an ANN, the method comprising:

creating the ANN as a plurality of paths,

where each path of the plurality of paths includes:

(a) a plurality of vertices each representing a neural unit within the ANN, and

(b) a plurality of edges each representing a weighted connection within the ANN, and

where each path of the plurality of paths is assigned a number;

for each path of the plurality of paths, initializing all weighted connections within the path with a same sign determined based on the number assigned to the path, where the sign is determined as either a positive sign or a negative sign; and

training the ANN, including for each weighted connection within the ANN:

keeping the sign of the weighted connection fixed during the training, and

training a magnitude of a value of the weighted connection,

where the training of the ANN is accelerated by training only the magnitude of the value of the weighted connection while keeping the sign of the weighted connection fixed.

2 . The method of claim 1 , wherein at least one layer of the ANN is at least one of a fully connected layer and a 1×1 convolution.

3 . The method of claim 1 , wherein the signs of the values of all weighted connections along a first predetermined number of the plurality of paths are positive, and the signs of the values of all weighted connections along a second predetermined number of the plurality of paths are negative.

4 . The method of claim 1 , wherein the ANN is trained utilizing labeled input training data.

5 . The method of claim 4 , wherein the training is semi-supervised.

6 . The method of claim 4 , wherein the training is unsupervised.

7 . The method of claim 1 , further comprising processing input data by the trained ANN to produce output data, the input data including one or more of image data, textual data, audio data, and video data, random numbers, pseudo-random numbers, and quasi-random numbers, and the output data including one or more of a classification, a categorization, a regression, a function approximation, a probability, and samples of data according to a learned distribution.

8 . The method of claim 1 , where each path of the plurality of paths of the ANN extends from an input layer of the ANN to an output layer of the ANN.

9 . The method of claim 1 , further comprising:

implementing the trained ANN; and

processing input data by the trained ANN to produce output data,

wherein during the processing of the input data to produce the output data, the same sign of all weighted connections within a path of the plurality of paths of the trained ANN is not explicitly considered to improve a performance of the trained ANN.

10 . A method to accelerate training of an ANN, the method comprising:

creating the ANN as a plurality of paths each connecting at least one input of the ANN to at least one output of the ANN,

where each of the plurality of paths are generated utilizing a low discrepancy sequence;

initializing a value of all weighted connections in a path of the plurality of paths with a same sign determined from a dimension of the low discrepancy sequence, where the sign is determined as either a positive sign or a negative sign; and

training the ANN, where a sign of a value of each weighted connection is fixed and only the magnitude of the value of each weighted connection is trained:

keeping the sign of the weighted connection fixed during the training, and

training a magnitude of a value of the weighted connection,

where the training of the ANN is accelerated by training only the magnitude of the value of the weighted connection while keeping the sign of the weighted connection fixed.

11 . The method of claim 10 , wherein dimensions of the low discrepancy sequences that cause coalescing links within the ANN are skipped.

12 . The method of claim 10 , wherein the low discrepancy sequence is at least one of a Sobol′, Niederreiter, and Niederreiter-Xing sequence.

13 . The method of claim 10 , wherein the low discrepancy sequence includes a scrambled low discrepancy sequence.

14 . The method of claim 10 , wherein successive dimensions of the low discrepancy sequence form (0,m,s)-nets in base b.

15 . The method of claim 10 , wherein each of the plurality of paths are sampled by a Cascaded Sobol′ sequence.

16 . The method of claim 10 , wherein connections between two layers of the ANN are created by referencing neurons by a corresponding component of the low discrepancy sequence and referencing inputs in their natural order or a permutation of their natural order.

17 . The method of claim 10 , wherein convolutional layers within the ANN are sampled by selecting one of the input channels of the ANN and one of a plurality of convolutional neurons of the ANN specified by an edge of a path.

18 . The method of claim 10 , wherein the ANN is trained utilizing labeled input training data.

19 . The method of claim 18 , wherein the training is semi-supervised.

20 . The method of claim 18 , wherein the training is unsupervised.

21 . The method of claim 10 , further comprising processing input data by the trained ANN to produce output data, the input data including one or more of image data, textual data, audio data, and video data, random numbers, pseudo-random numbers, and quasi-random numbers, and the output data including one or more of a classification, a categorization, a regression, a function approximation, a probability, and samples of data according to a learned distribution.

22 . The method of claim 10 , further comprising:

implementing the trained ANN; and

processing input data by the trained ANN to produce output data,

wherein during the processing of the input data to produce the output data, the same sign of all weighted connections within a path of the plurality of paths of the trained ANN is not explicitly considered to improve a performance of the trained ANN.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2022
From: KELLER, ALEXANDER; VAN KEIRSBILCK, MATTHIJS JULES
To: NVIDIA CORPORATION
Reel/Frame 059259/0914 →
Continuity (2)
Provisional Application 63156819 · Mar 4, 2021
Related Publication 20220284294A1 · Sep 8, 2022
References Cited (58)
US 20050149463A1 · Bolt · 2005 [cited by examiner]
US 20190122106A1 · Lee · 2019 [cited by examiner]
US 20190171936A1 · Karras et al. · 2019 [cited by applicant]
US 20190294972A1 · Keller · 2019 [cited by examiner]
US 20220319039A1 · Filip · 2022 [cited by examiner]
CN 110909756A · 2020 [cited by applicant]
DE 102018108314A1 · 2018 [cited by applicant]
DE 102018123518A1 · 2019 [cited by applicant]
G. Bellec et al., “Deep Rewiring: Training very sparse deep networks,” Aug. 7, 2018, https://arxiv.org/abs/1711.05136 (Year: 2018). [cited by examiner]
M. Mehmet et al., “Neural Networks for Shortest Path Computation and Routing in Computer Networks,” published in Nov. 1993 (Year: 1993). [cited by examiner]
Bellec, G., Scherr, F., Subramoney, A. et al. A solution to the learning dilemma for recurrent networks of spiking neurons. Nat Commun 11, 3625 (2020). https://doi.org/10.1038/s41467-020-17236-y (Year: 2020). [cited by examiner]
“An Explanation of Xavier Initialization,” published on Feb. 14, 2015 (Jones). (Year: 2015). [cited by examiner]
Lois Paulin, David Coeurjolly, Jean-Claude Iehl, Nicolas Bonneel, Alexander Keller, et al.. Cascaded Sobol′ Sampling. ACM Transactions on Graphics, 2021, Proceedings of SIGGRAPH Asia 2021, 40 (6), pp. 274:1--274:13. 10.… [cited by examiner]
Simons Foundation article, “When Less is More: Sparse Networks Outperform Dense Ones,” published on May 5, 2017 (“Simons”). <URL = https://www.simonsfoundation.org/2017/05/05/when-less-is-more-sparse-networks-can-outper… [cited by examiner]
Changpinyo et al., “The power of sparsity in convolutional neural networks,” arXiv, 2017, pp. 1-13, retrieved from https://arxiv.org/abs/1702.06257. [cited by applicant]
Child et al., “Generating Long Sequences with Sparse Transformers,” arXiv, 2019, 10 pages, retrieved from https://arxiv.org/abs/1904.10509. [cited by applicant]
Dettmers et al., “Sparse Networks from Scratch: Faster Training without Losing Performance,” arXiv, 2019, 14 pages, retrieved from https://arxiv.org/abs/1907.04840. [cited by applicant]
Dey et al., “Interleaver Design for Deep Neural Networks,” IEEE 51st Asilomar Conference on Signals, Systems, and Computers, 2017, 5 pages. [cited by applicant]
Dey et al., “Characterizing Sparse Connectivity Patterns in Neural Networks,” IEEE Information Theory and Applications Workshop, 2018, 8 pages, retrieved from https://arxiv.org/abs/1711.02131v4. [cited by applicant]
Dey et al., “Pre-Defined Sparse Neural Networks with Hardware Acceleration,” arXiv, 2018 pp. 1-14, retrieved from http://arxiv.org/abs/1812.01164. [cited by applicant]
Dey et al., “Accelerating Training of Deep Neural Networks via Sparse Edge Processing,” 26th International Conference on Artificial Neural Networks (ICANN), 2017, pp. 1-8, retrieved from https://arxiv.org/abs/1711.01343. [cited by applicant]
Dick et al., “Digital Nets and Sequences. Discrepancy Theory and Quasi-Monte Carlo Integration,” Cambridge University Press, 2010, 629 pages. [cited by applicant]
Farhat et al., “Optical implementation of the Hopfield model,” Applied Optics, vol. 24, No. 10, May 15, 1985, pp. 1469-1475. [cited by applicant]
Frankie et al., “The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks,” ICLR, 2019, pp. 1-42. [cited by applicant]
Glorot et al., “Deep Sparse Rectifier Neural Networks,” Proceedings of the 14th International Conference on Artificial Intelligence and Statistics, 2011, pp. 315-323. [cited by applicant]
Gray et al., “GPU Kernels for Block-Sparse Weights,” Semantic Scholar, 2017, 12 pages, retrieved from https://www.semanticscholar.org/paper/GPU-Kernels-for-Block-Sparse-Weights-Gray-Radford/a07609c2ed39d049d3e59b61408fb… [cited by applicant]
He et al., “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification,” IEEE International Conference on Computer Vision, Feb. 6, 2015, pp. 1-11, retrieved from https://www.semanticscho… [cited by applicant]
Heitz et al., “A Low-Discrepancy Sampler that Distributes Monte Carlo Errors as a Blue Noise in Screen Space,” SIGGRAPH'19 Talks, 2019, 3 pages. [cited by applicant]
Huang et al., “Densely Connected Convolutional Networks,” IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 4700-4708. [cited by applicant]
Ioffe et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” Proceedings of the 32nd International Conference on Machine Learning, 2015, 9 pages. [cited by applicant]
Jayakumar et al., “Top-KAST: Top-K Always Sparse Training,” 34th Conference on Neural Information Processing Systems (NeurIPS), 2020, pp. 1-11. [cited by applicant]
Joe et al., “Notes on generating Sobol′ sequences,” Technical report, School of Mathematics and Statistics, University of New South Wales, Aug. 2008, 3 pages, retrieved from https://web.maths.unsw.edu.au/˜fkuo/sobol/joe… [cited by applicant]
Keller, A., “Myths of computer graphics,” in Monte Carlo and Quasi-Monte Carlo Methods 2004, Springer, 2006, pp. 217-243, retrieved from https://ur.booksc.me/book/21518084/92cddb. [cited by applicant]
Keller, A., “Quasi-Monte Carlo Image Synthesis in a Nutshell,” NVIDIA Research, 2013, 37 pages, retrieved from https://www.semanticscholar.org/paper/Quasi-Monte-Carlo-Image-Synthesis-in-a-Nutshell-Keller/401844288ae3a1c… [cited by applicant]
Keller et al., “Parallel Quasi-Monte Carlo Integration by Partitioning Low Discrepancy Sequences,” NVIDIA, 2012, 12 pages, retrieved from https://www.semanticscholar.org/paper/Parallel-Quasi-Monte-Carlo-Integration-by-L… [cited by applicant]
Kriman et al., “QuartzNet: Deep Automatic Speech Recognition with 1D Time-Channel Separable Convolutions,” 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, 5 pages, retrieved… [cited by applicant]
Kundu et al., “Pre-defined Sparsity for Low-Complexity Convolutional Neural Networks,” arXiv, 2020, pp. 1-14, retrieved from https://arxiv.org/abs/2001.10710. [cited by applicant]
Kundu et al., “pSConv: A Pre-defined Sparse Kernel Based Convolution for Deep CNNs,” 57th Annual Allerton Conference on Communication, Control, and Computing, 2019, pp. 100-107. [cited by applicant]
Kung et al., “Systolic arrays (for VLSI),” Department of Computer Science, Carnegie-Mellon University, Apr. 1978, 31 pages. [cited by applicant]
Lecun et al., “Gradient-Based Learning Applied to Document Recognition,” Proceedings of the IEEE, Nov. 1998, pp. 1-46. [cited by applicant]
Molchanov et al., “Importance Estimation for Neural Network Pruning,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 11264-11272. [cited by applicant]
Mordido et al., “Instant Quantization of Neural Networks using Monte Carlo Methods,” 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing (NeurIPS 2019 EMC2), 2019, pp. 1-5. [cited by applicant]
Mordido et al., “Monte Carlo Gradient Quantization,” CVPR 2020 Joint Workshop on Efficient Deep Learning in Computer Vision, 2020, 9 pages. [cited by applicant]
Niederreiter, H., “Random Number Generation and Quasi-Monte Carlo Methods,” Society for Industrial and Applied Mathematics, 1992, 243 pages. [cited by applicant]
Paulin et al., “Cascaded Sobol′ Sampling,” ACM Transactions on Graphics, vol. 1, No. 1, Sep. 2021, pp. 1-13. [cited by applicant]
Rumelhartet al., “Learning representations by back-propagating errors,” Nature, vol. 323, Oct. 9, 1986, pp. 533-536. [cited by applicant]
Sobol, I. M., “On the Distribution of points in a cube and the approximate evaluation of integrals,” USSR Computational Mathematics and Mathematical Physics, pp. 86-112. [cited by applicant]
De Sousa, C., “An overview on weight initialization methods for feedforward neural networks,” IEEE International Joint Conference on Neural Networks, Jul. 2016, 9 pages. [cited by applicant]
Wachter, C., “Quasi-Monte Carlo Light Transport Simulation by Efficient Ray Tracing,” Ph.D. Thesis, Universität Ulm, 2007, 136 pages. [cited by applicant]
Zhou et al., “Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask,” 33rd Conference on Neural Information Processing Systems (NeurIPS), 2019, 18 pages. [cited by applicant]
Zhu et al., “Trained Ternary Quantization,” arXiv, 2016, 9 pages, retrieved from https://arxiv.org/pdf/1612.01064v1.pdf. [cited by applicant]
Owen, A., “Randomly permuted (t,m,s)-Nets and (t,s)-Sequences,” Department of Statistics, Standford University, Sep. 1994, 20 pages. [cited by applicant]
Joe et al., “Remark on Algorithm 659: Implementing Sobol's Quasirandom Sequence Generator,” ACM Transactions on Mathematical Software, vol. 29, No. 1, Mar. 2003, pp. 49-57. [cited by applicant]
Rui et al., “A perfect shuffle type of interpattern association optical neural network model,” Act A Photonica Sinica, vol. 29, No. 1, Jan. 2000, 7 pages. [cited by applicant]
Stone, H., “Parallel processing with the perfect shuffle,” IEEE Transactions on Computers, vol. c-20, No. 2, Feb. 1971, pp. 153-161. [cited by applicant]
Office Action from Chinese Patent Application No. 202210156811.5, dated Apr. 22, 2025, 11 pages. [cited by applicant]
Bellec et al., “Deep Rewiring: Training very sparse deep networks,” ICLR, 2018, pp. 1-24, retrieved from https://arxiv.org/pdf/1711.05136. [cited by applicant]
Araujo et al., “A Neural Network for Shortest Path Computation,” IEEE Transactions on Neural Networks, Feb. 2001, 16 pages. [cited by applicant]