IP Library › Granted Patent US 12,566,955
Granted Patent B2
US 12,566,955 · App. 17/815,025 · Granted Mar 3, 2026

Method for training a neural network

Inventor: Akos Utasi (Göd, HU)
Assignee: CONTINENTAL AUTONOMOUS MOBILITY GERMANY GMBH
G06N3/08G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,955
App. No.
17/815,025
Granted
Mar 3, 2026
Kind
B2
Abstract

The present disclosure relates to a method for training an artificial neural network, the method including providing a neural network to be trained, wherein after training the neural network is to be operated based on a first activation function. An initial training of the neural network is executed based on the at least a second activation function, the at least one second activation function being different to the first activation function. In a transition phase, further training steps are executed using a combination of the first activation function and the at least one second activation function. The combination of training functions is changed over time such that an overweighting of the second activation function at the beginning of the transition phase changes towards an overweighting of first activation function at the end of the transition phase. A final training step is based on the first activation function.

Claims (28)

1 . A method for training an artificial neural network, the method comprising:

providing a neural network to be trained, wherein after training the neural network to be operated based on a first activation function;

executing an initial training of the neural network based on at least one second activation function, the at least one second activation function being different from the first activation function;

executing further training steps in a transition phase in which the neural network is trained using a combination of the first activation function and the at least one second activation function, wherein the combination of training functions is changed over time by a tuning factor, wherein a scheduler adapts the tuning factor such that the first activation function is overweighted increasingly within the transition phase, so that the transition phase includes at least a beginning transition step in which a weight of the at least one second activation function in the combination is greater than a weight of the first activation function and at least an ending transition step in which a weight of the first activation function in the combination is greater than a weight of the at least one second activation function; and

executing a final training step based on the first activation function, the final training step being the last step of training the neural network before the neural network is operated based on the first activation function.

2 . The method according to claim 1 , wherein the combination is a linear combination of the first activation function and the at least one second activation function.

3 . The method according to claim 2 , wherein an overall activation function providing a linear combination of the first activation function and the second activation function is built according to the equation:

ƒ act,overall =α·ƒ act,1 +(1−α)·ƒ act,2

wherein:

a is the tuning factor, wherein aϵ[0,1];

f act,1 is the first activation function; and

f act,2 is the second activation function.

4 . The method according to claim 1 , wherein the tuning factor is changed linearly or nonlinearly in the transition phase.

5 . The method according to claim 1 , wherein the combination is a random selection of the first activation function or the at least one second activation function.

6 . The method according to claim 5 , wherein a probability for randomly selecting the first activation function or the at least one second activation function is changeable by means of the tuning factor.

7 . The method according to claim 6 , wherein the probability for randomly selecting the first activation function or the at least one second activation function is 1 or essentially 1 in selecting the at least one second activation function at a beginning of the transition phase and is lowered towards 0 at an end of the transition phase.

8 . The method according to claim 6 , wherein the tuning factor is changed linearly or nonlinearly in the transition phase.

9 . The method according to claim 5 , wherein random selection is performed based on a random number generator and a scheduler which provides a changeable decision threshold.

10 . The method according to claim 1 , wherein the neural network comprises multiple layers and the combination of the first activation function and the at least one second activation function is applied to each layer of the neural network.

11 . The method according to claim 10 , wherein the combination of activation functions is implemented as a random selection of the first activation function or the at least one second activation function and is performed for each layer independently from the other layers.

12 . The method according to claim 1 , wherein the first activation function is a Rectified Linear Unit (RELU) activation function which is described by the following formula:

y ( x )=max(0, x )

and the at least one second activation function is selected out of the list of the following activation functions: Swish, Mish, gaussian-error linear unit (GELU), exponential linear unit (ELU).

13 . A non-transitory computer readable medium comprising instructions which, when executed by a computer, cause the computer to carry out a method for training an artificial neural network comprising:

receiving information regarding a neural network to be trained, wherein after training the neural network is to be operated based on a first activation function;

executing an initial training of the neural network based on at least one second activation function, the at least one second activation function being different from the first activation function;

executing further training steps in a transition phase in which the neural network is trained using a combination of the first activation function and the at least one second activation function, wherein the combination of training functions is changed over time by a tuning factor, wherein a scheduler adapts the tuning factor such that the first activation function is overweighted increasingly within the transition phase, so that the transition phase includes at least a beginning transition step in which a weight of the at least one second activation function in the combination is greater than a weight of the first activation function and at least an ending transition step in which a weight of the first activation function in the combination is greater than a weight of the at least one second activation function; and

executing a final training based on the first activation function, the final training step being the last step of training the neural network before the neural network is operated based on the first activation function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2022
From: UTASI, AKOS
To: CONTINENTAL AUTONOMOUS MOBILITY GERMANY GMBH
Reel/Frame 061071/0968 →
Priority Claims (1)
EP 21187954 · Jul 27, 2021 · regional
Continuity (1)
Related Publication 20230035615A1 · Feb 2, 2023
References Cited (43)
US 11651230B2 · Li · 2023 [cited by applicant]
US 20180225564A1 · Haiut · 2018 [cited by examiner]
US 20190114531A1 · Torkamani et al. · 2019 [cited by applicant]
US 20190139622A1 · Osthege · 2019 [cited by applicant]
US 20200343985A1 · O'shea et al. · 2020 [cited by applicant]
US 20210142171A1 · Jung et al. · 2021 [cited by applicant]
US 20210232091A1 · Hong · 2021 [cited by examiner]
US 20210248473A1 · Shazeer · 2021 [cited by examiner]
US 20210370993A1 · Qian · 2021 [cited by examiner]
US 20220138562A1 · Biryukova · 2022 [cited by applicant]
US 20230035069A1 · Utasi · 2023 [cited by applicant]
US 20230394304A1 · Zhu et al. · 2023 [cited by applicant]
BR 112019027609B1 · 2025 [cited by examiner]
CN 107122825A · 2017 [cited by applicant]
CN 107516128A · 2017 [cited by applicant]
CN 108388941A · 2018 [cited by applicant]
CN 112906866A · 2021 [cited by applicant]
Oostwal et al. Phase Transitions in Layered Neural Networks: The Role of The Activation Function, Dec. 2020, University of Groningen (Year: 2020). [cited by examiner]
Manessi et al., Learning Combinations of Activation Functions, Apr. 25, 2019, Published as a conference paper at ICPR. (Year: 2019). [cited by examiner]
Jie Renlong et al: “Regularized Flexible Activation Function Combination for Deep Neural Networks”, 2020 25th International Conference on Pattern Recognition (ICPR), IEEE, Jan. 10, 2021 (Jan. 10, 2021), pp. 2001-2008, X… [cited by applicant]
Brosnan Yuen et al: “Universal Activation Function for Machine Learning”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 7, 2020 (Nov. 7, 2020), XP081808784. [cited by applicant]
Zhang Yichi et al, “FracBNN: Accurate and FPGA-Efficient Binary Neural Networks with Fractional Activations”, The 2021 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays, ACMPUB27, New York, NY, USA, Fe… [cited by applicant]
Nandi et al., Improving the Performance of Neural Networks with an Ensemble of Activation Functions, 2020 International Joint Conference on Neural Networks (IJCNN), Jul. 19-24, 2020, pp. 1-7, DOI: 10.1109/CNN48605.2020.… [cited by applicant]
Diganta Misra, “A Self Regularized Non-Monotonic Activation Function”, the 31st British Machine Vision Virtual Conference, Sep. 7-10, 2020, https://www.bmvc2020-conference.com/assets/papers/0928.pdf., arXiv preprint arX… [cited by applicant]
Jang et al, Neural Networks with Activation Networks, arXiv:1811.08618 [cs.CV]. 2018. [cited by applicant]
Shridhar et al., “A Probabilistic Activation Function for Deep Neural Networks”, https://arxiv.org/pdf/1905.10761.pdf, Jun. 16, 2020. [cited by applicant]
Ramachandran et al., “Searching for Activation Functions”, arXiv: 1710.05941v2, Oct. 27, 2017, https://arxiv.org/pdf/1710.05941.pdf. [cited by applicant]
Harmon et al., “Activation Ensembles for Deep Neural Networks”, arXiv: 1702.07790v1, https://arxiv.org/pdf/1702.07790.pdf, Feb. 24, 2017. [cited by applicant]
Eurpoean Search Report for counterpart EPO application EP 21 187 954.9, Feb. 4, 2022. [cited by applicant]
Hendrycks et al., “Gaussian Error Linear Units (GELUs)”, arXiv:1606.08415v4, Jul. 8, 2020. [cited by applicant]
Nair et al., “Rectified Linear Units Improve Restricted Boltzmann Machines,” ICML '10: Proceedings of the 27 International Conference on International Conference on Machine Learning, Jun. 2010, pp. 807-814. [cited by applicant]
Clevert et al., “Fast and Accurate Deep Network Learning by Exponential Linear units (ELUs),” 4th Intl Conf on Learning Representations 2016, May 2016, pp. 1-8, arXiv 1511.07289. [cited by applicant]
Tianhe Yu et al., “Gradient Surgery for Multi-Task Learning,” arXiv:2001.067824v4, Dec. 22, 2020. [cited by applicant]
Chen et al, “Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout,” 34th Conf. on Neural Information Processing Systems, 2020. [cited by applicant]
Garrett Bingham et al: “Discovering Parametric Activation Functions”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jan. 30, 2021 (Jan. 30, 2021), XP081870742. [cited by applicant]
Shuanglong Liu et al: “Optimizing Fully Spectral Convolutional Neural Networks on FPGA”, 2020 International Conference on Field-Programmable Technology (ICFPT), IEEE, Dec. 9, 2020 (Dec. 9, 2020), pp. 39-47, XP033910218. [cited by applicant]
Samba Raju Chiluveru et al: “Efficient Hardware Implementation of DNN-Based Speech Enhancement Algorithm With Precise Sigmoid Activation Function”, IEEE Transactions on Circuits and Systems II: Express Briefs, IEEE, USA… [cited by applicant]
Dugas et al. “Incorporating Second-Order Functional Knowledge for Better Option Pricing”, Proceedings of the 13th International Conference on Neural Information Processing Systems, 2000. [cited by applicant]
Extended European Search Report issued Feb. 1, 2022, by the European Patent Office in European Patent Application No. 21187958.0-1203. (8 pages). [cited by applicant]
Office Action issued by the U.S. Patent and Trademark Office in the U.S. Appl. No. 17/815,075, mailed May 28, 2025, U.S. Patent and Trademark Office, Alexandria, VA. (14 pages). [cited by applicant]
Gulcehre et al., “Mollifying Networks”, arXiv:1608.04980v1 [cs.LG] (Aug. 17, 2016), pp. 1-11. [cited by applicant]
Office Action (The First Office Action) issued Nov. 4, 2025, by the State Intellectual Property Office of People's Republic of China in corresponding Chinese Patent Application No. 202210863417.5 and an English translat… [cited by applicant]
Office Action (The First Office Action) issued Nov. 14, 2025, by the State Intellectual Property Office of People's Republic of China in Chinese Patent Application No. 202210859174.8 and an English translation of the Of… [cited by applicant]