IP Library › Granted Patent US 12,400,117
Granted Patent B2
US 12,400,117 · App. 17/960,694 · Granted Aug 26, 2025

Parameter-efficient method for training neural networks

Inventors: Yuhuang Hu (Zürich, CH); Shih-Chii Liu (Zürich, CH)
Assignee: UNIVERSITÄT ZÜRICH
G06N3/08G06N3/045G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,117
App. No.
17/960,694
Granted
Aug 26, 2025
Kind
B2
Abstract

A computer-implemented method is used to adapt a first artificial neural network for data classification tasks. The first artificial neural network is characterized by a first number of first weight parameters, and includes a set of first network layers. The method includes freezing at least some of the first weight parameters of the first neural network to obtain frozen first weight parameters and duplicating the frozen first weight parameters to obtain duplicated first weight parameters. A second artificial neural network is applied to the duplicated first weight parameters to obtain modulated first weight parameters. The second artificial neural network is characterized by a second number of second weight parameters, the second number being smaller than the first number. The frozen first weight parameters are replaced in the first neural network with the modulated first weight parameters to obtain a modulated first artificial neural network adapted for a data classification task.

Claims (26)

1. A computer-implemented method of adapting a first artificial neural network for one or more data classification tasks, the first artificial neural network being characterised by a first number of first weight parameters, and comprising a set of first network layers, the first artificial neural network being configured to receive input data samples, and output data samples indicative of results of data classification tasks, the method comprising the steps of:

freezing at least some of the first weight parameters of the first artificial neural network to obtain frozen first weight parameters;

duplicating the frozen first weight parameters to obtain duplicated first weight parameters;

applying a second artificial neural network to the duplicated first weight parameters to obtain modulated first weight parameters, the second artificial neural network being characterised by a second number of second weight parameters, the second number being smaller than the first number; and

replacing the frozen first weight parameters in the first artificial neural network with the modulated first weight parameters to obtain a modulated first artificial neural network adapted for a given data classification task,

wherein the method further comprises training the second artificial neural network with a task-specific training data set prior to applying the second artificial neural network to the duplicated first weight parameters.

2. The method according to claim 1 , wherein the method further comprises providing a computing apparatus with the first artificial neural network, and with one or more of the second artificial neural networks, and carrying out the steps of claim 1 on the computing apparatus.

3. The method according to claim 1 , wherein the first artificial neural network is a convolutional neural network, wherein the first network layers are convolutional layers and the frozen first weight parameters are convolution weights, and wherein the first artificial neural network comprises a set of second network layers comprising one or more classification layers.

4. The method according to claim 3 , wherein all the convolution weights of all the convolutional layers of the first artificial neural network are frozen and untrainable.

5. The method according to claim 1 , wherein the second artificial neural network comprises a set of third layers, which are fully connected layers.

6. The method according to claim 1 , wherein the second artificial neural network comprises a set of non-linear activation functions.

7. The method according to claim 1 , wherein the number of the frozen first weight parameters equals the number of the modulated first weight parameters.

8. The method according to claim 1 , wherein the second artificial neural network comprises a multilayer perceptron network.

9. The method according to claim 1 , wherein the second artificial neural network is applied to the duplicated first weight parameters layer-wise such that the second artificial neural network is applied to the duplicated first weight parameters of a respective first network layer before applying the second artificial neural network to the duplicated first weight parameters of a subsequent first network layer.

10. The method according to claim 1 , wherein the frozen first weight parameters are arranged in kernels in the first artificial neural network so that a respective kernel comprises a given number of channels with a given spatial dimension, and wherein the method further comprises flattening the duplicated first weight parameters to obtain flattened first weight parameters, and applying the second artificial neural network to the flattened first weight parameters.

11. The method according to claim 10 , wherein the flattened first weight parameters form two-dimensional sample files with a given number of rows and a given number of columns, wherein the number of columns per sample file of a respective first network layer equals the number of kernels in the respective first network layer multiplied by the number of channels in a respective kernel, while the number of rows per sample file of the respective first network layer equals a kernel height dimension multiplied by a kernel width dimension, or vice versa.

12. The method according to claim 10 , wherein the method further comprises de-flattening the modulated first weight parameters prior to replacing the frozen first weight parameters with the modulated first weight parameters.

13. The method according to claim 1 , wherein the second number is denoted by K, and the first number is denoted by N, and wherein K equals at most 0.5×N.

14. The method according to claim 1 , wherein the method further comprises initialising the second artificial neural network with an initialisation function which is the sum of an identity matrix and a matrix in which the entries are drawn from a zero-mean Gaussian distribution.

15. A computer program product comprising instructions for implementing the steps of the method according to claim 1 when loaded and run on an electronic device.

16. A computing apparatus for adapting a first artificial neural network for one or more data classification tasks, the first artificial neural network being characterised by a first number of first weight parameters, and comprising a set of first network layers, the first artificial neural network being configured to receive input data samples, and output data samples indicative of results of data classification tasks, the computing apparatus being configured to perform operations comprising:

freeze at least some of the first weight parameters of the first artificial neural network to obtain frozen first weight parameters;

duplicate the frozen first weight parameters to obtain duplicated first weight parameters;

apply a second artificial neural network to the duplicated first weight parameters to obtain modulated first weight parameters, the second artificial neural network being characterised by a second number of second weight parameters, the second number being smaller than the first number; and

replace the frozen first weight parameters in the first artificial neural network with the modulated first weight parameters to obtain a modulated first artificial neural network adapted for a given data classification task,

wherein the computing apparatus is further configured to train the second artificial neural network with a task-specific training data set prior to applying the second artificial neural network to the duplicated first weight parameters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2022
From: HU, YUHUANG; LIU, SHIH-CHII
To: UNIVERSITÄT ZÜRICH
Reel/Frame 061325/0963 →
Priority Claims (1)
EP 21201085 · Oct 5, 2021 · regional
Continuity (1)
Related Publication 20230107228A1 · Apr 6, 2023
References Cited (10)
US 20210264271A1 · Gebre · 2021 [cited by examiner]
US 20220164934A1 · Gao · 2022 [cited by examiner]
US 20220215656A1 · Dong · 2022 [cited by examiner]
US 20240103920A1 · Khayiguian · 2024 [cited by examiner]
CN 113837116A · 2021 [cited by examiner]
Kingma, Diederik P. “Adam: A method for stochastic optimization.” arXiv preprint arXiv:1412.6980 (2014). (Year: 2014). [cited by examiner]
Donti, Priya, Brandon Amos, and J. Zico Kolter. “Task-based end-to-end model learning in stochastic optimization.” Advances in neural information processing systems 30 (2017). (Year: 2017). [cited by examiner]
Xudong, et al., “Context-Gated Convolution,” Computer Vision, pp. 701-718 (Dec. 4, 2020). [cited by applicant]
Zhengjue, et al., “MetaSCI: Scalable and Adaptive Reconstruction for Video Compressive Sensing,” CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, pp. 2083-2092 (Jun. 20, 2021). [cited by applicant]
European Search Report dated Apr. 4, 2022 as received in Application No. 21201085.4. [cited by applicant]
Cited By (1)
US 12,682,236