IP Library › Granted Patent US 12,632,733
Granted Patent B2
US 12,632,733 · App. 16/981,018 · Granted May 19, 2026

Methods, systems, articles of manufacture and apparatus to train a neural network

Inventors: Anbang Yao (Beijing, CN); Dawei Sun (Beijing, CN); Aojun Zhou (Beijing, CN); Hao Zhao (Beijing, CN); Yurong Chen (Beijing, CN)
Assignee: Intel Corporation
G06N3/082G06N5/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,733
App. No.
16/981,018
Granted
May 19, 2026
Kind
B2
Abstract

Methods, systems, apparatus, and articles of manufacture are disclosed to train a neural network. An example apparatus includes an architecture evaluator to determine an architecture type of a neural network, a knowledge branch implementor to select a quantity of knowledge branches based on the architecture type, and a knowledge branch inserter to improve a training metric by appending the quantity of knowledge branches to respective layers of the neural network.

Claims (55)

1 . An apparatus to train a neural network, the apparatus comprising:

interface circuitry;

machine-readable instructions; and

at least one processor circuit to be programmed by the machine-readable instructions to:

determine an architecture type of a neural network, the neural network including intermediate layers between a first layer and a last layer;

select a quantity of network classifiers based on the architecture type; and

improve a training metric of the neural network by:

determining a middle one of the intermediate layers; and

attaching a first one of the network classifiers to the middle one of the intermediate layers of the neural network; and

remove the first one of the network classifiers from a backbone of the neural network prior to deployment of the neural network

wherein a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers includes a knowledge interaction matrix generated based on a binary indicator function, the binary indicator function representative of a set of layer indices of the neural network identifying one or more activation locations of the knowledge interaction.

2 . The apparatus as defined in claim 1 , wherein one or more of the at least one processor circuit is to calculate the quantity of network classifiers based on a quantity of layers associated with the neural network.

3 . The apparatus as defined in claim 2 , wherein one or more of the at least one processor circuit is to calculate the quantity of network classifiers based on dividing the quantity of layers associated with the neural network by a branch factor.

4 . The apparatus of claim 2 , wherein a current class probability output from the first one of the network classifiers represents a soft label used to align a probabilistic prediction output from a second one of the network classifiers with the current class probability output.

5 . The apparatus as defined in claim 1 , wherein one or more of the at least one processor circuit is to identify candidate insertion locations of the neural network.

6 . The apparatus as defined in claim 5 , wherein one or more of the at least one processor circuit is to insert one of the quantity of network classifiers at one of the candidate insertion locations.

7 . The apparatus as defined in claim 1 , wherein one or more of the at least one processor circuit is to attach a second one of the network classifiers to a second layer adjacent to the middle one of the intermediate layers.

8 . The apparatus of claim 1 , wherein a soft cross-entropy loss function defines a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers.

9 . At least one non-transitory computer readable medium comprising instructions to cause at least one processor circuit to:

determine an architecture type of a neural network, the neural network including intermediate layers between a first layer and a last layer;

select a quantity of network classifiers based on the architecture type; and

improve a training metric of the neural network by:

determining a middle one of the intermediate layers; and

attaching a first one of the network classifiers to the middle one of the intermediate layers of the neural network; and

remove the first one of the network classifiers from a backbone of the neural network prior to deployment of the neural network

wherein a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers includes a knowledge interaction matrix generated based on a binary indicator function, the binary indicator function representative of a set of layer indices of the neural network identifying one or more activation locations of the knowledge interaction.

10 . The at least one non-transitory computer readable medium as defined in claim 9 , wherein the computer readable instructions are to cause one or more of the at least one processor circuit to calculate the quantity of network classifiers based on a quantity of layers associated with the neural network.

11 . The at least one non-transitory computer readable medium as defined in claim 10 , wherein the computer readable instructions are to cause one or more of the at least one processor circuit to calculate the quantity of network classifiers based on dividing the quantity of layers associated with the neural network by a branch factor.

12 . The at least one non-transitory computer readable medium as defined in claim 9 , wherein the computer readable instructions are to cause one or more of the at least one processor circuit to identify candidate insertion locations of the neural network.

13 . The at least one non-transitory computer readable medium as defined in claim 12 , wherein the computer readable instructions are to cause the one or more of the at least one processor circuit to insert one of the quantity of network classifiers at one of the candidate insertion locations.

14 . The at least one non-transitory computer readable medium as defined in claim 9 , wherein the computer readable instructions are to cause one or more of the at least one processor circuit to attach a second one of the network classifiers to a second layer adjacent to the middle one of the intermediate layers.

15 . The at least one non-transitory computer readable medium as defined in claim 9 , wherein the computer readable instructions are to cause the one or more of the at least one processor circuit to select a knowledge interaction framework for the quantity of network classifiers.

16 . The at least one non-transitory computer readable medium as defined in claim 15 , wherein the computer readable instructions are to cause the one or more of the at least one processor circuit to implement the knowledge interaction framework as at least one of a top-down knowledge interaction framework, a bottom-up knowledge interaction framework, or a bi-directional knowledge interaction framework.

17 . The at least one non-transitory computer readable medium as defined in claim 15 , wherein the computer readable instructions are to cause the one or more of the at least one processor circuit to define an optimization goal, the optimization goal to include the selected knowledge interaction framework.

18 . A computer implemented method to train a neural network, the method comprising:

determining, by executing an instruction with at least one processor, an architecture type of a neural network, the neural network including intermediate layers between a first layer and a last layer;

selecting, by executing an instruction with one or more of the at least one processor, a quantity of network classifiers based on the architecture type; and

improving, by executing an instruction with one or more of the at least one processor, a training metric of the neural network by:

determining a middle one of the intermediate layers; and

attaching a first one of the network classifiers to the middle one of the intermediate layers of the neural network; and

removing, by executing an instruction with one or more of the at least one processor, the first one of the network classifiers from a backbone of the neural network prior to deployment of the neural network,

wherein a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers includes a knowledge interaction matrix generated based on a binary indicator function, the binary indicator function representative of a set of layer indices of the neural network identifying one or more activation locations of the knowledge interaction.

19 . The method as defined in claim 18 , further including selecting a knowledge interaction framework for the quantity of network classifiers.

20 . The method as defined in claim 19 , further including applying at least one of a top-down knowledge interaction framework, a bottom-up knowledge interaction framework, or a bi-directional knowledge interaction framework.

21 . The method as defined in claim 19 , further including defining an optimization goal, the optimization goal to include the selected knowledge interaction framework.

22 . A system to train a neural network, the system comprising:

means for determining an architecture type of a neural network, the neural network including intermediate layers between a first layer and a last layer;

means for selecting a quantity of network classifiers based on the architecture type; and

means for improving a training metric of the neural network by:

determining a middle one of the intermediate layers; and

attaching a first one of the network classifiers to the middle one of the intermediate layers of the neural network; and

means for implementing a knowledge branch to remove the first one of the network classifiers from a backbone of the neural network prior to deployment of the neural network

wherein a knowledge interaction between the first one of the network classifiers and a second one of the network classifiers includes a knowledge interaction matrix generated based on a binary indicator function, the binary indicator function representative of a set of layer indices of the neural network identifying one or more activation locations of the knowledge interaction.

23 . The system as defined in claim 22 , further including means for calculating the quantity of network classifiers based on a quantity of layers associated with the neural network.

24 . The system as defined in claim 23 , wherein the means for calculating the quantity of network classifiers is based on dividing the quantity of layers associated with the neural network by a branch factor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2020
From: YAO, ANBANG; SUN, DAWEI; ZHOU, AOJUN; ZHAO, HAO; CHEN, YURONG
To: INTEL CORPORATION
Reel/Frame 054518/0320 →
Continuity (1)
Related Publication 20210019628A1 · Jan 21, 2021
References Cited (51)
US 5867397A · Koza · 1999 [cited by examiner]
US 6601052B1 · Lee · 2003 [cited by examiner]
US 10853691B1 · Webb · 2020 [cited by examiner]
US 11551026B2 · Kim · 2023 [cited by examiner]
US 20020059154A1 · Rodvold · 2002 [cited by examiner]
US 20050114278A1 · Saptharishi · 2005 [cited by examiner]
US 20110029471A1 · Chakradhar · 2011 [cited by examiner]
US 20140079297A1 · Tadayon · 2014 [cited by examiner]
US 20140201126A1 · Zadeh · 2014 [cited by examiner]
US 20150141123A1 · Callaway · 2015 [cited by examiner]
US 20160117587A1 · Yan · 2016 [cited by examiner]
US 20170140240A1 · Socher · 2017 [cited by examiner]
US 20180203439A1 · Hattori · 2018 [cited by examiner]
US 20180365557A1 · Kobayashi · 2018 [cited by examiner]
US 20190034785A1 · Murray · 2019 [cited by examiner]
US 20190080456A1 · Song · 2019 [cited by examiner]
US 20190114540A1 · Lee · 2019 [cited by examiner]
US 20190180187A1 · Rawal · 2019 [cited by examiner]
US 20190244358A1 · Shi · 2019 [cited by examiner]
US 20190294931A1 · Risser · 2019 [cited by examiner]
US 20190303762A1 · Sui · 2019 [cited by examiner]
US 20190354837A1 · Zhou · 2019 [cited by examiner]
US 20200090045A1 · Baker · 2020 [cited by examiner]
US 20200193332A1 · Zhang · 2020 [cited by examiner]
US 20210019122A1 · Kobayashi · 2021 [cited by examiner]
US 20210192357A1 · Sinha · 2021 [cited by examiner]
CN 104751228 · 2015 [cited by applicant]
CN 104751842 · 2015 [cited by applicant]
CN 106355248 · 2017 [cited by applicant]
CN 106779064 · 2017 [cited by applicant]
Y. Lu, A. Kumar, S. Zhai, Y. Cheng, T. Javidi and R. Feris, “Fully-Adaptive Feature Sharing in Multi-Task Networks with Applications in Person Attribute Classification,” 2017 IEEE Conf. on Computer Vision and Pattern Re… [cited by examiner]
X. Jin, Y. Chen, J. Dong, J. Feng, and S. Yan, “Collaborative Layer-wise Discriminative Learning in Deep Neural Networks,” In: Leibe, B., Matas, J., Sebe, N., Welling, M. (eds) Computer Vision—ECCV 2016. ECCV 2016. Lect… [cited by examiner]
International Searching Authority, “International Search Report,” mailed in connection with International Patent Application No. PCT/CN2018/096599, on Apr. 28, 2019, 4 pages. [cited by applicant]
International Searching Authority, “Written Opinion,” mailed in connection with International Patent Application No. PCT/CN2018/096599, on Apr. 28, 2019, 4 pages. [cited by applicant]
International Searching Authority, “International Preliminary Report on Patentability,” issued in connection with International Patent Application No. PCT/CN2018/096599, mailed on Feb. 4, 2021, 6 pages. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” Advances in Neural Information Processing Systems 25 (NIPS 2012), 2012, 9 pages. [cited by applicant]
Taigman et al., “DeepFace: Closing the Gap to Human-Level Performance in Face Verification,” 2014 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 23-28, 2014, 8 pages. [cited by applicant]
Long et al., “Fully convolutional networks for semantic segmentation,” 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 2015, 10 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, 12 pages. [cited by applicant]
Huang et al., “Densely Connected Convolutional Networks,” 2017 IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, Jul. 21-26, 2017, 9 pages. [cited by applicant]
Zagoruyko et al., “Wide Residual Networks,” arXiv:1605.07146v4 [cs.CV], Jun. 14, 2017, 15 pages. [cited by applicant]
Silver et al., “Mastering the Game of Go with Deep Neural Networks and Tree Search,” Nature 529, 2016, 38 pages. [cited by applicant]
Xie et al., “Aggregated Residual Transformations for Deep Neural Networks,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, 10 pages. [cited by applicant]
Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” arXiv:1704.04861v1 [cs.CV] Apr. 17, 2017, 9 pages. [cited by applicant]
Zoph et al., “Neural Architecture Search With Reinforcement Learning” ICLR, 2017, 16 pages. [cited by applicant]
Lee et al., “Deeply-Supervised Nets,” Artificial Intelligence and Statistics, 2015, 10 pages. [cited by applicant]
Szegedy et al., “2015 IEEE Conference on Computer Going deeper with convolutions,” Vision and Pattern Recognition (CVPR), Boston, MA, USA, 2015 12 pages. [cited by applicant]
Huang et al., “Multi-Scale Dense Networks for Resource Efficient Image Classification,” International Conference on Learning Representations, 2017, 14 pages. [cited by applicant]
Mets, “Microsoft Neural Net Shows Deep Learning Can Get Way Deeper,” Wired, Jan. 14, 2016, 11 pages. [cited by applicant]
Nielsen, “Neural Networks and Deep Learning,” Determination Press, 2015, 63 pages. [cited by applicant]
Kozyrkov, “The simplest explanation of machine learning you'll ever need” Medium, May 24, 2018, 6 pages. [cited by applicant]