IP Library › Granted Patent US 10,438,112
Granted Patent B2
US 10,438,112 · App. 14/836,901 · Granted Oct 8, 2019

Method and apparatus of learning neural network via hierarchical ensemble learning

Inventors: Qiang Zhang (Pasadena, CA); Zhengping Ji (Temple City, CA); Lilong Shi (Pasadena, CA); Ilia Ovsiannikov (Studio City, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06N3/0454G06N3/082G06N3/084G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,438,112
App. No.
14/836,901
Filed
Aug 26, 2015
Granted
Oct 8, 2019
Kind
B2
Art Unit
2122
USPC
706/25
Abstract

A method for configuring a neural network is provided. The method includes: selecting a neural network including a plurality of layers, each of the layers including a plurality of neurons for processing an input and providing an output; and, incorporating at least one switch configured to randomly select and disable at least a portion of the neurons in each layer. Another method in the computer program product is disclosed.

Claims (54)

1. A method to train a neural network, the method comprising:

selecting, using a processor, a first subset of feature detectors from a plurality of feature detectors of a first layer of a neural network, the first subset of feature detectors and a second subset of feature detectors forming the plurality of feature detectors of the first layer;

disabling each feature detector of the first subset of feature detectors;

enabling each feature detector of the second subset of feature detectors;

selecting a third subset of feature detectors from a plurality of feature detectors of a second layer of the neural network, the second layer being closer to an output of the neural network than the first layer, the third subset of feature detectors and a fourth subset of feature detectors forming the plurality of feature detectors of the second layer, the feature detectors of the third subset of feature detectors being connected to each disabled feature detector of the first subset of feature detectors;

disabling each feature detector of the third subset of feature detectors;

enabling each feature detector of the fourth subset of feature detectors;

inputting training data into the neural network with the first subset and the third subset of feature detectors disabled and the second subset and the fourth subset of feature detectors enabled;

backpropagating data output from the neural network in response to inputting the training data into the neural network with the first subset and the third subset of feature detectors disabled and the second subset and the third subset of feature detectors enabled; and

updating at least one feature detector of at least one of the second subset and the fourth subset of feature detectors based on an optimization of a loss function.

2. The method of claim 1 , wherein the third subset of feature detectors further comprises the feature detectors that are connected to each disabled feature detector of the first subset of feature detectors and at least one feature detector selected from a remaining feature detector of the plurality of feature detectors of the second layer.

3. The method of claim 1 , wherein updating the at least one feature detector further comprises updating the at least one feature detector of at least one of the second subset and the fourth subset of feature detectors based on a gradient descent associated with each updated feature detector.

4. The method of claim 3 , wherein updating the at least one feature detector further comprises:

updating at least one feature detector of the second subset of feature detectors based on a gradient descent of the updated feature detector of the second subset of feature detectors; and

updating at least one feature detector of the fourth subset of feature detectors based on a gradient descent of the updated feature detector of the fourth subset of feature detectors.

5. The method of claim 1 , wherein selecting the first subset of feature detector comprises randomly selecting the first subset of feature detectors.

6. The method of claim 1 , wherein the feature detectors of at least one of the first layer and the second layer comprises a plurality of channels.

7. The method of claim 6 , wherein the feature detectors of the first layer comprises a plurality of channels and the feature detectors of the second layer comprises a plurality of channels.

8. A system, comprising:

a processor programmed to initiate executable operations to train a neural network comprising:

selecting a first subset of feature detectors from a plurality of feature detectors of a first layer of the neural network, the first subset of feature detectors and a second subset of feature detectors forming the plurality of feature detectors of the first layer;

disabling each feature detector of the first subset of feature detectors;

enabling each feature detector of the second subset of feature detectors;

selecting a third subset of feature detectors from a plurality of feature detectors of a second layer of the neural network, the second layer being closer to an output of the neural network than the first layer, the third subset of feature detectors and a fourth subset of feature detectors forming the plurality of feature detectors of the second layer, the feature detectors of the third subset of feature detectors being connected to each disabled feature detector of the first subset of feature detectors;

disabling each feature detector of the third subset of feature detectors;

enabling each feature detector of the fourth subset of feature detectors;

inputting training data into the neural network with the first subset and the third subset of feature detectors disabled and the second subset and the fourth subset of feature detectors enabled;

backpropagating data output from the neural network in response to inputting the training data into the neural network with the first subset and the third subset of feature detectors disabled and the second subset and the third subset of feature detectors enabled; and

updating at least one feature detector of at least one of the second subset and the fourth subset of feature detectors based on an optimization of a loss function.

9. The system of claim 8 , wherein the third subset of feature detectors further comprises the feature detectors that are connected to each disabled feature detector of the first subset of feature detectors and at least one feature detector selected from a remaining feature detector of the plurality of feature detectors of the second layer.

10. The system of claim 8 , wherein updating the at least one feature detector further comprises updating the at least one feature detector of at least one of the second subset and the fourth subset of feature detectors based on a gradient descent associated with each updated feature detector.

11. The system of claim 10 , wherein updating the at least one feature detector further comprises:

updating at least one feature detector of the second subset of feature detectors based on a gradient descent of the updated feature detector of the second subset of feature detectors; and

updating at least one feature detector of the fourth subset of feature detectors based on a gradient descent of the updated feature detector of the fourth subset of feature detectors.

12. The system of claim 8 , wherein selecting the first subset of feature detector comprises randomly selecting the first subset of feature detectors.

13. The system of claim 8 , wherein the feature detectors of at least one of the first layer and the second layer comprises a plurality of channels.

14. The system of claim 13 , wherein the feature detectors of the first layer comprises a plurality of channels and the feature detectors of the second layer comprises a plurality of channels.

15. A non-transitory computer-readable medium having stored thereon instructions that, if executed by a processor, result in at least the following:

selecting, a first subset of feature detectors from a plurality of feature detectors of a first layer of a neural network, the first subset of feature detectors and a second subset of feature detectors forming the plurality of feature detectors of the first layer;

disabling each feature detector of the first subset of feature detectors;

enabling each feature detector of the second subset of feature detectors;

selecting a third subset of feature detectors from a plurality of feature detectors of a second layer of the neural network, the second layer being closer to an output of the neural network than the first layer, the third subset of feature detectors and a fourth subset of feature detectors forming the plurality of feature detectors of the second layer, the feature detectors of the third subset of feature detectors being connected to each disabled feature detector of the first subset of feature detectors;

disabling each feature detector of the third subset of feature detectors;

enabling each feature detector of the fourth subset of feature detectors;

inputting training data into the neural network with the first subset and the third subset of feature detectors disabled and the second subset and the fourth subset of feature detectors enabled;

backpropagating data output from the neural network in response to inputting the training data into the neural network with the first subset and the third subset of feature detectors disabled and the second subset and the third subset of feature detectors enabled; and

updating at least one feature detector of at least one of the second subset and the fourth subset of feature detectors based on an optimization of a loss function.

16. The non-transitory computer-readable medium of claim 15 , wherein the third subset of feature detectors further comprises the feature detectors that are connected to each disabled feature detector of the first subset of feature detectors and at least one feature detector selected from a remaining feature detector of the plurality of feature detectors of the second layer.

17. The non-transitory computer-readable medium of claim 15 , wherein updating the at least one feature detector further comprises updating the at least one feature detector of at least one of the second subset and the fourth subset of feature detectors based on a gradient descent associated with each updated feature detector.

18. The non-transitory computer-readable medium of claim 17 , wherein updating the at least one feature detector further comprises:

updating at least one feature detector of the second subset of feature detectors based on a gradient descent of the updated feature detector of the second subset of feature detectors; and

updating at least one feature detector of the fourth subset of feature detectors based on a gradient descent of the updated feature detector of the fourth subset of feature detectors.

19. The non-transitory computer-readable medium of claim 15 , wherein selecting the first subset of feature detector comprises randomly selecting the first subset of feature detectors.

20. The non-transitory computer-readable medium of claim 15 , wherein the feature detectors of at least one of the first layer and the second layer comprises a plurality of channels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2015
From: ZHANG, QIANG; JI, ZHENGPING; SHI, LILONG; OVSIANNIKOV, ILIA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 036431/0465 →
Continuity (2)
Provisional Application 62166627 · May 26, 2015
Related Publication 20160350649A1 · Dec 1, 2016