IP Library Granted Patent US 8,234,228
Granted Patent B2
US 8,234,228 · App. 12/367,278 · Granted Jul 31, 2012

Method for training a learning machine having a deep multi-layered network with labeled and unlabeled training data

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,234,228
App. No.
12/367,278
Granted
Jul 31, 2012
Kind
B2
Abstract

The invention includes a method for training a learning machine having a deep multi-layered network, with labeled and unlabeled training data. The deep multi-layered network is a network having multiple layers of non-linear mapping. The method generally includes applying unsupervised embedding to any one or more of the layers of the deep network. The unsupervised embedding is operative as a semi-supervised regularizer in the deep network.

Claims (31)

1. A method for training a learning machine comprising a deep network including a plurality of layers, the method comprising the steps of:

applying a regularizer to one or more of the layers of the deep network in a first computer process;

training the regularizer with unlabeled data in a second computer process; and

training the deep network with labeled data in a third computer process;

wherein the regularizer comprises a layer separate from the one or more layers of the network.

2. The method of claim 1 , wherein the regularizer training step includes the steps of:

labeling similar pairs of the unlabeled data as positive pairs; and

labeling dis-similar pairs of the unlabeled data as negative pairs.

3. The method of claim 2 , wherein the regularizer training step further includes the steps of:

calculating a first distance metric which pushes the positive pairs of data together; and

calculating a second distance metric which pulls the negative pairs of data apart.

4. The method of claim 1 , wherein the regularizer training step is performed by stochastic gradient descent.

5. The method of claim 1 , wherein the one or more layers with the applied regularizer comprises a hidden layer of the network.

6. The method of claim 1 , wherein the one or more layers with the applied regularizer comprises an output layer of the network.

7. The method of claim 1 , wherein the one or more layers with the applied regularizer comprises an input layer of the network.

8. An apparatus for use in discriminative classification and regression, the apparatus comprising:

an input device for inputting unlabeled and labeled data associated with a phenomenon of interest;

a processor; and

a memory communicating with the processor, the memory comprising instructions executable by the processor for implementing a learning machine comprising a deep network including a plurality of layers and training the learning machine by:

applying a regularizer to one or more of the layers of the deep network;

training the regularizer with unlabeled data; and

training the deep network with labeled data;

wherein the regularizer applying, the regularizer training and the deep network training are performed in separate computer processes.

9. The apparatus of claim 8 , wherein at least two of the regularizer applying, regularizer training, and the deep network training are performed as a single computer process.

10. The method of claim 8 , wherein the regularizer training step includes the step of assigning weights to pairs of unlabeled data to indicate whether they are similar or dis-similar.

11. The method of claim 10 , wherein weights are in a binary form.

12. A method for training a learning machine comprising a deep network including a plurality of layers, the method comprising the steps of:

applying a regularizer to one or more of the layers of the deep network in a first computer process;

training the regularizer with unlabeled data in a second computer process; and

training the deep network with labeled data in a third computer process;

wherein at least one of the regularizers comprises a layer separate from the one or more layers of the network.

Assignments (1)
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE 8223797 ADD 8233797 PREVIOUSLY RECORDED ON REEL 030156 FRAME 0037. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 30, 2017
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 042587/0845 →