IP Library Granted Patent US 8,438,120
Granted Patent B2
US 8,438,120 · App. 12/597,257 · Granted May 7, 2013

Machine learning hyperparameter estimation

Inventor: Stephan Alexander Raaijmakers (Amsterdam, NL)
Assignee: Nederlandse Organisatie voor toegepast-natuurwetenschappelijk Onderzoek TNO
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,438,120
App. No.
12/597,257
Granted
May 7, 2013
Kind
B2
Abstract

A method of determining hyperparameters (HP) of a classifier ( 1 ) in a machine learning system ( 10 ) iteratively produces an estimate of a target hyperparameter vector. The method comprises the steps of selecting from the random sample the hyperparameter vector producing the best result in the present and any previous iterations, and updating the estimate of the target hyperparameter vector by using said selected hyperparameter vector. The random sample may be restricted by using the hyperparameter vector producing the best result in the present and any previous iterations.

Claims (156)

1. A method of determining hyperparameters of a classifier in a machine learning system by iteratively producing an estimate of a target hyperparameter vector, each iteration comprising the steps of:

drawing a random sample of hyperparameter vectors from a set of possible hyperparameter vectors,

updating the estimate of the target hyperparameter vector by using the random sample, and

selecting, from the random sample of hyperparameter vectors, a hyperparameter vector producing a best result in the present and any previous iterations, and wherein the step of updating the estimate of the target hyperparameter vector uses said hyperparameter vector producing the best result.

2. The method according to claim 1 , further comprising the steps of:

selecting a further hyperparameter vector producing the best result in any previous iterations, and

restricting the random sample of hyperparameter vectors by using the further selected hyperparameter vector.

3. The method according to claim 2 , wherein the step of restricting the random sample of hyperparameter vectors involves using an interval surrounding the further selected hyperparameter vector.

4. The method according to claim 3 , wherein the step of restricting the random sample of hyperparameter vectors is carried out prior to the step of selecting the hyperparameter vector producing the best result in the present and any previous iterations.

5. The method according to claim 3 , wherein the step of restricting the random sample of hyperparameter vectors is carried out during the step of updating the estimate of the target hyperparameter vector.

6. The method according to claim 2 , wherein the step of restricting the random sample of hyperparameter vectors is carried out prior to the step of selecting the hyperparameter vector producing the best result in the present and any previous iterations.

7. The method according to claim 2 , wherein the step of restricting the random sample of hyperparameter vectors is carried out during the step of updating the estimate of the target hyperparameter vector.

8. The method according to claim 1 , wherein the step of updating the estimate of the hyperparameter vector uses a weighting function.

9. The method according to claim 1 , wherein E t is the selected hyperparameter vector X t i at iteration t producing the best result S(X t i ), the selected hyperparameter vector X t i having elements X t ij , wherein the step of updating the estimate of the target hyperparameter vector comprises the step of determining the hyperparameter v t j , where

v

j

t

=

i

=

1

n

I

{

S

(

X

i

t

)

γ

t

}

W

(

X

i

t

;

E

t

)

X

ij

t

i

=

1

n

I

{

S

(

X

i

t

)

γ

t

}

W

(

X

i

t

;

E

t

)

,

wherein γ t is a threshold value and W is a weighting function.

10. The method according to claim 9 , wherein the weighting function W is given by

W

(

X

i

t

;

E

t

)

=

1

-

j

=

1

m

(

X

ij

t

-

E

j

T

)

2

j

=

1

m

(

X

ij

t

)

2

j

=

1

m

(

E

j

t

)

2

.

11. The method according to claim 1 , wherein the steps are carried out by a programmed computer apparatus including a processor and a computer readable medium including computer executable instructions.

12. A classifier for use in a machine learning system, the classifier using hyperparameters as control parameters, wherein the hyperparameters are determined according to the set of steps recited in claim 1 .

13. A non-transitory computer readable medium product including computer-executable instructions for carrying out a method of determining hyperparameters of a classifier in a machine learning system by iteratively producing an estimate of a target hyperparameter vector, each iteration comprising the steps of:

drawing a random sample of hyperparameter vectors from a set of possible hyperparameter vectors,

updating the estimate of the target hyperparameter vector by using the random sample, and

selecting, from the random sample of hyperparameter vectors, a hyperparameter vector producing a best result in the present and any previous iterations, and wherein the step of updating the estimate of the target hyperparameter vector uses said hyperparameter vector producing the best result.

14. A device for determining hyperparameters of a classifier in a machine learning system, the device comprising a processor arranged for iteratively producing an estimate of a target hyperparameter vector, each iteration comprising the steps of:

drawing a random sample of hyperparameter vectors from a set of possible hyperparameter vectors,

updating the estimate of the hyperparameter vector by using the random sample, and

selecting, from the random sample of hyperparameter vectors, a hyperparameter vector producing a best result in the present and any previous iterations, and wherein the step of updating the estimate of the target hyperparameter vector uses said hyperparameter vector producing the best result.

15. The device according to claim 14 , wherein the processor is further arranged for:

selecting a further hyperparameter vector producing the best result in any previous iterations, and

restricting the random sample by using the hyperparameter vector producing the best result in any previous iteration.

16. A machine learning system, comprising the device for determining hyperparameters of a classifier defined in claim 14 .

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2020
From: DATASERVE TECHNOLOGIES LLC
To: K.MIZRA LLC
Reel/Frame 053579/0590 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 5, 2020
From: NEDERLANDSE ORGANISATIE VOOR TOEGEPAST-NATUURWETENSCHAPPELIJK ONDERZOEK (TNO)
To: DATASERVE TECHNOLOGIES LLC
Reel/Frame 052113/0431 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2010
From: RAAIJMAKERS, STEPHAN ALEXANDER
To: NEDERLANDSE ORGANISATIE VOOR TOEGEPAST-NATUURWETENSCHAPPELIJK ONDERZOEK TNO
Reel/Frame 024660/0481 →
Priority Claims (2)
EP 07106963 · Apr 25, 2007 · regional
EP 07112037 · Jul 9, 2007 · regional
Continuity (1)
Related Publication 20100280979A1 · Nov 4, 2010