IP Library Granted Patent US 10,839,315
Granted Patent B2
US 10,839,315 · App. 15/607,718 · Granted Nov 17, 2020

Method and system of selecting training features for a machine learning algorithm

Inventors: Anastasiya Aleksandrovna Bezzubtseva (Lipetsk, RU); Alexandr Leonidovich Shishkin (Kirovo-Chepetsk, RU); Gleb Gennadievich Gusev (Moscow, RU); Aleksey Valyerevich Drutsa (Moscow, RU)
Assignee: YANDEX EUROPE AG
G06N20/00G06F16/93G06F16/24578G06F16/334G06F16/90335
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,839,315
App. No.
15/607,718
Granted
Nov 17, 2020
Kind
B2
Abstract

Methods and systems for selecting a selected-sub-set of features from a plurality of features for training a machine learning module, the training of the machine learning module to enable classification of an electronic document to a target label, the plurality of features associated with the electronic document. In one embodiment, the method comprises analyzing a given training document to extract the plurality of features, and for a given not-yet-selected feature of the plurality of features: generating a set of relevance parameters iteratively, generating a set of redundancy parameters iteratively and determining a feature significance score based on the set of relevance parameters and the set of redundancy parameters. The method further comprises selecting a feature associated with a highest value of the feature significance score and adding the selected feature to the selected-sub-set of features.

Claims (104)

1. A computer-implemented method for selecting a selected-sub-set of features from a plurality of features for training a machine learning module, the machine learning module executable by an electronic device,

the training of machine learning module to enable classification of an electronic document to a target label,

the plurality of features associated with the electronic document,

the method executable by the electronic device,

the method comprising:

analyzing, by the electronic device, a given training document to extract the plurality of features associated therewith, the given training document having a pre-assigned target label;

generating a set of relevance parameters by iteratively executing, for a given not-yet-selected feature of the plurality of features:

determining, by the electronic device, a respective relevance parameter of the given not-yet-selected feature to the pre-assigned target label, the relevance parameter indicative of a level of synergy of the given not-yet-selected feature, together with the set of relevance parameters including one or more already-selected features of the plurality of features, to determination of the pre-assigned target label, the respective relevance parameter is determined using:

h j :=argmax I ( c: b|h 1 , . . . ,h j−1 ,h )

wherein I is a mutual information;

wherein c is the pre-assigned target label; and

wherein b is the given not-yet-selected feature;

adding, by the electronic device, the respective relevance parameter to the set of relevance parameters;

generating a set of redundancy parameters by iteratively executing, for the given not-yet-selected feature of the plurality of features:

determining, by the electronic device, a respective redundancy parameter of the given not-yet-selected feature to the pre-assigned target label, the redundancy parameter indicative of a level of redundancy of the given not-yet-selected feature, together with a sub-set of relevance parameters and the set of redundancy parameters, including one or more already-selected features of the plurality of features to determination of the pre-assigned target label;

adding, by the electronic device, the respective redundancy parameter to the set of redundancy parameters;

analyzing, by the electronic device, the given not-yet-selected feature to determine a feature significance score based on the set of relevance parameters and the set of redundancy parameters;

selecting, by the electronic device, a given selected feature, the given selected feature associated with a highest value of the feature significance score and adding the given selected feature to the selected-sub-set of features; and

storing, by the machine learning module, the selected-sub-set of features.

2. The method of claim 1 , wherein the method further comprises, after the analyzing, by the electronic device, the given training document to extract the plurality of features associated therewith, binarizing the plurality of features and using a set of binarized features as the plurality of features.

3. The method of claim 2 , wherein the selected-sub-set of features comprises a pre-determined number of k selected features and wherein the generating the set of relevance parameters iteratively, the generating the set of redundancy parameters iteratively, the analyzing the given not-yet-selected features and selecting the given selected feature are repeated for a total of k times.

4. The method of claim 3 , wherein the method further comprises, prior to generating the set of relevance parameters, determining a parameter t specifying a number of features taken into account in the set of relevance parameters, and wherein the determining the respective relevance parameter is iteratively executed for t−1 steps, and wherein the determining the respective redundancy parameters is iteratively executed for t steps.

5. The method of claim 4 , wherein the parameter t is at least 3.

6. The method of claim 1 , wherein the respective redundancy parameter is determined using:

g j :=argmin I ( c;b,h 1 , . . . ,h j−1 |g 1 , . . . ,g j−1 ,g )

wherein I is the mutual information;

wherein c is the pre-assigned target label;

wherein b is the given not-yet-selected feature; and

wherein h 1 , . . . h j−1 is the sub-set of relevance parameters.

7. The method of claim 6 , wherein the analyzing, by the electronic device, the given not-yet-selected feature to determine the feature significance score based on the sub-set of relevance parameters and the sub-set of redundancy parameters is determined using:

J i [ f ]: =max b∈B[f] I ( c;b;h 1 , . . . ,h t−1 |g 1 , . . . ,g t )

wherein J i is the feature significance score;

wherein b is the given unselected binarized feature;

wherein B[f] is the set of binarized features;

wherein I is the mutual information;

wherein c is the pre-assigned target label;

wherein h 1 , . . . , h t−1 is the set of relevance parameters; and

wherein g 1 , . . . , g t is the set of redundancy parameters.

8. The method of claim 7 , wherein the given selected feature associated with a highest value of the feature significance score is determined using:

f best :=argmax f∈F\S J i [ f ]

wherein f is a given selected feature of the plurality of features; and

wherein F\S is a set of not-yet-selected features.

9. The method of claim 8 , wherein the method further comprises, prior to generating the second set of relevance parameters:

analyzing, by the electronic device, each feature of the plurality of features, to determine an individual relevance parameter of a given feature of the plurality of features to the pre-assigned target label, the individual relevance parameter indicative of a degree of relevance of the given feature to the determination of the pre-assigned target label; and

selecting, by the electronic device, from the plurality of features a first selected feature, the first selected feature associated with a highest value of the individual relevance parameter and adding the first selected feature to the selected-sub-set of features.

10. The method of claim 9 , wherein the individual relevance parameter is determined using:

f

best

:=

argmax

f

F

max

b

B

[

f

]

I

(

c

;

b

)

wherein f is the given feature of the plurality of features;

wherein F is the plurality of features;

wherein I is the mutual information;

wherein c is the pre-assigned target label;

wherein b is the given feature; and

wherein B[f] is the set of binarized features.

11. The method of claim 8 , wherein the generating the respective relevance parameter is further based on the plurality of features and wherein the generating the respective redundancy parameter is based on the selected-sub-set of features.

12. The method of claim 11 , wherein adding the given selected feature to the selected-sub-set of features comprises adding the set of relevance parameters to the selected-sub-set of features.

13. A server for selecting a selected-sub-set of features from a plurality of features for training a machine learning module, the training of the machine learning module to enable classification of an electronic document to a target label, the plurality of features associated with the electronic document, the server comprising:

a memory storage;

a processor coupled to the memory storage, the processor configured to:

analyze a given training document to extract the plurality of features associated therewith, the given training document having a pre-assigned target label;

generate a set of relevance parameters by iteratively executing, for a given not-yet-selected feature of the plurality of features:

determine a respective relevance parameter of the given not-yet-selected feature to the pre-assigned target label, the relevance parameter indicative of a level of synergy of the given not-yet-selected feature, together with the set of relevance parameters including one or more already-selected features of the plurality of features, to determination of the pre-assigned target label, the respective relevance parameter is determined using:

h h :=argmax I ( c: b|h 1 , . . . ,h j−1 ,h )

wherein I is a mutual information;

wherein c is the pre-assigned target label; and

wherein b is the given not-yet-selected feature add the respective relevance parameter to the set of relevance parameters;

generate a set of redundancy parameters by iteratively executing, for the given not-yet-selected feature of the plurality of features:

determine a respective redundancy parameter of the given not-yet-selected feature to the pre-assigned target label, the redundancy parameter indicative of a level of redundancy of the given not-yet-selected feature, together with a sub-set of relevance parameters and the set of redundancy parameters, including one or more already-selected features of the plurality of features to determination of the pre-assigned target label;

add the respective redundancy parameter to the set of redundancy parameters;

analyze the given not-yet-selected feature determine a feature significance score based on the set of relevance parameters and the set of redundancy parameters;

select a given selected feature, the given selected feature associated with a highest value of the feature significance score and adding the given selected feature to the selected-sub-set of features; and

store, at the memory storage, the selected-sub-set of features.

14. The server of claim 13 , wherein the processor is further configured to, after analyzing the given training document to extract a plurality of features associated therewith, the given training document having a pre-assigned target label, binarize the plurality of features and use a set of binarized features as the plurality of features.

15. The server of claim 14 , wherein the selected-sub-set of features comprises a pre-determined number of k selected features and wherein the generating the set of relevance parameters iteratively, the generating the set of redundancy parameters iteratively, the analyzing the given not-yet-selected features and selecting the the given selected feature are repeated for a total of k times.

16. The server of claim 15 , wherein the processor is further configured to, prior to generating the set of relevance parameters, determine a parameter t specifying a number of features taken into account in the set of relevance parameters, and wherein the determining the respective relevance parameter is iteratively executed for t−1 steps, and wherein the determining the respective redundancy parameters is iteratively executed for t steps.

17. The server of claim 16 , wherein the parameter t is superior to 3.

18. The server of claim 13 , wherein the respective redundancy parameter is determined using:

g j :=argmin I ( c;b,h 1 , . . . ,h j−1 |g 1 , . . . ,g j−1 ,g ),

wherein I is the mutual information;

wherein c is the pre-assigned target label;

wherein b is the binarized given not-yet-selected feature; and

wherein h 1 , . . . , h j−1 is the sub-set of relevance parameters.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068525/0349 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2017
From: BEZZUBTSEVA, ANASTASIYA ALEKSANDROVNA; SHISHKIN, ALEXANDR LEONIDOVICH; GUSEV, GLEB GENNADIEVICH; DRUTSA, ALEKSEY VALYEREVICH
To: YANDEX LLC
Reel/Frame 042524/0082 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2017
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 042524/0183 →
Priority Claims (1)
RU 2016132425 · Aug 5, 2016 · national
Continuity (1)
Related Publication 20180039911A1 · Feb 8, 2018
Cited By (2)
US 12,361,098 US 12,412,094