IP Library Granted Patent US 11,416,765
Granted Patent B2
US 11,416,765 · App. 15/892,699 · Granted Aug 16, 2022

Methods and systems for evaluating training objects by a machine learning algorithm

Inventor: Pavel Aleksandrovich Burangulov (Moscow, RU)
Assignee: YANDEX EUROPE AG
G06N20/00G06F16/951G06F16/953G06F17/17G06N5/003G06N5/04G06N20/20H04L67/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,416,765
App. No.
15/892,699
Granted
Aug 16, 2022
Kind
B2
Abstract

Methods and systems for training a machine learning algorithm (MLA) comprising: acquiring a first set of training samples having a plurality of features, iteratively training a first predictive model based on the plurality of features and generating a respective first prediction error indicator. Analyzing the respective first prediction error indicator for each iteration to determine an overfitting point, and determining at least one evaluation starting point. Acquiring an indication of a new set of training objects, and iteratively retraining the first predictive model with at least one training object from the at least one evaluation starting point to obtain a plurality of retrained first predictive models and generating a respective retrained prediction error indicator. Based on a plurality of retrained prediction error indicators and a plurality of the associated first prediction error indicators, selecting one of the first set of training samples and the at least one training object.

Claims (41)

1. A computer-implemented method for training a machine learning algorithm (MLA), the MLA executable by a server, the method comprising:

acquiring, by the MLA, a first set of training samples, the first set of training samples having a plurality of features;

iteratively training, by the MLA, a first predictive model based on at least a portion of the plurality of features, the training including, for each first training iteration:

generating a respective first prediction error indicator, the respective first prediction error indicator being at least partially indicative of a prediction error associated with the first predictive model at an associated first training iteration;

analyzing, by the MLA, the respective first prediction error indicator for each first training iteration to determine an overfitting point, the overfitting point corresponding to a given first training iteration after which a trend in the first prediction error indicator changes from decreasing to increasing;

determining, by the MLA, at least one evaluation starting point, the at least one evaluation starting point being positioned at a number of iterations before the overfitting point;

acquiring, by the MLA, an indication of a new set of training objects;

iteratively retraining, by the MLA, the first predictive model being in a respective trained state associated with the at least one evaluation starting point with:

at least one training object of the new set of training objects to obtain a plurality of retrained first predictive models;

for each one of the plurality of retrained first predictive models:

generating a respective retrained prediction error indicator for at least one retraining iteration corresponding to at least one first training iteration, the respective retrained prediction error indicator being at least partially indicative of a prediction error associated with the retrained first predictive model;

based on a plurality of retrained prediction error indicators associated with the plurality of retrained first predictive models and a plurality of the associated first prediction error indicators, selecting, by the MLA, one of the first set of training samples and the at least one training object of the new set of training objects.

2. The method of claim 1 , wherein the new set of training objects is one of a new set of features or a new set of training samples.

3. The method of claim 2 , wherein the training and the retraining the first predictive model are executed by applying a gradient boosting technique.

4. The method of claim 3 , wherein selecting, by the MLA, the at least one training object of the new set of training objects comprises comparing the plurality of retrained prediction error indicators with the plurality of the associated first prediction error indicators by applying a statistical hypothesis test.

5. The method of claim 4 wherein the first prediction error indicator and the respective retrained first predictive model prediction error indicator are one of a mean squared error (MSE) or a mean absolute error (MAE).

6. The method of claim 5 , wherein the statistical hypothesis test is a Wilcoxon signed-rank test.

7. The method of claim 6 , wherein the at least one evaluation starting point is a plurality of evaluation starting points.

8. The method of claim 7 , wherein each one of the plurality of evaluation starting points is associated with a respective plurality of retrained first predictive models.

9. A system for training a machine learning algorithm (MLA), the system comprising:

a processor;

a non-transitory computer-readable medium comprising instructions;

the processor, upon executing the instructions, being configured to:

acquire, by the MLA, a first set of training samples, the first set of training samples having a plurality of features;

iteratively train, by the MLA, a first predictive model based on at least a portion of the plurality of features, the training including, for each first training iteration:

generate a respective first prediction error indicator, the respective first prediction error indicator being at least partially indicative of a prediction error associated with the first predictive model at an associated first training iteration;

analyze, by the MLA, the respective first prediction error indicator for each first training iteration to determine an overfitting point, the overfitting point corresponding to a given first training iteration after which a trend in the first prediction error indicator changes from decreasing to increasing;

determine, by the MLA, at least one evaluation starting point, the at least one evaluation starting point being positioned at a number of iterations before the overfitting point;

acquire, by the MLA, an indication of a new set of training objects;

iteratively retrain, by the MLA, the first predictive model being in a respective trained state associated with the at least one evaluation starting point with:

at least one training object of the new set of training objects to obtain a plurality of retrained first predictive models;

for each one of the plurality of retrained first predictive models:

generate a respective retrained prediction error indicator for at least one retraining iteration corresponding to at least one first training iteration, the respective retrained prediction error indicator being at least partially indicative of a prediction error associated with the retrained first predictive model;

based on a plurality of retrained prediction error indicators associated with the plurality of retrained first predictive models and a plurality of the associated first prediction error indicators, select, by the MLA, one of the first set of training samples and the at least one training object of the new set of training objects.

10. The system of claim 9 , wherein the new set of training objects is one of a new set of features or a new set of training samples.

11. The system of claim 10 , wherein to execute the training and the retraining the first predictive model, the processor is configured to apply a gradient boosting technique.

12. The system of claim 11 , wherein to execute the selecting, by the MLA, the at least one training object of the new set of training objects, the processor is configured to compare the plurality of retrained prediction error indicators with the plurality of the associated first prediction error indicators by applying a statistical hypothesis test.

13. The system of claim 12 , wherein the first prediction error indicator and the respective retrained prediction error indicator are one of a mean squared error (MSE) or a mean absolute error (MAE).

14. The system of claim 13 , wherein the statistical hypothesis test is a Wilcoxon signed-rank test.

15. The system of claim 14 , wherein the at least one evaluation starting point is a plurality of evaluation starting points.

16. The system of claim 15 , wherein each one of the plurality of evaluation starting points is associated with a respective plurality of retrained first predictive models.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068525/0349 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2018
From: BURANGULOV, PAVEL ALEKSANDROVICH
To: YANDEX LLC
Reel/Frame 044880/0309 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 9, 2018
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 044880/0375 →
Priority Claims (1)
RU RU2017126674 · Jul 26, 2017 · national
Continuity (1)
Related Publication 20190034830A1 · Jan 31, 2019