IP Library › Granted Patent US 12,555,024
Granted Patent B2
US 12,555,024 · App. 16/458,924 · Granted Feb 17, 2026

Dynamic data selection for a machine learning model

Inventors: Someshwar Maroti Kale (Karnataka, IN); Vijayalakshmi Krishnamurthy (Sunnyvale, CA); Utkarsh Milind Desai (Bangalore, IN)
Assignee: Oracle International Corporation
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,024
App. No.
16/458,924
Granted
Feb 17, 2026
Kind
B2
Abstract

Embodiments implement a machine learning prediction model with dynamic data selection. A number of data predictions generated by a trained machine learning model can be accessed, where the data predictions include corresponding observed data. An accuracy for the machine learning model can be calculated based on the accessed number of data predictions and the corresponding observed data. The accessing and calculating can be iterated using a variable number of data predictions, where the variable number of data predictions is adjusted based on an action taken during a previous iteration, and, when the calculated accuracy fails to meet an accuracy criteria during a given iteration, a training for the machine learning model can be triggered.

Claims (52)

1 . A method for implementing a machine learning prediction model with dynamic data selection, the method comprising:

accessing a number of data predictions comprising predictions of data points, the data predictions being generated by a trained machine learning model deployed at a deployment time, wherein actual data values that correspond to the data points of the data predictions are stored;

calculating an accuracy for the machine learning model based on the accessed number of data predictions and the corresponding actual data values;

iterating the accessing and calculating using a variable number of data predictions, wherein,

a defined algorithm automatically adjusts the variable number of data predictions that are used to calculate the accuracy of the machine learning model between successive iterations, the variable number of data predictions between successive iterations being adjusted by a delta value that is calculated by applying a factor value;

when the iterating begins the factor value is a numeric value greater than one, and the factor value per iteration is systematically decreased number of iterations increases relative to the deployment time;

for a given iteration of the iterations, the defined algorithm applies a factor value for the given iteration using a first application type to calculate the delta value when training was triggered for an iteration prior to the given iteration, and the defined algorithm applies the factor value for the given iteration using a second application type to calculate the delta value when training was not triggered for the iteration prior to the given iteration; and

when the calculated accuracy during the given iteration fails to meet an accuracy criteria, a training for the machine learning model is triggered;

terminating, in response to the factor value meeting a criteria value, the iterating of the accessing and calculating such that a configured number of data predictions is determined based on the iterations, wherein training of the machine learning model is triggered by multiple of the iterations; and

generating, using the machine learning model after the terminating, additional predictions, wherein the configured number of data predictions is used to calculate an accuracy for the additional predictions and the calculated accuracy of the additional predictions is used to trigger training of the machine learning model.

2 . The method of claim 1 , wherein the iteration prior to the given iteration comprises an iteration directly preceding the given iteration.

3 . The method of claim 1 , wherein the triggered training comprises a retraining or updated training for the trained machine learning model.

4 . The method of claim 1 , wherein, when training is triggered, next iterations of the accessing and calculating use data predictions generated by the machine learning model generated trained via by the triggered training.

5 . The method of claim 1 , wherein the iterating is performed according to a predetermined period, and the predetermined period is a predetermined period of time or a predetermined amount of data predictions that have corresponding actual data values.

6 . The method of claim 1 , wherein the variable number of data predictions for the given iteration is increased when training was triggered during the iteration prior to the given iteration and decreased when training was not triggered during the iteration prior to the given iteration.

7 . The method of claim 1 , wherein,

the variable number of data predictions for the given iteration is based on the variable number of data predictions for the iteration prior to the given iteration;

the variable number of data predictions for the given iteration comprises the variable number of data predictions for the iteration prior to the given iteration multiplied by the factor value for the given iteration when training was triggered during the iteration prior to the given iteration; and

the variable number of data predictions for the given iteration comprises the variable number of data predictions for the iteration prior to the given iteration divided by the factor value for the given iteration when a training was not triggered during the iteration prior to the given iteration.

8 . A system for implementing a machine learning prediction model with dynamic data selection, the system comprising:

a processor; and

a memory storing instructions for execution by the processor, the instructions configuring the processor to:

access a number of data predictions comprising predictions of data points, the data predictions being generated by a trained machine learning model deployed at a deployment time, wherein actual data values that correspond to the data points of the data predictions are stored;

calculate an accuracy for the machine learning model based on the accessed number of data predictions and the corresponding actual data values;

iterate the accessing and calculating using a variable number of data predictions, wherein,

a defined algorithm automatically adjusts the variable number of data predictions that are used to calculate the accuracy of the machine learning model between successive iterations, the variable number of data predictions between successive iterations being adjusted by a delta value that is calculated by applying a factor value;

when the iterating begins the factor value is a numeric value greater than one, and the factor value per iteration is systematically decreased as a number of iterations increases relative to the deployment time;

for a given iteration of the iterations, the defined algorithm applies a factor value for the given iteration using a first application type to calculate the delta value when training was triggered for an iteration prior to the given iteration, and the defined algorithm applies the factor value for the given iteration using a second application type to calculate the delta value when training was not triggered for the iteration prior to the given iteration; and

when the calculated accuracy during the given iteration fails to meet an accuracy criteria, a training for the machine learning model is triggered;

terminate, in response to the factor value meeting a criteria value, the iterating of the accessing and calculating such that a configured number of data predictions is determined based on the iterations, wherein training of the machine learning model is triggered by multiple of the iterations; and

generate, using the machine learning model after the terminating, additional predictions, wherein the configured number of data predictions is used to calculate an accuracy for the additional predictions and the calculated accuracy of the additional predictions is used to trigger training of the machine learning model.

9 . The system of claim 8 , wherein the iteration prior to the given iteration comprises an iteration directly preceding the given iteration.

10 . The system of claim 9 , wherein the iterating is performed according to a predetermined period, and the predetermined period is a predetermined period of time or a predetermined amount of data predictions that have corresponding actual data values.

11 . The system of claim 10 , wherein the variable number of data predictions for the given iteration is increased when training was triggered during the iteration prior to the given iteration and the variable number of data predictions is decreased when training was not triggered during the iteration prior to the given iteration.

12 . A non-transitory computer readable medium having instructions stored thereon that, when executed by a processor, cause the processor to implement a machine learning prediction model with dynamic data selection, wherein, when executed, the instructions cause the processor to:

access a number of data predictions comprising predictions of data points, the data predictions being generated by a trained machine learning model deployed at a deployment time, wherein the data predictions comprise corresponding actual data values of the data points;

calculate an accuracy for the machine learning model based on the accessed number of data predictions and the corresponding actual data values;

iterate the accessing and calculating using a variable number of data predictions, wherein,

a defined algorithm automatically adjusts the variable number of data predictions that are used to calculate the accuracy of the machine learning model between successive iterations, the variable number of data predictions between successive iterations being adjusted by a delta value that is calculated by applying a factor value;

when the iterating begins the factor value is a numeric value greater than one, and the factor value per iteration is systematically decreased as a number of iterations increases relative to the deployment time; and

for a given iteration of the iterations, the defined algorithm applies a factor value for the given iteration using a first application type to calculate the delta value when training was triggered for an iteration prior to the given iteration, and the defined algorithm applies the factor value for the given iteration using a second application type to calculate the delta value when training was not triggered for the iteration prior to the given iteration,

when the calculated accuracy during the given iteration fails to meet an accuracy criteria, a training for the machine learning model is triggered;

terminate, in response to the factor value meeting a criteria value, the iterating of the accessing and calculating such that a configured number of data predictions is determined based on the iterations, wherein training of the machine learning model is triggered by multiple of the iterations; and

generate, using the machine learning model after the terminating, additional predictions, wherein the configured number of data predictions is used to calculate an accuracy for the additional predictions and the calculated accuracy of the additional predictions is used to trigger training of the machine learning model.

13 . The method of claim 1 , wherein applying the factor value for the given iteration using the first application type calculates a delta value that increases the variable number of data points between the iteration prior to the given iteration and the given iteration, and applying the factor value for the given iteration using the second application type calculates a delta value that decreases the variable number of data points between the iteration prior to the given iteration and the given iteration.

14 . The system of claim 8 , wherein applying the factor value for the given iteration using the first application type calculates a delta value that increases the variable number of data points between the iteration prior to the given iteration and the given iteration, and applying the factor value for the given iteration using the second application type calculates a delta value that decreases the variable number of data points between the iteration prior to the given iteration and the given iteration.

15 . The method of claim 1 , wherein the factor value per iteration is systematically decreased as the number of iterations increases relative to the deployment time by applying a step value to the factor value between successive iterations.

16 . The method of claim 15 , wherein applying the step value to the factor value over the iterations decreases the factor value to the criteria value or below the criteria value by a time when the iterating is terminated.

17 . The method of claim 15 , wherein the step value is a numeric value less than one and applying the step value to the factor value comprises multiplying the factor value by the step value.

18 . The method of claim 1 , wherein the machine learning model comprises a demand forecast model, the generated additional predictions comprise demand forecasts, and product shipments are executed in response to the generated additional predictions.

19 . The system of claim 8 , wherein the factor value per iteration is systematically decreased as the number of iterations increases relative to the deployment time by applying a step value to the factor value between successive iterations.

20 . The system of claim 19 , wherein applying the step value to the factor value over the iterations decreases the factor value to the criteria value or below the criteria value by a time when the iterating is terminated.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2020
From: KRISHNAMURTHY, VIJAYALAKSHMI; DESAI, UTKARSH MILIND
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 051543/0284 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2019
From: KALE, SOMESHWAR MAROTI
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 049648/0894 →
Priority Claims (1)
IN 201941003803 · Jan 30, 2019 · national
Continuity (1)
Related Publication 20200242511A1 · Jul 30, 2020
References Cited (40)
US 8849736B2 · Miranda et al. · 2014 [cited by applicant]
US 8868516B2 · Laredo et al. · 2014 [cited by applicant]
US 9703823B2 · Daly · 2017 [cited by examiner]
US 9898467B1 · Pitzel et al. · 2018 [cited by applicant]
US 10339468B1 · Johnston · 2019 [cited by examiner]
US 10657457B1 · Jeffery · 2020 [cited by examiner]
US 20040158562A1 · Caulfield et al. · 2004 [cited by applicant]
US 20100179930A1 · Teller et al. · 2010 [cited by applicant]
US 20150310055A1 · Derstadt et al. · 2015 [cited by applicant]
US 20160062493A1 · Beissinger · 2016 [cited by examiner]
US 20160188574A1 · Homma · 2016 [cited by examiner]
US 20160371601A1 · Grove · 2016 [cited by examiner]
US 20170323216A1 · Fano · 2017 [cited by applicant]
US 20180197087A1 · Luo et al. · 2018 [cited by applicant]
US 20190012575A1 · Xiao et al. · 2019 [cited by applicant]
US 20190130303A1 · Bigaj · 2019 [cited by examiner]
US 20200167669A1 · Lei · 2020 [cited by examiner]
CN 104679868B · 2017 [cited by applicant]
CN 108375808A · 2018 [cited by applicant]
CN 109191030A · 2019 [cited by applicant]
JP 2018535492A · 2018 [cited by applicant]
RU 2672394C1 · 2018 [cited by applicant]
Koychev et al., “Tracking Drifting Concepts by Time Window Optimisation”, 2005, Research and Development in Intelligent Systems XXII SGAI 2005, pp. 46-59 (Year: 2005). [cited by examiner]
Klinkenberg et al., “Detecting Concept Drift with Support Vector Machines”, ICML '00: Proceedings of the Seventeenth International Conference on Machine Learning, 2000, vol. 7 (2000), pp. 487-494 (Year: 2000). [cited by examiner]
Hassani, “Concept Drift Detection of Event Streams Using an Adaptive Window”, Jan. 6, 2019, retrieved from https://pure.tue.nl/ws/files/149896665/0230_dsm_ecms2019_0073.pdf (Year: 2019). [cited by examiner]
Seeliger et al., “Detecting Concept Drift in Processes using Graph Metrics on Process Graphs”, 2017, Proceedings of 9th International Conference on Subject-oriented Business Process Management, vol. 9 (2017), pp. 1-10 (… [cited by examiner]
Bifet et al., “Learning from Time-Changing Data with Adaptive Windowing”, 2007, Proceedings of the 2007 SIAM International Conference on Data Mining, vol. 2007, pp. 443-448 (Year: 2007). [cited by examiner]
Estrada et al., “NSC: A New Progressive Sampling Algorithm”, 2004, workshop on Machine Learning for Scienific data Analysis (IBERAMIA 2004), pp. 1-10 (Year: 2004). [cited by examiner]
Vansteenwegen et al., “An iterated local search algorithm for the single-vehicle cyclic inventory routing problem”, 2014, European Journal of Operational Research, vol. 237 No. 3, pp. 802-813 (Year: 2014). [cited by examiner]
Choudhary et al., “On the Runtime-Efficacy Trade-off of Anomaly Detection Techniques for Real-Time Streaming Data”, 2017, arXiv, v1, pp. 1-14 (Year: 2017). [cited by examiner]
Widmer et al., “Learning in the Presence of Concept Drift and Hidden Contexts”, 1996, Machine Learning, vol. 23, pp. 69-101 (Year: 1996). [cited by examiner]
Provost et al., “Efficient Progressive Sampling”, 1999, Proceedings of the fifth ACM SIGKDD international conference on Knowledge discovery and data mining, vol. 5 (1999), pp. 23-32 (Year: 1999). [cited by examiner]
International Search Report issued in the corresponding International Application No. PCT/US2019/040693 mailed on Oct. 15, 2019. [cited by applicant]
Kumar et al., “Self-Paced Learning for Latent Variable Models”, Advances in Neural Information Processing Systems 23, Jan. 1, 2010 (Jan. 1, 2010), pp. 1189-1197, XP055239017. [cited by applicant]
Losing et al., “Incremental on-line learning: A review and comparison of state of the art algorithms”, Neurocomputing, Elsevier, Amsterdam, NL, vol. 275, Sep. 28, 2017 (Sep. 28, 2017), pp. 1261-1274, XP085310253. [cited by applicant]
Raza et al., Learning with Covariate Shift-Detection and Adaptation in Non-Stationary Environments: Application to Brain-Computer Interface, 2015 International Joint Conference on Neural Networks (IJCNN), IEEE, Jul. 12,… [cited by applicant]
Dai et al., “Improving Data Quality Through Deep Learning and Statistical Models”, Advances in Intelligent Systems and Computing, Jul. 2018. [cited by applicant]
Gosain et al., “Neural Network Approach to Predict Quality of Data Warehouse Multidimensional Model”, Proc. of Int. Conf. on Advances in Computer Science, pp. 241-244, 2010. [cited by applicant]
Gudivada et al., “Data Quality Considerations for Big Data and Machine Learning: Going Beyond Data Cleaning and Transformations”, International Journal on Advances in Software 10.1 (2017), pp. 1-20. [cited by applicant]
Mahankali, “The machine learning approach to data quality”, Data, Analytics & Artificial Intelligence, Wipro, retrieved from https://www.wipro.com/en-IN/analytics/the-machine-learning-approach-to-data-quality/ on May 16… [cited by applicant]