IP Library › Granted Patent US 12,619,598
Granted Patent B2
US 12,619,598 · App. 17/453,330 · Granted May 5, 2026

Data allocation with user interaction in a machine learning system

Inventors: Bei Chen (Blanchardstown, IE); Massimiliano Mattetti (Dublin, IE); Rahul Nair (Dublin, IE); Elizabeth Daly (Dublin, IE); Oznur Alkan (Dublin, IE)
Assignee: International Business Machines Corporation
G06F16/2379
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,598
App. No.
17/453,330
Granted
May 5, 2026
Kind
B2
Abstract

Various embodiments are provided for providing enhanced data allocation for machine learning operations in a computing environment by one or more processors in a computing system. One or more data sampling strategies may be determined based on a dataset. One or more enhanced training data allocations may be suggested for machine learning operations in a cloud computing environment based on the one or more data sampling strategies.

Claims (53)

1 . A method for improving machine learning operations within a cloud computing environment when training a machine learning model for execution within the cloud computing environment, the method comprising:

determining one or more training data sampling strategies based on a dataset hosted using a cloud computing service in a cloud computing environment;

predicting a cost in resource utilization of each of the one or more training data sampling strategies in terms of a cost of data storage in a cloud object and in terms of a cost of training a machine learning model in the cloud computing environment to consume data provided by each of the one or more data sampling strategies;

predicting a degree of impact to accuracy of the one or more training data sampling strategies;

suggesting as part of a data pre-processing step of an automated machine learning model prediction service one or more enhanced training data allocations for machine learning operations in a cloud computing environment based on the one or more data sampling strategies, the one or more enhanced training data allocations each having one or more portions of data removed from the cloud computing environment to minimize cost of training the machine learning model while adhering to constraints of accuracy and run time; and

training the machine learning model utilizing the suggested enhanced training data allocations from the dataset, the machine learning model trained with the one or more portions of data removed from the dataset prior to training of the machine learning model to minimize the cost of training the machine learning model while adhering to the constraints of accuracy and run time.

2 . The method of claim 1 , further including receiving, as the dataset, a plurality of data types and data features, wherein the plurality of data types includes at least tabular data and timeseries data and the data features include at least a change point, seasonality data, and clustered data.

3 . The method of claim 1 , further including:

applying a forward allocation as a data sampling strategy for tabular data;

applying a backward allocation as a data sampling strategy for time series data;

applying a stratified sampling as a data sampling strategy for clustered data;

applying constraint sampling to include a defined time period as a data sampling strategy for seasonal data; and

using a change point detection as a data sampling strategy for abnormal data.

4 . The method of claim 1 , further including collecting feedback data based on the one or more data sampling strategies.

5 . The method of claim 1 , further including providing one or more t-shirt size options, data storage options, and the one or more data sampling strategies for suggesting one or more enhanced training data allocations.

6 . The method of claim 1 , further including providing a projected learning curve for the machine learning operations and benefit tradeoffs for each of the one or more enhanced training data allocations.

7 . The method of claim 1 , further including predicting a degree of impact on the dataset for each of the one or more enhanced training data allocations based on a training accuracy, training time, the dataset, computing hardware configurations, and the one or more data sampling strategies.

8 . A system for improving machine learning operations within a cloud computing environment when training a machine learning model for execution within the cloud computing environment, the system comprising:

one or more computers with executable instructions that when executed cause the system to:

determine one or more training data sampling strategies based on a dataset hosted using a cloud computing service in a cloud computing environment;

predicting a cost in resource utilization of each of the one or more training data sampling strategies in terms of a cost of data storage in a cloud object and in terms of a cost of training a machine learning model in the cloud computing environment to consume data provided by each of the one or more data sampling strategies;

predicting a degree of impact to accuracy of the one or more training data sampling strategies;

suggest as part of a data pre-processing step of an automated machine learning model prediction service one or more enhanced training data allocations for machine learning operations in a cloud computing environment based on the one or more data sampling strategies, the one or more enhanced training data allocations each having one or more portions of data removed from the cloud computing environment to minimize cost of training the machine learning model while adhering to constraints of accuracy and run time; and

training the machine learning model utilizing the suggested enhanced training data allocations from the dataset, the machine learning model trained with the one or more portions of data removed from the dataset prior to training of the machine learning model to minimize the cost of training the machine learning model while adhering to constraints of accuracy and run time.

9 . The system of claim 8 , wherein the executable instructions when executed cause the system to receive, as the dataset, a plurality of data types and data features, wherein the plurality of data types includes at least tabular data and timeseries data and the data features include at least a change point, seasonality data, and clustered data.

10 . The system of claim 8 , wherein the executable instructions when executed cause the system to:

apply a forward allocation as a data sampling strategy for tabular data;

apply a backward allocation as a data sampling strategy for time series data;

apply a stratified sampling as a data sampling strategy for clustered data;

apply constraint sampling to include a defined time period as a data sampling strategy for seasonal data; and

use a change point detection as a data sampling strategy for abnormal data.

11 . The system of claim 8 , wherein the executable instructions when executed cause the system to collect feedback data based on the one or more data sampling strategies.

12 . The system of claim 8 , wherein the executable instructions when executed cause the system to provide one or more t-shirt size options, data storage options, and the one or more data sampling strategies for suggesting one or more enhanced training data allocations.

13 . The system of claim 8 , wherein the executable instructions when executed cause the system to provide a projected learning curve for the machine learning operations and benefit tradeoffs for each of the one or more enhanced training data allocations.

14 . The system of claim 8 , wherein the executable instructions when executed cause the system to predict a degree of impact on the dataset for each of the one or more enhanced training data allocations based on a training accuracy, training time, the dataset, computing hardware configurations, and the one or more data sampling strategies.

15 . A computer program product for improving machine learning operations within a cloud computing environment when training a machine learning model for execution within the cloud computing environment, the computer program product comprising:

one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions comprising:

program instructions to determine one or more training data sampling strategies based on a dataset hosted using a cloud computing service in a cloud computing environment;

program instructions to predict a cost in resource utilization of each of the one or more training data sampling strategies in terms of a cost of data storage in a cloud object and in terms of a cost of training a machine learning model in the cloud computing environment to consume data provided by each of the one or more data sampling strategies;

program instructions to suggest as part of a data pre-processing step of an automated machine learning model prediction service one or more enhanced training data allocations for machine learning operations in a cloud computing environment based on the one or more data sampling strategies, the one or more enhanced training data allocations each having one or more portions of data removed from the cloud computing environment to minimize cost of training the machine learning model while adhering to constraints of accuracy and run time; and

program instructions to train the machine learning model utilizing the suggested enhanced training data allocations from the dataset, the machine learning model trained with the one or more portions of data removed from the dataset prior to training of the machine learning model to minimize the cost of training the machine learning model while adhering to the constraints of accuracy and run time.

16 . The computer program product of claim 15 , further including program instructions to receive, as the dataset, a plurality of data types and data features, wherein the plurality of data types includes at least tabular data and timeseries data and the data features include at least a change point, seasonality data, and clustered data.

17 . The computer program product of claim 15 , further including program instructions to:

apply a forward allocation as a data sampling strategy for tabular data;

apply a backward allocation as a data sampling strategy for time series data;

apply a stratified sampling as a data sampling strategy for clustered data;

apply constraint sampling to include a defined time period as a data sampling strategy for seasonal data; and

use a change point detection as a data sampling strategy for abnormal data.

18 . The computer program product of claim 15 , further including program instructions to collect feedback data based on the one or more data sampling strategies.

19 . The computer program product of claim 15 , further including program instructions to:

provide one or more t-shirt size options, data storage options, and the one or more data sampling strategies for suggesting one or more enhanced training data allocations; and

provide a projected learning curve for the machine learning operations and benefit tradeoffs for each of the one or more enhanced training data allocations.

20 . The computer program product of claim 15 , further including program instructions to predict a degree of impact on the dataset for each of the one or more enhanced training data allocations based on a training accuracy, training time, the dataset, computing hardware configurations, and the one or more data sampling strategies.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2021
From: CHEN, BEI; MATTETTI, MASSIMILIANO; NAIR, RAHUL; DALY, ELIZABETH; ALKAN, OZNUR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057999/0881 →
Continuity (1)
Related Publication 20230136461A1 · May 4, 2023
References Cited (31)
US 8359223B2 · Chi et al. · 2013 [cited by applicant]
US 10572819B2 · Klinger et al. · 2020 [cited by applicant]
US 10719301B1 · Dasgupta · 2020 [cited by examiner]
US 20120030157A1 · Tsuchida · 2012 [cited by examiner]
US 20160132787A1 · Drevo et al. · 2016 [cited by applicant]
US 20200272825A1 · Zhang · 2020 [cited by examiner]
US 20210117868A1 · Sriharsha · 2021 [cited by examiner]
US 20210256407A1 · Therani · 2021 [cited by examiner]
US 20210271934A1 · White · 2021 [cited by examiner]
US 20210374518A1 · Zhu · 2021 [cited by examiner]
US 20230005023A1 · Zhang · 2023 [cited by examiner]
Meek, Christopher, Bo Thiesson, and David Heckerman. “The learning-curve sampling method applied to model-based clustering.” Journal of Machine Learning Research 2.Feb. 2002: 397-418. (Year: 2002). [cited by examiner]
Guan et al., “Cost-Sensitive Elimination of Mislabeled Training Data,” in 402 Info. Sci. 170-81 (2017). (Year: 2017). [cited by examiner]
Yadwadkar, Neeraja, “Machine Learning for Automatic Resource Management in the Datacenter and the Cloud”, Dissertation, Electrical Engineering and Computer Sciences University of California at Berkeley, Aug. 2018 (pp. 1… [cited by applicant]
Zhang et al., “MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving”, Proceedings of the 2019 USENIX Annual Technical Conference, Jul. 2019 (pp. 15). [cited by applicant]
Meinardi et al., “How to Manage and Optimize Costs of Public Cloud IaaS and PaaS”, Gartner Information Technology Research, Published Mar. 23, 2020, (pp. 50) https://www.gartner.com/en/documents/3982411/how-to-manage-an… [cited by applicant]
Garcia et al., “A cloud-based framework for machine learning workloads and applications”, IEEE access, 8, 18681-18692, vol. 4, 2016, (pp. 12). [cited by applicant]
Ranganathan, S., Haribara, Y., “Cinnamon AI saves 70% on ML model training costs with Amazon SageMaker Managed Spot Training”, Retrieved from https://aws.amazon.com/blogs/machine-learning/cinnamon-ai-saves-70-on-ml-mode… [cited by applicant]
Apply machine type recommendations to VM instances, Retrieved from: https://cloud.google.com/compute/docs/instances/apply-machine-type-recommendations-for-instances, dated Apr. 9, 2025, 10 pages. [cited by applicant]
AWS Compute Optimizer FAQs, Retrieved from web https://aws.amazon.com/compute-optimizer/faqs/, dated Mar. 30, 2025, 10 pages. [cited by applicant]
Chaurasiya, BK., Optimizing costs for machine learning with Amazon SageMaker, Retrieved from: https://aws.amazon.com/tr/blogs/machine-learning/optimizing-costs-for-machine-learning-with-amazon-sagemaker/, Oct. 27, 2020 … [cited by applicant]
Cloud Recommendation Engine, Retrieved from web https://docs.device42.com/reports/cloud-recommendation-engine/, dated Mar. 30, 2025, 9 pages. [cited by applicant]
Han, et al., Efficient Service Recommendation System for Cloud Computing Market, Retrieved from: https://link.springer.com/chapter/10.1007%2F978-3-642-10549-4_14, Nov. 2009, vol. 63, pp. 117-118. [cited by applicant]
Hossain, et al., A Data Compression and Storage Optimization Framework for IoT Sensor Data in Cloud Storage, 21st International Conference of Computer and Information Technology (lCClT), Dec. 21-23, 2018, 6 pages. [cited by applicant]
Kolhar, et al., Storage Allocation Scheme for Virtual Instances of Cloud Computing, Neural Computing and Applications,Jan. 5, 2016, vol. 28, pp. 1397-1404. [cited by applicant]
MLOps: Continuous Delivery and Automation Pipelines in Machine Learning, Retrieved from: https://cloud.google.com/solutions/machine-learning/best-practices-for-ml-performance-cost, Aug. 28, 2024, 19 pages. [cited by applicant]
What is Active Assist, Retrieved from: https://cloud.google.com/recommender/docs/whatis-activeassist, dated Mar. 30, 2025, 5 pages. [cited by applicant]
No Author, “IBM Cloud Object Storage: Pricing”, IBM, Oct. 24, 2020, 9 Pages. [cited by applicant]
No Author, “Recommender”, Google Cloud, Oct. 16, 2021, 7 Pages. [cited by applicant]
No Author, “VM Instance Sizing Recommender”, Google Cloud, Dec. 5, 2020, 6 Pages. [cited by applicant]
Yashchin Emmanuel. “Discussion: A Review of Some Sampling and Aggregation Strategies for Basic Statistical Process Monitoring (I. M. Zwetsloot and W. H. Woodall)”, Journal of Quality Technology, Jul. 12, 2019, 7 Pages. [cited by applicant]