IP Library Granted Patent US 12,675,505
Granted Patent B1
US 12,675,505 · App. 19/013,257 · Granted Jul 7, 2026

Model exploration for grouped datasets

Inventors: Dhavalkumar C. Patel (White Plains, NY); Jayant R. Kalagnanam (Briarcliff Manor, NY); Robert Jeffrey Baseman (Brewster, NY); Chandrasekhara K. Reddy (Kinnelon, NJ); Fateh A. Tipu (Wappingers Falls, NY); Dung Tien Phan (Pleasantville, NY); Nam H. Nguyen (Pleasantville, NY)
Assignee: International Business Machines Corporation
G06F16/285G06F9/542
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,505
App. No.
19/013,257
Granted
Jul 7, 2026
Kind
B1
Abstract

An embodiment includes creating, by a system, a first dataset and a second dataset from a group of datasets by performing a skyline operation of a feature extracted from the group of datasets. The embodiment includes updating a pipeline metric in a pipeline store by orchestrating an automated modeler based on the first dataset where responsive to the updating an event notification message is transmitted. The embodiment includes detecting that the pipeline metric has been stored in the pipeline store based on the event notification message. The embodiment also includes responsive to the event notification message, creating a search space according to a criterion applied to the pipeline store and orchestrating the automated modeler according to the search space and the second dataset where the system models the first dataset and the second dataset individually and the second dataset is modelled based in part on the pipeline metric.

Claims (32)

1 . A computer-implemented method comprising:

creating, by a system, a first dataset and a second dataset from a group of datasets by performing a skyline operation of a machine learning feature extracted from the group of datasets, the skyline operation creating the first dataset and the second dataset wherein the first dataset and the second dataset each comprises data that are not dominated by another machine learning feature;

updating a pipeline metric, wherein the pipeline metric is a measurement of a process of an automated machine learning modeler, in a pipeline store by orchestrating the automated machine learning modeler based on the first dataset wherein responsive to the updating an event notification message is transmitted;

detecting that the pipeline metric has been stored in the pipeline store based on the event notification message; and

responsive to the event notification message, creating a search space according to a criterion applied to the pipeline store wherein the criterion is based on ranking the pipeline metric together with historical pipeline metrics and orchestrating the automated machine learning modeler according to the search space and the second dataset wherein the system models the first dataset and the second dataset individually and the second dataset is modelled based in part on the pipeline metric.

2 . The computer-implemented method of claim 1 , wherein creating the search space further comprises extracting historical pipeline metrics from the pipeline store and creating the search space from the historical pipeline metrics.

3 . The computer-implemented method of claim 1 , wherein orchestrating the automated machine learning modeler with a scoring model wherein the scoring model is derived from the pipeline metric.

4 . The computer-implemented method of claim 1 , wherein the search space comprises a weighted graph wherein the criterion applied to the pipeline store comprises weighing a node and an edge of the weighted graph according to a machine learning estimator performance and a machine learning transformer performance.

5 . The computer-implemented method of claim 1 , wherein the event notification message comprises a binary word of a shared memory buffer.

6 . The computer-implemented method of claim 1 , further comprising establishing a dimensional space by applying a data centric approach, a model centric approach, a joint approach or a user-controlled approach.

7 . The computer-implemented method of claim 1 , wherein the first dataset is smaller than the second dataset and comprises optimized dimensions of the group of datasets.

8 . A computer program product comprising one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by a processor to cause the processor to perform operations comprising:

creating, by a system, a first dataset and a second dataset from a group of datasets by performing a skyline operation of a machine learning feature extracted from the group of datasets, the skyline operation creating the first dataset and the second dataset wherein the first dataset and the second dataset each comprises data that are not dominated by another machine learning feature;

updating a pipeline metric, wherein the pipeline metric is a measurement of a process of an automated machine learning modeler, in a pipeline store by orchestrating the automated machine learning modeler based on the first dataset wherein responsive to the updating an event notification message is transmitted;

detecting that the pipeline metric has been stored in the pipeline store based on the event notification message; and

responsive to the event notification message, creating a search space according to a criterion applied to the pipeline store wherein the criterion is based on ranking the pipeline metric together with historical pipeline metrics and orchestrating the automated machine learning modeler according to the search space and the second dataset wherein the system models the first dataset and the second dataset individually and the second dataset is modelled based in part on the pipeline metric.

9 . The computer program product of claim 8 , wherein creating the search space further comprises extracting historical pipeline metrics from the pipeline store and creating the search space from the historical pipeline metrics.

10 . The computer program product of claim 8 , wherein orchestrating the automated machine learning modeler with a scoring model wherein the scoring model is derived from the pipeline metric.

11 . The computer program product of claim 8 , wherein the search space comprises a weighted graph wherein the criterion applied to the pipeline store comprises weighing a node and an edge of the weighted graph according to a machine learning estimator performance and a machine learning transformer performance.

12 . The computer program product of claim 8 , wherein the event notification message comprises a binary word of a shared memory buffer.

13 . The computer program product of claim 8 , further comprising establishing a dimensional space by applying a data centric approach, a model centric approach, a joint approach or a user-controlled approach.

14 . The computer program product of claim 8 , wherein the first dataset is smaller than the second dataset and comprises optimized dimensions of the group of datasets.

15 . A computer system comprising a processor and one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by the processor to cause the processor to perform operations comprising:

creating, by a system, a first dataset and a second dataset from a group of datasets by performing a skyline operation of a machine learning feature extracted from the group of datasets, the skyline operation creating the first dataset and the second dataset wherein the first dataset and the second dataset each comprises data that are not dominated by another machine learning feature;

updating a pipeline metric, wherein the pipeline metric is a measurement of a process of an automated machine learning modeler, in a pipeline store by orchestrating the automated machine learning modeler based on the first dataset wherein responsive to the updating an event notification message is transmitted;

detecting that the pipeline metric has been stored in the pipeline store based on the event notification message; and

responsive to the event notification message, creating a search space according to a criterion applied to the pipeline store wherein the criterion is based on ranking the pipeline metric together with historical pipeline metrics and orchestrating the automated machine learning modeler according to the search space and the second dataset wherein the system models the first dataset and the second dataset individually and the second dataset is modelled based in part on the pipeline metric.

16 . The computer system of claim 15 , wherein creating the search space further comprises extracting historical pipeline metrics from the pipeline store and creating the search space from the historical pipeline metrics.

17 . The computer system of claim 15 , wherein orchestrating the automated machine learning modeler with a scoring model wherein the scoring model is derived from the pipeline metric.

18 . The computer system of claim 15 , wherein the search space comprises a weighted graph wherein the criterion applied to the pipeline store comprises weighing a node and an edge of the weighted graph according to a machine learning estimator performance and a machine learning transformer performance.

19 . The computer system of claim 15 , wherein the event notification message comprises a binary word of a shared memory buffer.

20 . The computer system of claim 15 , further comprising establishing a dimensional space by applying a data centric approach, a model centric approach, a joint approach or a user-controlled approach.