IP Library › Granted Patent US 11,205,138
Granted Patent B2
US 11,205,138 · App. 16/419,116 · Granted Dec 21, 2021

Model quality and related models using provenance data

Inventors: Samiulla Zakir Hussain Shaikh (Bangalore, IN); Himanshu Gupta (New Delhi, IN); Rajmohan Chandrahasan (Kanchipuram, IN); Sameep Mehta (New Delhi, IN); Manish Anand Bhide (Hyderabad, IN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,205,138
App. No.
16/419,116
Granted
Dec 21, 2021
Kind
B2
Abstract

A method, computer system, and a computer program product for utilizing provenance data to improve machine learning is provided. Embodiments of the present invention may include collecting provenance data. Embodiments of the present invention may include identifying model quality improvements based on the collected provenance data. Embodiments of the present invention may include identifying related models based on the collected provenance data. Embodiments of the present invention may include recommending model quality improvements to a user.

Claims (48)

1. A method for improving machine learning, the method comprising:

collecting, via a first computer, provenance data comprising at least one member selected from a group consisting of data origins and data movements;

identifying, via the first computer, a first machine learning model, the provenance data relating to a first dataset that has been input to the first machine learning model, the first machine learning model having been trained by a first user;

identifying, via the first computer, a second machine learning model having been trained by a second user, the identifying being based on the collected provenance data and based on a comparison of a user profile of the first user and a user profile of the second user;

recommending, via the first computer and to the first user, a second dataset related to the first dataset, the second dataset being stored in a provenance store; and

training, via the first computer, the identified second machine learning model by using the recommended second dataset.

2. The method of claim 1 , wherein the provenance data is collected from the provenance store.

3. The method of claim 1 , wherein the provenance data further comprises machine learning data that relates to the first machine learning model and that includes at least one member selected from a group consisting of: model metrics, training datasets, feature sets, a usage history, reviews, ratings, feedback, machine learning algorithms, model metadata, and deployment metadata.

4. The method of claim 1 , wherein the provenance data relates to the first dataset and further includes at least one member selected from a group consisting of: a usage history, a forward lineage, a backward lineage, transformations applied for data preparation and data cleaning, column classifications, reviews, ratings, feedback, and a data distribution.

5. The method of claim 1 , further comprising identifying model quality improvements for the first machine learning model based on the collected provenance data, wherein the identifying the model quality improvements includes analyzing at least one member selected from a group consisting of:

datasets,

transformations of data,

machine learning algorithms, and

parameters.

6. The method of claim 1 , wherein the identifying the second machine learning model further includes analyzing at least one member selected from a group consisting of provenance data-based relationships, context-based relationships, usage-based relationships, and data profile-based relationships.

7. The method of claim 1 , wherein the second dataset was used to execute a same task as the first dataset.

8. The method of claim 1 , wherein the identifying the second machine learning model further comprises identifying that the first dataset was also used to train the second machine learning model.

9. The method of claim 1 , wherein the second machine learning model is identified from a machine learning model catalog, and the second machine learning model was listed in the machine learning model catalog before the first machine learning model is identified.

10. A computer system for improving machine learning, the computer system comprising:

one or more processors, one or more computer-readable memories, one or more computer readable storage mediums, and program instructions stored on at least one of the one or more computer readable storage mediums for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories, wherein the computer system is configured to perform a method comprising:

collecting provenance data comprising at least one member selected from a group consisting of data origins and data movements;

identifying a first machine learning model, the provenance data relating to a first dataset that has been input to the first machine learning model, the first machine learning model having been trained by a first user;

identifying a second machine learning model having been trained by a second user, the identifying being based on the collected provenance data and based on a comparison of a user profile of the first user and a user profile of the second user;

recommending, to the first user, a second dataset related to the first dataset, the second dataset being stored in a provenance store; and

training the identified second machine learning model by using the recommended second dataset.

11. The computer system of claim 10 , wherein the provenance data is collected from the provenance store.

12. The computer system of claim 10 , wherein the provenance data further comprises machine learning data that relates to the first machine learning model and that includes at least one member selected from a group consisting of: model metrics, training datasets, feature sets, a usage history, reviews, ratings, feedback, machine learning algorithms, model metadata, and deployment metadata.

13. The computer system of claim 10 , wherein the provenance data relates to the first dataset and further includes at least one member selected from a group consisting of: a usage history, a forward lineage, a backward lineage, transformations applied for data preparation and data cleaning, column classifications, reviews, ratings, feedback, and a data distribution.

14. The computer system of claim 10 , wherein the method further comprises identifying model quality improvements for the first machine learning model based on the collected provenance data, and wherein the identifying the model quality improvements further includes analyzing at least one member selected from a group consisting of:

datasets,

transformations of data,

machine learning algorithms, and

parameters.

15. The computer system of claim 10 , wherein the identifying the second machine learning model further includes analyzing at least one member selected from a group consisting of provenance data-based relationships, context-based relationships, usage- based relationships, and data profile-based relationships.

16. A computer program product for improving machine learning, the computer program product comprising:

one or more computer readable storage mediums and program instructions stored on at least one of the one or more computer readable storage mediums, the program instructions being executable by a processor to cause the processor to perform a method comprising:

collecting provenance data comprising at least one member selected from a group consisting of data origins and data movements;

identifying a first machine learning model, the provenance data relating to a first dataset that has been input to the first machine learning model, the first machine learning model having been trained by a first user;

identifying a second machine learning model having been trained by a second user, the identifying being based on the collected provenance data and based on a comparison of a user profile of the first user and a user profile of the second user;

recommending to the first user a second dataset related to the first dataset, the second dataset being stored in a provenance store; and

training the identified second machine learning model by using the recommended second dataset.

17. The computer program product of claim 16 , wherein the provenance data further comprises machine learning data that relates to the first machine learning model and that includes at least one member selected from a group consisting of: model metrics, training datasets, feature sets, a usage history, reviews, ratings, feedback, machine learning algorithms, model metadata, and deployment metadata.

18. The computer program product of claim 16 , wherein the provenance data relates to the first dataset and further includes at least one member selected from a group consisting of: a usage history, a forward lineage, a backward lineage, transformations applied for data preparation and data cleaning, column classifications, reviews, ratings, feedback, and a data distribution.

19. The computer program product of claim 16 , wherein the method further comprises identifying model quality improvements for the first machine learning model based on the collected provenance data, and wherein the identifying the model quality improvements further includes analyzing at least one member selected from a group consisting of:

datasets,

transformations of data,

machine learning algorithms, and

parameters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2019
From: SHAIKH, SAMIULLA ZAKIR HUSSAIN; GUPTA, HIMANSHU; CHANDRAHASAN, RAJMOHAN; MEHTA, SAMEEP; BHIDE, MANISH ANAND
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 049250/0935 →
Continuity (1)
Related Publication 20200372398A1 · Nov 26, 2020
Cited By (5)
US 12,248,488 US 12,277,129 US 12,619,907 US 12,694,333 US 12,701,051