IP Library › Granted Patent US 11,475,374
Granted Patent B2
US 11,475,374 · App. 16/893,073 · Granted Oct 18, 2022

Techniques for automated self-adjusting corporation-wide feature discovery and integration

Inventors: Alberto Polleri (London, GB); Larissa Cristina Dos Santos Romualdo Suzuki (Wokingham, GB); Sergio Aldea Lopez (London, GB); Marc Michiel Bron (London, GB); Dan David Golding (London, GB); Alexander Ioannides (London, GB); Maria del Rosario Mestre (London, GB); Hugo Alexandre Pereira Monteiro (London, GB); Oleg Gennadievich Shevelev (London, GB); Xiaoxue Zhao (London, GB); Matthew Charles Rowe (Milton Keynes, GB)
Assignee: ORACLE INTERNATIONAL CORPORATION
G06N20/20G06F8/75G06F8/77G06F11/3409G06F11/3466G06F16/211G06F16/2365G06F16/24573G06F16/24578G06F16/285G06F16/367G06F16/907G06F16/9024G06F16/9035G06K9/6231G06K9/6232G06K9/6259G06K9/6298G06N5/003G06N5/025G06N20/00H04L9/088H04L9/3236
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,475,374
App. No.
16/893,073
Filed
Jun 4, 2020
Granted
Oct 18, 2022
Kind
B2
Art Unit
2167
USPC
707/737
Abstract

The present disclosure relates to systems and methods for a self-adjusting corporation-wide discovery and integration feature of a machine learning system that can review a client's data store, review the labels for the various data schema, and effectively map the client's data schema to classifications used by the machine learning model. The various techniques can automatically select the features that are predictive for each individual use case (i.e., one client), effectively making a machine learning solution client-agnostic for the application developer. A weighted list of common representations of each feature for a particular machine learning solution can be generated and stored. When new data is added to the data store, a matching service can automatically detect which features should be fed into the machine-learning solution based at least in part on the weighted list. The weighted list can be updated as new data is made available to the model.

Claims (110)

1. A method for automatically reconciling a schema for labeling a data set for a machine learning application for use in a production environment, the method comprising:

receiving a first input identifying one or more locations of the data set;

receiving an input of a problem to generate a solution using the machine learning application;

analyzing the data to extract one or more labels for the schema of the data set, wherein the one or more labels describe a type of data that is contained in that portion of the data set;

generating a first list of common categories for each of the one or more labels for the schema by:

accessing a library of terms stored in a memory, wherein the terms correspond to categories known by a machine learning model; and

correlating the one or more labels with the categories based at least in part by identifying a category for each of the one or more labels;

generating a mapping of the one or more labels with the categories of the machine learning model, wherein the mapping identifies a location in the data set for each of the categories of the machine learning model;

analyzing the data to extract one or more features described by the data set, wherein the features are incorporated into the machine learning model and correspond to a subset of the one or more labels for the schema;

assigning weights to the one or more features based at least in part on an influence of the one or more features to the solution using the machine learning application; and

storing the mapping and the weights in the memory.

2. The method of claim 1 , wherein the assigning weights to the one or more features comprises:

generating a second list, wherein the second list identifies the one or more features of the data set;

determining a ranking of each the one or more features in the second list based at least in part on the influence of the one or more features to the solution using the machine learning application; and

assigning the weights to the one or more features in the second list based at least in part on the ranking of the features to the solution of the machine learning application.

3. The method of claim 2 , wherein the determining the ranking comprises:

determining a machine learning algorithm from a plurality of algorithms stored in a library wherein the algorithm incorporates the one or more features to calculate a result;

modifying the machine learning algorithm by removing a first feature of the one or more features;

calculating a first result of the modified machine learning algorithm;

comparing the first result of the modified machine learning algorithm with ground truth data; and

calculating a ranking for the first feature based at least in part on the comparing the first result with the ground truth data, wherein the first feature is ranked higher in importance for a decreased difference between the first result and the ground truth data as compared with one or more other results.

4. The method of claim 2 , further comprising:

identifying a new location of additional data;

analyzing the additional data to identify one or more new features, wherein the new features are not identified on the second list of the one or more features in the memory;

generating a revised list of the one or more features in the memory that includes the one or more new features;

determining a revised ranking of each the one or more features and the one or more new features in the revised list based at least in part on an influence of the one or more new features to the solution using the machine learning application; and

assigning weights to each of the ranked features in the revised list based at least in part on the revised ranking of the new feature for the solution generated by the machine learning application.

5. The method of claim 1 , further comprising:

presenting the mapping of the one or more labels with the categories of the machine learning model; and

receiving a second input, wherein the second input correlates a label of the one or more labels for the schema of the data with a category of the one or more categories known by the machine learning model.

6. The method of claim 1 , further comprising:

extracting the data set stored at the one or more locations;

storing the extracted data set in the memory; and

renaming the one or more labels of the data set to match the mapping of the one or more labels.

7. The method of claim 1 , further comprising:

identifying a new label of the one or more labels, wherein the new label does not correlate to the categories of the machine learning data; and

adding the new label and associated metadata to the library of terms stored in the memory.

8. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause a data processing apparatus to perform operations for reconciling a schema for labeling a data set for a machine learning application for use in a production environment, the operations comprising:

receiving a first input identifying one or more locations of the data set;

receiving an input of a problem to generate a solution using the machine learning application;

analyzing the data to extract one or more labels for the schema of the data set, wherein the one or more labels describe a type of data that is contained in that portion of the data set;

generating a first list of common categories for each of the one or more labels for the schema by:

accessing a library of terms stored in a memory, wherein the terms correspond to categories known by a machine learning model; and

correlating the one or more labels with the categories based at least in part by identifying a category for each of the one or more labels;

generating a mapping of the one or more labels with the categories of the machine learning model, wherein the mapping identifies a location in the data set for each of the categories of the machine learning model;

analyzing the data to extract one or more features described by the data set, wherein the features are incorporated into the machine learning model and correspond to a subset of the one or more labels for the schema;

assigning weights to the one or more features based at least in part on an influence of the one or more features to the solution using the machine learning application; and

storing the mapping and the weights in the memory.

9. The computer-program product of claim 8 , wherein the operation of assigning weights to the one or more features comprises:

generating a second list, wherein the second list identifies the one or more features of the data set;

determining a ranking of each the one or more features in the second list based at least in part on the influence of the one or more features to the solution using the machine learning application; and

assigning the weights to the one or more features in the second list based at least in part on the ranking of the features to the solution of the machine learning application.

10. The computer-program product of claim 9 , including instructions configured to cause a data processing apparatus to perform further operations comprising:

determining a machine learning algorithm from a plurality of algorithms stored in a library wherein the algorithm incorporates the one or more features to calculate a result;

modifying the machine learning algorithm by removing a first feature of the one or more features;

calculating a first result of the modified machine learning algorithm;

comparing the first result of the modified machine learning algorithm with ground truth data; and

calculating a ranking for the first feature based at least in part on the comparing the first result with the ground truth data, wherein the first feature is ranked higher in importance for a decreased difference between the first result and the ground truth data as compared with one or more other results.

11. The computer-program product of claim 9 , including instructions configured to cause a data processing apparatus to perform further operations comprising:

identifying a new location of additional data;

analyzing the additional data to identify one or more new features, wherein the new features are not identified on the first second list of the one or more features in the memory;

generating a revised list of the one or more features in the memory that includes the one or more new features;

determining a revised ranking of each the one or more features and the one or more new features in the revised list based at least in part on an influence of the one or more new features to the solution; and

assigning weights to each of the ranked features in the revised list based at least in part on the revised ranking of the new feature for the solution generated by the machine learning application.

12. The computer-program product of claim 8 , including instructions configured to cause a data processing apparatus to perform further operations comprising:

presenting the mapping of the one or more labels with the categories of the machine learning model; and

receiving a second input, wherein the second input correlates a label of the one or more labels for the schema of the data with a category of the one or more categories known by the machine learning model.

13. The computer-program product of claim 8 , including instructions configured to cause a data processing apparatus to perform further operations comprising:

extracting the data set stored at the one or more locations;

storing the extracted data set in the memory; and

renaming the one or more labels of the data set to match the mapping of the one or more labels.

14. The computer-program product of claim 8 , including instructions configured to cause a data processing apparatus to perform further operations comprising:

identifying a new label of the one or more labels, wherein the new label does not correlate to the categories of the machine learning data; and

adding the new label and associated metadata to the library of terms stored in the memory.

15. A system for automatically reconciling a schema for labeling a data set for a machine learning application for use in a production environment, comprising:

one or more data processors; and

a non-transitory computer-readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform operations comprising:

receiving a first input identifying one or more locations of the data set;

receiving an input of a problem to generate a solution using the machine learning application;

analyzing the data to extract one or more labels for the schema of the data set, wherein the one or more labels describe a type of data that is contained in that portion of the data set;

generating a first list of common categories for each of the one or more labels for the schema by:

accessing a library of terms stored in a memory, wherein the terms correspond to categories known by a machine learning model; and

correlating the one or more labels with the categories based at least in part by identifying a category for each of the one or more labels;

generating a mapping of the one or more labels with the categories of the machine learning model, wherein the mapping identifies a location in the data set for each of the categories of the machine learning model;

analyzing the data to extract one or more features described by the data set, wherein the features are incorporated into the machine learning model and correspond to a subset of the one or more labels for the schema;

assigning weights to the one or more features based at least in part on an influence of the one or more features to the solution using the machine learning application; and

storing the mapping and the weights in the memory.

16. The system of claim 15 , wherein the operation of assigning weights to the one or more features comprises:

generating a second list, wherein the second list identifies the one or more features of the data set;

determining a ranking of each the one or more features in the second list based at least in part on the influence of the one or more features to the solution; and

assigning the weights to the one or more features in the second list based at least in part on the ranking of the features to the solution of the machine learning application.

17. The system of claim 16 , wherein the non-transitory computer-readable storage medium includes further instructions which, when executed on the one or more data processors, cause the one or more data processors to perform further operations comprising:

determining a machine learning algorithm from a plurality of algorithms stored in a library wherein the algorithm incorporates the one or more features to calculate a result;

modifying the machine learning algorithm by removing a first feature of the one or more features;

calculating a first result of the modified machine learning algorithm;

comparing the first result of the modified machine learning algorithm with ground truth data; and

calculating a ranking for the first feature based at least in part on the comparing the first result with the ground truth data, wherein the first feature is ranked higher in importance for a decreased difference between the first result and the ground truth data as compared with one or more other results.

18. The system of claim 16 , wherein the non-transitory computer-readable storage medium includes further instructions which, when executed on the one or more data processors, cause the one or more data processors to perform further operations comprising:

identifying a new location of additional data;

analyzing the additional data to identify one or more new features, wherein the new features are not identified on the second list of the one or more features in the memory;

generating a revised list of the one or more features in the memory that includes the one or more new features;

determining a revised ranking of each the one or more features and the one or more new features in the revised list based at least in part on an influence of the one or more new features to the solution; and

assigning weights to each of the ranked features in the revised list based at least in part on the revised ranking of the new feature for the solution generated by the machine learning application.

19. The system of claim 15 , wherein the non-transitory computer-readable storage medium includes further instructions which, when executed on the one or more data processors, cause the one or more data processors to perform further operations comprising:

presenting the mapping of the one or more labels with the categories of the machine learning model; and

receiving a second input, wherein the second input correlates a label of the one or more labels for the schema of the data with a category of the one or more categories known by the machine learning model.

20. The system of claim 15 , wherein the non-transitory computer-readable storage medium includes further instructions which, when executed on the one or more data processors, cause the one or more data processors to perform further operations comprising:

extracting the data set stored at the one or more locations;

storing the extracted data set in the memory; and

renaming the one or more labels of the data set to match the mapping of the one or more labels.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE SURNAMES OF THE SECOND AND EIGHTH NAMED INVENTORS PREVIOUSLY RECORDED ON REEL 052874 FRAME 0908. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNOR'S INTEREST. Recorded Jun 16, 2020
From: POLLERI, ALBERTO; SUZUKI, LARISSA CRISTINA DOS SANTOS ROMUALDO; LOPEZ, SERGIO ALDEA; BRON, MARC MICHIEL; GOLDING, DAN DAVID; IOANNIDES, ALEXANDER; MESTRE, MARIA DEL ROSARIO; MONTEIRO, HUGO ALEXANDRE PEREIRA; SHEVELEV, OLEG GENNADIEVICH; ZHAO, XIAOXUE; ROWE, MATTHEW CHARLES
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 052947/0117 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2020
From: POLLERI, ALBERTO; ROMUALDO SUZUKI, LARISSA CRISTINA DOS SANTOS; LOPEZ, SERGIO ALDEA; BRON, MARC MICHIEL; GOLDING, DAN DAVID; IOANNIDES, ALEXANDER; MESTRE, MARIA DEL ROSARIO; PEREIRA MONTEIRO, HUGO ALEXANDRE; SHEVELEV, OLEG GENNADIEVICH; ZHAO, XIAOXUE; ROWE, MATTHEW CHARLES
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 052874/0908 →
Continuity (2)
Provisional Application 62900537 · Sep 14, 2019
Related Publication 20210081377A1 · Mar 18, 2021
Cited By (4)
US 12,328,232 US 12,489,762 US 12,609,832 US 12,730,977