IP Library › Granted Patent US 11,921,681
Granted Patent B2
US 11,921,681 · App. 17/344,489 · Granted Mar 5, 2024

Machine learning techniques for predictive structural analysis

Inventors: Vijaychandar Natesan (Bangalore, IN); Ramesh R. Ganesan (Bangalore, IN); Rakesh P A (Bengaluru, IN); Rahul Singh (Bengaluru, IN); Sarath C Varma Kutcharlapati (Vizianagaram, IN); Varunkumar Akula (Karimnagar, IN)
Assignee: Optum Technology, Inc.
G06F16/211G06F16/2264G06F16/2282G06F16/25G06F16/285G06N5/04G06N20/00G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,921,681
App. No.
17/344,489
Granted
Mar 5, 2024
Kind
B2
Abstract

Various embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for performing predictive structural analysis. Certain embodiments of the present invention utilize systems, methods, and computer program products that perform predictive structural analysis using at least one of table column classification machine learning models, table column clustering machine learning models, structural variance generation machine learning models, and emergence report generation machine learning models.

Claims (74)

1. A computer-implemented method comprising:

identifying, by one or more processors, a reference table data object associated with a table data object, wherein (i) the table data object comprises a plurality of table columns and (ii) the reference table data object comprises a plurality of reference table columns;

extracting, by the one or more processors, a reference table column feature of a plurality of reference table column features for each of the plurality of reference table columns, wherein at least one reference table column feature of the plurality of reference table column features comprises a sparsity feature of the corresponding reference table column;

for each table column pair that comprises a table column of the table data object and a reference table column of the reference table data object, determining, by the one or more processors, a table column pair similarity measure based at least in part on a table column mapping and a reference table column mapping, wherein: (i) the reference table column mapping maps the corresponding reference table column to a multi-dimensional clustering space based at least in part on a defined set of table column features and (ii) the defined set of table column features comprises at least a sparsity feature of the corresponding table column;

determining, by the one or more processors and based at least in part on each table column pair similarity measure, a variance report for the table data object, wherein the variance report describes at least one table column that does not achieve a similarity threshold associated with its table column pair; and

initiating, by the one or more processors, the performance of one or more prediction-based actions based at least in part on the variance report.

2. The computer-implemented method of claim 1 , further comprising:

identifying, by the one or more processors, an unidentified table column set of the plurality of table columns, wherein each overall column type prediction for a table column in the unidentified table column set describes that the table column is not associated with a candidate table column type; and

generating, by the one or more processors, an overall unidentified table column report that describes one or more unidentified table column groupings as determined based at least in part on the unidentified table column set.

3. The computer-implemented method of claim 2 , wherein determining the one or more unidentified table column groupings comprises:

for each unidentified table column of a plurality of unidentified table columns, determining, by the one or more processors, a plurality of unidentified table column features; and

determining, by the one or more processors and based at least in part on each plurality of unidentified table column features for an unidentified table column, the one or more unidentified table column groupings of the plurality of unidentified table columns.

4. The computer-implemented method of claim 1 , further comprising:

for each table column:

generating, by the one or more processors, using a header-based table classification machine learning model of a plurality of classification machine learning models and based at least in part on a table column name set for the table column, a predicted header-based column type of a plurality of predicted column types for the table column and a header-based column type voting weight of a plurality of column type voting weights for the predicted header-based column type;

generating, by the one or more processors, using a data-based table classification machine learning model of the plurality of classification machine learning models and based at least in part on a table column value set for the table column, a predicted data-based column type of the plurality of predicted column types for the table column and a data-based column type voting weight of the plurality of column type voting weights for the predicted data-based column type;

generating, by the one or more processors, using an entity recognition classification machine learning model of the plurality of classification machine learning models and based at least in part on the table column value set, a predicted entity-recognition-based column type of the plurality of predicted column types for the table column and an entity-recognition-based column type voting weight of the plurality of column type voting weights for the predicted entity-recognition-based column type;

generating, by the one or more processors, using a pattern matching classification machine learning model of the plurality of classification machine learning models and based at least in part on the table column name set, a predicted pattern-machine-based column type of the plurality of predicted column types for the table column and a pattern-matching-based column type voting weight of the plurality of column type voting weights for the predicted entity-recognition-based column type; and

generating, by the one or more processors, using a voting machine learning model and based at least in part on the plurality of predicted column types and the plurality of column type voting weights, an overall column type prediction for the table column; and

initiating, by the one or more processors, the performance of one or more second prediction-based actions based at least in part on each overall column type prediction for a table column.

5. The computer-implemented method of claim 4 , wherein generating each overall column type prediction for a table column comprises:

for each candidate column type of a plurality of candidate column types:

identifying, by the one or more processors, a predicted column type set of the plurality of predicted column types for the table column that correspond to the candidate column type;

identifying, by the one or more processors, a column type voting weight set of the plurality of column type voting weights that correspond to the predicted column type set; and

determining, by the one or more processors, a candidate column type voting value for the candidate column type with respect to the table column based at least in part on the column type voting weight set; and

generating, by the one or more processors, the overall column type prediction based at least in part on each candidate column type voting value for a candidate column type with respect to the table column.

6. The computer-implemented method of claim 1 , further comprising:

for each table column:

determining, by the one or more processors, using a table column clustering machine learning model and based at least in part on a plurality of table column features of the table column, a related table column cluster set for the table column; and

determining, by the one or more processors, a functional grouping of the table column based at least in part on the related table column cluster set for the table column.

7. The computer-implemented method of claim 6 , wherein the plurality of table column features comprises at least one of a data type feature of the table column, a data pattern feature of the table column, a most frequent entity type feature of the table column, the sparsity feature of the table column, or an adjacent column name feature of the table column.

8. A system comprising memory and one or more processors communicatively coupled to the memory, the one or more processors configured to:

identify a reference table data object associated with a table data object, wherein (i) the table data object comprises a plurality of table columns and (ii) the reference table data object comprises a plurality of reference table columns;

extract a reference table column feature of a plurality of reference table column features for each of the plurality of reference table columns, wherein at least one reference table column feature of the plurality of reference table column features comprises a sparsity feature of the corresponding reference table column;

for each table column pair that comprises a table column of the table data object and a reference table column of the reference table data object, determine a table column pair similarity measure based at least in part on a table column mapping and a reference table column mapping, wherein: (i) the reference table column mapping maps the corresponding reference table column to a multi-dimensional clustering space based at least in part on a defined set of table column features and (ii) the defined set of table column features comprises at least a sparsity feature of the corresponding table column;

determine, based at least in part on each table column pair similarity measure, a variance report for the table data object, wherein the variance report describes at least one table column that does not achieve a similarity threshold associated with its table column pair; and

initiate the performance of one or more prediction-based actions based at least in part on the variance report.

9. The system of claim 8 , wherein the one or more processors, are further configured to:

identify an unidentified table column set of the plurality of table columns, wherein each overall column type prediction for a table column in the unidentified table column set describes that the table column is not associated with a candidate table column type; and

generate an overall unidentified table column report that describes one or more unidentified table column groupings as determined based at least in part on the unidentified table column set.

10. The system of claim 9 , wherein determining the one or more unidentified table column groupings comprises:

for each unidentified table column of a plurality of unidentified table columns, determining a plurality of unidentified table column features; and

determining, based at least in part on each plurality of unidentified table column features for an unidentified table column, the one or more unidentified table column groupings of the plurality of unidentified table columns.

11. The system of claim 8 , wherein the one or more processors, are further configured to:

for each table column:

generate, using a header-based table classification machine learning model of a plurality of classification machine learning models and based at least in part on a table column name set for the table column, a predicted header-based column type of a plurality of predicted column types for the table column and a header-based column type voting weight of a plurality of column type voting weights for the predicted header-based column type;

generate, using a data-based table classification machine learning model of the plurality of classification machine learning models and based at least in part on a table column value set for the table column, a predicted data-based column type of the plurality of predicted column types for the table column and a data-based column type voting weight of the plurality of column type voting weights for the predicted data-based column type;

generate, using an entity recognition classification machine learning model of the plurality of classification machine learning models and based at least in part on the table column value set, a predicted entity-recognition-based column type of the plurality of predicted column types for the table column and an entity-recognition-based column type voting weight of the plurality of column type voting weights for the predicted entity- recognition-based column type;

generate, using a pattern matching classification machine learning model of the plurality of classification machine learning models and based at least in part on the table column name set, a predicted pattern-machine-based column type of the plurality of predicted column types for the table column and a pattern-matching-based column type voting weight of the plurality of column type voting weights for the predicted entity- recognition-based column type; and

generate, using a voting machine learning model and based at least in part on the plurality of predicted column types and the plurality of column type voting weights, an overall column type prediction for the table column; and

initiate the performance of one or more second prediction-based actions based at least in part on each overall column type prediction for a table column.

12. The system of claim 11 , wherein generating each overall column type prediction for a table column comprises:

for each candidate column type of a plurality of candidate column types:

identifying a predicted column type set of the plurality of predicted column types for the table column that correspond to the candidate column type;

identifying a column type voting weight set of the plurality of column type voting weights that correspond to the predicted column type set; and

determining a candidate column type voting value for the candidate column type with respect to the table column based at least in part on the column type voting weight set; and

generating the overall column type prediction based at least in part on each candidate column type voting value for a candidate column type with respect to the table column.

13. The system of claim 8 , wherein the one or more processors are further configured to:

for each table column:

determine, using a table column clustering machine learning model and based at least in part on a plurality of table column features of the table column, a related table column cluster set for the table column; and

determine a functional grouping of the table column based at least in part on the related table column cluster set for the table column.

14. The system of claim 13 , wherein the plurality of table column features comprises at least one of a data type feature of the table column, a data pattern feature of the table column, a most frequent entity type feature of the table column, the sparsity feature of the table column, or an adjacent column name feature of the table column.

15. One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processors, cause the one or more processors to:

identify a reference table data object associated with a table data object, wherein (i) the table data object comprises a plurality of table columns and (ii) the reference table data object comprises a plurality of reference table columns;

extract a reference table column feature of a plurality of reference table column features for each of the plurality of reference table columns, wherein at least one reference table column feature of the plurality of reference table column features comprises a sparsity feature of the corresponding reference table column;

for each table column pair that comprises a table column of the table data object and a reference table column of the reference table data object, determine a table column pair similarity measure based at least in part on a table column mapping and a reference table column mapping, wherein: (i) the reference table column mapping maps the corresponding reference table column to a multi-dimensional clustering space based at least in part on a defined set of table column features and (ii) the defined set of table column features comprises at least a sparsity feature of the corresponding table column;

determine, based at least in part on each table column pair similarity measure, a variance report for the table data object, wherein the variance report describes at least one table column that does not achieve a similarity threshold associated with its table column pair; and

initiate the performance of one or more prediction-based actions based at least in part on the variance report.

16. The one or more non-transitory computer-readable storage media of claim 15 , wherein the instructions further cause the one or more processors to:

identify an unidentified table column set of the plurality of table columns, wherein each overall column type prediction for a table column in the unidentified table column set describes that the table column is not associated with a candidate table column type; and

generate an overall unidentified table column report that describes one or more unidentified table column groupings as determined based at least in part on the unidentified table column set.

17. The one or more non-transitory computer-readable storage media of claim 16 , wherein determining the one or more unidentified table column groupings comprises:

for each unidentified table column of a plurality of unidentified table columns, determining a plurality of unidentified table column features; and

determining, based at least in part on each plurality of unidentified table column features for an unidentified table column, the one or more unidentified table column groupings of the plurality of unidentified table columns.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 10, 2021
From: NATESAN, VIJAYCHANDAR; GANESAN, RAMESH R.; AMBITI, KISHOR KUMAR K.; RAMANATHAN, SIVAKUMAR; T S, GIRISH KUMAR; P A, RAKESH; SINGH, RAHUL; KUTCHARLAPATI, SARATH C VARMA; AKULA, VARUNKUMAR
To: OPTUM TECHNOLOGY, INC.
Reel/Frame 056503/0843 →
Priority Claims (1)
IN 202111018632 · Apr 22, 2021 · national
Continuity (1)
Related Publication 20220342857A1 · Oct 27, 2022