IP Library › Granted Patent US 12,111,797
Granted Patent B1
US 12,111,797 · App. 18/371,931 · Granted Oct 8, 2024

Schema inference system

Inventors: Kirsten Rae Lum (Issaquah, WA); Wing Yew Lum (Issaquah, WA); Christopher John Gutierrez (Taos, NM)
Assignee: Storytellers.ai LLC
G06F16/211G06F16/221G06F16/24544G06F16/2456
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,111,797
App. No.
18/371,931
Granted
Oct 8, 2024
Kind
B1
Abstract

Concrete data types for raw data organized in one or more columns of one or more tables may be determined. Functional data types for the one or more columns may be determined based on the raw data and the concrete data types such that a portion of the one or more columns may be associated with an identifier data type. Relationships between the one or more tables may be determined based on the portion of the one or more columns associated with the identifier data type. A schema representing the raw data may be generated based on the one or more relationships and the one or more tables.

Claims (92)

1. A method for managing data in a network using one or more processors to execute instructions that are configured to cause actions, comprising:

employing raw data to determine one or more tables with one or more columns, wherein the raw data includes metadata and is organized in the one or more columns of one or more tables;

employing the raw data to determine one or more concrete data types that correspond to the raw data organized in the one or more columns of the one or more tables;

determining one or more functional data types for the one or more columns of the one or more tables based on the raw data and correspondence with the one or more concrete data types, wherein a portion of the one or more columns are associated with an identifier data type;

determining one or more existing relationships associated with the one or more tables based on one or more of a metric or a statistical feature of the metadata for the raw data;

determining one or more inferred relationships associated with the one or more tables based on the portion of the one or more columns associated with the identifier data type and one or more common values associated with the one or more functional data types for the one or more columns;

employing one or more large language models to predict one or more inferred relationships associated with the one or more tables, wherein the one or more large language models are trained by one or more semantic relationship evaluators to generate the one or more predicted inferred relationships between the one or more tables;

executing one or more join expressions on two or more predicted inferred relationships for validation, wherein each predicted inferred relationship associated with one or more join execution errors is invalidated; and

generating a schema representing the raw data based on the one or more existing relationships, the one or more inferred relationships, each predicted inferred relationship that is validated, and the one or more tables.

2. The method of claim 1 , further comprising:

providing one or more queries that include one or more join expressions that are based on the one or more relationships; and

validating the schema based on an execution of the one or more queries, wherein the validation is based on an absence of errors associated with the execution of the one or more queries.

3. The method of claim 1 , further comprising:

employing one or more data sources to provide the raw data, wherein the raw data includes one or more of a comma separated value file, a spreadsheet, an extensible markup language file, a database export file, an application export file, or a network connection to data source.

4. The method of claim 1 , wherein determining the one or more concrete data types, further comprises:

determining one or more evaluators that declare one or more operations to infer the one or more concrete data types; and

inferring the one or more concrete data types based on the one or more operations, wherein the one or more concrete data types include one or more of an integer type, a floating point type, a character type, or a string type.

5. The method of claim 1 , wherein determining the one or more functional data types, further comprises:

determining one or more evaluators that declare one or more operations to infer the one or more functional data types; and

inferring the one or more functional data types based on the one or more operations, wherein the one or more functional data types include one or more of the identifier data type, a category data type, a text data type, or a numeric data type.

6. The method of claim 1 , further comprising:

determining one or more metrics associated with the one or more columns based on the raw data, wherein the one or more metrics include one or more of a row count, a median value, a mean value, a cardinality, or a distribution of values; and

providing a profile for each column, wherein the one or more metrics for each column are included in the profile; and

employing the profile for each column to further determine the one or more concrete data types or the one or more functional data types.

7. The method of claim 1 , further comprising:

determining a portion of the one or more relationships based on an evaluation of header information associated with the one or more columns, wherein the portion of the one or more relationships are associated with a portion of the one or more columns that are associated with related header information.

8. The method of claim 1 , further comprising:

determining a portion of the one or more relationships based on an evaluation of one or more statistical features of a portion of the raw data associated with the portion of the one or more relationships.

9. The method of claim 1 , further comprising:

determining a portion of the one or more relationships based on an evaluation of one or more semantic characteristics of the one or more columns, wherein the evaluation of one or more semantic characteristics is based on a response from a large language model that is trained by a prompt that includes information associated with the one or more columns.

10. A network computer for managing data, comprising:

a memory that stores at least instructions; and

one or more processors that execute instructions that are configured to cause actions, including:

employing raw data to determine one or more tables with one or more columns, wherein the raw data includes metadata and is organized in the one or more columns of one or more tables;

employing the raw data to determine one or more concrete data types that correspond to the raw data organized in the one or more columns of the one or more tables;

determining one or more functional data types for the one or more columns of the one or more tables based on the raw data and correspondence with the one or more concrete data types, wherein a portion of the one or more columns are associated with an identifier data type;

determining one or more existing relationships associated with the one or more tables based on one or more of a metric or a statistical feature of the metadata for the raw data;

determining one or more inferred relationships associated with the one or more tables based on the portion of the one or more columns associated with the identifier data type and one or more common values associated with the one or more functional data types for the one or more columns;

employing one or more large language models to predict one or more inferred relationships associated with the one or more tables, wherein the one or more large language models are trained by one or more semantic relationship evaluators to generate the one or more predicted inferred relationships between the one or more tables;

executing one or more join expressions on two or more predicted inferred relationships for validation, wherein each predicted inferred relationship associated with one or more join execution errors is invalidated; and

generating a schema representing the raw data based on the one or more existing relationships, the one or more inferred relationships, each predicted inferred relationship that is validated, and the one or more tables.

11. The network computer of claim 10 , further comprising:

providing one or more queries that include one or more join expressions that are based on the one or more relationships; and

validating the schema based on an execution of the one or more queries, wherein the validation is based on an absence of errors associated with the execution of the one or more queries.

12. The network computer of claim 10 , further comprising:

employing one or more data sources to provide the raw data, wherein the raw data includes one or more of a comma separated value file, a spreadsheet, an extensible markup language file, a database export file, an application export file, or a network connection to data source.

13. The network computer of claim 10 , wherein determining the one or more concrete data types, further comprises:

determining one or more evaluators that declare one or more operations to infer the one or more concrete data types; and

inferring the one or more concrete data types based on the one or more operations, wherein the one or more concrete data types include one or more of an integer type, a floating point type, a character type, or a string type.

14. The network computer of claim 10 , wherein determining the one or more functional data types, further comprises:

determining one or more evaluators that declare one or more operations to infer the one or more functional data types; and

inferring the one or more functional data types based on the one or more operations, wherein the one or more functional data types include one or more of the identifier data type, a category data type, a text data type, or a numeric data type.

15. The network computer of claim 10 , further comprising:

determining one or more metrics associated with the one or more columns based on the raw data, wherein the one or more metrics include one or more of a row count, a median value, a mean value, a cardinality, or a distribution of values; and

providing a profile for each column, wherein the one or more metrics for each column are included in the profile; and

employing the profile for each column to further determine the one or more concrete data types or the one or more functional data types.

16. The network computer of claim 10 , further comprising:

determining a portion of the one or more relationships based on an evaluation of header information associated with the one or more columns, wherein the portion of the one or more relationships are associated with a portion of the one or more columns that are associated with related header information.

17. The network computer of claim 10 , further comprising:

determining a portion of the one or more relationships based on an evaluation of one or more statistical features of a portion of the raw data associated with the portion of the one or more relationships.

18. The network computer of claim 10 , further comprising:

determining a portion of the one or more relationships based on an evaluation of one or more semantic characteristics of the one or more columns, wherein the evaluation of one or more semantic characteristics is based on a response from a large language model that is trained by a prompt that includes information associated with the one or more columns.

19. A processor readable non-transitory storage media that includes instructions configured for managing data over a network, wherein execution of the instructions by one or more processors on one or more network computers performs actions, comprising:

employing raw data to determine one or more tables with one or more columns, wherein the raw data includes metadata and is organized in the one or more columns of one or more tables;

employing the raw data to determine one or more concrete data types that correspond to the raw data organized in the one or more columns of the one or more tables;

determining one or more functional data types for the one or more columns of the one or more tables based on the raw data and correspondence with the one or more concrete data types, wherein a portion of the one or more columns are associated with an identifier data type;

determining one or more existing relationships associated with the one or more tables based on one or more of a metric or a statistical feature of the metadata for the raw data;

determining one or more inferred relationships associated with the one or more tables based on the portion of the one or more columns associated with the identifier data type and one or more common values associated with the one or more functional data types for the one or more columns;

employing one or more large language models to predict one or more inferred relationships associated with the one or more tables, wherein the one or more large language models are trained by one or more semantic relationship evaluators to generate the one or more predicted inferred relationships between the one or more tables;

executing one or more join expressions on two or more predicted inferred relationships for validation, wherein each predicted inferred relationship associated with one or more join execution errors is invalidated; and

generating a schema representing the raw data based on the one or more existing relationships, the one or more inferred relationships, each predicted inferred relationship that is validated, and the one or more tables.

20. The media of claim 19 , further comprising:

providing one or more queries that include one or more join expressions that are based on the one or more relationships; and

validating the schema based on an execution of the one or more queries, wherein the validation is based on an absence of errors associated with the execution of the one or more queries.

21. The media of claim 19 , further comprising:

employing one or more data sources to provide the raw data, wherein the raw data includes one or more of a comma separated value file, a spreadsheet, an extensible markup language file, a database export file, an application export file, or a network connection to data source.

22. The media of claim 19 , wherein determining the one or more concrete data types, further comprises:

determining one or more evaluators that declare one or more operations to infer the one or more concrete data types; and

inferring the one or more concrete data types based on the one or more operations, wherein the one or more concrete data types include one or more of an integer type, a floating point type, a character type, or a string type.

23. The media of claim 19 , wherein determining the one or more functional data types, further comprises:

determining one or more evaluators that declare one or more operations to infer the one or more functional data types; and

inferring the one or more functional data types based on the one or more operations, wherein the one or more functional data types include one or more of the identifier data type, a category data type, a text data type, or a numeric data type.

24. The media of claim 19 , further comprising:

determining one or more metrics associated with the one or more columns based on the raw data, wherein the one or more metrics include one or more of a row count, a median value, a mean value, a cardinality, or a distribution of values; and

providing a profile for each column, wherein the one or more metrics for each column are included in the profile; and

employing the profile for each column to further determine the one or more concrete data types or the one or more functional data types.

25. The media of claim 19 , further comprising:

determining a portion of the one or more relationships based on an evaluation of header information associated with the one or more columns, wherein the portion of the one or more relationships are associated with a portion of the one or more columns that are associated with related header information.

26. The media of claim 19 , further comprising:

determining a portion of the one or more relationships based on an evaluation of one or more statistical features of a portion of the raw data associated with the portion of the one or more relationships.

27. The media of claim 19 , further comprising:

determining a portion of the one or more relationships based on an evaluation of one or more semantic characteristics of the one or more columns, wherein the evaluation of one or more semantic characteristics is based on a response from a large language model that is trained by a prompt that includes information associated with the one or more columns.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2023
From: LUM, KIRSTEN RAE; LUM, WING YEW; GUTIERREZ, CHRISTOPHER JOHN
To: STORYTELLERS.AI LLC
Reel/Frame 065000/0508 →
Cited By (2)
US 12,423,064 US 12,645,694