IP Library Granted Patent US 12,639,652
Granted Patent B2
US 12,639,652 · App. 18/134,999 · Granted May 26, 2026

Automatically building business intelligence models

Inventors: Yeye He (Bellevue, WA); Yiming Lin (Irvine, CA); Surajit Chaudhuri (Kirkland, WA)
Assignee: Microsoft Technology Licensing, LLC
G06Q10/067G06F16/212G06F16/24544G06F16/2465
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,652
App. No.
18/134,999
Filed
Apr 14, 2023
Granted
May 26, 2026
Kind
B2
Examiner
JAMI, HARES
Art Unit
2164
USPC
707/606
Abstract

The present disclosure relates to methods and systems that automatically predict a business intelligence model for tables of data provided as input. The methods and systems automatically generate a graph representing the business intelligence model and provide the graph as output. The graph provides a visual representation of the business intelligence model with nodes of the graph representing each input table and edges of the graph representing weighted edges joining pairs of tables together.

Claims (32)

1 . A method, comprising:

receiving tables of data;

using a machine learning model to automatically predict a business intelligence model for the tables by predicting a probability of joinability of each pair of table columns in the tables, wherein the business intelligence model defines relationships between the data and the probability of joinability of a pair of table columns is used in creating edges of a graph;

outputting the graph for the business intelligence model using the probability of joinability, wherein the graph is constructed in a general snowflake structure using an equation that uses the probability of joinability in identifying nodes and edges of the graph and the nodes of the graph represent each input table of the tables of data and the edges of the graph represent weighted edges joining pairs of tables together;

using the machine learning model to perform an optimization leveraging graph properties and a general shape of business intelligence models to generate accurate predictions of join relationships between the tables in the business intelligence model, wherein the optimization enforces the snowflake structure and identifies and removes improper edges from the predicted probability of joinability in the graph; and

outputting the graph based on modifications from the optimization.

2 . The method of claim 1 , wherein the machine learning model is trained offline to predict the probability of joinability of each pair of tables in the tables of data.

3 . The method of claim 2 , wherein the graph further includes the probability of joinability presented on the edges of the graph.

4 . The method of claim 1 , further comprising:

performing an optimization of the graph using graph properties.

5 . The method of claim 4 , wherein the graph properties enforce a structure on the graph to perform the optimization of the graph.

6 . The method of claim 5 , wherein the graph properties include a minimum cost arborescence, an edge maximizing schema, or a cardinality constraint.

7 . The method of claim 4 , wherein the optimization includes adding edges to the graph.

8 . The method of claim 4 , wherein the optimization includes removing edges from the graph.

9 . The method of claim 1 , further comprising:

performing a recall mode optimization on the graph identifying missing joins from the predicted probability of joins in the graph; and

adding the missing joins as edges to the graph.

10 . A device, comprising:

a processor;

memory in electronic communication with the processor; and

instructions stored in the memory, the instructions being executable by the processor to:

receive tables of data;

use a machine learning model to automatically predict a business intelligence model for the tables by predicting a probability of joinability of each pair of table columns in the tables, wherein the business intelligence model defines relationships between the data and the probability of joinability of a pair of table columns is used in creating edges of a graph;

output the graph for the business intelligence model using the probability of joinability, wherein the graph is constructed in a general snowflake structure using an equation that uses the probability of joinability in identifying nodes and edges of the graph and the nodes of the graph represent each input table of the tables of data and edges of the graph represent weighted edges joining pairs of tables together;

use the machine learning model to perform an optimization leveraging graph properties and a general shape of business intelligence models to generate accurate predictions of join relationships between the tables in the business intelligence model, wherein the optimization enforces the snowflake structure and identifies and removes improper edges from the predicted probability of joinability in the graph; and

output the graph based on modifications from the optimization.

11 . The device of claim 10 , wherein the machine learning model is trained offline to predict a probability of joinability of each pair of tables in the tables of data and the probability of joinability is used in creating the edges of the graph.

12 . The device of claim 11 , wherein the graph further includes the probability of joinability presented on the edges of the graph.

13 . The device of claim 10 , wherein the instructions are further executable by the processor to perform an optimization of the graph using graph properties.

14 . The device of claim 13 , wherein the graph properties enforce a structure on the graph to perform the optimization of the graph.

15 . The device of claim 14 , wherein the graph properties include a minimum cost arborescence, an edge maximizing schema, or a cardinality constraint.

16 . The device of claim 13 , wherein the instructions are further executable by the processor to perform the optimization on the graph by adding edges to the graph or removing edges from the graph.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 17, 2023
From: HE, YEYE; LIN, YIMING; CHAUDHURI, SURAJIT
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063341/0588 →
Continuity (1)
Related Publication 20240346427A1 · Oct 17, 2024
References Cited (42)
US 10795895B1 · Taig · 2020 [cited by examiner]
US 20100325054A1 · Currie · 2010 [cited by examiner]
US 20120136684A1 · Pulido De Los Reyes · 2012 [cited by examiner]
US 20180173750A1 · Dumant · 2018 [cited by examiner]
US 20190378074A1 · Mcphatter · 2019 [cited by examiner]
US 20210158176A1 · Wan · 2021 [cited by applicant]
US 20240311623A1 · Rossi · 2024 [cited by examiner]
Pavlo, et al., “A Comparison of Approaches to Large-Scale Data Analysis”, In Proceedings of the ACM SIGMOND International Conference on Management of data, Jun. 29, 2009, 14 Pages. [cited by applicant]
Rostin, et al., “A Machine Learning Approach to Foreign Key Discovery”, In Proceedings of WebDB, Jun. 28, 2009, 6 Pages. [cited by applicant]
Karger, et al., “A Randomized Linear-Time Algorithm to Find Minimum Spanning Trees”, In Journal of the ACM, vol. 42, Issue 2, Mar. 1, 1995, pp. 321-328. [cited by applicant]
Land, et al., “An Automatic Method for Solving Discrete Programming Problems”, In Journal of 50 Years of Integer Programming 1958-2008: From the Early Years to the State-of-the-Art, Nov. 1, 2010, pp. 105-132. [cited by applicant]
Chaudhuri, et al., “An Overview of Business intelligence Technology”, In Journal of Communications of the ACM, vol. 54 , Issue 8, Aug. 1, 2011, pp. 88-98. [cited by applicant]
Li, et al., “Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples”, In Proceedings of the 2021 International Conference on Management of Data, Jun. 20, 2021, pp. 1064-1076. [cited by applicant]
Zhu, et al., “Auto-Join: Joining Tables by Leveraging Transformations”, In Journal of Proceedings of the VLDB Endowment, vol. 10, Issue 10, Jun. 1, 2017, pp. 1034-1045. [cited by applicant]
Yan, et al., “Auto-Suggest: Learning-to-Recommend Data Preparation Steps Using Data Science Notebooks”, In Proceedings of the ACM SIGMOD International Conference on Management of Data, Jun. 14, 2020, pp. 1539-1554. [cited by applicant]
Guruswami, et al., “Beating The Random Ordering Is Hard: Every Ordering Csp Is Approximation Resistant”, In Journal of SIAM Journal on Computing, vol. 40, Issue 3, Jun. 1, 2011, pp. 878-914. [cited by applicant]
Funke, et al., “Benchmarking Hybrid OLTP&OLAP Database Systems”, In Journal of Datenbanksysteme für Business, Technologie und Web (BTW), Jan. 1, 2011, pp. 390-409. [cited by applicant]
Negash, et al., “Business Intelligence”, In Handbook on decision support systems 2. Springer, Sep. 22, 2008, pp. 175-193. [cited by applicant]
Escoffier, et al., “Completeness in approximation classes beyond APX”, In Journal of Theoretical computer science, vol. 359, Issue 1-3, Aug. 14, 2006, pp. 369-377. [cited by applicant]
Hartmanis, Juris, “Computers and Intractability: A Guide to the Theory of NP-Completeness (Michael R. Garey and David S. Johnson)”, In Journal of Siam Review, vol. 24, Issue 1, Jan. 1, 1982, pp. 90-91. [cited by applicant]
DeWitt, Davidj. , “DeWitt Clause”, Retrieved From: https://en.wikipedia.org/wiki/David_DeWitt#DeWitt_Clause, Feb. 27, 2022, 2 Pages. [cited by applicant]
Rosen, Kennethh. , “Discrete Mathematics And Its Applications”, In Publication of Taylor & Francis Group, Apr. 3, 2008, 524 Pages. [cited by applicant]
Marchi, et al., “Efficient Algorithms for Mining Inclusion Dependencies”, In Proceedings of the International Conference on Extending Database Technology, Springer, Mar. 25, 2002, pp. 464-476. [cited by applicant]
Chen, et al., “Fast Foreign-Key Detection in Microsoft SQL Server PowerPivot for Excel”, In Journal of Proceedings of the VLDB Endowment, vol. 7, Issue 13, Sep. 1, 2014, pp. 1417-1428. [cited by applicant]
Jiang, et al., “Holistic Primary Key and Foreign Key Detection”, In Journal of Intelligent Information Systems, vol. 54, Issue 3, Jun. 1, 2020, pp. 439-461. [cited by applicant]
Idoine, Carlie, “How to Enable Self-Service Analytics”, Retrieved From: https://www.gartner.com/en/documents/3957086, Sep. 9, 2019, 4 Pages. [cited by applicant]
Casanova, et al., “Inclusion Dependencies and Their Interaction with Functional Dependencies”, In Proceedings of the 1st ACM SIGACT-SIGMOD symposium on Principles of database systems, Mar. 29, 1982, pp. 171-176. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/022916, Jun. 17, 2024, 15 pages. [cited by applicant]
Lin, et al., “Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph,” retrieve from: arXiv preprint arXiv:2306.12515, Jun. 21, 2023, 18pages. [cited by applicant]
Zhang, et al., “On Multi-Column Foreign Key Discovery”, In Journal of Proceedings of the VLDB Endowment, vol. 3, Issue 1, Sep. 1, 2010, pp. 805-814. [cited by applicant]
Edmonds, Jack, “Optimum Branchings”, In Journal of Research of the national Bureau of Standards, vol. 71B, Issue 4, Oct. 1, 1967, pp. 233-240. [cited by applicant]
“Power BI”, Retrieved From: https://powerbi.microsoft.com/en-us/, Retrieved on: Jan. 16, 2023, 12 Pages. [cited by applicant]
Mizil, et al., “Predicting Good Probabilities With Supervised Learning”, In Proceedings of the 22nd international conference on Machine learning, Aug. 7, 2005, 8 Pages. [cited by applicant]
Reimers, et al., “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”, In repository of arXiv:1908.10084v1, Aug. 27, 2019, 11 Pages. [cited by applicant]
“SentenceTransformers Documentation”, Retrieved From: https://www.sbert.net/, Retrieved on: Jan. 16, 2023, 7 Pages. [cited by applicant]
Mackinlay, et al., “Show Me: Automatic Presentation for Visual Analysis”, In Journal of IEEE transactions on visualization and computer graphics, Nov. 5, 2007, pp. 1137-1144. [cited by applicant]
Yang, et al., “Summarizing Relational Databases”, In Journal of Proceedings of the VLDB Endowment, vol. 2, Issue 1, Aug. 24, 2009, pp. 634-645. [cited by applicant]
“Tableau”, Retrieved From: https://www.tableau.com/, Retrieved on: Jan. 16, 2023, 13 Pages. [cited by applicant]
Kimball, et al., “The Data Warehouse Toolkit: The Complete Guide to Dimensional Modeling”, Published by John Wiley & Sons, Aug. 8, 2011, 447 Pages. [cited by applicant]
Kronz, et al., “The Gartner 2022 Analytics & BI Platforms Magic Quadrant Highlights”, Retrieved From: https://www.gartner.com/en/webinar/453533/1068121, Retrieved on: Jan. 16, 2023, 2 Pages. [cited by applicant]
Hanrahan, Pat, “VizQL: a language for query, analysis and visualization”, In Proceedings of the ACM SIGMOD international conference on Management of data, Jun. 27, 2006, pp. 721-721. [cited by applicant]
Schmidt, et al., “XMark: A Benchmark for XML Data Management”, In Proceedings of the 28th International Conference on Very Large Databases, Aug. 20, 2002, 13 Pages. [cited by applicant]