IP Library › Granted Patent US 12,373,415
Granted Patent B2
US 12,373,415 · App. 18/207,050 · Granted Jul 29, 2025

Synthesizing transformations to relationalize data tables

Inventors: Yeye He (Bellevue, WA); Cong Yan (Issaquah, WA); Yue Wang (Redmond, WA); Surajit Chaudhuri (Kirkland, WA); Peng Li (Atlanta, GA)
Assignee: Microsoft Technology Licensing, LLC
G06F16/2282G06F16/21
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,415
App. No.
18/207,050
Filed
Jun 7, 2023
Granted
Jul 29, 2025
Kind
B2
Art Unit
2153
USPC
707/609
Abstract

This document relates to relational databases and corresponding data tables. Non-conforming data tables can be automatically transformed into conforming relational data tables. One example can obtain conforming relational data tables and can generate training data without human labelling by identifying a transformational operator that will transform an individual conforming relational data table to a non-conforming data table and an inverse transformational operator that will transform the non-conforming data table back to the individual conforming relational data table. The example can train a model with the training data. The trained model can synthesize programs to transform other non-conforming data tables to conforming relational data tables.

Claims (43)

1. A method comprising:

obtaining a conforming relational data table that conforms to a relational format;

generating a training example by performing a first transformational operation on the conforming relational data table, the first transformational operation resulting in a non-conforming data table that is formatted differently than the relational format;

identifying a second transformational operation that will transform the non-conforming data table back to the conforming relational data table;

training a machine learning model with the training example to obtain a trained machine learning model, the training involving evaluating a loss function to determine weights of the trained machine learning model, the loss function being based at least on whether the machine learning model correctly predicts the second transformational operation that will transform the non-conforming data table back to the conforming relational data table;

synthesizing, with the trained machine learning model, a program for a different non-conforming data table that is formatted differently than the relational format; and,

transforming, with the synthesized program, the different non-conforming data table into another conforming relational data table that conforms to the relational format.

2. The method of claim 1 , further comprising:

selecting the first transformational operation from a set of transformational operators.

3. The method of claim 1 , wherein selecting the first transformational operation comprises selecting a single transformational operator or multiple serially performed transformational operators.

4. The method of claim 3 , wherein identifying the second transformational operation comprises identifying a single inverse transformational operator or multiple serially performed inverse transformational operators.

5. The method of claim 1 , wherein the obtaining comprises obtaining multiple conforming relational data tables and training examples are generated from each of the multiple conforming relational data tables.

6. The method of claim 5 , wherein the training the machine learning model comprises training the machine learning model utilizing the training examples generated from the multiple conforming relational data tables.

7. The method of claim 1 , wherein training the machine learning model comprises training the machine learning model without a human-provided label for the training example.

8. The method of claim 1 , wherein the machine learning model is trained prior to receiving the different non-conforming data table.

9. The method of claim 8 , further comprising:

generating a user interface; and

receiving input identifying the different non-conforming data table through the user interface.

10. The method of claim 9 , further comprising:

presenting the transforming of the different non-conforming data table into the another conforming relational data table on the user interface.

11. The method of claim 1 , the trained machine learning model comprising an embedding layer and a convolutional layer.

12. A system, comprising:

a processor; and,

a storage resource storing computer-readable instructions which, when executed by the processor, cause the processor to:

obtain conforming relational data tables that are formatted in a relational format;

generate training data without human labelling by identifying:

a transformational operator that will transform an individual conforming relational data table to a non-conforming data table that is formatted differently than the relational format, and

an inverse transformational operator that will transform the non-conforming data table back to the individual conforming relational data table that is formatted in the relational format; and,

train a machine learning model with the training data by evaluating a loss function to determine weights of the trained machine learning model, the loss function being based at least on whether the machine learning model correctly predicts the inverse transformational operation that will transform the individual non-conforming data table back to the conforming relational data table.

13. The system of claim 12 , wherein the processor is further configured to synthesize programs with the trained machine learning model for other individual conforming relational data tables.

14. The system of claim 13 , wherein the processor is further configured to rank the synthesized programs.

15. The system of claim 14 , wherein the processor is further configured to re-rank the synthesized programs with input-output re-ranking.

16. The system of claim 15 , wherein the processor is further configured to receive an additional data table and utilize the trained machine learning model to synthesize a program to transform the additional data table into a conforming relational data table.

17. The system of claim 16 , wherein the processor is further configured to cause a user interface to be generated and to receive the additional data table via the user interface.

18. A computing device, comprising:

hardware; and,

an Auto-Tables component configured to utilize a trained machine learning model to synthesize a program to transform a particular input data table into a particular conforming relational data table,

wherein, prior to synthesizing the program, the trained machine learning model has been trained by evaluating a loss function to determine weights of the trained machine learning model, the loss function being based at least on whether the machine learning model correctly predicts operations that will transform other non-conforming data tables into other conforming relational data tables,

wherein the particular conforming relational data table is in a relational format, and

wherein the particular input data table is in a different format than the relational format.

19. The computing device of claim 18 , wherein the Auto-Tables component is further configured to cause a user interface to be generated and to receive input identifying the particular input data table via the user interface.

20. The computing device of claim 19 , wherein the Auto-Tables component is further configured to cause the transformation of the particular input data table into the particular conforming relational data table to be presented on the user interface.

21. The computing device of claim 18 , wherein the Auto-Tables component is further configured to recognize that another input data table is already a conforming relational data table and to not transform the another input data table.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2023
From: HE, YEYE; YAN, CONG; WANG, YUE; CHAUDHURI, SURAJIT; LI, PENG
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065232/0809 →
Continuity (1)
Related Publication 20240411740A1 · Dec 12, 2024
References Cited (82)
US 6714943B1 · Ganesh · 2004 [cited by examiner]
US 7028057B1 · Vasudevan · 2006 [cited by examiner]
US 7287216B1 · Lee · 2007 [cited by examiner]
US 7801886B1 · Gabriel · 2010 [cited by examiner]
US 8417715B1 · Bruckhaus · 2013 [cited by examiner]
US 11256485B2 · Shi · 2022 [cited by examiner]
US 20040024790A1 · Everett · 2004 [cited by applicant]
US 20160179869A1 · Hutchins · 2016 [cited by examiner]
US 20190014151A1 · Firke · 2019 [cited by examiner]
US 20190258942A1 · Gu et al · 2019 [cited by applicant]
US 20200175390A1 · Conti · 2020 [cited by examiner]
US 20210019125A1 · Shi · 2021 [cited by examiner]
US 20220067048A1 · Mohan · 2022 [cited by applicant]
US 20220343444A1 · Chan · 2022 [cited by examiner]
US 20230236587A1 · Yang · 2023 [cited by examiner]
CN 112052414A · 2020 [cited by applicant]
EP 1686498A2 · 2006 [cited by applicant]
“Auto Tables-Supp-Materials-for-Review”, Retrieved from: https://1drv.ms/u/s!AkvY8ho1gepOicQ5-x0l4_-Smn5CuQ?e=6F8ue7, Retrieved Date: Feb. 28, 2023, 1 Page. [cited by applicant]
“Data Restructuring Using Excel”, Retrieved from: https://techcommunity.microsoft.com/t5/excel/data-restructuring-using-excel/m-p/287547, Retrieved Date: Feb. 28, 2023, 9 Pages. [cited by applicant]
“Foofah: programming-by-example data transformation program synthesizer”, Retrieved from: https://github.com/umich-dbgroup/foofah, Apr. 23, 2018, 4 Pages. [cited by applicant]
“Pandas”, Retrieved from: https://pandas.pydata.org/, Retrieved Date: Feb. 28, 2023, 2 Pages. [cited by applicant]
“Pandas Melt with Multi Index Data Set and Resetting Index—Why is this working?”, Retrieved from: https://stackoverflow.com/questions/53917303/pandas-melt-with-multi-index-data-set-and-resetting-index-why-is-this-workin… [cited by applicant]
“Pandas Melt with Multiple Value Vars”, Retrieved from: https://stackoverflow.com/questions/45066873/pandas-melt-with-multiple-value-vars, Retrieved Date: Feb. 28, 2023, 4 Pages. [cited by applicant]
“pandas.DataFrame.explode”, Retrieved from: https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.explode.html, Retrieved Date: Feb. 28, 2023, 2 Pages. [cited by applicant]
“pandas.DataFrame.pivot”, Retrieved from: https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.pivot.html, Retrieved Date: Feb. 28, 2023, 3 Pages. [cited by applicant]
“pandas.DataFrame.stack”, Retrieved from: https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.stack.html, Retrieved Date: Feb. 28, 2023, 3 Pages. [cited by applicant]
“pandas.DataFrame.transpose”, Retrieved from: https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.transpose.html, Retrieved Date: Feb. 28, 2023, 2 Pages. [cited by applicant]
“pandas.melt”, Retrieved from: https://pandas.pydata.org/docs/reference/api/pandas.melt.html, Retrieved Date: Feb. 28, 2023, 2 Pages. [cited by applicant]
“pandas.wide_to_long”, Retrieved from: https://pandas.pydata.org/docs/reference/api/pandas.wide_to_long.html, Retrieved Date: Feb. 28, 2023, 4 Pages. [cited by applicant]
“Pivot data from wide to long”, Retrieved from: https://tidyr.tidyverse.org/reference/pivot_longer.html, Retrieved Date: Feb. 28, 2023, 8 Pages. [cited by applicant]
“Pivot table issue”, Retrieved from: https://techcommunity.microsoft.com/t5/excel/pivot-table-issue/m-p/3015448, Retrieved Date: Feb. 28, 2023, 7 Pages. [cited by applicant]
“Reshape wide to long in pandas”, Retrieved from: https://stackoverflow.com/questions/36537945/reshape-wide-to-long-in-pandas, Retrieved Date: Feb. 28, 2023, 5 Pages. [cited by applicant]
“Scythe: Synthesizing SQL queries from input output examples.”, Retrieved from: https://github.com/Mestway/Scythe, Retrieved Date: Feb. 28, 2023, 2 Pages. [cited by applicant]
“Simultaneously melt multiple columns in Python Pandas”, Retrieved from: https://stackoverflow.com/questions/51519101/simultaneously-melt-multiple-columns-in-python-pandas, Retrieved Date: Feb. 28, 2023, 6 Pages. [cited by applicant]
“Transposing data for better analysis”, Retrieved from: https://techcommunity.microsoft.com/t5/excel/transposing-data-for-better-analysis/m-p/1297106, Retrieved Date: Feb. 28, 2023, 6 Pages. [cited by applicant]
Barowy, et al., “FlashRelate: extracting relational data from semi-structured spreadsheets using examples”, In ACM SIGPLAN Notices 50, Issue 6, Jun. 13, 2015, pp. 218-228. [cited by applicant]
Bojanowski, et al., “Enriching Word Vectors with Subword Information”, In Journal of Transactions of the Association for Computational Linguistics, vol. 5, Jun. 2017, pp. 135-146. [cited by applicant]
Codd, E. F., “The relational model for database management: version 2”, In Publication of Addison-Wesley Longman Publishing Co., Inc, Jan. 1, 1990, 567 Pages. [cited by applicant]
Deng, et al., “Imagenet: A Large-Scale Hierarchical Image Database”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 20, 2009, pp. 248-255. [cited by applicant]
Deng, et al., “Turl: Table understanding through representation learning”, In Proceedings of SIGMOD Record, vol. 51, No. 1, Mar. 2022, pp. 33-40. [cited by applicant]
Dong, et al., “TableSense: Spreadsheet Table Detection with Convolutional Neural Networks”, In Proceedings of the AAAI Conference on Artificial Intelligence, Jul. 17, 2019, pp. 69-76. [cited by applicant]
Gulwani, et al., “Spreadsheet Data Manipulation Using Examples”, In Magazine Communications of the ACM, vol. 55, Issue 8, Aug. 1, 2012, pp. 97-105. [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 27, 2016, pp. 770-778. [cited by applicant]
He, et al., “Transform-data-by-example (TDE) an extensible search engine for data transformations”, In Proceedings of the VLDB Endowment, vol. 11, Issue 10, Jun. 1, 2018, pp. 1165-1177. [cited by applicant]
Herzig, et al., “TAPAS: Weakly Supervised Table Parsing via Pre-training”, In Repository of arXiv:2004.02349v1, Apr. 5, 2020, 14 Pages. [cited by applicant]
Ishio, et al., “PATSQL—SQL Synthesizer”, Retrieved from: https://github.com/takashi-ishio, Retrieved Date: Feb. 28, 2023, 8 Pages. [cited by applicant]
Jin, et al., “Foofah: Transforming Data by Example”, In Proceedings of the ACM International Conference on Management of Data, May 14, 2017, pp. 683-698. [cited by applicant]
Kent, William, “A simple guide to five normal forms in relational database theory”, In Journal of Communications of the ACM, vol. 26, No. 2, Feb. 1983, pp. 120-125. [cited by applicant]
Krizhevsky, et al., “Imagenet classification with deep convolutional neural networks”, In journal of Communications of the ACM, vol. 60, Issue 6, Jun. 2017, pp. 84-90. [cited by applicant]
Maddock, David, “Pivot Chart 4 columns, set responses to 4 questions”, Retrieved from: https://techcommunity.microsoft.com/t5/user/viewprofilepage/user-id/949996#profile, May 5, 2021, 6 Pages. [cited by applicant]
Manning, et al., “An Introduction to Information Retrieval”, In Publication of Cambridge University Press, Jul. 12, 2008, 581 Pages. [cited by applicant]
Murphy, Kevin, “Machine Learning: A Probabilistic Perspective”, In Publication MIT Press, Sep. 18, 2012, 1098 Pages. [cited by applicant]
Paszke, et al., “Automatic Differentiation in PyTorch”, In Proceedings of 31st Conference on Neural Information Processing Systems, Oct. 28, 2017, 4 Pages. [cited by applicant]
Pennington, et al., “GloVe: Global Vectors for Word Representation”, In Proceedings of Conference on Empirical Methods in Natural Language Processing, Oct. 25, 2014, pp. 1532-1543. [cited by applicant]
Reimers, et al., “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”, In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on … [cited by applicant]
Shorten, et al., “A survey on image data augmentation for deep learning”, In Journal of Journal of big data vol. 6, Issue 1, Dec. 2019, 48 Pages. [cited by applicant]
Simonyan, et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition”, In repository of arXiv:1409.1556v5, Dec. 23, 2014, 13 Pages. [cited by applicant]
Takenouchi, et al., “PATSQL: effcient synthesis of SQL queries from example tables with quick inference of projected columns”, In Repository of arXiv:2010.05807v1, Oct. 12, 2020, 11 Pages. [cited by applicant]
Tran, et al., “Query by Output”, In Proceedings of the ACM SIGMOD International Conference on Management of Data, Jun. 29, 2009, pp. 535-548. [cited by applicant]
Wang, et al., “Synthesizing Highly Expressive SQL Queries from Input-Output Examples”, In Proceedings of the 38th ACM SIGPLAN Conference on Programming Language Design and Implementation, Jun. 18, 2017, pp. 452-466. [cited by applicant]
Wang, et al., “TUTA: Tree-based Transformers for Generally Structured Table Pre-training”, In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, Aug. 2021, pp. 1780-1790. [cited by applicant]
Yan, et al., “Auto-Suggest: Learning-to-Recommend Data Preparation Steps using Data Science Notebooks”, In Proceedings of the ACM SIGMOD International Conference on Management of Data, Jun. 11, 2020, pp. 1539-1554. [cited by applicant]
Yin, et al., “TABERT: Pretraining for Joint Understanding of Textual and Tabular Data”, In Repository of arXiv:2005.08314v1, May 17, 2020, 15 Pages. [cited by applicant]
Zhang, “Automatically synthesizing SQL queries from input-output examples”, In Proceedings of 28th IEEE/ACM International Conference on Automated Software Engineering (ASE), Nov. 11, 2013, pp. 224-234. [cited by applicant]
“Bee-Synth/Bee”, Retrieved from internet URL: https://github.com/BEE-Synth/Bee/tree/291a824622e36fccfa43461e85be3f836e3f4eff/Eval/Benchmarks/Spreadsheet/flashrelate-01, 2 pages. [cited by applicant]
“Power Query—Data Cleaning (Unpivot, Transpose,etc)”, Microsoft 365, Retrieved from internet URL: https://techcommunity.microsoft.com/t5/excel/power-query-data-cleaning-unpivot-transpose-etc/m-p/2400300, May 31, 2021, 7… [cited by applicant]
“Unpivot grouped data”, Microsoft 365, Retrieved from Internet URL: https://techcommunity.microsoft.com/t5/excel/unpivot-grouped-data/m-p/3686239, Nov. 29, 2022, 5 pages. [cited by applicant]
“UnPivot Monthly Data”, Microsoft 365, Retrieved from internet URL: https://techcommunity.microsoft.com/t5/excel/unpivot-monthly-data/m-p/1867836, Nov. 9, 2020, 8 pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners”, arXiv preprint arXiv:2005.14165, Jul. 22, 2020, 75 pages. [cited by applicant]
Gao, et al., “Navigating the Data Lake with Datamaran: Automatically Extracting Structure from Log Datasets”, SIGMOD '18: Proceedings of the 2018 International Conference on Management of Data, May 27, 2018, pp. 943-958. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/031519, Sep. 6, 2024, 15 pages. [cited by applicant]
Jin, et al., “Auto-transform: learning-to-transform by patterns”, Proceedings of the VLDB Endowment, vol. 13, Issue 12, Jul. 1, 2020, pp. 2368-2381. [cited by applicant]
John CNL, “Excel—Ideas feature! ”, Microsoft, Retrieved from internet URL: https://answers.microsoft.com/en-us/msoffice/forum/all/excel-ideas-feature/c9574cf9-dccc-4356-95d3-07d268e39d82, Jul. 12, 2024, 7 pages. [cited by applicant]
Koehler, et al., “Incorporating Data Context to Cost-Effectively Automate End-to-End Data Wrangling”, IEEE Transactions on Big Data, vol. 7, Issue 1, Apr. 15, 2019, pp. 169-186. [cited by applicant]
Li, et al., “A Survey of Convolutional Neural Networks: Analysis, Applications, and Prospects”, IEEE Transactions on Neural Networks and Learning Systems, vol. 33, Issue 12, Jun. 10, 2021, pp. 6999-7019. [cited by applicant]
Li, et al., “Auto-Tables: Relationalize Tables without Using Examples”, ACM SIGMOD Record, vol. 53, Issue 1, May 14, 2024, pp. 76-85. [cited by applicant]
Li, et al., “Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples”, arXiv preprint arXiv:2307.14565, Jul. 27, 2023, 15 pages. [cited by applicant]
Li, et al., “Auto-Tables: Synthesizing Multi-Step Transformations to Relationalize Tables without Using Examples”, arXiv:2307.14565v2, Aug. 9, 2023, 15 pages. [cited by applicant]
Lin, et al., “Auto-BI: Automatically Build BI-Models Leveraging Local Join Prediction and Global Schema Graph”, arXiv preprint arXiv:2306.12515, Jun. 21, 2023, 18 pages. [cited by applicant]
Nobari, et al., “Efficiently Transforming Tables for Joinability”, IEEE 38th International Conference on Data Engineering (ICDE), May 9, 2022, 14 pages. [cited by applicant]
Yang, et al., “Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search”, arXiv preprint arXiv:2106.13861, Aug. 3, 2021, 16 pages. [cited by applicant]
Zhu, et al., “Auto-join: joining tables by leveraging transformations”, Proceedings of the VLDB Endowment, vol. 10, Issue 10, Jun. 1, 2017, pp. 1034-1045. [cited by applicant]