IP Library › Granted Patent US 11,880,344
Granted Patent B2
US 11,880,344 · App. 17/321,138 · Granted Jan 23, 2024

Synthesizing multi-operator data transformation pipelines

Inventors: Yeye He (Redmond, WA); Surajit Chaudhuri (Kirkland, WA); Junwen Yang (Chicago, IL)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G06F16/211G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,880,344
App. No.
17/321,138
Granted
Jan 23, 2024
Kind
B2
Abstract

Methods and systems for generating multi-operator data transformation pipelines. An example method includes accessing raw data for transformation; receiving a selection of a target table or target visualization, wherein the target table or target visualization is for data other than the raw data; extracting table properties and target constraints; and based on the extracted table properties and target constraints, synthesizing one or more multi-operator data transformation pipelines for transforming the raw data to a generated table or generated visualization.

Claims (77)

1. A computing device, comprising:

at least one processing unit; and

system memory encoding instructions that, when executed by the at least one processing unit, cause the computing device to perform operations comprising:

receive raw data for transformation;

provide a user interface for receiving a selection of a target database table or target visualization dashboard;

receive, by way of the user interface, a selection of the target database table or the target visualization dashboard, wherein the target database table or target visualization dashboard is for data other than the raw data;

extract table properties and target constraints from the selected target database table or target visualization dashboard;

based on the table properties and target constraints extracted from the selected target database table or target visualization dashboard, generate a plurality of multi-operator data transformation pipelines;

receive a selection of one of the plurality of multi-operator data transformation pipelines; and

responsive to receiving the selection, transform the raw data to a generated database table or generated visualization dashboard having a format matching the selected target database table or target visualization dashboard using the selected one of the plurality of multi-operator data transformation pipelines.

2. The system of claim 1 , wherein the raw data includes multiple input tables.

3. The system of claim 1 , wherein operators in the plurality of multi-operator data transformation pipelines include at least two or more table-reshaping operators or string transformation operators.

4. The system of claim 3 , wherein:

the table-reshaping operators include at least one of a join operator, a union operator, a groupby operator, an agg operator, a pivot operator, an unpivot operator, or an explode operator; and

the string transformation operators include at least one of a split operator, a substring operator, a concatenate operator, casing operator, or an index operator.

5. The system of claim 1 , wherein the target constraints include at least one of a key-column constraint or a functional-dependency constraint.

6. The system of claim 1 , wherein the plurality of multi-operator data transformation pipelines includes a first multi-operator data transformation pipeline and a second multi-operator data transformation pipeline, and wherein the operations further comprise:

concurrently display:

operators of the first multi-operator data transformation pipeline as selectable visual indicators; and

operators of the second multi-operator data transformation pipeline as selectable visual indicators.

7. The system of claim 1 , wherein the operations further comprise:

generate single-operator partial pipelines;

for each single-operator partial pipeline, determine a likelihood probability and a constraint-matching criteria, wherein the constraint-matching criteria is based on the target constraints;

based on the determined likelihood probabilities and constraint matching criteria for the single-operator partial pipelines, select a subset of the single-operator partial pipelines;

generate, from the subset of the single-operator partial pipelines, double-operator partial pipelines;

for each double-operator partial pipeline, determine a likelihood probability and a constraint-matching criteria, wherein the constraint-matching criteria is based on the target constraints; and

based on the determined likelihood probabilities and constraint matching criteria for the double-operator partial pipelines, select a subset of the double-operator partial pipelines;

wherein the plurality of multi-operator data transformation pipelines are based on the subset of the double-operator pipelines.

8. The system of claim 7 , wherein selecting the subset of single-operator partial pipelines and the subset of double-operator partial pipelines includes using at least one reinforcement learning model.

9. A method for generating a multi-operator data transformation pipeline, the method comprising:

accessing raw data for transformation;

providing a user interface for receiving a selection of a target database table or a target visualization dashboard;

receiving a selection of the target database table or target visualization dashboard by way of the user interface, wherein the target database table or target visualization dashboard is for data other than the raw data;

extracting table properties and target constraints from the selected target database table or target visualization dashboard;

based on the table properties and target constraints extracted from the selected target database table or target visualization dashboard, generating a plurality of multi-operator data transformation pipelines for transforming the raw data to a generated database table or generated visualization dashboard matching the selected target database table or target visualization dashboard;

receiving a selection of one of the plurality of multi-operator data transformation pipelines; and

responsive to receiving the selection, transforming the raw data to the generated database table or generated visualization dashboard having a format matching the selected target database table or target visualization dashboard using the selected one of the plurality of multi-operator data transformation pipelines.

10. The method of claim 9 , wherein the raw data includes multiple input tables.

11. The method of claim 9 , wherein operators in the plurality of one or multi-operator data transformation pipelines include at least two or more table-reshaping operators or string transformation operators.

12. The method of claim 11 , wherein:

the table-reshaping operators include at least one of a join operator, a union operator, a groupby operator, an agg operator, a pivot operator, an unpivot operator, or an explode operator; and

the string transformation operators include at least one of a split operator, a substring operator, a concatenate operator, casing operator, or an index operator.

13. The method of claim 9 , wherein the target constraints include at least one of a key-column constraint or a functional-dependency constraint.

14. The method of claim 9 , wherein the multi-operator data transformation pipelines includes a first multi-operator data transformation pipeline and a second multi-operator data transformation pipeline, and wherein the method further comprises:

concurrently displaying:

operators of the first multi-operator data transformation pipeline as selectable visual indicators; and

operators of the second multi-operator data transformation pipeline as selectable visual indicators.

15. The method of claim 9 , further comprising:

generating single-operator partial pipelines;

for each single-operator partial pipeline, determining a likelihood probability and a constraint-matching criteria, wherein the constraint-matching criteria is based on the target constraints;

based on the determined likelihood probabilities and constraint matching criteria for the single-operator partial pipelines, selecting a subset of the single-operator partial pipelines;

generating, from the subset of the single-operator partial pipelines, double-operator partial pipelines;

for each double-operator partial pipeline, determining a likelihood probability and a constraint-matching criteria, wherein the constraint-matching criteria is based on the target constraints; and

based on the determined likelihood probabilities and constraint matching criteria for the double-operator partial pipelines, selecting a subset of the double-operator partial pipelines;

wherein the plurality of multi-operator data transformation pipelines are based on the subset of the double-operator pipelines.

16. The method of claim 15 , wherein selecting the subset of single-operator partial pipelines and the subset of double-operator partial pipelines includes using at least one reinforcement learning model.

17. Computer storage media storing instructions that cause one or more processors to perform operations comprising:

receiving raw data for transformation;

providing a user interface for receiving a selection of a target database table or target visualization dashboard;

receiving a selection of a target database table or target visualization dashboard by way of the user interface, wherein the target database table or target visualization dashboard is for data other than the raw data;

extracting table properties and target constraints from the selected target database table or target visualization dashboard;

based on the table properties and target constraints extracted from the selected target database table or target visualization dashboard, generating a plurality of one or more multi-operator data transformation pipelines for transforming the raw data to a generated database table or generated visualization dashboard matching the selected target table or target visualization;

receiving a selection of one of the plurality of multi-operator data transformation pipelines; and

responsive to receiving the selection, transforming the raw data to a generated database table or generated visualization dashboard having a format matching the selected target database table or target visualization dashboard using the selected one of the plurality of multi-operator data transformation pipelines.

18. The computer storage media storing instructions of claim 17 , wherein:

operators in the plurality of multi-operator data transformation pipelines include at least two or more table-reshaping operators or string transformation operators;

the table-reshaping operators include at least one of a join operator, a union operator, a groupby operator, an agg operator, a pivot operator, an unpivot operator, or an explode operator; and

the string transformation operators include at least one of a split operator, a substring operator, a concatenate operator, casing operator, or an index operator.

19. The computer storage media of claim 17 , wherein the target constraints include at least one of a key-column constraint or a functional-dependency constraint.

20. The computer storage media of claim 17 , having further instructions stored thereupon that cause the one or more processors to perform operations comprising:

generating single-operator partial pipelines;

for each single-operator partial pipeline, determining a likelihood probability and a constraint-matching criteria, wherein the constraint-matching criteria is based on the target constraints;

based on the determined likelihood probabilities and constraint matching criteria for the single-operator partial pipelines, selecting a subset of the single-operator partial pipelines;

generating, from the subset of the single-operator partial pipelines, double-operator partial pipelines;

for each double-operator partial pipeline, determining a likelihood probability and a constraint-matching criteria, wherein the constraint-matching criteria is based on the target constraints; and

based on the determined likelihood probabilities and constraint matching criteria for the double-operator partial pipelines, selecting a subset of the double-operator partial pipelines;

wherein the plurality of multi-operator data transformation pipelines are based on the subset of the double-operator pipelines; and wherein the subset of single-operator partial pipelines and the subset of double-operator partial pipelines are selected using at least one reinforcement learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2021
From: HE, YEYE; CHAUDHURI, SURAJIT; YANG, JUNWEN
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 056256/0401 →
Continuity (1)
Related Publication 20220365910A1 · Nov 17, 2022
Cited By (1)
US 12,210,981