IP Library › Granted Patent US 12,265,828
Granted Patent B2
US 12,265,828 · App. 18/159,968 · Granted Apr 1, 2025

Automated runtime configuration for dataflows

Inventors: Abhishek Uday Kumar Shah (Seattle, WA); Anudeep Sharma (Bothell, WA); Mark A. Kromer (Snohomish, WA); Jikai Ma (Redmond, WA)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G06F9/328G06F9/30036G06F11/3006G06F16/164G06N20/20G06F9/5044G06F9/5066G06F9/5072G06N3/044G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,828
App. No.
18/159,968
Filed
Jan 26, 2023
Granted
Apr 1, 2025
Kind
B2
Art Unit
2192
USPC
712/220
Abstract

Methods, systems and computer program products are provided for automated runtime configuration for dataflows to automatically select or adapt a runtime environment or resources to a dataflow plan prior to execution. Metadata generated for dataflows indicates dataflow information, such as numbers and types of sources, sinks and operations, and the amount of data being consumed, processed and written. Weighted dataflow plans are created from unweighted dataflow plans based on metadata. Weights that indicate operation complexity or resource consumption are generated for data operations. A runtime environment or resources to execute a dataflow plan is/are selected based on the weighted dataflow and/or a maximum flow. Preferences may be provided to influence weighting and runtime selections.

Claims (62)

1. A method in a server, comprising:

receiving a dataflow plan comprising representations of data operations in a dataflow pipeline;

determining, based on the received dataflow plan, metadata of the received dataflow plan;

determining, based on the received dataflow plan, a feature set comprising a plurality of features;

selecting a machine learning (ML) model from a plurality of ML models based on the metadata of the received dataflow plan and metadata of a training set the selected ML model was trained on;

providing the feature set to the selected ML model;

receiving, from the selected ML model, the weighted dataflow plan associating a weight associated with a data operation of the data operations; and

causing execution of the received dataflow plan by resources allocated based on the weight.

2. The method of claim 1 , wherein the weight is selected from a weight range corresponding to an execution complexity range for the data operation in terms of consumption of at least one of computing, memory, storage or network resources.

3. The method of claim 1 , wherein the weight comprises a combined weight or a weight vector indicating at least two of computing, memory, storage, and network resources.

4. The method of claim 1 , wherein the ML model determines the weight based at least on a portion of the metadata of the received dataflow plan corresponding to the data operation.

5. The method of claim 1 , further comprising:

determining a maximum flow for the dataflow plan having the weight applied thereto,

wherein said causing execution of the received dataflow plan comprises:

causing allocation of the resources based on the maximum flow.

6. The method of claim 5 , wherein said determining the maximum dataflow comprises:

applying a max-flow min-cut theorem to the dataflow plan having the weight applied thereto.

7. The method of claim 1 , wherein said causing execution of the received dataflow plan comprises:

causing allocation of the resources based on an indication of a preference related to at least one of execution time to execute the received dataflow plan or execution cost to execute the received dataflow plan.

8. The method of claim 1 , wherein said causing execution of the received dataflow plan comprises:

providing feedback to a user interface indicating a plurality of execution environments and associated costs; and

prompting user input to select an execution environment from among the plurality of execution environments to execute the received dataflow plan.

9. The method of claim 1 , wherein said causing execution of the received dataflow plan comprises:

providing the weight to a second ML model; and

causing allocation of the resources based at least on an output generated by the second ML model based at least on the weight.

10. A system, comprising:

a processor; and

memory that stores program code executable by the processor to perform a method, the method comprising:

receiving a dataflow plan comprising representations of data operations in a dataflow pipeline;

determining, based on the received dataflow plan, metadata of the received dataflow plan;

determining, based on the received dataflow plan, a feature set comprising a plurality of features;

selecting a machine learning (ML) model from a plurality of ML models based on metadata of the received dataflow plan and metadata of a training set the selected ML model was trained on;

providing the feature set to the selected ML model;

receiving, from the selected ML model, a weight associated with a data operation of the data operations; and

causing execution of the received dataflow plan by resources allocated based on the weight.

11. The system of claim 10 , wherein the weight is selected from a weight range corresponding to an execution complexity range for the data operation in terms of consumption of at least one of computing, memory, storage or network resources.

12. The system of claim 10 , wherein the weight comprises a combined weight or a weight vector indicating at least two of computing, memory, storage, and network resources.

13. The system of claim 10 , wherein the ML model determines the weight based at least on the corresponding metadata.

14. The system of claim 10 , wherein the method further comprises:

determining a maximum flow for the dataflow plan having the weight applied thereto,

wherein said causing execution of the received dataflow plan comprises:

causing allocation of the resources based on the maximum flow.

15. The system of claim 14 , wherein said determining the maximum dataflow comprises:

applying a max-flow min-cut theorem to the dataflow plan having the weight applied thereto.

16. The system of claim 10 , wherein said causing execution of the received dataflow plan comprises:

causing allocation of the resources based on an indication of a preference related to at least one of execution time to execute the received dataflow plan or execution cost to execute the received dataflow plan.

17. The system of claim 10 , wherein said causing execution of the received dataflow plan comprises:

providing feedback to a user interface indicating a plurality of execution environments and associated costs; and

prompting user input to select an execution environment from among the plurality of execution environments to execute the received dataflow plan.

18. The system of claim 10 , wherein said causing execution of the received dataflow plan comprises:

providing the weight to a second ML model; and

causing allocation of the resources based at least on an output generated by the second ML model based at least on the weight.

19. A method in a server, comprising:

receiving a dataflow plan comprising representations of data operations in a dataflow pipeline;

determining, based on the received dataflow plan, metadata of the received dataflow plan that comprises information regarding the data operations;

selecting a machine learning (ML) model from a plurality of ML models based on the metadata of the dataflow plan and metadata of a training set the selected ML model was trained on;

providing the generated metadata to the selected ML model;

receiving, from the selected ML model, a weight associated with a data operation of the data operations; and

causing execution of the received dataflow plan by resources allocated based on the weight.

20. The method of claim 19 , wherein said causing execution of the received dataflow plan comprises:

determining a maximum flow for the dataflow plan having the weight applied thereto; and

causing allocation of the resources based on the maximum flow.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2023
From: SHAH, ABHISHEK UDAY KUMAR; SHARMA, ANUDEEP; KROMER, MARK A.; MA, JIKAI
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062507/0958 →
Continuity (2)
Continuation 16811448 · Mar 6, 2020
Related Publication 20230168895A1 · Jun 1, 2023
References Cited (12)
US 10728120B2 · Beyer · 2020 [cited by examiner]
US 11269911B1 · Jones · 2022 [cited by examiner]
US 20160189034A1 · Shakeri · 2016 [cited by examiner]
US 20180218295A1 · Hasija · 2018 [cited by examiner]
US 20180232702A1 · Dialani · 2018 [cited by examiner]
US 20190147076A1 · Yang · 2019 [cited by examiner]
US 20200042362A1 · Cui · 2020 [cited by examiner]
Yung-Chuan Jiang et al., Temporal Partitioning Data Flow Graphs for Dynamically Reconfigurable Computing, 2007 IEEE, [Retrieved on May 23, 2023]. Retrieved from the internet: <URL: https://ieeexplore.ieee.org/stamp/stam… [cited by examiner]
Shanjiang Tang et al., A Survey on Spark Ecosystem: Big Data Processing Infrastructure, Machine Learning, and Applications, Jan. 2022, [Retrieved on Sep. 18, 2024]. Retrieved from the internet: <URL: https://ieeexplore.… [cited by examiner]
“Machine Learning”, Retrieved From: https://en.wikipedia.org/w/index.php?title=Machine_learning&oldid=909989905, Aug. 8, 2019, 21 Pages. [cited by applicant]
“Office Action Issued in European Patent Application No. 21707578.7”, Mailed Date: Oct. 24, 2023, 9 Pages. [cited by applicant]
Communication pursuant to Article 94(3) received in European Application No. 21707578.7, mailed on May 24, 2024, 4 pages. [cited by applicant]