IP Library Granted Patent US 11,379,140
Granted Patent B2
US 11,379,140 · App. 17/494,542 · Granted Jul 5, 2022

System and method for model training orchestration

Inventors: Luis Capelo (New York, NY); Williams Falcon (New York, NY); Karolis Rusenas (New York, NY); Luca Antiga (New York, NY); Neven Miculinic (New York, NY)
Assignee: Grid.ai, Inc.
G06F3/0644G06F3/0604G06F3/067G06F3/0659G06F9/4881G06N20/00G06F2209/505G06F2209/508
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,379,140
App. No.
17/494,542
Granted
Jul 5, 2022
Kind
B2
Abstract

A system for large-scale machine learning experiment execution, including: a platform configured to determine an experiment set from a run specification and schedule a run to one or more clusters; and a set of agents configured to receive the experiment set from the platform and facilitate individual experiment execution through a cluster orchestrator.

Claims (38)

1. A system, comprising:

a plurality of agents, each installed on a different computing cluster, wherein each computing cluster is controlled by a different cluster orchestrator; and

a processing system, configured to:

receive cluster telemetry, comprising metrics for experiments running on a cluster, from the agent for the respective cluster; and

send control instructions determined based on the cluster telemetry to the agent, wherein the agent controls the respective cluster orchestrator according to the control instructions, wherein the control instructions are in a platform-standard protocol, wherein the agent controls the respective cluster orchestrator using a cluster orchestrator-specific protocol.

2. The system of claim 1 , wherein the processing system facilitates storage of artifacts generated by the experiments within an object store connected to the respective cluster, wherein the processing system cannot directly access the artifacts.

3. The system of claim 1 , wherein the experiments running on a cluster are part of a shared run, wherein the processing system determines the experiments for each run and schedules runs to clusters, wherein different experiments are scheduled to different sets of nodes within the cluster.

4. The system of claim 3 , wherein the cluster orchestrator of the respective cluster schedules the experiment to the different sets of nodes within the cluster.

5. The system of claim 3 , wherein all experiments of the shared run are specified by a single-line run request comprising: a set of model identifiers, a set of datastore identifiers, and a set of hyperparameter values, wherein the experiments are automatically scheduled responsive to receipt of the single line run request without human intervention.

6. The system of claim 1 , further comprising an experiment reconciler, configured to reconcile experiments across different clusters.

7. A system for multi-cluster machine learning experiment orchestration, comprising:

a plurality of agents, each installed on a different computing cluster, wherein each agent is configured to control a cluster orchestrator of the computing cluster, and wherein each cluster orchestrator controls a plurality of nodes within the respective computing cluster; and

a processing system, comprising a multitenant platform, configured to:

determine a set of runs, each run comprising a set of machine learning experiments; and

schedule runs to clusters, wherein the agent associated with the cluster controls run execution on the cluster;

wherein the processing system concurrently controls execution of different sets of runs on different sets of clusters for different users.

8. The system of claim 7 , wherein a run is executed on a heterogeneous computing architecture.

9. The system of claim 7 , wherein the runs are scheduled to clusters based on an optimization over an estimated execution time and estimated execution cost for all experiments within the run to execute on the respective cluster.

10. The system of claim 7 , wherein each run is associated with a run specification comprising a set of model identifiers, a set of dataset identifiers, and a set of hyperparameter values, wherein each experiment within the run is associated with an experiment configuration comprising at least one model identifier from the set of model identifiers, at least one dataset identifier from the set of dataset identifiers, and a combination of hyperparameter values.

11. The system of claim 10 , wherein the processing system determines a set of experiment configurations associated with a run and sends the experiment configurations to the agent associated with the cluster scheduled for the run.

12. The system of claim 10 , wherein each experiment is assigned to a set of nodes within the cluster by the cluster orchestrator.

13. The system of claim 10 , wherein the experiment configuration is further determined based on run telemetry from a previously-executed run.

14. The system of claim 7 , wherein the processing system is configured to initialize a set of clusters, determine the set of runs, schedule the runs, and dynamically manage the runs without user intervention, responsive to receipt of a single line request comprising a run specification.

15. The system of claim 7 , wherein the processing system initializes a cluster on behalf of a user using authorization credentials provided to the processing system by the user.

16. The system of claim 7 , wherein the agent sends run telemetry to the processing system, wherein the processing system schedules another run of the set to another agent based on the run telemetry.

17. The system of claim 7 , wherein each experiment generates a trained machine learning model associated with model metrics, wherein the processing system accesses the model metrics via the agent and cannot access the trained machine learning model.

18. The system of claim 7 , wherein the cluster orchestrator comprises a Kubernetes™ deployment.

19. A system, comprising:

a plurality of agents, each installed on a different computing cluster, wherein each computing cluster is controlled by a different cluster orchestrator; and

a processing system, configured to:

receive cluster telemetry, comprising metrics for experiments running on a cluster, from the agent for the respective cluster, wherein the experiments running on the cluster are part of a shared run, wherein the processing system determines the experiments for each run and schedules runs to clusters, wherein different experiments are scheduled to different sets of nodes within the cluster; and

send control instructions determined based on the cluster telemetry to the agent, wherein the agent controls the respective cluster orchestrator according to the control instructions.

20. A system, comprising:

a plurality of agents, each installed on a different computing cluster, wherein each computing cluster is controlled by a different cluster orchestrator;

a processing system, configured to:

receive cluster telemetry, comprising metrics for experiments running on a cluster, from the agent for the respective cluster;

send control instructions determined based on the cluster telemetry to the agent, wherein the agent controls the respective cluster orchestrator according to the control instructions; and

reconcile experiments across different clusters.

Assignments (2)
SECURITY INTEREST Recorded Feb 3, 2026
From: GRID.AI, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 073679/0233 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2021
From: CAPELO, LUIS; FALCON, WILLIAMS; RUSENAS, KAROLIS; ANTIGA, LUCA; MICULINIC, NEVEN
To: GRID.AI, INC.
Reel/Frame 058005/0027 →
Continuity (9)
Provisional Application 63182218 · Apr 30, 2021
Provisional Application 63173666 · Apr 12, 2021
Provisional Application 63173657 · Apr 12, 2021
Provisional Application 63173674 · Apr 12, 2021
Provisional Application 63168667 · Mar 31, 2021
Provisional Application 63088908 · Oct 7, 2020
Provisional Application 63088888 · Oct 7, 2020
Provisional Application 63087406 · Oct 5, 2020
Related Publication 20220107744A1 · Apr 7, 2022
Cited By (2)
US 12,254,295 US 12,287,990