IP Library Granted Patent US 12,117,983
Granted Patent B2
US 12,117,983 · App. 18/512,028 · Granted Oct 15, 2024

Model ML registry and model serving

Inventors: Aaron Daniel Davidson (Berkeley, CA); Clemens Mewald (Lafayette, CA); Tomas Nykodym (San Francisco, CA)
Assignee: Databricks, Inc.
G06F16/219G06F16/955G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,117,983
App. No.
18/512,028
Granted
Oct 15, 2024
Kind
B2
Abstract

A system includes an interface, a processor, and a memory. The interface is configured to receive a version of a model from a model registry. The processor is configured to store the version of the model, start a process running the version of the model, and update a proxy with version information associated with the version of the model, wherein the updated proxy indicates to redirect an indication to invoke the version of the model to the process. The memory is coupled to the processor and configured to provide the processor with instructions.

Claims (58)

1. A method comprising:

receiving, by a model server system, an application programming interface (API) request to execute a machine learning model hosted by the model server system, the API request including execution data and a version indicator identifying a first version of the machine learning model to execute the API request, wherein the model server system hosts at least two different versions of the machine learning model and the API request was initiated by an application executing on a client device that is remote to the model server system;

identifying the first version of the machine learning model based on the version indicator included in the API request;

generating an input based on the execution data included in the API request;

generating an output by providing the input to the first version of the machine learning model; and

providing, to the client device, a response to the API request, the response having been generated based on the output.

2. The method of claim 1 , wherein the application is a chat application.

3. The method of claim 1 , further comprising:

receiving, by the model server system, a subsequent API request to execute the machine learning model, the subsequent API request including a version indicator identifying a second version of the machine learning model, the second version of the machine learning model being different than the first version of the machine learning model;

identifying the second version of the machine learning model based on the version indicator included in the subsequent API request;

generating a subsequent input based on execution data included in the subsequent API request;

generating a subsequent output by providing the subsequent input to the second version of the machine learning model; and

providing a response to the subsequent API request, the response to the subsequent API request having been generated based on the subsequent output.

4. The method of claim 1 , wherein identifying the first version of the machine learning model based on the version indicator included in the API request comprises:

querying a proxy server redirect table based on the version indicator included in the API request to determine routing information for the first version of the machine learning model.

5. The method of claim 4 , wherein providing the input to the first version of the machine learning model comprises routing the input to an endpoint based on the routing information for the first version of the machine learning model.

6. The method of claim 1 , wherein the API request further includes an authentication token.

7. The method of claim 6 , further comprising:

authenticating the API request based on the authentication token.

8. A computer system comprising:

one or more computer processors; and

one or more computer-readable mediums storing instruction that, when executed by the one or more computer processors, cause the computer system to perform operations comprising:

receiving an application programming interface (API) request to execute a machine learning model hosted by the model server system, the API request including execution data and a version indicator identifying a first version of the machine learning model to execute the API request, wherein the model server system hosts at least two different versions of the machine learning model and the API request was initiated by an application executing on a client device that is remote to the model server system;

identifying the first version of the machine learning model based on the version indicator included in the API request;

generating an input based on the execution data included in the API request;

generating an output by providing the input to the first version of the machine learning model; and

providing, to the client device, a response to the API request, the response having been generated based on the output.

9. The computer system of claim 8 , wherein the application is a chat application.

10. The computer system of claim 8 , wherein the operations further comprise:

receiving a subsequent API request to execute the machine learning model, the subsequent API request including a version indicator identifying a second version of the machine learning model, the second version of the machine learning model being different than the first version of the machine learning model;

identifying the second version of the machine learning model based on the version indicator included in the subsequent API request;

generating a subsequent input based on execution data included in the subsequent API request;

generating a subsequent output by providing the subsequent input to the second version of the machine learning model; and

providing a response to the subsequent API request, the response to the subsequent API request having been generated based on the subsequent output.

11. The computer system of claim 8 , wherein identifying the first version of the machine learning model based on the version indicator included in the API request comprises:

querying a proxy server redirect table based on the version indicator included in the API request to determine routing information for the first version of the machine learning model.

12. The computer system of claim 11 , wherein providing the input to the first version of the machine learning model comprises routing the input to an endpoint based on the routing information for the first version of the machine learning model.

13. The computer system of claim 8 , wherein the API request further includes an authentication token.

14. The computer system of claim 13 , wherein the operations further comprise:

authenticating the API request based on the authentication token.

15. A non-transitory computer-readable medium storing instructions that, when executed by one or more computer processors of a computer system, cause the computer system to perform operations comprising:

receiving an application programming interface (API) request to execute a machine learning model hosted by the model server system, the API request including execution data and a version indicator identifying a first version of the machine learning model to execute the API request, wherein the model server system hosts at least two different versions of the machine learning model and the API request was initiated by an application executing on a client device that is remote to the model server system;

identifying the first version of the machine learning model based on the version indicator included in the API request;

generating an input based on the execution data included in the API request;

generating an output by providing the input to the first version of the machine learning model; and

providing, to the client device, a response to the API request, the response having been generated based on the output.

16. The non-transitory computer-readable medium of claim 15 , wherein the application is a chat application.

17. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

receiving a subsequent API request to execute the machine learning model, the subsequent API request including a version indicator identifying a second version of the machine learning model, the second version of the machine learning model being different than the first version of the machine learning model;

identifying the second version of the machine learning model based on the version indicator included in the subsequent API request;

generating a subsequent input based on execution data included in the subsequent API request;

generating a subsequent output by providing the subsequent input to the second version of the machine learning model; and

providing a response to the subsequent API request, the response to the subsequent API request having been generated based on the subsequent output.

18. The non-transitory computer-readable medium of claim 15 , wherein identifying the first version of the machine learning model based on the version indicator included in the API request comprises:

querying a proxy server redirect table based on the version indicator included in the API request to determine routing information for the first version of the machine learning model.

19. The non-transitory computer-readable medium of claim 18 , wherein providing the input to the first version of the machine learning model comprises routing the input to an endpoint based on the routing information for the first version of the machine learning model.

20. The non-transitory computer-readable medium of claim 15 , wherein the operations further comprise:

authenticating the API request based on an authentication token included in the API request.

Assignments (3)
SECURITY INTEREST Recorded Jan 6, 2025
From: DATABRICKS, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 069825/0419 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME FROM "DATABRICKS INC." TO --DATABRICKS, INC.-- PREVIOUSLY RECORDED ON REEL 65612 FRAME 667. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT DOCUMENT. Recorded Sep 13, 2024
From: DAVIDSON, AARON DANIEL; NYKODYM, TOMAS; MEWALD, CLEMENS
To: DATABRICKS, INC.
Reel/Frame 068982/0160 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 19, 2023
From: DAVIDSON, AARON DANIEL; MEWALD, CLEMENS; NYKODYM, TOMAS
To: DATABRICKS INC.
Reel/Frame 065612/0667 →
Continuity (4)
Continuation 18162579 · Jan 31, 2023
Continuation 17324907 · May 19, 2021
Provisional Application 63080569 · Sep 18, 2020
Related Publication 20240152496A1 · May 9, 2024