IP Library Granted Patent US 12,541,491
Granted Patent B2
US 12,541,491 · App. 18/885,322 · Granted Feb 3, 2026

Model ML registry and model serving

Inventors: Aaron Daniel Davidson (Berkeley, CA); Clemens Mewald (Lafayette, CA); Tomas Nykodym (San Francisco, CA)
Assignee: Databricks, Inc.
G06F16/219G06F16/955G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,491
App. No.
18/885,322
Granted
Feb 3, 2026
Kind
B2
Abstract

A system includes an interface, a processor, and a memory. The interface is configured to receive a version of a model from a model registry. The processor is configured to store the version of the model, start a process running the version of the model, and update a proxy with version information associated with the version of the model, wherein the updated proxy indicates to redirect an indication to invoke the version of the model to the process. The memory is coupled to the processor and configured to provide the processor with instructions.

Claims (50)

1 . A method comprising:

receiving, by a model server system, an application programming interface (API) request to execute a machine learning model, the API request identifying a first machine learning model to execute the API request, wherein the model server system hosts at least two different machine learning models;

generating an input based on execution data included in the API request;

generating an output by providing the input to the first machine learning model identified by the API request data; and

returning the output in response to the API request.

2 . The method of claim 1 , wherein the API request was initiated by an application executing on a client device that is remote to the model server system.

3 . The method of claim 1 , wherein the application is a chat application.

4 . The method of claim 1 , further comprising:

receiving, by the model server system, a subsequent API request identifying a second machine learning model, the second machine learning model being different than the first machine learning model;

generating a subsequent input based on execution data included in the subsequent API request;

generating a subsequent output by providing the subsequent input to the second machine learning model identified by the subsequent API request data; and

returning the subsequent output in response to the subsequent API request.

5 . The method of claim 1 , further comprising:

querying a proxy server redirect table based on a version indicator included in the API request to determine routing information for the first machine learning model.

6 . The method of claim 5 , wherein providing the input to the first machine learning model comprises routing the input to an endpoint based on the routing information for the first version of the machine learning model.

7 . The method of claim 1 , wherein the API request further includes an authentication token.

8 . The method of claim 7 , further comprising:

authenticating the API request based on the authentication token.

9 . A model server system comprising:

one or more computer processors; and

one or more computer-readable mediums storing instruction that, when executed by the one or more computer processors, cause the model server system to perform operations comprising:

receiving an application programming interface (API) request to execute a machine learning model, the API request identifying a first machine learning model to execute the API request, wherein the model server system hosts at least two different machine learning models;

generating an input based on execution data included in the API request;

generating an output by providing the input to the first machine learning model identified by the API request data; and

returning the output in response to the API request.

10 . The model serving system of claim 9 , wherein the API request was initiated by an application executing on a client device that is remote to the model server system.

11 . The model serving system of claim 9 , wherein the application is a chat application.

12 . The model serving system of claim 9 , the operations further comprising:

receiving a subsequent API request identifying a second machine learning model, the second machine learning model being different than the first machine learning model;

generating a subsequent input based on execution data included in the subsequent API request;

generating a subsequent output by providing the subsequent input to the second machine learning model identified by the subsequent API request data; and

returning the subsequent output in response to the subsequent API request.

13 . The model serving system of claim 9 , the operations further comprising:

querying a proxy server redirect table based on a version indicator included in the API request to determine routing information for the first machine learning model.

14 . The model serving system of claim 13 , wherein providing the input to the first machine learning model comprises routing the input to an endpoint based on the routing information for the first version of the machine learning model.

15 . The model serving system of claim 9 , wherein the API request further includes an authentication token.

16 . The model serving system of claim 15 , the operations further comprising:

authenticating the API request based on the authentication token.

17 . A non-transitory computer-readable medium storing instruction that, when executed by one or more computer processors of a model server system, cause the model server system to perform operations comprising:

receiving an application programming interface (API) request to execute a machine learning model, the API request identifying a first machine learning model to execute the API request, wherein the model server system hosts at least two different machine learning models;

generating an input based on execution data included in the API request;

generating an output by providing the input to the first machine learning model identified by the API request data; and

returning the output in response to the API request.

18 . The non-transitory computer-readable medium of claim 17 , wherein the API request was initiated by an application executing on a client device that is remote to the model server system.

19 . The non-transitory computer-readable medium of claim 17 , wherein the application is a chat application.

20 . The non-transitory computer-readable medium of claim 17 , the operations further comprising:

receiving a subsequent API request identifying a second machine learning model, the second machine learning model being different than the first machine learning model;

generating a subsequent input based on execution data included in the subsequent API request;

generating a subsequent output by providing the subsequent input to the second machine learning model identified by the subsequent API request data; and

returning the subsequent output in response to the subsequent API request.

Assignments (2)
SECURITY INTEREST Recorded Jan 6, 2025
From: DATABRICKS, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 069825/0419 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2024
From: DAVIDSON, AARON DANIEL; MEWALD, CLEMENS; NYKODYM, TOMAS
To: DATABRICKS, INC.
Reel/Frame 068955/0404 →
Continuity (5)
Continuation 18512028 · Nov 17, 2023
Continuation 18162579 · Jan 31, 2023
Continuation 17324907 · May 19, 2021
Provisional Application 63080569 · Sep 18, 2020
Related Publication 20250021536A1 · Jan 16, 2025
References Cited (36)
US 6212560B1 · Fairchild · 2001 [cited by applicant]
US 7506334B2 · Curtis · 2009 [cited by examiner]
US 7698398B1 · Lai · 2010 [cited by applicant]
US 8069435B1 · Lai · 2011 [cited by applicant]
US 8346929B1 · Lai · 2013 [cited by applicant]
US 9516053B1 · Muddu · 2016 [cited by examiner]
US 9723110B2 · Yang et al. · 2017 [cited by applicant]
US 10270886B1 · Postelnik et al. · 2019 [cited by applicant]
US 10380500B2 · Miao et al. · 2019 [cited by applicant]
US 10482069B1 · Barnes et al. · 2019 [cited by applicant]
US 10880347B1 · Krishnan et al. · 2020 [cited by applicant]
US 11468369B1 · Wilson et al. · 2022 [cited by applicant]
US 11562180B2 · Nushi · 2023 [cited by examiner]
US 11681819B1 · Surazski · 2023 [cited by examiner]
US 20110153957A1 · Gao et al. · 2011 [cited by applicant]
US 20110283256A1 · Raundahl Gregersen et al. · 2011 [cited by applicant]
US 20150019825A1 · Gao et al. · 2015 [cited by applicant]
US 20150213134A1 · Nie et al. · 2015 [cited by applicant]
US 20170091651A1 · Miao et al. · 2017 [cited by applicant]
US 20190250898A1 · Yang · 2019 [cited by applicant]
US 20190384640A1 · Swamy · 2019 [cited by examiner]
US 20200134484A1 · Hazard · 2020 [cited by examiner]
US 20200177960A1 · Rakshit · 2020 [cited by examiner]
US 20200401696A1 · Ringlein · 2020 [cited by examiner]
US 20210097125A1 · Khanna · 2021 [cited by examiner]
US 20210174253A1 · Moore et al. · 2021 [cited by applicant]
US 20210263978A1 · Banipal · 2021 [cited by examiner]
US 20210264321A1 · Xiang · 2021 [cited by applicant]
US 20220014963A1 · Yeh · 2022 [cited by examiner]
US 20220092043A1 · Davidson et al. · 2022 [cited by applicant]
US 20220263843A1 · Aslam · 2022 [cited by examiner]
US 20220303680A1 · Ahmed · 2022 [cited by examiner]
US 20220414455A1 · Collins et al. · 2022 [cited by applicant]
US 20230136939A1 · Kanzelberger · 2023 [cited by examiner]
US 20230177031A1 · Davidson · 2023 [cited by examiner]
United States Office Action, U.S. Appl. No. 17/324,907, filed Dec. 7, 2022, 9 pages. [cited by applicant]