IP Library Patent Application 18527111
Patent Application
App. No. 18/527,111

SELECTING OPTIMAL HARDWARE CONFIGURATIONS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/527,111
Abstract

A data processing service builds a container for a customer to run a trained large language model (LLM). The data processing service receives a trained LLM and a desired configuration from a user of a client device. Based on the desired configuration, the data processing service selects a hardware configuration and structures weights of the trained LLM based on the hardware configuration. The data processing service generates a container image reflecting the hardware configuration, registers the container image to a container registry, and generates a container from the container image as well as an application programming interface (API) endpoint for the container. The data processing service deploys the trained LLM in the API endpoint using the container such that the trained LLM is accessible through API calls.

Claims (50)

1 . A method of building a container for a client to run a trained large language model (LLM) comprising:

receiving the trained LLM and a desired configuration, the trained LLM including a set of weights;

selecting a hardware configuration based on the desired configuration;

structuring the set of weights of the trained LLM based on the hardware configuration;

generating a container image reflecting the hardware configuration;

registering the container image to a container registry;

generating the container from the container image to deploy the trained LLM in the container;

generating an application programming interface (API) endpoint for the container; and

deploying the trained LLM in the API endpoint using the container, the trained LLM accessible through API calls.

2 . The method of claim 1 , wherein structuring the set of weights of the trained LLM comprises splitting the weights using tensor parallelism.

3 . The method of claim 1 , wherein selecting the hardware configuration comprises selecting the hardware configuration based on a queries per second (QPS) of the trained LLM.

4 . The method of claim 1 , further comprising selecting a batching configuration for the trained LLM.

5 . The method of claim 1 , further comprising quantizing the trained LLM.

6 . The method of claim 1 , wherein selecting the hardware configuration comprises:

determining a particular hardware configuration in a hardware configuration table that has a highest throughput for a model type of the trained LLM;

computing an expected price per hour of the determined particular hardware configuration; and

comparing the expected price per hour to a cost threshold.

7 . The method of claim 1 , wherein selecting the hardware configuration comprises determining a particular hardware configuration in a hardware configuration table that has a lowest latency for a model type of the trained LLM and has an expected price per hour that does not exceed a cost threshold.

8 . The method of claim 7 , wherein determining the hardware configuration in the hardware configuration table that has the lowest latency for the model type of the trained LLM comprises simulating an expected latency of at least one hardware configuration.

9 . A non-transitory computer readable storage medium comprising stored program code, the program code comprising instructions, the instructions when executed cause a processor system to:

receive a trained LLM and a desired configuration, the trained LLM including a set of weights;

select a hardware configuration based on the desired configuration;

structure the set of weights of the trained LLM based on the hardware configuration;

generate a container image reflecting the hardware configuration and registering the container image to a container registry;

generate a container from the container image to deploy the trained LLM in the container;

generate an application programming interface (API) endpoint for the container; and

deploy the trained LLM in the API endpoint using the container, the trained LLM accessible through API calls.

10 . The non-transitory computer readable storage medium of claim 9 , wherein the instructions for structuring the set of weights of the trained LLM comprise instructions that, when executed, cause the processor system to split the weights using tensor parallelism.

11 . The non-transitory computer readable storage medium of claim 9 , wherein the instructions for selecting the hardware configuration comprise instructions that, when executed, cause the processor system to select the hardware configuration based on a queries per second (QPS) of the trained LLM.

12 . The non-transitory computer readable storage medium of claim 9 , wherein the instructions further comprise instructions that, when executed, cause the processor system to select a batching configuration for the trained LLM.

13 . The non-transitory computer readable storage medium of claim 9 , wherein the instructions further comprise instructions that, when executed, cause the processor system to quantize the trained LLM.

14 . The non-transitory computer readable storage medium of claim 9 , wherein the instructions for selecting the hardware configuration comprise instructions that, when executed, cause the processor system to:

determine a particular hardware configuration in a hardware configuration table that has a highest throughput for a model type of the trained LLM; and

compute an expected price per hour of the determined particular hardware configuration;

compare the expected price per hour to a cost threshold.

15 . The method of claim 1 , wherein the instructions for selecting the hardware configuration comprise instructions that, when executed, cause the processor system to determine a particular hardware configuration in a hardware configuration table that has a lowest latency for a model type of the trained LLM and has an expected price per hour that does not exceed a cost threshold.

16 . The method of claim 15 , wherein the instructions for determining the hardware configuration in the hardware configuration table that has the lowest latency for the model type of the trained LLM comprise instructions that, when executed, cause the processor system to simulate an expected latency of at least one hardware configuration.

17 . A computer system, comprising:

a computer processor; and

a non-transitory computer readable storage medium comprising stored instructions that when executed by the computer processor, cause the computer system to:

receive a trained LLM and a desired configuration, the trained LLM including a set of weights;

select a hardware configuration based on the desired configuration;

structure the set of weights of the trained LLM based on the hardware configuration;

generate a container image reflecting the hardware configuration and registering the container image to a container registry;

generate a container from the container image to deploy the trained LLM in the container;

generate an application programming interface (API) endpoint for the container; and

deploy the trained LLM in the API endpoint using the container, the trained LLM accessible through API calls.

18 . The computer system of claim 17 , wherein the instructions for structuring the set of weights of the trained LLM comprise instructions causing the computer system to split the weights using tensor parallelism.

19 . The computer system of claim 17 , wherein the instructions for selecting the hardware configuration comprise instructions causing the computer system to select a hardware configuration based on a queries per second (QPS) of the trained LLM.

20 . The computer system of claim 17 , wherein the instructions further comprise instructions causing the computer system to select a batching configuration for the trained LLM.

Assignments (2)
SECURITY INTEREST Recorded Jan 6, 2025
From: DATABRICKS, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 069825/0419 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2024
From: BILAL, AHMED; CHEN, STEVEN YIKUN; FONTAINE, BRUCE LAURENT; KHUDIA, DAYA SHANKER; LI, CHENRAN; MATHUR, ANKIT
To: DATABRICKS, INC.
Reel/Frame 066131/0950 →