IP Library Granted Patent US 12,626,070
Granted Patent B2
US 12,626,070 · App. 18/436,105 · Granted May 12, 2026

Serverless functional routing for large language model inference service

Inventors: Bo Wen (Chappaqua, NY); Chen Wang (Chappaqua, NY); Huamin Chen (Newton, MA)
Assignee: International Business Machines Corporation
G06F40/40G06F16/3347
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,070
App. No.
18/436,105
Granted
May 12, 2026
Kind
B2
Abstract

A computer-implemented method for serving a large language model (LLM) application via a serverless function router communicative with multiple endpoints that each have a set of subject matter expert models stored thereon is provided. The computer-implemented method includes receiving a prompt, querying a database comprising multiple datasets for an indication as to which one of the multiple datasets has a highest level of similarity with the prompt, recognizing one of the multiple endpoints as having the set of the expert models stored thereon which have a closest match with the one of the multiple datasets and routing the prompt to the one of the multiple endpoints having the set of the expert models stored thereon which have the closest match with the one of the multiple datasets.

Claims (51)

1 . A computer-implemented method for serving a large language model (LLM) application, the computer-implemented method comprising:

disposing a serverless function router in communication with multiple discrete servers serving as multiple endpoints;

grouping subject matter expert models into sets of subject matter expert models and storing each set of subject matter expert models on one of the multiple endpoints;

receiving a prompt at the serverless function router;

querying, by the serverless function router, a database comprising multiple datasets for an indication as to which one of the multiple datasets has a highest level of similarity with the prompt;

recognizing, by the serverless function router, one of the multiple endpoints as having the set of the expert models stored thereon which have a closest match with the one of the multiple datasets; and

routing, by the serverless function router, the prompt to the one of the multiple endpoints having the set of the expert models stored thereon which have the closest match with the one of the multiple datasets.

2 . The computer-implemented method according to claim 1 , wherein the database comprises a vector database.

3 . The computer-implemented method according to claim 1 , wherein each of the multiple datasets has a closest match with the subject matter expert models stored on one of the multiple endpoints.

4 . The computer-implemented method according to claim 3 , wherein:

the set of subject matter expert models of a first one of the multiple endpoints is configured to handle prompts relating to medical subject matter,

the set of subject matter expert models of a second one of the multiple endpoints is configured to handle prompts relating to financial subject matter, and

the set of subject matter expert models of a third one of the multiple endpoints is configured to handle prompts relating to technical subject matter.

5 . The computer-implemented method according to claim 3 , wherein each of the sets of subject matter expert models of each of the first, second and third ones of the multiple endpoints comprises one or more foundation models and one or more fine-tuned models.

6 . The computer-implemented method according to claim 1 , wherein the serverless function router is a prompt aware router and the routing comprises prompt aware routing.

7 . A computer program product for serving a large language model (LLM) application, the computer program product comprising one or more computer readable storage media having computer readable program code collectively stored on the one or more computer readable storage media, the computer readable program code being executed by a processor of a computer system to cause the computer system to perform a method comprising:

disposing a serverless function router in communication with multiple discrete servers serving as multiple endpoints;

grouping subject matter expert models into sets of subject matter expert models and storing each set of subject matter expert models on one of the multiple endpoints;

receiving a prompt at the serverless function router;

receiving a prompt at the serverless function router;

querying, by the serverless function router, a database comprising multiple datasets for an indication as to which one of the multiple datasets has a highest level of similarity with the prompt;

recognizing, by the serverless function router, one of the multiple endpoints as having the set of the expert models stored thereon which have a closest match with the one of the multiple datasets; and

routing, by the serverless function router, the prompt to the one of the multiple endpoints having the set of the expert models stored thereon which have the closest match with the one of the multiple datasets.

8 . The computer program product according to claim 7 , wherein the database comprises a vector database.

9 . The computer program product according to claim 7 , wherein each of the multiple datasets has a closest match with the subject matter expert models stored on one of the multiple endpoints.

10 . The computer program product according to claim 9 , wherein:

the set of subject matter expert models of a first one of the multiple endpoints is configured to handle prompts relating to medical subject matter,

the set of subject matter expert models of a second one of the multiple endpoints is configured to handle prompts relating to financial subject matter, and

the set of subject matter expert models of a third one of the multiple endpoints is configured to handle prompts relating to technical subject matter.

11 . The computer program product according to claim 9 , wherein each of the sets of subject matter expert models of each of the first, second and third ones of the multiple endpoints comprises one or more foundation models and one or more fine-tuned models.

12 . The computer program product according to claim 7 , wherein the serverless function router is a prompt aware router and the routing comprises prompt aware routing.

13 . A computing system comprising:

a processor;

a memory coupled to the processor; and

one or more computer readable storage media coupled to the processor, the one or more computer readable storage media collectively containing instructions that are executed by the processor via the memory to implement a method for serving a large language model (LLM) application comprising:

disposing a serverless function router in communication with multiple discrete servers serving as multiple endpoints;

grouping subject matter expert models into sets of subject matter expert models and storing each set of subject matter expert models on one of the multiple endpoints;

receiving a prompt at the serverless function router;

receiving a prompt at the serverless function router;

querying, by the serverless function router, a database comprising multiple datasets for an indication as to which one of the multiple datasets has a highest level of similarity with the prompt;

recognizing, by the serverless function router, one of the multiple endpoints as having the set of the expert models stored thereon which have a closest match with the one of the multiple datasets; and

routing, by the serverless function router, the prompt to the one of the multiple endpoints having the set of the expert models stored thereon which have the closest match with the one of the multiple datasets.

14 . The computing system according to claim 13 , wherein:

the database comprises a vector database, and

each of the multiple datasets has a closest match with the subject matter expert models stored on one of the multiple endpoints.

15 . The computing system according to claim 14 , wherein:

the set of subject matter expert models of a first one of the multiple endpoints is configured to handle prompts relating to medical subject matter,

the set of subject matter expert models of a second one of the multiple endpoints is configured to handle prompts relating to financial subject matter, and

the set of subject matter expert models of a third one of the multiple endpoints is configured to handle prompts relating to technical subject matter.

16 . The computing system according to claim 14 , wherein each of the sets of subject matter expert models of each of the first, second and third ones of the multiple endpoints comprises one or more foundation models and one or more fine-tuned models.

17 . The computing system according to claim 13 , wherein the serverless function router is a prompt aware router and the routing comprises prompt aware routing.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2024
From: RED HAT, INC.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 069550/0992 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2024
From: CHEN, HUAMIN
To: RED HAT, INC.
Reel/Frame 066414/0541 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2024
From: WEN, BO; WANG, CHEN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 066414/0156 →
Continuity (1)
Related Publication 20250259008A1 · Aug 14, 2025
References Cited (19)
US 20230197067A1 · Robert Jose · 2023 [cited by examiner]
US 20230281510A1 · Royer · 2023 [cited by examiner]
US 20240296315A1 · Singh · 2024 [cited by examiner]
US 20240311405A1 · Kim · 2024 [cited by examiner]
US 20240411751A1 · Radmilac · 2024 [cited by examiner]
US 20250117411A1 · Mohammed · 2025 [cited by examiner]
CN 115204143A · 2022 [cited by applicant]
CN 115292484A · 2022 [cited by applicant]
CN 116127046A · 2023 [cited by applicant]
CN 116402164A · 2023 [cited by applicant]
CN 116629345A · 2023 [cited by applicant]
CN 116703454A · 2023 [cited by applicant]
Schnitzer et al. “Large Language Model Routing with Benchmark Datasets”. arXiv:2309.15789v1 [cs.CL] Sep. 27, 2023 (Year: 2023). [cited by examiner]
Fu et al. “ServerlessLLM: Locality-Enhanced Serverless Inference for Large Language Models”. arXiv:2401.14351v1 [cs.LG] Jan. 25, 2024 (Year: 2024). [cited by examiner]
Lu et al. “Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models”. arXiv:2311.08692v1 [cs.CL] Nov. 15, 2023 (Year: 2023). [cited by examiner]
Hari et al., “Tryage: Real-time, intelligent routing of user prompts to large language model.” arXiv preprint arXiv:2308.11601 (2023): 11 pages. [cited by applicant]
Lin et al. “Predictive Prompts with Joint Training of Large Language Models for Explainable Recommendation.” Mathematics 11.4230 (2023): 12 pages. [cited by applicant]
Zhou et al., “Mixture of Experts with Expert Choice Routing” :https://blog.research.google/2022/11/mixture-of-experts-with-expert-choice.html?m=1 (Retrieved Dec. 23, 2023), 6 pages. [cited by applicant]
Google Research, “Mixture-of-Experts with Expert Choice Routing, ”https://blog.research.google/2022/11/mixture-of-experts-with-expert-choice.html?m=1(Retrieved Feb. 7, 2024), 6 pages. [cited by applicant]