IP Library › Granted Patent US 12,314,825
Granted Patent B2
US 12,314,825 · App. 18/800,900 · Granted May 27, 2025

Prompt routing system and method

Inventors: Shriyash K. Upadhyay (North Potomac, MD); Etan J. Ginsberg (North Potomac, MD); Dory Zidon (North Potomac, MD); Luka Samkharadze (North Potomac, MD)
Assignee: Martian Learning, Inc.
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,825
App. No.
18/800,900
Granted
May 27, 2025
Kind
B2
Abstract

In variants, the method can include determining training data, determining a router, and using the router. In variants, using the router can include receiving a runtime prompt, predicting performance scores for the runtime prompt for each of a set of candidate models, optionally predicting operational metrics for responding to the runtime prompt for each of the set of candidate models, selecting a candidate model based on the predicted performance scores and optionally the predicted operational metrics, and optionally determining a response based on the runtime prompt.

Claims (52)

1. A method for routing machine learning model prompts, comprising:

training a scoring model comprising a neural network to predict candidate model response scores based on neural network prompts;

extracting an encoder from the scoring model, wherein the encoder is a strict subset of layers of the scoring model;

receiving a prompt;

determining a candidate model response score for each of a set of candidate models based on the prompt, by:

determining a prompt encoding from the prompt using the encoder extracted from the scoring model; and

determining the candidate model scores based on the prompt encoding;

selecting a runtime model from the set of candidate models based on the set of candidate model response scores; and

facilitating runtime model response determination for the prompt using the runtime model.

2. The method of claim 1 , wherein the runtime model is selected before the prompt is run through any of the set of candidate models.

3. The method of claim 1 , wherein determining the candidate model response scores comprises:

determining a set of stored encodings generated using the encoder, wherein each stored encoding is associated with a stored prompt; and

comparing the prompt encoding to the set of stored encodings.

4. The method of claim 3 , wherein determining the candidate model response scores further comprises:

determining a set of similar encodings to the prompt encoding from the set of stored encodings based on the comparison; and

using response scores associated with the set of similar encodings as the candidate model response scores.

5. The method of claim 4 , wherein the set of similar encodings is determined using cosine similarity.

6. The method of claim 1 , further comprising determining a set of Quality of Service (QoS) metrics for each candidate model, and wherein selecting the runtime model comprises a multivariate optimization of candidate model response scores and QoS metrics.

7. The method of claim 6 , wherein each set of QoS metrics is determined based on the prompt.

8. The method of claim 6 , wherein a QoS metric within each set of QoS metrics comprises a resource allocation.

9. The method of claim 1 , wherein the scoring model is trained by:

determining a set of multiple performance scores for a prompt-response pair, wherein each performance score is determined by a different reward model;

determining a target performance score from the set of multiple performance scores; and

training the scoring model to predict the target performance score given the prompt in the prompt-response pair.

10. The method of claim 1 , wherein each candidate model is associated with multiple candidate model response scores, wherein the runtime model is selected using the multiple candidate response scores for each of the set of candidate models.

11. The method of claim 1 , wherein the prompt is received from a third party.

12. The method of claim 11 , further comprising: receiving a set of operation metric preferences from the third party, wherein the runtime model is selected based on the operation metric preferences, wherein the operation metric preferences comprise at least one of latency, cost, or computational efficiency.

13. A method for model routing, comprising:

training a routing model comprising a neural network to predict response scores for candidate models based on a prompt, wherein the routing model is trained based on:

a set of prompts within a set of training prompt-response pairs, each training response determined using a candidate model within a set of candidate models; and

a set of training response scores, each training response score evaluating a training prompt-response pair from the set of training prompt-response pairs;

extracting an encoder from the routing model, wherein the encoder is a strict subset of layers of the routing model;

receiving a new prompt from a user;

determining a response score for each candidate model within the set of candidate models based on the new prompt by:

determining a prompt encoding of the new prompt using the encoder extracted from the routing model: and

non-probabilistically determining the response scores for the candidate models using the prompt encoding;

selecting a runtime model from the set of candidate models for the new prompt based on the response scores; and

facilitating determination of a response to the new prompt using the selected runtime model, wherein the response is returned to the user.

14. The method of claim 13 , wherein each training response score in the set of training response scores is determined by a reward model, and wherein the method further comprises:

after facilitating determination of the response, receiving a response evaluation from the user; and

retraining the reward model using the response evaluation.

15. The method of claim 13 , wherein the training response scores are generated by a scoring model trained to determine training response scores for training prompt-response pairs.

16. The method of claim 13 , wherein each predicted response score comprises a continuous value.

17. The method of claim 13 , wherein the routing model is part of a plurality of routing models, the method further comprising determining a plurality of response scores for each candidate model, wherein each response score in the plurality of response scores is determined by a different routing model; wherein the runtime model is selected based on the plurality of response scores for the runtime model.

18. The method of claim 13 , further comprising:

receiving a set of operational metric constraints from the user; and

predicting a set of operational metrics for each candidate model;

wherein the runtime model is further selected based on the predicted set of operational metrics satisfying the set of operational metric constraints.

19. The method of claim 13 , wherein facilitating determination of the response comprises transmitting the new prompt to a third party hosting the selected runtime model, after selecting the runtime model.

20. The method of claim 13 , wherein selecting the runtime model comprises:

identifying a subset of the set of training prompt-response pairs that are similar to the new prompt, using the routing model; and

selecting the runtime model using a heuristic applied to the training response score for each of the subset of training prompt-response pairs.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2025
From: UPADHYAY, SHRIYASH K.; GINSBERG, ETAN J.; ZIDON, DORY; SAMKHARADZE, LUKA; ZVERIANSKII, ALEKSANDR; BAKUTA, ARTEM; ROMANOV, ANDREI
To: MARTIAN LEARNING, INC.
Reel/Frame 072120/0820 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 19, 2024
From: SAMKHARADZE, LUKA
To: MARTIAN LEARNING, INC.
Reel/Frame 069645/0183 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2024
From: ZIDON, DORY
To: MARTIAN LEARNING, INC.
Reel/Frame 068872/0863 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2024
From: UPADHYAY, SHRIYASH K.; GINSBERG, ETAN J.
To: MARTIAN LEARNING, INC.
Reel/Frame 068361/0055 →
Continuity (4)
Provisional Application 63532199 · Aug 11, 2023
Provisional Application 63588591 · Oct 6, 2023
Provisional Application 63598879 · Nov 14, 2023
Related Publication 20250053876A1 · Feb 13, 2025
References Cited (19)
US 20200202256A1 · Chaudhari · 2020 [cited by examiner]
US 20200380418A1 · Strope · 2020 [cited by examiner]
US 20210150155A1 · Kim · 2021 [cited by examiner]
US 20210209412A1 · Quader · 2021 [cited by examiner]
US 20220318689A1 · Li-Bland · 2022 [cited by examiner]
US 20220400159A1 · Chi et al. · 2022 [cited by applicant]
US 20230222344A1 · Chai · 2023 [cited by examiner]
US 20240273345A1 · Bharadwaj · 2024 [cited by examiner]
CN 116226334A · 2023 [cited by examiner]
Bai, Yuntao , et al., “Constitutional AI: Harmlessness from AI Feedback”, arXiv:2212.08073, https://arxiv.org/abs/2212.08073, Dec. 15, 2022. [cited by applicant]
Chen, Lingjiao , et al., “FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance”, arXiv:2305.05176, https://arxiv.org/abs/2305.05176, May 9, 2023. [cited by applicant]
Ding, Dujian , et al., “Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing”, International Conference on Learning Representations, Apr. 22, 2024. [cited by applicant]
Kag, Anil , et al., “Efficient Edge Inference by Selective Query”, International Conference on Learning Representations, published May 1, 2023. [cited by applicant]
Lu, Keming , et al., “Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models”, arXiv:2311.08692, https://arxiv.org/abs/2311.08692, Nov. 25, 2023. [cited by applicant]
Ouyang, Long , et al., “Training language models to follow instructions with human feedback”, arXiv:2203.02155, https://arxiv.org/abs/2203.02155, Mar. 4, 2022. [cited by applicant]
Šakota, Marija , et al., “Fly-Swat or Cannon? Cost-Effective Language Model Choice via Meta-Modeling”, arXiv:2308.06077, https://arxiv.org/abs/2308.06077, Aug. 11, 2023. [cited by applicant]
Shazeer, Noam , et al., “Outrageously Large Neural Networks: the Sparsely-Gated Mixture-of-Experts Layer”, arXiv:1701.06538, https://arxiv.org/abs/1701.06538, Jan. 23, 2017. [cited by applicant]
Wang, Yiding , et al., “Tabi: An Efficient Multi-Level Inference System for Large Language Models”, EuroSys '23: Proceedings of the Eighteenth European Conference on Computer Systems, pp. 233-248, https://doi.org/10.114… [cited by applicant]
Zhao, Xu , et al., “Automatic Model Selection with Large Language Models for Reasoning”, arXiv:2305.14333, https://arxiv.org/abs/2305.14333, May 23, 2023. [cited by applicant]