IP Library › Granted Patent US 12,321,862
Granted Patent B1
US 12,321,862 · App. 18/830,573 · Granted Jun 3, 2025

Latency-, accuracy-, and privacy-sensitive tuning of artificial intelligence model selection parameters and systems and methods of the same

Inventors: Avi Levin (Tel Aviv, IL); Miriam Silver (Tel Aviv, IL); Payal Jain (London, GB); Biraj Krushna Rath (London, GB); Stuart Murray (London, GB); Nimrod Barak (New York, NY)
G06N3/091
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,862
App. No.
18/830,573
Filed
Sep 11, 2024
Granted
Jun 3, 2025
Kind
B1
Art Unit
2499
USPC
706/15
Abstract

The disclosed data generation platform enables generation of an output in response to an output generation request based on tuning a routing model that enables model selection in a dynamic, system-sensitive manner. For example, the disclosed data generation platform receives an output generation request for a user device and generates a risk indicator associated with the output generation request. The platform can determine a current system state and generate a set of performance indicators and associated weighting values based on the risk indicator and the system state. The data generation platform can select a first routing model based on the weighting values. The data generation platform can provide the output generation request to the first routing model to generate an indication of a model with which to generate a model output responsive to the input. The data generation platform can enable access to the generated model output.

Claims (118)

1. A non-transitory computer-readable storage medium comprising instructions thereon, wherein the instructions when executed by at least one data processor of a system, cause the system to:

receive an output generation request for a user device,

wherein the output generation request includes an input for generation of a text-based output and a user identifier associated with the user device;

provide the user identifier and the input to a risk evaluation model to generate a risk indicator associated with (1) the input and (2) a user associated with the user identifier,

wherein the risk indicator includes a composite value associated with (1) a first estimated security risk associated with the input and (2) a second estimated security risk associated with the user;

dynamically monitor one or more system resource measurements to determine a current system state,

wherein the current system state indicates a real-time computational resource usage of a computing ecosystem;

provide the current system state and the risk indicator to a performance parameter determination model to generate a set of performance indicators and associated weighting values,

wherein the performance indicators include performance metrics associated with the computing ecosystem, and

wherein generating the set of performance indicators and associated weighting values comprises:

generating a first weighting value associated with a first performance indicator associated with latency requirements,

generating a second weighting value associated with a second performance indicator associated with accuracy requirements, and

generating a third weighting value associated with a third performance indicator associated with privacy requirements;

provide the set of performance indicators, the associated weighting values, and the input to an evaluation model to identify a first routing model of a set of routing models;

provide the output generation request to the first routing model to generate an indication of a large-language model;

provide the input to the large-language model to generate a model output responsive to the input; and

transmit the model output to a server system enabling access to the generated model output by the user device.

2. The non-transitory computer-readable storage medium of claim 1 , wherein the instructions for generating the risk indicator cause the system to:

determine an input classification associated with the output generation request,

wherein the input classification includes an indication that the input is associated with (1) security information, (2) a high urgency level, or (3) a high accuracy requirement;

in response to determining the input classification associated with the output generation request, provide the input classification to the risk evaluation model to generate an input risk value associated with the input; and

generate the risk indicator according to the first estimated security risk comprising the input risk value.

3. The non-transitory computer-readable storage medium of claim 1 , wherein the instructions for generating the risk indicator cause the system to:

retrieve, from a user activity database, a user activity history associated with the user associated with the user identifier,

wherein the user activity history includes previous output generation requests associated with the user identifier;

provide the user activity history to a user risk determination model to generate a user activity risk value associated with the user; and

generate the risk indicator according to the second estimated security risk comprising the user activity risk value.

4. The non-transitory computer-readable storage medium of claim 1 , wherein the instructions for dynamically monitoring the one or more system resource measurements cause the system to:

monitor the one or more system resource measurements to determine an updated system state, wherein the updated system state indicates real-time computational resource usage of the computing ecosystem during generation of the model output responsive to the input;

determine that a first system resource measurement of the updated system state includes a first value;

compare the first value with a threshold measurement value; and

responsive to comparing the first value with the threshold measurement value, cause termination of the generation of the model output.

5. The non-transitory computer-readable storage medium of claim 1 , wherein the instructions for generating the set of performance indicators and the associated weighting values cause the system to:

detect that the output generation request includes a request for prioritization of the input in a queue of inputs; and

in response to detecting that the output generation request includes the request for prioritization, generate the first weighting value such that the first weighting value is greater than the second weighting value and the third weighting value.

6. The non-transitory computer-readable storage medium of claim 1 ,

wherein the instructions for generating the model output responsive to the input cause the system to:

in response to providing the output generation request to the first routing model, generate routing instructions comprising an input modification indicator signaling whether to generate a modified input;

determine that the routing instructions comprise an indication to apply one or more instructions associated with a protocol to generate the modified input;

responsive to determining that the routing instructions comprise the protocol to generate the modified input, execute the one or more instructions associated with the protocol to generate the modified input,

wherein the protocol modifies at least one attribute of the input to generate the modified input; and

provide the modified input to the large-language model to generate an updated model output responsive to the modified input.

7. A system comprising:

at least one hardware processor; and

at least one non-transitory memory storing instructions, which, when executed by the at least one hardware processor, cause the system to:

receive an output generation request for a user device,

wherein the output generation request includes an input for generation of an output and a user identifier;

provide the user identifier and the input to a risk evaluation model to generate a risk indicator associated with (1) the input and (2) a user associated with the user identifier,

wherein the risk indicator includes a composite value associated with (1) a first estimated security risk associated with the input and (2) a second estimated security risk associated with the user;

dynamically monitor one or more system resource measurements to determine a current system state,

wherein the current system state indicates a real-time computational resource usage of a computing ecosystem;

provide the current system state and the risk indicator to a performance parameter determination model to generate a set of performance indicators and associated weighting values,

wherein the performance indicators include performance metrics associated with the computing ecosystem, and

wherein generating the set of performance indicators and associated weighting values comprises:

generating a first weighting value associated with a first performance indicator associated with latency requirements,

generating a second weighting value associated with a second performance indicator associated with accuracy requirements, and

generating a third weighting value associated with a third performance indicator associated with privacy requirements;

provide the set of performance indicators, the associated weighting values, and the input to an evaluation model to identify a first routing model of a set of routing models;

provide the output generation request to the first routing model to generate an indication of an artificial intelligence model;

provide the input to the artificial intelligence model to generate a model output responsive to the input; and

enable access to the generated model output by the user device.

8. The system of claim 7 , wherein the instructions for generating the risk indicator cause the system to:

determine an input classification associated with the output generation request,

wherein the input classification includes an indication that the input is associated with (1) security information, (2) a high urgency level, or (3) a high accuracy requirement;

in response to determining the input classification associated with the output generation request, provide the input classification to the risk evaluation model to generate an input risk value associated with the input; and

generate the risk indicator according to the first estimated security risk comprising the input risk value.

9. The system of claim 7 , wherein the instructions for generating the risk indicator cause the system to:

retrieve, from a user activity database, a user activity history associated with the user associated with the user identifier;

provide the user activity history to a user risk determination model to generate a user activity risk value associated with the user; and

generate the risk indicator according to the second estimated security risk comprising the user activity risk value.

10. The system of claim 7 , wherein the instructions for dynamically monitoring the one or more system resource measurements cause the system to:

monitor the one or more system resource measurements to determine an updated system state, wherein the updated system state indicates real-time computational resource usage of the computing ecosystem during generation of the model output responsive to the input;

determine that a first system resource measurement of the updated system state includes a first value;

compare the first value with a threshold measurement value; and

responsive to comparing the first value with the threshold measurement value, cause termination of the generation of the model output.

11. The system of claim 7 , wherein the instructions for generating the set of performance indicators and the associated weighting values cause the system to:

detect that the output generation request includes a request for prioritization of the input in a queue of inputs; and

in response to detecting that the output generation request includes the request for prioritization, generate the first weighting value such that the first weighting value is greater than the second weighting value and the third weighting value.

12. The system of claim 7 , wherein the instructions for generating the model output responsive to the input cause the system to:

in response to providing the output generation request to the first routing model, generate routing instructions comprising an input modification indicator signaling whether to generate a modified input;

determine that the routing instructions comprise an indication to apply one or more instructions associated with a protocol to generate the modified input;

responsive to determining that the routing instructions comprise the protocol to generate the modified input, execute the one or more instructions associated with the protocol to generate the modified input,

wherein the protocol modifies at least one attribute of the input to generate the modified input; and

provide the modified input to the artificial intelligence model to generate an updated model output responsive to the modified input.

13. A method comprising:

receiving an output generation request for a user device,

wherein the output generation request includes an input for generation of an output and a user identifier;

providing the user identifier and the input to a risk evaluation model to generate a risk indicator associated with (1) the input and (2) a user associated with the user identifier,

wherein the risk indicator includes a composite value associated with (1) a first estimated security risk associated with the input and (2) a second estimated security risk associated with the user, and

wherein generating the risk indicator comprises:

retrieving, from a user activity database, a user activity history associated with the user associated with the user identifier,

providing the user activity history to a user risk determination model to generate a user activity risk value associated with the user, and

generating the risk indicator according to the second estimated security risk comprising the user activity risk value;

dynamically monitoring one or more system resource measurements to determine a current system state,

wherein the current system state indicates a real-time computational resource usage of a computing ecosystem;

providing the current system state and the risk indicator to a performance parameter determination model to generate a set of performance indicators and associated weighting values,

wherein the performance indicators include performance metrics associated with the computing ecosystem;

providing the set of performance indicators, the associated weighting values, and the input to an evaluation model to identify a first routing model of a set of routing models;

providing the output generation request to the first routing model to generate an indication of an artificial intelligence model;

providing the input to the artificial intelligence model to generate a model output responsive to the input; and

enabling access to the generated model output by the user device.

14. The method of claim 13 , wherein generating the risk indicator comprises:

determining an input classification associated with the output generation request,

wherein the input classification includes an indication that the input is associated with (1) security information, (2) a high urgency level, or (3) a high accuracy requirement;

in response to determining the input classification associated with the output generation request, providing the input classification to the risk evaluation model to generate an input risk value associated with the input; and

generating the risk indicator according to the first estimated security risk comprising the input risk value.

15. The method of claim 13 , wherein dynamically monitoring the one or more system resource measurements comprises:

monitoring the one or more system resource measurements to determine an updated system state, wherein the updated system state indicates real-time computational resource usage of the computing ecosystem during generation of the model output responsive to the input;

determining that a first system resource measurement of the updated system state includes a first value;

comparing the first value with a threshold measurement value; and

responsive to comparing the first value with the threshold measurement value, causing termination of the generation of the model output.

16. The method of claim 13 , wherein generating the set of performance indicators and the associated weighting values comprises:

generating a first weighting value associated with a first performance indicator associated with latency requirements;

generating a second weighting value associated with a second performance indicator associated with accuracy requirements; and

generating a third weighting value associated with a third performance indicator associated with privacy requirements.

17. The method of claim 16 , wherein generating the set of performance indicators and the associated weighting values comprises:

detecting that the output generation request includes a request for prioritization of the input in a queue of inputs; and

in response to detecting that the output generation request includes the request for prioritization, generating the first weighting value such that the first weighting value is greater than the second weighting value and the third weighting value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 8, 2025
From: LEVIN, AVI; SILVER, MIRIAM; JAIN, PAYAL; RATH, BIRAJ KRUSHNA; MURRAY, STUART; BARAK, NIMROD
To: CITIBANK, N.A.
Reel/Frame 071066/0472 →
Continuity (4)
Continuation In Part 18821880 · Aug 30, 2024
Continuation In Part 18661532 · May 10, 2024
Continuation In Part 18661519 · May 10, 2024
Continuation In Part 18633293 · Apr 11, 2024
References Cited (25)
US 9842045B2 · Heorhiadi · 2017 [cited by examiner]
US 11573848B2 · Linck · 2023 [cited by examiner]
US 11656852B2 · Mazurskiy · 2023 [cited by examiner]
US 11750717B2 · Walsh · 2023 [cited by examiner]
US 11875123B1 · Ben David et al. · 2024 [cited by applicant]
US 11875130B1 · Bosnjakovic et al. · 2024 [cited by applicant]
US 11924027B1 · Mysore · 2024 [cited by examiner]
US 11947435B2 · Boulineau · 2024 [cited by examiner]
US 11960515B1 · Pallakonda et al. · 2024 [cited by applicant]
US 11983806B1 · Ramesh · 2024 [cited by examiner]
US 11990139B1 · Sandrew · 2024 [cited by examiner]
US 11995412B1 · Mishra · 2024 [cited by applicant]
US 12001463B1 · Pallakonda et al. · 2024 [cited by applicant]
US 12026599B1 · Lewis et al. · 2024 [cited by applicant]
US 20170262164A1 · Jain et al. · 2017 [cited by applicant]
US 20220311681A1 · Palladino · 2022 [cited by examiner]
US 20220318654A1 · Lin · 2022 [cited by examiner]
US 20220414536A1 · M L · 2022 [cited by examiner]
US 20240020538A1 · Socher et al. · 2024 [cited by applicant]
US 20240095077A1 · Singh · 2024 [cited by examiner]
US 20240129345A1 · Kassam · 2024 [cited by examiner]
WO 2024020416A1 · 2024 [cited by applicant]
Generative machine learning models; IPCCOM000272835D, Aug. 17, 2023. (Year: 2023). [cited by examiner]
Hu, Q., J., et al., “Routerbench: A Benchmark for Multi-LLM Routing System,” arXiv:2403.12031v2 [cs.LG] Mar. 28, 2024, 16 pages. [cited by applicant]
Peers, M., “What California AI Bill Could Mean,” The Briefing, published and retrieved Aug. 30, 2024, 8 pages, https://www.theinformation.com/articles/what-california-ai-bill-could-mean. [cited by applicant]
Cited By (3)
US 12,561,335 US 12,681,999 US 12,694,343