SYSTEMS AND METHODS FOR MONITORING AND EVALUATING LANGUAGE PERFORMANCE ACCORDING TO REAL-TIME DATA IN A DISTRIBUTED NETWORKING ENVIRONMENT
Described herein are systems and methods for monitoring and evaluating language performance according to real-time data in a distributed networking environment. The system can maintain an evaluation dataset for language models. The evaluation dataset can include an evaluation example. The evaluation example can include a respective input prompt, indicating an intent relating to wager opportunities, and a respective output message, identifying information associated with the intent and a corresponding wager recommendation. The system can generate candidate outputs using input prompts and the respective input prompt. The system can determine evaluation scores for the language models based on the respective output message of the evaluation example and the candidate outputs, with a first score corresponding to a first language model. The system can update, based on the first score satisfying an assignment criterion, a data structure to assign the first language model to input prompts identifying the intent.
1 . A system, comprising:
one or more processors coupled to non-transitory memory, the one or more processors configured to:
maintain an evaluation dataset for a plurality of language models, the evaluation dataset comprising a first evaluation example including:
(i) a respective input prompt indicating an intent relating to wager opportunities, and
(ii) a respective output message identifying information associated with the intent and a corresponding wager recommendation;
generate, using the plurality of language models, a plurality of candidate outputs using a plurality of input prompts and the respective input prompt of the first input example;
determine a plurality of evaluation scores for the plurality of language models based on the respective output message of the first evaluation example and the plurality of candidate outputs, a first score of the plurality of scores corresponding to a first language model of the plurality of language models; and
update, based on the first score satisfying an assignment criterion, a data structure to assign the first language model to input prompts identifying the intent.
2 . The system of claim 1 , wherein the one or more processors are further configured to:
receive, from a client device, a prompt corresponding to the intent relating to wager opportunities;
select the first language model of the plurality of language models based on the intent of the prompt; and
generate, using the prompt and the first language model, output identifying at least one wager recommendation corresponding to the prompt.
3 . The system of claim 1 , wherein the one or more processors are further configured to:
determine the plurality of evaluation scores based on a semantic similarity between the respective output message of the first evaluation example and the plurality of candidate outputs.
4 . The system of claim 1 , wherein the respective input prompt of the first evaluation example comprises a plurality of wager recommendations and the respective output message comprises a first wager recommendation of the plurality of wager recommendations.
5 . The system of claim 4 , wherein the one or more processors are further configured to:
maintain a plurality of historical wager opportunities; and
generate the plurality of wager recommendations based on a search operation of the plurality of wager opportunities using the respective input prompt.
6 . The system of claim 4 , wherein the one or more processors are further configured to:
determine a first evaluation score of the plurality of evaluation scores corresponding to the first language model based on a comparison of the first wager recommendation and a corresponding wager recommendation included in a respective candidate output generated by the first language model.
7 . The system of claim 1 , wherein the one or more processors are further configured to:
determine a second plurality of evaluation scores for the plurality of language models using a second evaluation example corresponding to a second intent; and
assign a second language model of the plurality of language models to prompts corresponding to the second intent.
8 . The system of claim 1 , wherein the intent identifies one or more of a wager type, a live event type, a team identifier, or an athlete identifier.
9 . The system of claim 1 , wherein the one or more processors are further configured to:
generate a plurality of candidate outputs using a combination of the plurality of input prompts and the plurality of historical wager opportunities; and
determine the plurality of evaluation scores based on the respective output message of a third evaluation example and the plurality of candidate outputs generated using the combination.
10 . The system of claim 1 , wherein the one or more processors are further configured to:
maintain a player profile associated with a client device; and
generate the plurality of candidate outputs based on the player profile and the respective input prompt.
11 . A method, comprising:
maintaining, by one or more processors coupled to non-transitory memory, an evaluation dataset for a plurality of language models, the evaluation dataset comprising a first evaluation example including:
(i) a respective input prompt indicating an intent relating to wager opportunities, and
(ii) a respective output message identifying information associated with the intent and a corresponding wager recommendation; and
generating, by the one or more processors, a plurality of candidate outputs using a plurality of input prompts and the respective input prompt of the first input example;
determining, by the one or more processors, a plurality of evaluation scores for the plurality of language models based on the respective output message of the first evaluation example and the plurality of candidate outputs, a first score of the plurality of scores corresponding to a first language model of the plurality of language models; and
updating, by the one or more processors, based on the first score satisfying an assignment criterion, a data structure to assign the first language model to input prompts identifying the intent.
12 . The method of claim 11 , further comprising:
receiving, by the one or more processors, from a client device, a prompt corresponding to the intent relating to wager opportunities;
selecting, by the one or more processors, the first language model of the plurality of language models based on the intent of the prompt; and
generating, by the one or more processors, using the prompt and the first language model, output identifying at least one wager recommendation corresponding to the prompt.
13 . The method of claim 11 , further comprising:
determining, by the one or more processors, the plurality of evaluation scores based on a semantic similarity between the respective output message of the first evaluation example and the plurality of candidate outputs.
14 . The method of claim 11 , wherein the respective input prompt of the first evaluation example comprises a plurality of wager recommendations and the respective output message comprises a first wager recommendation of the plurality of wager recommendations.
15 . The method of claim 14 , further comprising:
maintaining, by the one or more processors, a plurality of historical wager opportunities; and
generating, by the one or more processors, the plurality of wager recommendations based on a search operation of the plurality of wager opportunities using the respective input prompt.
16 . The method of claim 14 , further comprising:
determining, by the one or more processors, a first evaluation score of the plurality of evaluation scores corresponding to the first language model based on a comparison of the first wager recommendation and a corresponding wager recommendation included in a respective candidate output generated by the first language model.
17 . The method of claim 11 , further comprising:
determining, by the one or more processors, a second plurality of evaluation scores for the plurality of language models using a second evaluation example corresponding to a second intent; and
assigning, by the one or more processors, a second language model of the plurality of language models to prompts corresponding to the second intent.
18 . The method of claim 11 , wherein the intent identifies one or more of a wager type, a live event type, a team identifier, or an athlete identifier.
19 . The method of claim 11 , further comprising:
generating, by the one or more processors, a plurality of candidate outputs using a combination of the plurality of input prompts and the plurality of historical wager opportunities; and
determining, by the one or more processors, the plurality of evaluation scores based on the respective output message of a third evaluation example and the plurality of candidate outputs generated using the combination.
20 . The method of claim 11 , further comprising:
maintaining, by the one or more processors, a player profile associated with a client device; and
generating, by the one or more processors, the plurality of candidate outputs based on the player profile and the respective input prompt.