Streamlined evaluation of simultaneously executed artificial intelligence agents
Systems and methods are described for comparing execution of two or more artificial intelligence (AI) AI agents. A platform can provide a user interface (UI) that allows for selection or creation of multiple AI agents. The AI agents can utilize different agent objects, such as different prompts, datasets, or models. The AI agents can be displayed on a single UI screen, where execution of the AI agents is simultaneously simulated. The same inputs can be provided to the multiple AI agents, and the corresponding outputs can display on screen. The platform can also vectorize and compare the semantic similarity of the outputs, presenting an indication of the semantic similarity on the same UI screen.
1 . A method of comparatively evaluating performance of artificial intelligence (“AI”) agents, comprising:
receiving, at a user interface (“UI”), a selection of a first AI agent;
causing simultaneous display, on a single screen of the UI (“UI screen”), of the first AI agent and a second AI agent, the first AI agent comprising first multiple agent objects, and the second AI agent comprising second multiple agent objects, the second AI agent differing from the first AI agent by at least one agent object, wherein the first multiple agent objects and the second multiple agent objects are simultaneously displayed in an agent topology;
simultaneously executing the first AI agent and the second AI agent with a same series of queries, wherein an agent executor executes each of the first multiple agent objects according to a first manifest file of the first AI agent and each of the second multiple agent objects according to a second manifest file of the second AI agent, wherein executing the first AI agent generates first outputs responsive to the same series of queries and executing the second AI agent generates second outputs responsive to the same series of queries;
collecting, by the agent executor, first execution metrics from each of the executed first multiple agent objects and second execution metrics from each of the executed second multiple agent objects; and
causing simultaneous display, on the UI screen, of the first execution metrics, the first multiple agent objects, the second execution metrics, and the second multiple agent objects, wherein each first execution metric of the first execution metrics is displayed with the respective agent object from which the first execution metric was collected in a first column of the UI screen and each second execution metric of the second execution metrics is displayed with the respective agent object from which the second execution metric was collected in a second column of the UI screen.
2 . The method of claim 1 , wherein the first and second execution metrics comprise performance metrics and similarity metrics.
3 . The method of claim 2 , wherein the performance metrics comprise latency or response time.
4 . The method of claim 2 , wherein the performance metrics comprise token usage.
5 . The method of claim 2 , wherein the performance metrics comprise graphical processing unit cost.
6 . The method of claim 2 , wherein the performance metrics comprise telemetry data from hosting environments for the first multiple agent objects and the second multiple agent objects.
7 . The method of claim 2 , wherein the similarity metrics comprise a value representing semantic similarity of vectorized outputs between respective agent objects of the first AI agent and the second AI agent.
8 . The method of claim 1 , wherein a different hosting region is selected for the first AI agent than for the second AI agent.
9 . The method of claim 1 , wherein the second AI agent differs from the first AI agent based on a hosting type, wherein the hosting type comprises one of user environment, AI platform hosted, and hyperscaler hosted.
10 . The method of claim 1 , wherein the simultaneous display of the first and second execution metrics identifies variations in execution metrics of corresponding agent objects across the first and second AI agents.
11 . The method of claim 1 , wherein the at least one agent object comprises a first dataset, a first AI model, or a first prompt package.
12 . The method of claim 1 , further comprising displaying a first total of the first execution metrics for the first AI agent, and a second total of the second execution metrics for the second AI agent.
13 . The method of claim 1 , wherein historical metrics are presented on the UI screen for comparison with the first and second execution metrics.
14 . The method of claim 1 , wherein the first and second execution metrics further comprise an execution duration for each of the first and second multiple agent objects, including a start time and an end time, and the UI screen displays each of the execution durations in association with the respective agent object from which it was collected.
15 . The method of claim 1 , wherein the UI screen displays a visual indicator of a divergence threshold being exceeded by comparing execution metrics of corresponding agent objects of the first AI agent and the second AI agent.
16 . A system for comparatively evaluating performance of artificial intelligence (“AI”) agents, comprising:
a memory storage including a non-transitory, computer-readable medium comprising instructions; and
at least one hardware-based processor that executes the instructions to carry out stages comprising:
receiving, at a user interface (“UI”), a selection of a first AI agent;
causing simultaneous display, on a single screen of the UI (“UI screen”), of the first AI agent and a second AI agent, the first AI agent comprising first multiple agent objects, and the second AI agent comprising second multiple agent objects, the second AI agent differing from the first AI agent by at least one agent object, wherein the first multiple agent objects and the second multiple agent objects are simultaneously displayed in an agent topology;
simultaneously executing the first AI agent and the second AI agent with a same series of queries, wherein an agent executor executes each of the first multiple agent objects according to a first manifest file of the first AI agent and each of the second multiple agent objects according to a second manifest file of the second AI agent, wherein the first AI agent generates first outputs responsive to the same series of queries and executing the second AI agent generates second outputs responsive to the same series of queries;
collecting, by the agent executor, first performance metrics from each of the executed first multiple agent objects and second performance metrics from each of the executed second multiple agent objects; and
causing simultaneous display, on the UI screen, of the first performance metrics, the first multiple agent objects, the second performance metrics, and the second multiple agent objects, wherein each first performance metric of the first performance metrics is displayed with the respective agent object from which the first performance metric was collected in a first column of the UI screen and each second performance metric of the second performance metrics is displayed with the respective agent object from which the second performance metric was collected in a second column of the UI screen.
17 . A non-transitory, computer-readable medium having instructions for comparatively evaluating performance of artificial intelligence (“AI”) agents, the instructions, when executed by a processor, causing the processor to perform stages comprising:
receiving, at a user interface (“UI”), a selection of a first AI agent;
causing simultaneous display, on a single screen of the UI (“UI screen”), of the first AI agent and a second AI agent, the first AI agent comprising first multiple agent objects, and the second AI agent comprising second multiple agent objects, the second AI agent differing from the first AI agent by at least one agent object, wherein the first multiple agent objects and the second multiple agent objects are simultaneously displayed in an agent topology;
simultaneously executing the first AI agent and the second AI agent with a same series of queries, wherein an agent executor executes each of the first multiple agent objects according to a first manifest file of the first AI agent and each of the second multiple agent objects according to a second manifest file of the second AI agent, wherein executing the first AI agent generates first outputs responsive to the same series of queries and executing the second AI agent generates second outputs responsive to the same series of queries;
collecting, by the agent executor, first performance metrics from each of the executed first multiple agent objects and second performance metrics from each of the executed second multiple agent objects; and
causing simultaneous display, on the UI screen, of the first performance metrics, the first multiple agent objects, the second performance metrics, and the second multiple agent objects, wherein each first performance metric of the first performance metrics is displayed with the respective agent object from which the first performance metric was collected in a first column of the UI screen and each second performance metric of the second performance metrics is displayed with the respective agent object from which the second performance metric was collected in a second column of the UI screen.