Data Flow Logic for Providing Artificial Intelligence Agents that Automate Multimodal Software Usage
A system for providing artificial intelligence agents that automate software usage includes training servers configured to train agents during training, production servers configured to execute the trained agents during inference, a plurality of training datasets, and data flow logic. The data flow logic is configured to, provide, during the training, the agents and the plurality of training datasets to the training servers to cause the training servers to train the agents on the plurality of training datasets and thereby produce the trained agents, configure the production servers with the trained agents for use during the inference, provide, during the inference, prompts issued by clients to the production servers to cause the production servers to translate the prompts into agent calls to the trained agents that in turn cause the trained agents to generate outputs that are responsive to the prompts, and make the outputs available to the clients.
1 . A system for providing artificial intelligence agents that automate software usage, comprising:
a plurality of training servers configured to train agents during training;
a plurality of production servers configured to execute the trained agents during inference;
a plurality of training datasets, wherein the plurality of datasets includes images of defined agentic trajectories for multimodal interface automation task workflows, each trajectory accompanied by detailed multimodal annotations describing actions and agentic decisions; and
a data flow logic configured to:
during the training, provide the agents and the plurality of training datasets to the training servers to cause the training servers to train the agents on the plurality of training datasets and thereby produce the trained agents;
configure the production servers with the trained agents for use during the inference;
during the inference, provide multimodal interface automation prompts issued by clients to the production servers to cause the production servers to translate the prompts into multimodal interface automation agent calls, wherein each multimodal interface automation agent call specifies interface actions to cause the trained agents to generate outputs that automate multimodal interface workflow in response to the multimodal prompts; and
make the outputs available to the clients.
2 . The system of claim 1 , wherein the data flow logic is further configured to use some training datasets in the plurality of training datasets for pre-training the agents, for post-training the agents, for finisher training the agents, for combined fine-tuning of the agents, and for agentic-fine tuning of the agents.
3 . The system of claim 1 , wherein the data flow logic is further configured to cause the plurality of training servers to periodically retrain the trained agents.
4 . The system of claim 3 , wherein the data flow logic is further configured to periodically reconfigure the production servers with the retrained agents responsive to reliability scores corresponding to multimodal task benchmark.
5 . The system of claim 1 , wherein the agent calls are multimodal interface automation agent calls.
6 . The system of claim 5 , wherein the data flow logic is further configured to periodically configure the clients with agent workflow logics that construct, based on the prompts, agent specifications that are configured to issue the multimodal interface automation agent calls to the trained agents.
7 . The system of claim 1 , wherein the plurality of training datasets includes:
a first training dataset including documents containing text interleaved with images;
a second training dataset including text embedded in images;
a third training dataset including recorded videos of software usage;
a fourth training dataset including portable document format (PDF) documents;
a fifth training dataset including recorded videos of software tool usage trajectories, wherein;
a sixth training dataset including images of open-domain web pages;
a seventh training dataset including images of specific-domain web pages; and/or
an eighth training dataset including images of agentic trajectories of the agent performing interface automation task workflows.
8 . The system of claim 7 , wherein images in the recorded videos of software tool usage trajectories are interleaved with text descriptions of tasks executed in the recorded videos through the software tool usage trajectories.
9 . The system of claim 8 , wherein the images in the recorded videos of software tool usage trajectories are further interleaved with text descriptions of actions executed on the images and image annotations resulting from execution of the actions.
10 . The system of claim 9 , wherein the actions include clicking, scrolling, and typing.
11 . The system of claim 7 , wherein the images of open-domain web pages are automatically crawled.
12 . The system of claim 11 , wherein the open-domain web pages are multimodal web pages.
13 . The system of claim 11 , wherein the open-domain web pages are part of software tools.
14 . The system of claim 11 , wherein the images of open-domain web pages are interleaved with text descriptions of synthetic tasks and image annotations resulting from execution of the synthetic tasks.
15 . The system of claim 14 , wherein the synthetic tasks include website-wise tasks, element-wise tasks, and action-wise tasks.
16 . The system of claim 15 , wherein the website-wise tasks include heading optical character recognition (OCR), captioning, and web question answering (WebQA).
17 . The system of claim 15 , wherein the element-wise tasks include element optical character recognition (OCR), element grounding/localization, and key-value pair identification.
18 . The system of claim 15 , wherein the action-wise tasks include action grounding and action prediction.
19 . A computer-implemented method for providing artificial intelligence agents that automate software usage, the computer-implemented method comprising:
training, with a plurality of training servers, agents;
executing, with a plurality of production servers, the trained agents during inference;
providing, during the training, the agents and a plurality of training datasets to the training servers to cause the training servers to train the agents on the plurality of training datasets and thereby produce the trained agents, wherein the plurality of datasets includes images of defined agentic trajectories for multimodal interface automation task workflows having each trajectory accompanied by detailed multimodal annotations describing actions and agentic decisions;
configuring the production servers with the trained agents for use during the inference;
providing, during the inference, multimodal interface automation prompts issued by clients to the production servers to cause the production servers to translate the prompts into multimodal interface automation agent calls, wherein each multimodal interface automation agent call specifies interface actions to cause the trained agents to generate outputs that automate multimodal interface workflow in response to the multimodal prompts; and
making the outputs available to the clients.
20 . A non-transitory computer readable storage medium impressed with computer program instructions for providing artificial intelligence agents that automate software usage, the instructions, when executed on a processor, implement a method comprising:
training, with a plurality of training servers, agents;
executing, with a plurality of production servers, the trained agents during inference;
providing, during the training, the agents and a plurality of training datasets to the training servers to cause the training servers to train the agents on the plurality of training datasets and thereby produce the trained agents, wherein the plurality of datasets includes images of defined agentic trajectories for multimodal interface automation task workflows having each trajectory accompanied by detailed multimodal annotations describing actions and agentic decisions;
configuring the production servers with the trained agents for use during the inference;
providing, during the inference, multimodal interface automation prompts issued by clients to the production servers to cause the production servers to translate the prompts into multimodal interface automation agent calls, wherein each multimodal interface automation agent call specifies interface actions to cause the trained agents to generate outputs that automate multimodal interface workflow in response to the multimodal prompts; and
making the outputs available to the clients.