IP Library Granted Patent US 12,585,862
Granted Patent B2
US 12,585,862 · App. 18/909,470 · Granted Mar 24, 2026

Training data for training artificial intelligence agents to automate multimodal software usage

Inventors: Sagnak Tasirlar (San Francisco, CA); David Abrahams (San Francisco, CA); Lina Lukyantseva (San Francisco, CA); Erich Elsen (San Francisco, CA); Maxwell Nye (San Francisco, CA); Augustus Odena (San Francisco, CA); Rohan Bavishi (San Francisco, CA); Vibhaa Sivaraman (San Francisco, CA); Adam Hoff (San Francisco, CA); Teddy Rothschild (San Francisco, CA); Shaya Zarkesh (San Francisco, CA); Deepak Moparthi (San Francisco, CA); Jacob van Gogh (San Francisco, CA); Claire Pajot (San Francisco, CA); Curtis Hawthorne (San Francisco, CA); Matt Elkherj (San Francisco, CA); Warut Vijitbenjaronk (San Francisco, CA); Arushi Somani (San Francisco, CA); Johnny Lee (San Francisco, CA); Joe Gershenson (San Francisco, CA); Jordyn Shuell (San Francisco, CA); Danielle Perszyk (San Francisco, CA)
Assignee: Anthropic, PBC
G06F40/166G06F3/0481G06F3/0484G06F16/951G06F40/174G06F40/284G06N3/0455G06N3/091G06N5/04G06N20/00G06V10/7715G06V10/774G06V10/803G06V10/82G06V20/40G06V30/19147G06V30/41G06F9/451
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,862
App. No.
18/909,470
Granted
Mar 24, 2026
Kind
B2
Abstract

A system for automating software usage includes an agent configured to automate. The agent is trained on one or more training data sets. The one or more training datasets include one or more of a first training dataset including documents containing text interleaved with images, a second training dataset including text embedded in images, a third training dataset including recorded videos of software usage, a fourth training dataset including portable document format (PDF) documents, a fifth training dataset including recorded videos of software tool usage trajectories, a sixth training dataset including images of open-domain web pages, a seventh training dataset including images of specific-domain web pages, and/or an eighth training dataset including images of agentic trajectories of the agent performing interface automation task workflows.

Claims (34)

1 . A system for automating actions on a software user interface, comprising:

a processor; and

a memory storing instructions which, when executed, cause the processor to perform steps comprising:

constructing an artificial intelligence agent based on a Transformer based vision language model configured to automate actions on a software user interface, including training the Transformer based vision language model on:

a first training dataset including documents containing text interleaved with images;

a second training dataset including text embedded in images;

a third training dataset including recorded videos of software usage;

a fourth training dataset including portable document format (PDF) documents;

a fifth training dataset including recorded videos of software tool usage trajectories;

a sixth training dataset including images of open-domain web pages;

a seventh training dataset including images of specific-domain web pages; and

an eighth training dataset including images of agentic trajectories of the agent performing interface automation task workflows;

receiving, via a user interface, a natural language input indicating a user intent to interact with the software user interface;

generating, by the Transformer based vision language model, a workflow indicating a past action and a next action relating to the user intent based on an input comprising a screenshot of the software user interface and the natural language input;

executing, by the artificial intelligence agent, at least the next action in a form of a machine-actuated command on the software user interface thereby resulting in a change in the software user interface; and

iteratively generating, by the Transformer based vision language model, next machine-actuated commands based on a screenshot of a changed software user interface; and

iteratively executing, by the artificial intelligence agent, the next machine-actuated commands thereby causing a task described by the natural language input to be completed by interacting with the software user interface.

2 . The system of claim 1 , wherein images in the recorded videos of software tool usage trajectories are interleaved with text descriptions of tasks executed in the recorded videos through the software tool usage trajectories.

3 . The system of claim 2 , wherein the images in the recorded videos of software tool usage trajectories are further interleaved with text descriptions of actions executed on the images and image annotations resulting from execution of the actions.

4 . The system of claim 3 , wherein the actions include clicking, scrolling, and typing.

5 . The system of claim 1 , wherein the images of open-domain web pages are automatically crawled.

6 . The system of claim 5 , wherein the open-domain web pages are multimodal web pages.

7 . The system of claim 5 , wherein the open-domain web pages are part of software tools.

8 . The system of claim 5 , wherein the images of open-domain web pages are interleaved with text descriptions of synthetic tasks and image annotations resulting from execution of the synthetic tasks.

9 . The system of claim 8 , wherein the synthetic tasks include website-wise tasks, element-wise tasks, and action-wise tasks.

10 . The system of claim 9 , wherein the website-wise tasks include heading optical character recognition (OCR), captioning, and web question answering (WebQA).

11 . The system of claim 9 , wherein the element-wise tasks include element optical character recognition (OCR), element grounding/localization, and key-value pair identification.

12 . The system of claim 9 , wherein the action-wise tasks include action grounding and action prediction.

13 . The system of claim 8 , wherein the images of open-domain web pages are further interleaved with unified resource locators (URLs) of the open-domain web pages.

14 . The system of claim 1 , wherein the specific-domain web pages are multimodal web pages.

15 . The system of claim 14 , wherein the specific-domain web pages are part of software tools.

16 . The system of claim 15 , wherein the images of specific-domain web pages are curated by the software tools.

17 . The system of claim 14 , wherein the images of specific-domain web pages are interleaved with text descriptions of synthetic tasks and image annotations resulting from execution of the synthetic tasks.

18 . The system of claim 17 , wherein the synthetic tasks include website-wise tasks, element-wise tasks, and action-wise tasks.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE NAME OF THE ASSIGNEE TO ANTHROPIC, PBC PREVIOUSLY RECORDED ON REEL 70785 FRAME 275. ASSIGNOR(S) HEREBY CONFIRMS THE THE ASSIGNMENT.. Recorded Apr 17, 2025
From: ADEPT AL LABS INC.
To: ANTHROPIC, PBC
Reel/Frame 071101/0374 →
CHANGE OF NAME Recorded Apr 12, 2025
From: PERSIMMON AI LABS, INC.
To: ADEPT AI LABS, INC.
Reel/Frame 070825/0499 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2025
From: ADEPT AL LABS INC.
To: ANTHROPIC, PBNC
Reel/Frame 070785/0275 →
Continuity (9)
Provisional Application 63638613 · Apr 25, 2024
Provisional Application 63638631 · Apr 25, 2024
Provisional Application 63638644 · Apr 25, 2024
Provisional Application 63567667 · Mar 20, 2024
Provisional Application 63567681 · Mar 20, 2024
Provisional Application 63567698 · Mar 20, 2024
Provisional Application 63567721 · Mar 20, 2024
Provisional Application 63567714 · Mar 20, 2024
Related Publication 20250299510A1 · Sep 25, 2025
References Cited (135)
US 6012030A · French-St. George et al. · 2000 [cited by applicant]
US 6226785B1 · Peterson et al. · 2001 [cited by applicant]
US 6859451B1 · Pasternack et al. · 2005 [cited by applicant]
US 8185544B2 · Oztekin et al. · 2012 [cited by applicant]
US 8493406B2 · Rubin et al. · 2013 [cited by applicant]
US 8855684B2 · Bellver et al. · 2014 [cited by applicant]
US 9218128B1 · Yuschik et al. · 2015 [cited by applicant]
US 9269048B1 · Chen · 2016 [cited by applicant]
US 10257225B1 · Sites · 2019 [cited by applicant]
US 10587708B2 · Laird-McConnell et al. · 2020 [cited by applicant]
US 11077320B1 · Hibbard · 2021 [cited by applicant]
US 11645564B2 · Wu et al. · 2023 [cited by applicant]
US 11809887B2 · Hinton et al. · 2023 [cited by applicant]
US 11907864B2 · Wu et al. · 2024 [cited by applicant]
US 12293272B1 · Poulis et al. · 2025 [cited by applicant]
US 20020062475A1 · Iborra et al. · 2002 [cited by applicant]
US 20030217054A1 · Bachman et al. · 2003 [cited by applicant]
US 20040054690A1 · Hillerbrand et al. · 2004 [cited by applicant]
US 20040078787A1 · Borek et al. · 2004 [cited by applicant]
US 20040215665A1 · Edgar et al. · 2004 [cited by applicant]
US 20050010418A1 · McNair et al. · 2005 [cited by applicant]
US 20060155954A1 · Haynes et al. · 2006 [cited by applicant]
US 20060161878A1 · Koh et al. · 2006 [cited by applicant]
US 20070233495A1 · Agapi et al. · 2007 [cited by applicant]
US 20080065388A1 · Cross et al. · 2008 [cited by applicant]
US 20080065390A1 · Ativanichayaphong et al. · 2008 [cited by applicant]
US 20080065453A1 · Settuducati · 2008 [cited by applicant]
US 20080118051A1 · Odinak et al. · 2008 [cited by applicant]
US 20080228494A1 · Cross · 2008 [cited by applicant]
US 20080301094A1 · Zhu et al. · 2008 [cited by applicant]
US 20110041140A1 · Harm et al. · 2011 [cited by applicant]
US 20130080641A1 · Lui et al. · 2013 [cited by applicant]
US 20130144682A1 · Dhara et al. · 2013 [cited by applicant]
US 20130226892A1 · Ehsani et al. · 2013 [cited by applicant]
US 20130268260A1 · Lundberg et al. · 2013 [cited by applicant]
US 20140157288A1 · Wong · 2014 [cited by applicant]
US 20140214404A1 · Kalia et al. · 2014 [cited by applicant]
US 20150339712A1 · Koutrika et al. · 2015 [cited by applicant]
US 20160162172A1 · Rathod · 2016 [cited by applicant]
US 20160335331A1 · Schnase et al. · 2016 [cited by applicant]
US 20170048170A1 · Smullen et al. · 2017 [cited by applicant]
US 20170091178A1 · Barbosa et al. · 2017 [cited by applicant]
US 20170255580A1 · Parekh et al. · 2017 [cited by applicant]
US 20170289305A1 · Liensberger et al. · 2017 [cited by applicant]
US 20180012141A1 · Chehreghani et al. · 2018 [cited by applicant]
US 20180060744A1 · Achin et al. · 2018 [cited by applicant]
US 20180137431A1 · Goldfarb et al. · 2018 [cited by applicant]
US 20180157739A1 · Khaitan et al. · 2018 [cited by applicant]
US 20180314943A1 · Liang et al. · 2018 [cited by applicant]
US 20190171984A1 · Irimie et al. · 2019 [cited by applicant]
US 20190187987A1 · Fauchère et al. · 2019 [cited by applicant]
US 20190332686A1 · Lee · 2019 [cited by applicant]
US 20190384807A1 · Dernoncourt et al. · 2019 [cited by applicant]
US 20200342316A1 · Shazeer et al. · 2020 [cited by applicant]
US 20210232992A1 · Disterheft et al. · 2021 [cited by applicant]
US 20220046129A1 · Clodore et al. · 2022 [cited by applicant]
US 20220048525A1 · Tsai et al. · 2022 [cited by applicant]
US 20220051219A1 · Sells et al. · 2022 [cited by applicant]
US 20220058981A1 · Neumann · 2022 [cited by applicant]
US 20220130013A1 · Pottorff et al. · 2022 [cited by applicant]
US 20220215262A1 · Narayanaswami et al. · 2022 [cited by applicant]
US 20220246257A1 · Priestas et al. · 2022 [cited by applicant]
US 20220291966A1 · Masood et al. · 2022 [cited by applicant]
US 20230031702A1 · Li et al. · 2023 [cited by applicant]
US 20230106716A1 · Xiong et al. · 2023 [cited by applicant]
US 20230156075A1 · Brewer et al. · 2023 [cited by applicant]
US 20230206913A1 · Vempaty et al. · 2023 [cited by applicant]
US 20230222285A1 · Zhang et al. · 2023 [cited by applicant]
US 20230222623A1 · Ke et al. · 2023 [cited by applicant]
US 20230281400A1 · Wang et al. · 2023 [cited by applicant]
US 20230306205A1 · Maeder et al. · 2023 [cited by applicant]
US 20230325693A1 · Wu et al. · 2023 [cited by applicant]
US 20230342167A1 · Radkoff et al. · 2023 [cited by applicant]
US 20230351149A1 · Yu et al. · 2023 [cited by applicant]
US 20230360388A1 · Singh · 2023 [cited by applicant]
US 20230386025A1 · Loddenkemper et al. · 2023 [cited by applicant]
US 20230419652A1 · Tiong et al. · 2023 [cited by applicant]
US 20240119257A1 · Guo et al. · 2024 [cited by applicant]
US 20240135232A1 · Betteridge et al. · 2024 [cited by applicant]
US 20240256835A1 · Dehghani et al. · 2024 [cited by applicant]
US 20240281472A1 · LaRhette et al. · 2024 [cited by applicant]
US 20240282084A1 · Serra et al. · 2024 [cited by applicant]
US 20240282094A1 · Tsimpoukelli et al. · 2024 [cited by applicant]
US 20240290065A1 · Park et al. · 2024 [cited by applicant]
US 20240303443A1 · Cheng et al. · 2024 [cited by applicant]
US 20240329943A1 · Sharma et al. · 2024 [cited by applicant]
US 20240338234A1 · Li · 2024 [cited by applicant]
US 20240362272A1 · Lee et al. · 2024 [cited by applicant]
US 20240370765A1 · Pierucci et al. · 2024 [cited by applicant]
US 20240404238A1 · Yu et al. · 2024 [cited by applicant]
US 20240412720A1 · Vasylyev · 2024 [cited by applicant]
US 20250028759A1 · Castillo et al. · 2025 [cited by applicant]
US 20250217170A1 · Shaw et al. · 2025 [cited by applicant]
US 20250245030A1 · Cyjon et al. · 2025 [cited by applicant]
US 20250272350A1 · Furuta et al. · 2025 [cited by applicant]
US 20250272506A1 · Sassak, Jr. et al. · 2025 [cited by applicant]
WO 2024049607A2 · 2024 [cited by applicant]
WO 2024146961A1 · 2024 [cited by applicant]
Deng, Xiang, et al. “Mind2web: Towards a generalist agent for the web.” Advances in Neural Information Processing Systems 36 (2023): 28091-28114. (Year: 2023). [cited by examiner]
Gur, Izzeddin, et al. “A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.” ICLR. 2024. (Year: 2024). [cited by examiner]
He, Hongliang, et al. “WebVoyager: Building an end-to-end web agent with large multimodal models.” arXiv preprint arXiv:2401.13919 (2024). (Year: 2024). [cited by examiner]
Koh, Jing Yu, et al. “Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.” arXiv preprint arXiv:2401.13649 (2024). (Year: 2024). [cited by examiner]
Liu, Junpeng, et al. “Visualwebbench: How far have multimodal Ilms evolved in web page understanding and grounding ?.” arXiv preprint arXiv:2404.05955 (2024). (Year: 2024). [cited by examiner]
Lu, Xing Han, Zdenek Kasner, and Siva Reddy. “Weblinx: Real-world website navigation with multi-turn dialogue.” arXiv preprint arXiv:2402.05930 (2024). (Year: 2024). [cited by examiner]
Zhang, Saizheng, et al. “Personalizing dialogue agents: I have a dog, do you have pets too ?.” arXiv preprint arXiv: 1801.07243 (2018). (Year: 2018). [cited by examiner]
Li et al, Demonstration+ Natural Language: Multi modal Interfaces for GUI-Based Interactive Task Learning Agents (Year: 2021) 43 pages. [cited by applicant]
Sethi, Pooja, et al. “Autonlu: Detecting, root-causing, and fixing nlu model errors.” arXiv preprint arXiv:2110.06384 (2021). (Year:2021) 10 pages. [cited by applicant]
Takebayashi et al., Multi modal Interface Agent for Enhancing Knowledge Sharing (Year: 1997) 4 pages. [cited by applicant]
U.S. Appl. No. 18/908,447 Non Final Office Action dated Dec. 13, 2024, 27 pages. [cited by applicant]
U.S. Appl. No. 18/909,068 Non-final Office Action dated Dec. 19, 2024, 34 pages. [cited by applicant]
U.S. Appl. No. 18/909,186 Non-final Rejection dated Dec. 9, 2024, 27 pages. [cited by applicant]
U.S. Appl. No. 18/909,455 Non-final Office Action dated Dec. 19, 2024, 30 pages. [cited by applicant]
U.S. Appl. No. 18/909,531 Non-final Rejection dated Jan. 3, 2025, 110 pages. [cited by applicant]
U.S. Appl. No. 18/909,588 Non-final Office Action dated Dec. 4, 2024, 41 pages. [cited by applicant]
Walker et al. “Neural semantic parsing with anonymization for command understanding in general-purpose service robots.” Robot World Cup. Cham: Springer International Publishing, 2019. 337-350. (Year: 2019) 14 pages. [cited by applicant]
Xie et al., OpenAgents: An Open Platform for Language Agents in the Wild, (Year: 2023) 34 pages. [cited by applicant]
Yin, Pengcheng. Learning Structured Neural Semantic Parsers. Diss. Carnegie Mellon University, 2021. (Year: 2021) 189 pages. [cited by applicant]
Zhou, Shuyan, et al. “Webarena: A realistic web environment for building autonomous agents.” arXiv preprint arXiv:2307.13854 (2023). (Year: 2023) 22 pages. [cited by applicant]
Adept Product Team, ‘Building Powerful Agents with Adept’, Aug. 23, 2024, 12 pages. [cited by applicant]
Adept Team, “Adept Fuyu-Heavy: A new multimodal model”, Jan. 24, 2024, 11 pages. [cited by applicant]
Erich Elsen, Augustus Odena, Maxwell Nye, Saǧnak Taşirlar, Tri Dao, Curtis Hawthorne, Deepak Moparthi, Arushi Somani, “Releasing Persimmon-8B”, Sep. 7, 2023, 7 pages. [cited by applicant]
Erich Elsen, Curtis Hawthorne, Arushi Somani, “The Adventure of the Errant Hardware”, Sep. 19, 2023, 14 pages. [cited by applicant]
Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, Saǧnak Taşirlar, “Fuyu-8B: A Multimodal Architecture for AI Agents”, Oct. 17, 2023, 22 pages. [cited by applicant]
Tri Dao, “FlashAttention: Fast Transformer training with long sequences”, Jan. 17, 2023, 9 pages. [cited by applicant]
International Search Report and Written Opinion mailed Jul. 1, 2025 in PCT Application No. PCT/US2025/020719 filed Mar. 20, 2025. [cited by applicant]
Lee, Yi-Lun, et al., “Multimodal Prompting with Missing Modalities for Visual Recognition,” 2023, 10 pages. [cited by applicant]
Moran, Douglas B., et al. “Multimodal User Interfaces in the Open Agent Architecture” 1997, 8 pages. [cited by applicant]
Song, Kisub, et al., “Generating multimodal user interfaces for Web services.” 2008, 11 pages. [cited by applicant]
Chen, Weihao, et al. “Miwa: Mixed-initiative web automation for better user control and confidence.” Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 2023. (Year: 2023). [cited by applicant]
Ortiz, Jose Javier Gonzalez, John Guttag, and Adrian Dalca. “Magnitude invariant parametrizations improve hypernetwork learning.” arXiv preprint arXiv:2304.07645 (2023). (Year: 2023). [cited by applicant]
Waibel, A., et al., “Multimodal Interfaces for Multimedia Information Agents,” 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing, Munich, Germany, 1997, pp. 167-170 vol. 1, doi: 10.1109/ICAS… [cited by applicant]
Zheng, L., et al., “AgentStudio: A Toolkit for Building General Virtual Agents,” Mar. 26, 2024, arXiv:2403.17918V1 [cs.AI], 12 pages. [cited by applicant]
Furuta, H., et al. “Multimodal Web Navigation with Instruction-Finetuned Foundation Models,” Feb. 25, 2024, arXiv:2305.11854v4 [cs.LG], 31 pages. [cited by applicant]
“Selenium Documentation Release 1.0,” Aug. 26, 2012, 201 pages. [cited by applicant]
“WebDriver W3C Recommendation,” World Wide Web Consortium, Jun. 5, 2018, 153 pages. Available at https://www.w3.org/TR/webdriver1/. [cited by applicant]