IP Library Granted Patent US 12,619,815
Granted Patent B2
US 12,619,815 · App. 18/909,558 · Granted May 5, 2026

Magnitude invariant multimodal agent for efficient image-text interface automation

Inventors: Curtis Hawthorne (San Francisco, CA); Erich Elsen (San Francisco, CA); Augustus Odena (San Francisco, CA); Maxwell Nye (San Francisco, CA); Arushi Somani (San Francisco, CA); Kyle Vigen (San Francisco, CA); Rohan Bavishi (San Francisco, CA); Sagnak Tasirlar (San Francisco, CA); Warut Vijitbenjaronk (San Francisco, CA); Ulas Kirazci (San Francisco, CA); Joe Gershenson (San Francisco, CA); Shaya Zarkesh (San Francisco, CA)
Assignee: Anthropic, PBC
G06F40/166G06F3/0481G06F3/0484G06F16/951G06F40/174G06F40/284G06N3/0455G06N3/091G06N5/04G06N20/00G06V10/7715G06V10/774G06V10/803G06V10/82G06V20/40G06V30/19147G06V30/41G06F9/451
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,815
App. No.
18/909,558
Granted
May 5, 2026
Kind
B2
Abstract

A system for magnitude-invariant image-text agentic interface automation is disclosed. A bit vectorization logic is configured to convert image patches in a plurality of image patches into magnitude-invariant bit vectors, and generate a plurality of lines of magnitude-invariant bit vectors. A tokenization logic is configured to translate the input text sequence into a sequence of input text tokens, and to translate the successive lines of magnitude-invariant bit vectors interleaved with a newline character into a sequence of input magnitude-invariant bit vector tokens. A linear projection logic is configured to linearly project a single token stream of the sequence of input text tokens and the sequence of input magnitude-invariant bit vector tokens into a decoder-only Transformer logic, wherein the linear projection of the single token stream bypasses any embedding lookup.

Claims (38)

1 . A system for magnitude-invariant image-text agentic interface automation, comprising:

memory storing an input image and an input text sequence;

patch extraction logic configured to extract image patches from the input image on a line-by-line basis, and generate a plurality of lines of image patches for the input image;

bit vectorization logic configured to convert image patches in the plurality of lines of image patches into magnitude-invariant bit vectors, and generate a plurality of lines of magnitude-invariant bit vectors;

newline insertion logic configured to interleave a newline character between successive lines of magnitude-invariant bit vectors in the plurality of lines of magnitude-invariant bit vectors, wherein the newline character specifies an end of a line in the input image;

tokenization logic configured to translate the input text sequence into a sequence of input text tokens, and to translate the successive lines of magnitude-invariant bit vectors interleaved with the newline character into a sequence of input magnitude-invariant bit vector tokens;

linear projection logic configured to linearly project a single token stream of the sequence of input text tokens and the sequence of input magnitude-invariant bit vector tokens into a decoder-only Transformer logic, wherein the linear projection of the single token stream bypasses any embedding lookup; and

the decoder-only Transformer logic configured to process the linearly projected, embedding lookup-bypassed single token stream to generate a sequence of output tokens that are responsive to the input image and the input text sequence.

2 . The system of claim 1 , wherein the bit vectorization logic is further configured to apply a RGB555 format compression to convert the image patches in the plurality of lines of image patches into the magnitude-invariant bit vectors, and generate the plurality of lines of magnitude-invariant bit vectors.

3 . The system of claim 2 , wherein the RGB555 format compression produces three 5-bit values, one for each of subpixel channels R (red), G (green), and B (blue).

4 . The system of claim 3 , wherein the three 5-bit values take either a 1 value or a −1 value.

5 . The system of claim 4 , wherein the three 5-bit values are magnitude-invariant to scale modification functions of the decoder-only Transformer logic.

6 . The system of claim 5 , wherein a layer normalization (LayerNorm) function is one of the scaling functions of the decoder-only Transformer logic.

7 . The system of claim 1 , wherein the bit vectorization logic is further configured to apply a RGB888 format compression to convert the image patches in the plurality of lines of image patches into the magnitude-invariant bit vectors, and generate the plurality of lines of magnitude-invariant bit vectors.

8 . The system of claim 7 , wherein the RGB888 format compression produces three 8-bit values, one for each of subpixel channels R (red), G (green), and B (blue).

9 . The system of claim 8 , wherein the three 8-bit values take either a 1 value or a −1 value.

10 . The system of claim 9 , wherein the three 8-bit values are magnitude-invariant to scale modification functions of the decoder-only Transformer logic.

11 . The system of claim 10 , wherein a layer normalization (LayerNorm) function is one of the scaling functions of the decoder-only Transformer logic.

12 . The system of claim 1 , wherein the bit vectorization logic is further configured to apply a RGB565 format compression to convert the image patches in the plurality of lines of image patches into the magnitude-invariant bit vectors, and generate the plurality of lines of magnitude-invariant bit vectors.

13 . The system of claim 12 , wherein the RGB565 format compression produces 5-bit values for R (red) and B (blue) subpixel channels and 6-bit values for G (green) subpixel channel.

14 . The system of claim 13 , wherein the 5-bit and the 6-bit values take either a 1 value or a −1 value.

15 . The system of claim 14 , wherein the 5-bit and the 6-bit values are magnitude-invariant to scale modification functions of the decoder-only Transformer logic.

16 . The system of claim 15 , wherein a layer normalization (LayerNorm) function is one of the scaling functions of the decoder-only Transformer logic.

17 . A system for magnitude-invariant image-text agentic interface automation, comprising:

memory storing an input image;

patch extraction logic configured to extract image patches from the input image on a line-by-line basis, and generate a plurality of lines of image patches for the input image;

bit vectorization logic configured to convert image patches in the plurality of lines of image patches into magnitude-invariant bit vectors, and generate a plurality of lines of magnitude-invariant bit vectors;

newline insertion logic configured to interleave a newline character between successive lines of magnitude-invariant bit vectors in the plurality of lines of magnitude-invariant bit vectors, wherein the newline character specifies an end of a line in the input image;

tokenization logic configured to translate the successive lines of magnitude-invariant bit vectors interleaved with the newline character into a sequence of input magnitude-invariant bit vector tokens;

linear projection logic configured to linearly project the sequence of input magnitude-invariant bit vector tokens into a decoder-only Transformer logic, wherein the linear projection of the sequence of input magnitude-invariant bit vector tokens bypasses any embedding lookup; and

the decoder-only Transformer logic configured to process the linearly projected, embedding lookup-bypassed sequence of input magnitude-invariant bit vector tokens to generate a sequence of output tokens that are responsive to the input image.

18 . A system for magnitude-invariant image-text agentic interface automation, comprising:

memory storing an input image;

patch extraction logic configured to extract image patches from the input image on a line-by-line basis, and generate a plurality of lines of image patches for the input image;

bit vectorization logic configured to convert image patches in the plurality of lines of image patches into magnitude-invariant bit vectors, and generate a plurality of lines of magnitude-invariant bit vectors;

tokenization logic configured to translate lines of the plurality of lines of magnitude-invariant bit vectors into a sequence of input magnitude-invariant bit vector tokens;

linear projection logic configured to linearly project the sequence of input magnitude-invariant bit vector tokens into a decoder-only Transformer logic; and

the decoder-only Transformer logic configured to process the linearly projected sequence of input magnitude-invariant bit vector tokens to generate a sequence of output tokens that are responsive to the input image.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE NAME OF THE ASSIGNEE TO ANTHROPIC, PBC PREVIOUSLY RECORDED ON REEL 70785 FRAME 275. ASSIGNOR(S) HEREBY CONFIRMS THE THE ASSIGNMENT.. Recorded Apr 17, 2025
From: ADEPT AL LABS INC.
To: ANTHROPIC, PBC
Reel/Frame 071101/0374 →
CHANGE OF NAME Recorded Apr 11, 2025
From: PERSIMMON AI LABS, INC.
To: ADEPT AI LABS, INC.
Reel/Frame 070820/0815 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2025
From: ADEPT AL LABS INC.
To: ANTHROPIC, PBNC
Reel/Frame 070785/0275 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2025
From: HAWTHORNE, CURTIS; ELSEN, ERICH; ODENA, AUGUSTUS; NYE, MAXWELL; SOMANI, ARUSHI; VIGEN, KYLE; BAVISHI, ROHAN; TASIRLAR, SAGNAK; VIJITBENJARONK, WARUT; KIRAZCI, ULAS; GERSHENSON, JOE; ZARKESH, SHAYA
To: ADEPT AI LABS INC.
Reel/Frame 070808/0236 →
Continuity (9)
Provisional Application 63638613 · Apr 25, 2024
Provisional Application 63638631 · Apr 25, 2024
Provisional Application 63638644 · Apr 25, 2024
Provisional Application 63567714 · Mar 20, 2024
Provisional Application 63567667 · Mar 20, 2024
Provisional Application 63567721 · Mar 20, 2024
Provisional Application 63567681 · Mar 20, 2024
Provisional Application 63567698 · Mar 20, 2024
Related Publication 20250299024A1 · Sep 25, 2025
References Cited (143)
US 6012030A · French-St. George et al. · 2000 [cited by applicant]
US 6226785B1 · Peterson et al. · 2001 [cited by applicant]
US 6859451B1 · Pasternack et al. · 2005 [cited by applicant]
US 8185544B2 · Oztekin et al. · 2012 [cited by applicant]
US 8493406B2 · Rubin et al. · 2013 [cited by applicant]
US 8855684B2 · Bellver et al. · 2014 [cited by applicant]
US 9218128B1 · Yuschik et al. · 2015 [cited by applicant]
US 9269048B1 · Chen · 2016 [cited by applicant]
US 10257225B1 · Sites · 2019 [cited by applicant]
US 10587708B2 · Laird-Mcconnell et al. · 2020 [cited by applicant]
US 11077320B1 · Hibbard · 2021 [cited by applicant]
US 11645564B2 · Wu et al. · 2023 [cited by applicant]
US 11809887B2 · Hinton et al. · 2023 [cited by applicant]
US 11907864B2 · Wu et al. · 2024 [cited by applicant]
US 12293272B1 · Poulis et al. · 2025 [cited by applicant]
US 20020062475A1 · Iborra et al. · 2002 [cited by applicant]
US 20030217054A1 · Bachman et al. · 2003 [cited by applicant]
US 20040054690A1 · Hillerbrand et al. · 2004 [cited by applicant]
US 20040078787A1 · Borek et al. · 2004 [cited by applicant]
US 20040215665A1 · Edgar et al. · 2004 [cited by applicant]
US 20050010418A1 · McNair et al. · 2005 [cited by applicant]
US 20060155954A1 · Haynes et al. · 2006 [cited by applicant]
US 20060161878A1 · Koh et al. · 2006 [cited by applicant]
US 20070233495A1 · Agapi et al. · 2007 [cited by applicant]
US 20080065388A1 · Cross et al. · 2008 [cited by applicant]
US 20080065390A1 · Ativanichayaphong et al. · 2008 [cited by applicant]
US 20080065453A1 · Settuducati · 2008 [cited by applicant]
US 20080118051A1 · Odinak et al. · 2008 [cited by applicant]
US 20080228494A1 · Cross · 2008 [cited by applicant]
US 20080301094A1 · Zhu et al. · 2008 [cited by applicant]
US 20110041140A1 · Harm et al. · 2011 [cited by applicant]
US 20130080641A1 · Lui et al. · 2013 [cited by applicant]
US 20130144682A1 · Dhara et al. · 2013 [cited by applicant]
US 20130226892A1 · Ehsani et al. · 2013 [cited by applicant]
US 20130268260A1 · Lundberg et al. · 2013 [cited by applicant]
US 20140157288A1 · Wong · 2014 [cited by applicant]
US 20140214404A1 · Kalia et al. · 2014 [cited by applicant]
US 20150339712A1 · Koutrika et al. · 2015 [cited by applicant]
US 20160162172A1 · Rathod · 2016 [cited by applicant]
US 20160335331A1 · Schnase et al. · 2016 [cited by applicant]
US 20170048170A1 · Smullen et al. · 2017 [cited by applicant]
US 20170091178A1 · Barbosa et al. · 2017 [cited by applicant]
US 20170255580A1 · Parekh et al. · 2017 [cited by applicant]
US 20170289305A1 · Liensberger et al. · 2017 [cited by applicant]
US 20180012141A1 · Chehreghani et al. · 2018 [cited by applicant]
US 20180060744A1 · Achin et al. · 2018 [cited by applicant]
US 20180137431A1 · Goldfarb et al. · 2018 [cited by applicant]
US 20180157739A1 · Khaitan et al. · 2018 [cited by applicant]
US 20180314943A1 · Liang et al. · 2018 [cited by applicant]
US 20190171984A1 · Irimie et al. · 2019 [cited by applicant]
US 20190187987A1 · Fauchère et al. · 2019 [cited by applicant]
US 20190332686A1 · Lee · 2019 [cited by applicant]
US 20190384807A1 · Dernoncourt et al. · 2019 [cited by applicant]
US 20200342316A1 · Shazeer et al. · 2020 [cited by applicant]
US 20210232992A1 · Disterheft et al. · 2021 [cited by applicant]
US 20220046129A1 · Clodore et al. · 2022 [cited by applicant]
US 20220048525A1 · Tsai et al. · 2022 [cited by applicant]
US 20220051219A1 · Sells et al. · 2022 [cited by applicant]
US 20220058981A1 · Neumann · 2022 [cited by applicant]
US 20220130013A1 · Pottorff et al. · 2022 [cited by applicant]
US 20220215262A1 · Narayanaswami et al. · 2022 [cited by applicant]
US 20220246257A1 · Priestas et al. · 2022 [cited by applicant]
US 20220291966A1 · Masood et al. · 2022 [cited by applicant]
US 20230031702A1 · Li et al. · 2023 [cited by applicant]
US 20230106716A1 · Xiong et al. · 2023 [cited by applicant]
US 20230156075A1 · Brewer et al. · 2023 [cited by applicant]
US 20230206913A1 · Vempaty et al. · 2023 [cited by applicant]
US 20230222285A1 · Zhang et al. · 2023 [cited by applicant]
US 20230222623A1 · Ke et al. · 2023 [cited by applicant]
US 20230281400A1 · Wang et al. · 2023 [cited by applicant]
US 20230306205A1 · Maeder et al. · 2023 [cited by applicant]
US 20230325693A1 · Wu et al. · 2023 [cited by applicant]
US 20230342167A1 · Radkoff et al. · 2023 [cited by applicant]
US 20230351149A1 · Yu et al. · 2023 [cited by applicant]
US 20230360388A1 · Singh · 2023 [cited by applicant]
US 20230386025A1 · Loddenkemper et al. · 2023 [cited by applicant]
US 20230419652A1 · Tiong et al. · 2023 [cited by applicant]
US 20240119257A1 · Guo et al. · 2024 [cited by applicant]
US 20240135232A1 · Betteridge et al. · 2024 [cited by applicant]
US 20240256835A1 · Dehghani et al. · 2024 [cited by applicant]
US 20240281472A1 · LaRhette et al. · 2024 [cited by applicant]
US 20240282084A1 · Serra et al. · 2024 [cited by applicant]
US 20240282094A1 · Tsimpoukelli et al. · 2024 [cited by applicant]
US 20240290065A1 · Park et al. · 2024 [cited by applicant]
US 20240303443A1 · Cheng et al. · 2024 [cited by applicant]
US 20240329943A1 · Sharma et al. · 2024 [cited by applicant]
US 20240338234A1 · Li · 2024 [cited by applicant]
US 20240362272A1 · Lee et al. · 2024 [cited by applicant]
US 20240370765A1 · Pierucci et al. · 2024 [cited by applicant]
US 20240404238A1 · Yu · 2024 [cited by examiner]
US 20240412720A1 · Vasylyev · 2024 [cited by applicant]
US 20250028759A1 · Castillo et al. · 2025 [cited by applicant]
US 20250217170A1 · Shaw et al. · 2025 [cited by applicant]
US 20250245030A1 · Cyjon et al. · 2025 [cited by applicant]
US 20250272350A1 · Furuta et al. · 2025 [cited by applicant]
US 20250272506A1 · Sassak, Jr. et al. · 2025 [cited by applicant]
WO 2024049607A2 · 2024 [cited by applicant]
WO 2024146961A1 · 2024 [cited by applicant]
Li et al., “Otter: Deep Diving Into Large Multi-Modality Models”, 2023, GitHub Repository, https://github.com/Luodian/Otter (Year: 2023). [cited by examiner]
Sravanthi, G., and M. GurunadhaBabu. “Design & Implementation of VGA Display System Based on CPLD and Dual Memory.” International Journal of VLSI System Design and Communication System 3.01 (2015): 0005-0009. (Year: 201… [cited by examiner]
Yu, Jiahui, et al. “Vector-quantized image modeling with improved vqgan.” arXiv preprint arXiv:2110.04627 (2021). (Year: 2021). [cited by examiner]
Humbertokramm, “Convert array RGB565 to RGB888 and then to PNG in python”, 2018, GitHub Repository, https://github.com/humbertokramm/RGB565toRGB888toPNG_-python- (Year: 2018). [cited by examiner]
Ortiz, Jose Javier Gonzalez, John Guttag, and Adrian Dalca. “Magnitude invariant parametrizations improve hypernetwork learning.” arXiv preprint arXiv:2304.07645 (2023). (Year: 2023). [cited by examiner]
Adept Product Team, ‘Building Powerful Agents with Adept’, Aug. 23, 2024, 12 pages. [cited by applicant]
Adept Team, “Adept Fuyu-Heavy: A new multimodal model”, Jan. 24, 2024, 11 pages. [cited by applicant]
Erich Elsen, Augustus Odena, Maxwell Nye, Sanak Tarlar, Tri Dao, Curtis Hawthorne, Deepak Moparthi, Arushi Somani, “Releasing Persimmon-8B”, Sep. 7, 2023, 7 pages. [cited by applicant]
Erich Elsen, Curtis Hawthorne, Arushi Somani, “The Adventure of the Errant Hardware”, Sep. 19, 2023, 14 pages. [cited by applicant]
Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, Saǧnak Tarlar, “Fuyu-8B: A Multimodal Architecture for AI Agents”, Oct. 17, 2023, 22 pages. [cited by applicant]
Tri Dao, “FlashAttention: Fast Transformer training with long sequences”, Jan. 17, 2023, 9 pages. [cited by applicant]
Chen, Delong, Samuel Cahyawijaya, Jianfeng Liu, Baoyuan Wang, and Pascale Fung. “Subobject-level Image Tokenization.” arXiv preprint arXiv:2402.14327 (2024). (Year: 2024). [cited by applicant]
F. Shi, R. Gao, W. Huang and L. Wang, “Dynamic MDETR: A Dynamic Multi modal Transformer Decoder for Visual Grounding,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, No. 2, pp. 1181-1198, Fe… [cited by applicant]
J. Wu, W. Gan, Z. Chen, S. Wan and p. S. Yu, “Multimodal Large Language Models: A Survey,” 2023 IEEE International Conference on Big Data (BigData), Sorrento, Italy, 2023, pp. 2247-2256 (Year: 2023). [cited by applicant]
Li, Bo, Peiyuan Zhang, Jingkang Yang, Yuanhan Zhang, Fanyi Pu, and Ziwei Liu. “Otterhd: A high-resolution multi-modality model.” arXiv preprint arXiv:2311.04219 (2023). (Year: 2023). [cited by applicant]
Lee, Yi-Lun, et al., “Multimodal Prompting with Missing Modalities for Visual Recognition,” 2023, 10 pages. [cited by applicant]
Moran, Douglas B., et al. “Multimodal User Interfaces in the Open Agent Architecture” 1997, 8 pages. [cited by applicant]
Song, Kisub, et al., “Generating multimodal user interfaces for Web services.” 2008, 11 pages. [cited by applicant]
Li et al, Demonstration+ Natural Language: Multi modal Interfaces for GUI-Based Interactive Task Learning Agents (Year: 2021) 43 pages. [cited by applicant]
Sethi, Pooja, et al. “Autonlu: Detecting, root-causing, and fixing nlu model errors.” arXiv preprint arXiv:2110.06384 (2021). (Year:2021) 10 pages. [cited by applicant]
Takebayashi et al., Multi modal Interface Agent for Enhancing Knowledge Sharing (Year: 1997) 4 pages. [cited by applicant]
U.S. Appl. No. 18/908,447 Non Final Office Action dated Dec. 13, 2024, 27 pages. [cited by applicant]
U.S. Appl. No. 18/909,068 Non-final Office Action dated Dec. 19, 2024, 34 pages. [cited by applicant]
U.S. Appl. No. 18/909,186 Non-final Rejection dated Dec. 9, 2024, 27 pages. [cited by applicant]
U.S. Appl. No. 18/909,455 Non-final Office Action dated Dec. 19, 2024, 30 pages. [cited by applicant]
U.S. Appl. No. 18/909,531 Non-final Rejection dated Jan. 3, 2025, 110 pages. [cited by applicant]
U.S. Appl. No. 18/909,588 Non-final Office Action dated Dec. 4, 2024, 41 pages. [cited by applicant]
Walker et al. “Neural semantic parsing with anonymization for command understanding in general-purpose service robots.” Robot World Cup. Cham: Springer International Publishing, 2019. 337-350. (Year: 2019) 14 pages. [cited by applicant]
Xie et al., OpenAgents: An Open Platform for Language Agents in the Wild, (Year: 2023) 34 pages. [cited by applicant]
Yin, Pengcheng. Learning Structured Neural Semantic Parsers. Diss. Carnegie Mellon University, 2021. (Year: 2021) 189 pages. [cited by applicant]
Zhou, Shuyan, et al. “Webarena: A realistic web environment for building autonomous agents.” arXiv preprint arXiv:2307.13854 (2023). (Year: 2023) 22 pages. [cited by applicant]
Chen, Weihao, et al. “Miwa: Mixed-initiative web automation for better user control and confidence.” Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. 2023. (Year: 2023). [cited by applicant]
Deng, Xiang, et al. “Mind2web: Towards a generalist agent for the web.” Advances in Neural Information Processing Systems 36 (2023): 28091-28114. (Year: 2023). [cited by applicant]
Gur, Izzeddin, et al. “A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis.” ICLR. 2024. (Year: 2024). [cited by applicant]
He, Hongliang, et al. “WebVoyager: Building an end-to-end web agent with large multi modal models.” arXiv preprint arXiv:2401.13919 (2024). (Year: 2024). [cited by applicant]
Koh, Jing Yu, et al. “Visualwebarena: Evaluating multi modal agents on realistic visual web tasks.” arXiv preprint arXiv:2401.13649 (2024). (Year: 2024). [cited by applicant]
Liu, Junpeng, et al. “Visualwebbench: How far have multi modal Ilms evolved in web page understanding and grounding ?. ” arXivpreprint arXiv:2404.05955 (2024). (Year: 2024). [cited by applicant]
Lu, Xing Han, Zdenek Kasner, and Siva Reddy. “Weblinx: Real-world website navigation with multi-turn dialogue.” arXiv preprint arXiv:2402.05930 (2024). (Year: 2024). [cited by applicant]
Zhang, Saizheng, et al. “Personalizing dialogue agents: I have a dog, do you have pets too ?.” arXiv preprint arXiv:1801.07243 (2018). (Year: 2018). [cited by applicant]
International Search Report and Written Opinion mailed Jul. 1, 2025 in PCT Application No. PCT/US2025/020719 filed Mar. 20, 2025. [cited by applicant]
Waibel, A., et al., “Multimodal Interfaces for Multimedia Information Agents,” 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing, Munich, Germany, 1997, pp. 167-170 vol. 1, doi: 10.1109/ICAS… [cited by applicant]
Zheng, L., et al., “AgentStufio: A Toolkit for Building General Virtual Agents,” Mar. 26, 2024, arXiv:2403.17918V1 [sc.AI], 12 pages. [cited by applicant]
Furuta, H., et al. “Multimodal Web Navigation with Instruction-Finetuned Foundation Models,” Feb. 25, 2024, arXiv:2305.11854v4 [cs.LG], 31 pages. [cited by applicant]
“Selenium Documentation Release 1.0,” Aug. 26, 2012, 201 pages. [cited by applicant]
“WebDriver W3C Recommendation,” World Wide Web Consortium, Jun. 5, 2018, 153 pages. Available at https://www.w3.org/TR/webdriver1/. [cited by applicant]