IP Library Granted Patent US 12,387,036
Granted Patent B1
US 12,387,036 · App. 18/909,588 · Granted Aug 12, 2025

Multimodal agent for efficient image-text interface automation

Inventors: Erich Elsen (San Francisco, CA); Curtis Hawthorne (San Francisco, CA); Augustus Odena (San Francisco, CA); Maxwell Nye (San Francisco, CA); Arushi Somani (San Francisco, CA); Kyle Vigen (San Francisco, CA); Rohan Bavishi (San Francisco, CA); Sagnak Tasirlar (San Francisco, CA); Warut Vijitbenjaronk (San Francisco, CA); Ulas Kirazci (San Francisco, CA); Joe Gershenson (San Francisco, CA); Shaya Zarkesh (San Francisco, CA)
Assignee: Anthropic, PBC
G06F40/166G06F3/0481G06F40/284G06V10/7715G06V10/803G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,036
App. No.
18/909,588
Granted
Aug 12, 2025
Kind
B1
Abstract

A system for image-text agentic interface automation is disclosed. A multimodal agent is configured to process arbitrary-length text sequences and arbitrary-resolution images. A newline insertion logic is configured to interleave a newline character between successive lines of image patches in a plurality of lines of image patches, wherein the newline character specifies an end of a line in an input image. A tokenization logic is configured to translate the input text sequence into a sequence of input text tokens, and to translate the successive lines of image patches interleaved with the newline character into a sequence of input image tokens. A linear projection logic is configured to linearly project a single token stream of the sequence of input text tokens and the sequence of input image tokens into a decoder-only Transformer logic, wherein the linear projection of the single token stream bypasses any embedding lookup.

Claims (38)

1. A system for image-text agentic interface automation, comprising:

a multimodal agent configured to process arbitrary-length text sequences and arbitrary-resolution images:

memory storing an input image and an input text sequence;

patch extraction logic configured to extract image patches from the input image on a line-by-line basis, and generate a plurality of lines of image patches for the input image;

newline insertion logic configured to interleave a newline character between successive lines of image patches in the plurality of lines of image patches, wherein the newline character specifies an end of a line in the input image;

tokenization logic configured to translate the input text sequence into a sequence of input text tokens, and to translate the successive lines of image patches interleaved with the newline character into a sequence of input image tokens;

linear projection logic configured to linearly project a single token stream of the sequence of input text tokens and the sequence of input image tokens into a decoder-only Transformer logic, wherein the linear projection of the single token stream bypasses any embedding lookup; and

the decoder-only Transformer logic configured to process the linearly projected, embedding lookup-bypassed single token stream to generate a sequence of output tokens that are responsive to the input image and the input text sequence.

2. The system of claim 1 , wherein the line in the input image is a row of image patches.

3. The system of claim 1 , wherein the line in the input image is a column of image patches.

4. The system of claim 1 , wherein the successive lines of image patches are arranged in a raster-scan order.

5. The system of claim 1 , wherein the decoder-only Transformer logic is further configured without any image-specific position embeddings.

6. The system of claim 5 , wherein the decoder-only Transformer logic is further configured to be trained on images of arbitrary size at training time, thereby obviating separate high and low-resolution training stages.

7. The system of claim 1 , wherein the decoder-only Transformer logic is further configured without a pooling logic.

8. The system of claim 1 , wherein the decoder-only Transformer logic is further configured without a causal attention logic.

9. The system of claim 1 , wherein the decoder-only Transformer logic is further configured to decouple input embeddings from output embeddings.

10. The system of claim 1 , wherein the decoder-only Transformer logic is further configured to use a squared rectified linear unit (ReLU) activation function.

11. The system of claim 1 , wherein the decoder-only Transformer logic is further configured to use a rotary positional embedding (RoPE).

12. The system of claim 1 , wherein the decoder-only Transformer logic is further configured to add a layer normalization (LayerNorm) function to Query (Q) and Key (K) embeddings before the Q and K embeddings enter attention calculations.

13. A system for image-text agentic interface automation, comprising:

a multimodal agent configured to process arbitrary-resolution images:

memory storing an input image;

patch extraction logic configured to extract image patches from the input image on a line-by-line basis, and generate a plurality of lines of image patches for the input image;

newline insertion logic configured to interleave a newline character between successive lines of image patches in the plurality of lines of image patches, wherein the newline character specifies an end of a line in the input image;

tokenization logic configured to translate the successive lines of image patches interleaved with the newline character into a sequence of input image tokens;

linear projection logic configured to linearly project the sequence of input image tokens into a decoder-only Transformer logic, wherein the linear projection of the sequence of input image tokens bypasses any embedding lookup; and

the decoder-only Transformer logic configured to process the linearly projected, embedding lookup-bypassed sequence of input image tokens to generate a sequence of output tokens that are responsive to the input image.

14. The system of claim 13 , wherein the line in the input image is a row of image patches.

15. The system of claim 13 , wherein the line in the input image is a column of image patches.

16. The system of claim 13 , wherein the decoder-only Transformer logic is further configured without any image-specific position embeddings.

17. The system of claim 16 , wherein the decoder-only Transformer logic is further configured to be trained on images of arbitrary size at training time, thereby obviating separate high and low-resolution training stages.

18. A computer-implemented method for image-text agentic interface automation, including:

storing an input image;

extracting image patches from the input image on a line-by-line basis, and generating a plurality of lines of image patches for the input image;

interleaving a newline character between successive lines of image patches in the plurality of lines of image patches, wherein the newline character specifies an end of a line in the input image;

translating the successive lines of image patches interleaved with the newline character into a sequence of input image tokens;

linearly projecting the sequence of input image tokens into a decoder-only Transformer logic, wherein the linear projection of the sequence of input image tokens bypasses any embedding lookup; and

processing the linearly projected, embedding lookup-bypassed sequence of input image tokens through the decoder-only Transformer logic to generate a sequence of output tokens that are responsive to the input image.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE NAME OF THE ASSIGNEE TO ANTHROPIC, PBC PREVIOUSLY RECORDED ON REEL 70785 FRAME 275. ASSIGNOR(S) HEREBY CONFIRMS THE THE ASSIGNMENT.. Recorded Apr 17, 2025
From: ADEPT AL LABS INC.
To: ANTHROPIC, PBC
Reel/Frame 071101/0374 →
CHANGE OF NAME Recorded Apr 11, 2025
From: PERSIMMON AI LABS, INC.
To: ADEPT AI LABS, INC.
Reel/Frame 070820/0815 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2025
From: ADEPT AL LABS INC.
To: ANTHROPIC, PBNC
Reel/Frame 070785/0275 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2025
From: ELSEN, ERICH; HAWTHORNE, CURTIS; ODENA, AUGUSTUS; NYE, MAXWELL; SOMANI, ARUSHI; VIGEN, KYLE; BAVISHI, ROHAN; TASIRLAR, SAGNAK; VIJITBENJARONK, WARUT; KIRAZCI, ULAS; GERSHENSON, JOE; ZARKESH, SHAYA
To: ADEPT AI LABS INC.
Reel/Frame 070808/0109 →
Continuity (8)
Provisional Application 63638613 · Apr 25, 2024
Provisional Application 63638644 · Apr 25, 2024
Provisional Application 63638631 · Apr 25, 2024
Provisional Application 63567667 · Mar 20, 2024
Provisional Application 63567681 · Mar 20, 2024
Provisional Application 63567698 · Mar 20, 2024
Provisional Application 63567721 · Mar 20, 2024
Provisional Application 63567714 · Mar 20, 2024
References Cited (73)
US 6226785B1 · Peterson et al. · 2001 [cited by applicant]
US 8185544B2 · Oztekin et al. · 2012 [cited by applicant]
US 8855684B2 · Bellver et al. · 2014 [cited by applicant]
US 9218128B1 · Yuschik et al. · 2015 [cited by applicant]
US 10257225B1 · Sites · 2019 [cited by applicant]
US 10587708B2 · Laird-McConnell et al. · 2020 [cited by applicant]
US 11645564B2 · Wu et al. · 2023 [cited by applicant]
US 11809887B2 · Hinton et al. · 2023 [cited by applicant]
US 11907864B2 · Wu et al. · 2024 [cited by applicant]
US 20020062475A1 · Iborra et al. · 2002 [cited by applicant]
US 20030217054A1 · Bachman et al. · 2003 [cited by applicant]
US 20040054690A1 · Hillerbrand et al. · 2004 [cited by applicant]
US 20040078787A1 · Borek et al. · 2004 [cited by applicant]
US 20040215665A1 · Edgar et al. · 2004 [cited by applicant]
US 20060155954A1 · Haynes et al. · 2006 [cited by applicant]
US 20060161878A1 · Koh et al. · 2006 [cited by applicant]
US 20080118051A1 · Odinak et al. · 2008 [cited by applicant]
US 20110041140A1 · Harm et al. · 2011 [cited by applicant]
US 20130226892A1 · Ehsani et al. · 2013 [cited by applicant]
US 20140157288A1 · Wong · 2014 [cited by applicant]
US 20140214404A1 · Kalia et al. · 2014 [cited by applicant]
US 20150339712A1 · Koutrika et al. · 2015 [cited by applicant]
US 20160162172A1 · Rathod · 2016 [cited by applicant]
US 20160335331A1 · Schnase et al. · 2016 [cited by applicant]
US 20170091178A1 · Barbosa et al. · 2017 [cited by applicant]
US 20170289305A1 · Liensberger et al. · 2017 [cited by applicant]
US 20180012141A1 · Chehreghani et al. · 2018 [cited by applicant]
US 20180060744A1 · Achin et al. · 2018 [cited by applicant]
US 20190171984A1 · Irimie et al. · 2019 [cited by applicant]
US 20190187987A1 · Fauchère et al. · 2019 [cited by applicant]
US 20190332686A1 · Lee · 2019 [cited by applicant]
US 20200342316A1 · Shazeer · 2020 [cited by examiner]
US 20230106716A1 · Xiong · 2023 [cited by examiner]
US 20230222623A1 · Ke · 2023 [cited by examiner]
US 20230281400A1 · Wang · 2023 [cited by examiner]
US 20230306205A1 · Maeder et al. · 2023 [cited by applicant]
US 20230325693A1 · Wu et al. · 2023 [cited by applicant]
US 20230342167A1 · Radkoff et al. · 2023 [cited by applicant]
US 20230351149A1 · Yu · 2023 [cited by examiner]
US 20230386025A1 · Loddenkemper et al. · 2023 [cited by applicant]
US 20230419652A1 · Tiong · 2023 [cited by examiner]
US 20240119257A1 · Guo · 2024 [cited by examiner]
US 20240256835A1 · Dehghani · 2024 [cited by examiner]
US 20240282084A1 · Serra · 2024 [cited by examiner]
US 20240282094A1 · Tsimpoukelli · 2024 [cited by examiner]
US 20240290065A1 · Park · 2024 [cited by examiner]
US 20240303443A1 · Cheng et al. · 2024 [cited by applicant]
US 20240362272A1 · Lee et al. · 2024 [cited by applicant]
US 20240370765A1 · Pierucci et al. · 2024 [cited by applicant]
US 20240412720A1 · Vasylyev · 2024 [cited by applicant]
WO WO2024146961A1 · 2024 [cited by examiner]
Li, Bo, Peiyuan Zhang, Jingkang Yang, Yuanhan Zhang, Fanyi Pu, and Ziwei Liu. “Otterhd: A high-resolution multi-modality model.” arXiv preprint arXiv:2311.04219 (2023). (Year: 2023). [cited by examiner]
F. Shi, R. Gao, W. Huang and L. Wang, “Dynamic MDETR: A Dynamic Multimodal Transformer Decoder for Visual Grounding,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, No. 2, pp. 1181-1198, Feb… [cited by examiner]
J. Wu, W. Gan, Z. Chen, S. Wan and P. S. Yu, “Multimodal Large Language Models: A Survey,” 2023 IEEE International Conference on Big Data (BigData), Sorrento, Italy, 2023, pp. 2247-2256 (Year: 2023). [cited by examiner]
Chen, Delong, Samuel Cahyawijaya, Jianfeng Liu, Baoyuan Wang, and Pascale Fung. “Subobject-level Image Tokenization.” arXiv preprint arXiv:2402.14327 (2024). (Year: 2024). [cited by examiner]
Adept Product Team, ‘Building Powerful Agents with Adept’, Aug. 23, 2024, 12 pages. [cited by applicant]
Adept Team, “Adept Fuyu-Heavy: A new multimodal model”, Jan. 24, 2024, 11 pages. [cited by applicant]
Erich Elsen, Augustus Odena, Maxwell Nye, Sanak Tarlar, Tri Dao, Curtis Hawthorne, Deepak Moparthi, Arushi Somani, “Releasing Persimmon-8B”, Sep. 7, 2023, 7 pages. [cited by applicant]
Erich Elsen, Curtis Hawthorne, Arushi Somani, “The Adventure of the Errant Hardware”, Sep. 19, 2023, 14 pages. [cited by applicant]
Rohan Bavishi, Erich Elsen, Curtis Hawthorne, Maxwell Nye, Augustus Odena, Arushi Somani, Sanak Tarlar, “Fuyu-8B: A Multimodal Architecture for AI Agents”, Oct. 17, 2023, 22 pages. [cited by applicant]
Tri Dao, “FlashAttention: Fast Transformer training with long sequences”, Jan. 17, 2023, 9 pages. [cited by applicant]
Li et al, Demonstration+ Natural Language: Multi modal Interfaces for GUI-Based Interactive Task Learning Agents (Year: 2021) 43 pages. [cited by applicant]
Sethi, Pooja, et al. “Autonlu: Detecting, root-causing, and fixing nlu model errors.” arXiv preprint arXiv:2110.06384 (2021). (Year:2021) 10 pages. [cited by applicant]
Takebayashi et al., Multi modal Interface Agent for Enhancing Knowledge Sharing (Year: 1997) 4 pages. [cited by applicant]
U.S. Appl. No. 18/908,447 Non Final Office Action dated Dec. 13, 2024, 27 pages. [cited by applicant]
U.S. Appl. No. 18/909,068 Non-final Office Action dated Dec. 19, 2024, 34 pages. [cited by applicant]
U.S. Appl. No. 18/909,186 Non-final Rejection dated Dec. 9, 2024, 27 pages. [cited by applicant]
U.S. Appl. No. 18/909,455 Non-final Office Action dated Dec. 19, 2024, 30 pages. [cited by applicant]
U.S. Appl. No. 18/909,531 Non-final Rejection dated Jan. 3, 2025, 110 pages. [cited by applicant]
Walker et al. “Neural semantic parsing with anonymization for command understanding in general-purpose service robots.” Robot World Cup. Cham: Springer International Publishing, 2019. 337-350. (Year: 2019) 14 pages. [cited by applicant]
Xie et al., OpenAgents: An Open Platform for Language Agents in the Wild, (Year: 2023) 34 pages. [cited by applicant]
Yin, Pengcheng. Learning Structured Neural Semantic Parsers. Diss. Carnegie Mellon University, 2021. (Year: 2021) 189 pages. [cited by applicant]
Zhou, Shuyan, et al. “Webarena: A realistic web environment for building autonomous agents.” arXiv preprint arXiv:2307.13854 (2023). (Year: 2023) 22 pages. [cited by applicant]
Cited By (3)
US 12,626,487 US 12,675,262 US 12,694,222