IP Library Patent Application 19418741
Patent Application
App. No. 19/418,741

Magnitude Invariant Multimodal Agent for Efficient Image-Text Interface Automation

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/418,741
Abstract

A system for magnitude-invariant image-text agentic interface automation is disclosed. A bit vectorization logic is configured to convert image patches in a plurality of image patches into magnitude-invariant bit vectors, and generate a plurality of lines of magnitude-invariant bit vectors. A tokenization logic is configured to translate the input text sequence into a sequence of input text tokens, and to translate the successive lines of magnitude-invariant bit vectors interleaved with a newline character into a sequence of input magnitude-invariant bit vector tokens. A linear projection logic is configured to linearly project a single token stream of the sequence of input text tokens and the sequence of input magnitude-invariant bit vector tokens into a decoder-only Transformer logic, wherein the linear projection of the single token stream bypasses any embedding lookup.

Claims (59)

1 . A method of conducting an artificial intelligent (AI) agent workflow on a computing environment, the method comprising:

receiving, at a server hosting an AI agent, at least one of:

(i) one or more screenshots of a computing environment,

(ii) an agent action history of one or more past action commands, and

(iii) a task description indicative of the AI agent workflow;

generating, at the server, an agent inference message including the one or more screenshots, the agent action history, the task description and/or an agent system prompt instructing the AI agent to perform a specific task;

generating, by the AI agent deployed at the server, an actuation function indicative of at least one action to fulfill the task description based on the agent inference message; and

transmitting the actuation function to the computing environment that executes a machine-actuated command form of the actuation function.

2 . The method of claim 1 , wherein the computing environment comprises a web browser application at a client device.

3 . The method of claim 1 , wherein the one or more screenshots comprise at least one of:

a current screenshot of the computing environment; and

a prior screenshot of the computing environment, wherein the computing environment evolves from the prior screenshot to the current screenshot after an actuation command is executed on the computing environment.

4 . The method of claim 2 , wherein the task description is received via a user interface of a client device.

5 . The method of claim 1 , wherein the actuation function is constructed in a domain-specific language indicative of one or more multimodal user interface interactions.

6 . The method of claim 1 , further comprising:

inferring one or more user intent indicators from the task description; and

combining the one or more user intent indicators into the agent inference message.

7 . The method of claim 6 , wherein the one or more user intent indicators comprise at least one of:

an indicator of a current state of the computing environment; and

a user-contemplated action for the computing environment.

8 . The method of claim 1 , further comprising:

training, using vision and language training data, a multimodal model of the AI agent to process the agent inference message including both image data corresponding to the one or more screenshots and text data corresponding to at least the task description.

9 . The method of claim 1 , further comprising:

causing the actuation function to be translated into and executed as the machine-actuated command on the computing environment even when a user interface of the computing environment has changed at runtime.

10 . A system of conducting an artificial intelligent (AI) agent workflow on a computing environment, the system comprising:

an AI agent implemented on one or more processors; and

an interface to receive at least one of:

(i) one or more screenshots of the computing environment,

(ii) an agent action history of one or more past action commands, and

(iii) a task description indicative of the AI agent workflow;

wherein the AI agent obtains an agent inference message including the one or more screenshots, the agent action history, the task description and/or an agent system prompt instructing the AI agent to perform a specific task, and

generates an actuation function indicative of at least one action to fulfill the task description based on the agent inference message, and

wherein the interface further transmits the actuation function to the computing environment that executes a machine-actuated command form of the actuation function.

11 . The system of claim 10 , wherein the computing environment comprises a web browser application at a client device.

12 . The system of claim 10 , wherein the one or more screenshots comprise at least one of:

a current screenshot of the computing environment; and

a prior screenshot of the computing environment, wherein the computing environment evolves from the prior screenshot to the current screenshot after an actuation command is executed on the computing environment.

13 . The system of claim 10 , wherein the task description is received via a user interface of a client device.

14 . The system of claim 10 , wherein the actuation function is constructed in a domain-specific language indicative of one or more multimodal user interface interactions.

15 . The system of claim 10 , wherein the one or more processors are configured to:

infer one or more user intent indicators from the task description; and

combine the one or more user intent indicators into the agent inference message.

16 . The system of claim 15 , wherein the one or more user intent indicators comprise at least one of:

an indicator of a current state of the computing environment; and

a user-contemplated action for the computing environment.

17 . The system of claim 10 , wherein the AI agent comprises a multimodal model trained on vision and language data to process the agent inference message including both image data corresponding to the one or more screenshots and text data corresponding to at least the task description.

18 . The system of claim 10 , wherein the actuation function is further translated into and executed as the machine-actuated command on the computing environment even when a user interface of the computing environment has changed at runtime.

19 . A non-transitory storage computer-readable medium storing a plurality of processor-executable instructions for conducting an artificial intelligent (AI) agent workflow on a computing environment, the instructions being executed by one or more processors to perform operations comprising:

receiving, at a server hosting the AI agent, at least one of:

(i) one or more screenshots of the computing environment,

(ii) an agent action history of one or more past action commands, and

(iii) a task description indicative of the AI agent workflow;

generating, at the server, an agent inference message including the one or more screenshots, the agent action history, the task description and/or an agent system prompt instructing the AI agent to perform a specific task;

generating, by the AI agent deployed at the server, an actuation function indicative of at least one action to fulfill the task description based on the agent inference message; and

transmitting the actuation function to the computing environment that executes a machine-actuated command form of the actuation function.

20 . The non-transitory storage computer-readable medium of claim 19 , wherein the computing environment comprises a web browser application at a client device, and

wherein the one or more screenshots comprise at least one of:

a current screenshot of the computing environment; and

a prior screenshot of the computing environment, wherein the computing environment evolves from the prior screenshot to the current screenshot after an actuation command is executed on the computing environment.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2026
From: HAWTHORNE, CURTIS; ELSEN, ERICH; SOMANI, ARUSHI; VIGEN, KYLE; BAVISHI, ROHAN; TASIRLAR, SAGNAK; VIJITBENJARONK, WARUT; KIRAZCI, ULAS; GERSHENSON, JOE; ZARKESH, SHAYA
To: ADEPT AI LABS INC.
Reel/Frame 074142/0687 →
CHANGE OF NAME Recorded Mar 20, 2026
From: PERSIMMON AI LABS, INC.
To: ADEPT AI LABS INC.
Reel/Frame 075265/0569 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2026
From: ADEPT AI LABS INC.
To: ANTHROPIC, PBC
Reel/Frame 075265/0699 →