IP Library Patent Application 18791986
Patent Application
App. No. 18/791,986

Adding Voice or Chat User interface to graphical user interface (gui)-based virtualized applications and desktops using large language and large action models

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/791,986
Abstract

Methods and systems for enhanced remote desktop interfaces are described. A computing system may train, using historical or live information, a LAM to execute, within a remote desktop application, textual actions with their parameters if any. A user declarative request (voice or chat) may be interpreted by a LLM to match a specific action (and potentially ask for the corresponding parameters in a conversational way). Subsequently, from the textual action and its parameters, the LAM may execute the action within a remote desktop application and report the result to the user via voice or chat.

Claims (55)

1 . A method comprising:

training, using historical remote desktop interaction information indicating user inputs and corresponding actions executed within historical remote desktop application sessions, a large action model (LAM), wherein training the LAM configures the LAM to execute, for a given textual input, one or more actions to perform within a given remote desktop application to complete a task requested by the given textual input;

deploying, to a remote desktop host server, a LAM agent, configured to access the LAM to identify the one or more actions;

receiving, during a remote desktop session, a textual input indicating a first task to perform;

identifying, based on the first task, a remote desktop application configured to perform the task and a list of actions that the remote desktop application is configured to perform;

identifying, using a large language model (LLM), at least one action of the list of actions to execute to perform the task;

executing, using the LAM, the at least one action to produce an action result; and

displaying the action result, wherein the action result comprises an indication that the task has been executed.

2 . The method of claim 1 , wherein training the LAM is further based on lists of actions corresponding to each remote desktop application of a plurality of remote desktop applications, wherein each list of actions is labelled based on the corresponding remote desktop application.

3 . The method of claim 1 , further comprising:

establishing, based on successful validation of authentication credentials provided at a client device, the remote desktop session, wherein establishing the remote desktop session comprises receiving, at the client device and from the remote desktop host server, an authentication token.

4 . The method of claim 3 , wherein establishing the remote desktop session further comprises:

identifying one or more applications corresponding to the remote desktop session and, for each of the one or more applications, a list of actions that the corresponding application is configured to performed.

5 . The method of claim 1 , wherein identifying the remote desktop application comprises applying a large language model to the textual input to identify the remote desktop application.

6 . The method of claim 1 , further comprising:

launching, after identifying the remote desktop application, before identifying the at least one action, and via communication with the remote desktop host server, the remote desktop application.

7 . The method of claim 6 , further comprising:

after launching the remote desktop application and prior to the identification of the at least one action, establishing a connection between a client device and the remote desktop host server.

8 . The method of claim 7 , wherein the connection comprises a remote desktop protocol connection, a websocket connection, or a LAM virtual channel (VC).

9 . The method of claim 1 , further comprising:

collecting feedback on the action result; and

updating, based on the feedback, the LAM agent.

10 . The method of claim 7 , wherein the client device comprises one of: smart glasses or a mobile device.

11 . A computing system comprising:

one or more processors;

memory storing computer executable instructions that, when executed by the one or more processors, cause the computing system to:

train, using historical remote desktop interaction information indicating user inputs and corresponding actions executed within historical remote desktop application sessions, a large action model (LAM), wherein training the LAM configures the LAM to execute, for a given textual input, one or more actions to perform within a given remote desktop application to complete a task requested by the given textual input;

deploy, to a remote desktop host server, a LAM agent, configured to access the LAM to identify the one or more actions;

receive, during a remote desktop session, a textual input indicating a first task to perform;

identify, based on the first task, a remote desktop application configured to perform the task and a list of actions that the remote desktop application is configured to perform;

identify, using a large language model (LLM), at least one action of the list of actions to execute to perform the task;

execute, using the LAM, the at least one action to produce an action result; and

display the action result, wherein the action result comprises an indication that the task has been executed.

12 . The computing system of claim 11 , wherein training the LAM is further based on lists of actions corresponding to each remote desktop application of a plurality of remote desktop applications, wherein each list of actions is labelled based on the corresponding remote desktop application.

13 . The computing system of claim 11 , wherein the memory stores additional computer executable instructions that, when executed by the one or more processors, further cause the computing system to:

establish, based on successful validation of authentication credentials provided at a client device, the remote desktop session, wherein establishing the remote desktop session comprises receiving, at the client device and from the remote desktop host server, an authentication token.

14 . The computing system of claim 13 , wherein establishing the remote desktop session further comprises:

identifying one or more applications corresponding to the remote desktop session and, for each of the one or more applications, a list of actions that the corresponding application is configured to performed.

15 . The computing system of claim 11 , wherein identifying the remote desktop application comprises applying a large language model to the textual input to identify the remote desktop application.

16 . The computing system of claim 11 , wherein the memory stores additional computer executable instructions that, when executed by the one or more processors, further cause the computing system to:

launch, after identifying the remote desktop application, before identifying the at least one action, and via communication with the remote desktop host server, the remote desktop application.

17 . The computing system of claim 16 , wherein the memory stores additional computer executable instructions that, when executed by the one or more processors, further cause the computing system to:

after launching the remote desktop application and prior to the identification of the at least one action, establish a connection between a client device and the remote desktop host server.

18 . The computing system of claim 17 , wherein the connection comprises a remote desktop protocol connection or a websocket connection.

19 . The computing system of claim 11 , wherein the memory stores additional computer executable instructions that, when executed by the one or more processors, further cause the computing system to:

collect feedback on the action result; and

update, based on the feedback, the LAM agent.

20 . One or more non-transitory computer-readable media storing instructions that, when executed by a computing system comprising at least one processor, a communication interface, and memory, cause the computing system to:

train, using historical remote desktop interaction information indicating user inputs and corresponding actions executed within historical remote desktop application sessions, a large action model (LAM), wherein training the LAM configures the LAM to execute, for a given textual input, one or more actions to perform within a given remote desktop application to complete a task requested by the given textual input;

deploy, to a remote desktop host server, a LAM agent, configured to access the LAM to identify the one or more actions;

receive, during a remote desktop session, a textual input indicating a first task to perform;

identify, based on the first task, a remote desktop application configured to perform the task and a list of actions that the remote desktop application is configured to perform;

identify, using a large language model (LLM), at least one action of the list of actions to execute to perform the task;

execute, using the LAM, the at least one action to produce an action result; and

display the action result, wherein the action result comprises an indication that the task has been executed.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 27, 2026
From: DIVOUX, HUBERT; INGALE, MUKUND
To: CITRIX SYSTEMS, INC.
Reel/Frame 073592/0362 →
PATENT SECURITY AGREEMENT Recorded Aug 15, 2025
From: CLOUD SOFTWARE GROUP, INC.; CITRIX SYSTEMS, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS NOTES COLLATERAL AGENT
Reel/Frame 072488/0172 →