IP Library › Granted Patent US 12,566,736
Granted Patent B2
US 12,566,736 · App. 18/758,868 · Granted Mar 3, 2026

Orchestrating dynamic actions through multi-modality inputs

Inventors: Waseem Akram Syed (San Jose, CA); Umamaheswara Rao Kothuri (Sunnyvale, CA)
Assignee: Intuit Inc.
G06F16/211G06F3/0481G06F16/215G06F16/248
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,736
App. No.
18/758,868
Granted
Mar 3, 2026
Kind
B2
Abstract

Certain aspects of the disclosure provide for processing multi-modality inputs to trigger dynamic actions. In examples, a method may include: receiving user input through one or more input modalities; converting the user input and one or more supported actions obtained from an actions repository into a unified schema; applying a set of moderation rules to the unified schema to generate a moderated input; generating a prompt based on the moderated input and one or more influencing examples; processing the prompt using one or more Large Language Models (LLMs) to obtain one or more matched actions and populated data models; augmenting the one or more matched actions with additional data; initiating one or more supported actions based on the one or more matched actions and populated data models; and generating a response to the user input based on results of executing the one or more supported actions.

Claims (64)

1 . A method for processing multi-modality inputs to trigger dynamic actions, the method comprising:

receiving user input through one or more input modalities, wherein the one or more input modalities include at least one of voice, image, or text;

converting the user input and one or more supported actions obtained from an actions repository into a unified schema, wherein the unified schema represents the user input and the one or more supported actions in a standardized format;

applying a set of moderation rules to the unified schema to generate a moderated input, including:

filtering irrelevant or unnecessary information from the unified schema; and

constraining the unified schema to adhere to specific formatting or structural requirements;

generating a prompt based on the moderated input and one or more influencing examples;

processing the prompt using one or more Large Language Models (LLMs) to obtain one or more matched actions and populated data models, wherein the one or more matched actions represents a mapping between the user input and the one or more supported actions;

augmenting the one or more matched actions with additional data, wherein the additional data provides context for executing the one or more matched actions;

initiating one or more supported actions based on the one or more matched actions and populated data models; and

generating a response to the user input based on results of executing the one or more supported actions.

2 . The method of claim 1 , wherein the one or more influencing examples includes at least one of: examples of desired response formats or structures; guidance on a tone, style, or level of formality to be used in the generated responses; or domain-specific terminology or conventions.

3 . The method of claim 1 , wherein generating the prompt based on the moderated input and the one or more influencing examples comprises:

combining the moderated input and the one or more influencing examples into a single input format; and

providing the combined moderated input to the one or more LLMs for processing.

4 . The method of claim 1 , wherein augmenting the one or more matched actions with additional data comprises:

identifying missing or incomplete information in the one or more matched actions; and

retrieving the missing or incomplete information from a data service.

5 . The method of claim 1 , wherein executing the one or more of the supported actions comprises executing the one or more of the supported actions through an adaptive user interface that dynamically adapts to specific requirements of a respective action of the one or more of the supported actions.

6 . The method of claim 5 , wherein the adaptive user interface comprises a set of reusable user interface components dynamically assembled based on requirements of each action.

7 . The method of claim 1 , wherein generating the response to the user input comprises presenting the response using an adaptive user interface.

8 . The method of claim 7 , wherein generating the response to the user input further comprises:

aggregating results of the one or more of the supported actions;

formatting the aggregated results into a user-friendly format; and

presenting the formatted aggregated results to a user through the adaptive user interface.

9 . The method of claim 8 , wherein the user-friendly format includes at least one of text, speech, visual elements, or interactive components that is specific to a respective action of the one or more of the supported actions.

10 . The method of claim 1 , wherein converting the user input and the one or more supported actions into the unified schema comprises:

identifying entities and intents associated with the user input; and

mapping the identified entities and intents to corresponding elements in the unified schema.

11 . The method of claim 1 , wherein initiating the one or more supported actions comprises:

prioritizing the one or more matched actions based on a set of criteria to obtain prioritized one or more matched actions; and

executing the prioritized one or more matched actions in a sequential order.

12 . The method of claim 11 , wherein the set of predefined criteria includes at least one of user preferences, action relevance, or action complexity.

13 . A processing system, comprising:

a memory comprising computer-executable instructions; and

a processor configured to execute the computer-executable instructions and cause the processing system to:

receive user input through one or more input modalities, wherein the one or more input modalities include at least one of voice, image, or text;

the user input and one or more supported actions obtained from an actions repository into a unified schema, wherein the unified schema represents the user input and the one or more supported actions in a standardized format;

apply a set of moderation rules to the unified schema to generate a moderated input, wherein to apply the set of moderation rules to the unified schema comprises:

to filter out irrelevant or unnecessary information from the unified schema, and

to constrain the unified schema to adhere to specific formatting or structural requirements;

generate a prompt based on the moderated input and one or more influencing examples;

process the prompt with one or more Large Language Models (LLMs) to obtain one or more matched actions and populated data models, wherein the one or more matched actions represents a mapping between the user input and the one or more supported actions;

augment the one or more matched actions with additional data, wherein the additional data provides context for executing the one or more matched actions;

initiate one or more of the one or more supported actions based on the one or more matched actions and populated data models; and

generate a response to the user input based on results of executing the one or more supported actions.

14 . The processing system of claim 13 , wherein to generate the prompt based on the moderated input and the one or more influencing examples, the computer-executable instructions are further executable by the processor to cause the processing system to:

combine the moderated input and the one or more influencing examples into a single input format; and

provide the combined moderated input to the one or more LLMs for processing.

15 . The processing system of claim 13 , wherein to augment the one or more matched actions with additional data, the computer-executable instructions are further executable by the processor to cause the processing system to:

identify missing or incomplete information in the one or more matched actions; and

retrieve the missing or incomplete information from a data service.

16 . The processing system of claim 13 , wherein to execute one or more of the supported actions, the computer-executable instructions are further executable by the processor to cause the processing system to execute one or more of the supported actions through an adaptive user interface that dynamically adapts to specific requirements of a respective action of the one or more of the supported actions.

17 . The processing system of claim 13 , wherein to generate the response to the user input, the computer-executable instructions are further executable by the processor to cause the processing system to present the response using an adaptive user interface.

18 . The processing system of claim 17 , wherein to generate the response to the user input, the computer-executable instructions are further executable by the processor to cause the processing system to:

aggregate results of the one or more of the supported actions;

format the aggregated results into a user-friendly format; and

present the formatted aggregated results to a user through the adaptive user interface.

19 . The processing system of claim 13 , wherein to convert the user input and the one or more supported actions into the unified schema, the computer-executable instructions are further executable by the processor to cause the processing system to:

identify entities and intents associated with the user input; and

map the identified entities and intents to corresponding elements in the unified schema.

20 . The processing system of claim 13 , wherein to initiate the one or more supported actions, the computer-executable instructions are further executable by the processor to cause the processing system to:

prioritize the one or more matched actions based on a set of criteria to obtain prioritized one or more matched actions; and

execute the prioritized one or more matched actions in a sequential order.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2024
From: SYED, WASEEM AKRAM; KOTHURI, UMAMAHESWARA RAO
To: INTUIT INC.
Reel/Frame 068129/0375 →
Continuity (1)
Related Publication 20260003833A1 · Jan 1, 2026
References Cited (23)
US 10686673B1 · Murugesan · 2020 [cited by examiner]
US 10915970B1 · Wang · 2021 [cited by examiner]
US 11070443B1 · Murugesan · 2021 [cited by examiner]
US 11647095B1 · Barsade · 2023 [cited by examiner]
US 12417250B1 · Sharma · 2025 [cited by examiner]
US 12436960B1 · Kulesza · 2025 [cited by examiner]
US 20100030525A1 · Dong · 2010 [cited by examiner]
US 20120116984A1 · Hoang · 2012 [cited by examiner]
US 20120158772A1 · Shafran · 2012 [cited by examiner]
US 20170060911A1 · Loscalzo · 2017 [cited by examiner]
US 20180068302A1 · Wang · 2018 [cited by examiner]
US 20190318413A1 · Schulz · 2019 [cited by examiner]
US 20220028020A1 · Wray · 2022 [cited by examiner]
US 20220392451A1 · Dai · 2022 [cited by examiner]
US 20230394478A1 · Fuentes · 2023 [cited by examiner]
US 20250225343A1 · Yu · 2025 [cited by examiner]
US 20250245092A1 · Gallagher, Jr. · 2025 [cited by examiner]
US 20250254240A1 · Greenberg · 2025 [cited by examiner]
US 20250265522A1 · Jagmohan · 2025 [cited by examiner]
US 20250307067A1 · Cho · 2025 [cited by examiner]
US 20250307688A1 · Li · 2025 [cited by examiner]
Unified user interface design: designing universally accessible interactions (Year: 2004). [cited by examiner]
Testing the Limits of Unified Sequence to Sequence LLM Pretraining on Diverse Table Data Tasks (Year: 2023). [cited by examiner]