IP Library › Granted Patent US 12,430,011
Granted Patent B2
US 12,430,011 · App. 18/619,127 · Granted Sep 30, 2025

Voice assistant-enabled client application with user view context and multi-modal input support

Inventors: Tudor Buzasu Klein (Yokohama, JP); Viktoriya Taranov (Kirkland, WA); Sergiy Gavrylenko (Issaquah, WA); Jaclyn Carley Knapp (Redmond, WA); Andrew Paul McGovern (Redmond, WA); Harris Syed (Redmond, WA); Chad Steven Estes (Redmond, WA); Jesse Daniel Eskes Rusak (Redmond, WA); David Ernesto Heekin Burkett (Redmond, WA); Allison Anne O'Mahony (Redmond, WA); Ashok Kuppusamy (Redmond, WA); Jonathan Reed Harris (Redmond, WA); Jose Miguel Rady Allende (Redmond, WA); Diego Hernan Carlomagno (Redmond, WA); Talon Edward Ireland (Redmond, WA); Michael Francis Palermiti, II (Redmond, WA); Richard Leigh Mains (Redmond, WA); Jayant Krishnamurthy (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F3/0484G06F3/167G10L15/08G10L15/22G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,011
App. No.
18/619,127
Granted
Sep 30, 2025
Kind
B2
Abstract

Various embodiments discussed herein enable client applications to be heavily integrated with a voice assistant in order to perform commands associated with voice utterances of users via voice assistant functionality and also seamlessly cause client applications to automatically perform native functions as part of executing the voice utterance. Such heavy integration also allows particular embodiments to support multi-modal input from a user for a single conversational interaction. In this way, client application user interface interactions, such as clicks, touch gestures, or text inputs are executed alternative or in addition to the voice utterances.

Claims (45)

1. A system comprising:

at least one computer processor; and

one or more computer storage media storing computer-useable instructions that, when used by the at least one computer processor, cause the at least one computer processor to perform operations comprising:

detecting a first user action of a first user;

subsequent to the detecting, capturing audio data comprising a voice utterance of the first user;

receiving, via a user interface, a manual user input of the first user;

receiving an indication that the manual user input of the first user was performed by the first user later in time than the voice utterance of the first user; and

based at least in part on the indication that the manual user input of the first user was performed by the first user later in time than the voice utterance of the first user, responding to only the manual user input and refraining from responding to the voice utterance.

2. The system of claim 1 , wherein the operations further comprise:

receiving a second voice utterance of the first user and a second manual input of the first user;

determining that the second voice utterance is associated with a first response type;

determining that the second voice utterance was received prior to receiving the second manual user input; and

based at least in part on the determining that the second voice utterance is associated with the first response type and the determining that the second voice utterance was received prior to the receiving of the second manual user input, causing a response to the second voice utterance and refraining from responding to the second manual user input.

3. The system of claim 2 , wherein the operations further comprise tagging the response type with an identifier (ID) and caching the ID in computer memory, wherein the causing of the response to the second voice utterance is based on the ID cached in the computer memory.

4. The system of claim 1 , wherein the manual user input comprises at least one of a touch gesture by the first user, text entry input by the first user, or a pointer click by the first user.

5. The system of claim 1 , wherein the user action includes at least one of: a wake word issued by the first user or an interaction with an element of the user interface.

6. The system of claim 1 , wherein the operations further comprise:

based on the receiving of the manual user input, deactivating a microphone; and

in response to the deactivating of the microphone and based on the receiving of the manual user input, responding to only the manual user input.

7. The system of claim 1 , wherein the operations further comprise:

receiving a second voice utterance later in time relative to a second manual user input; and

responding to only the second voice utterance based on the second voice utterance being received later in time relative to the second manual user input, wherein the second manual user input is excluded from being responded to.

8. The system of claim 1 , wherein the operations further comprise:

receiving one of a second voice utterance of the first user or a second manual input of the first user;

determining that one of the second voice utterance or the second manual input conflicts with the manual user input; and

based at least in part on the determining, refraining from responding to both the voice utterance and one of the second voice utterance or the second manual input.

9. A computer-implemented method comprising:

capturing audio data comprising a voice utterance of a first user;

receiving, via a user interface, a manual user input of the first user;

receiving an indication that the voice utterance of the first user was issued later in time than the manual user input of the first user; and

based at least in part on the indication that the voice utterance of the first user was issued later in time than the manual user input of the first user, responding to only the voice utterance of the first user and refraining from responding to the manual user input of the first user.

10. The computer-implemented method of claim 9 , further comprising:

receiving a second voice utterance of the first user and a second manual input of the first user;

determining that the second voice utterance is associated with a first response type;

determining that the second voice utterance was received prior to receiving the second manual user input; and

based at least in part on the determining that the second voice utterance is associated with the first response type and the determining that the second voice utterance was received prior to the receiving of the second manual user input, causing a response to the second voice utterance and refraining from responding to the second manual user input.

11. The computer-implemented method of claim 10 , further comprising tagging the response type with an identifier (ID) and caching the ID in computer memory, wherein the causing of the response to the second voice utterance is based on the ID cached in the computer memory.

12. The computer-implemented method of claim 9 , wherein the manual user input comprises at least one of a touch gesture by the first user, text entry input by the first user, or a pointer click by the first user.

13. The computer-implemented method of claim 9 , further comprising: detecting a user action of a first user, and wherein the user action includes at least one of: a wake word issued by the first user or an interaction with an element of the user interface.

14. The computer-implemented method of claim 9 , further comprising:

based on receiving second manual user input, deactivating a microphone; and

in response to the deactivating of the microphone and based on the receiving of the second manual user input, responding to only the second manual user input.

15. The computer-implemented method of claim 9 , further comprising:

receiving a second manual input later in time relative to a second voice utterance; and

responding to only the second manual input based on the second manual input being received later in time relative to the voice utterance, wherein the second voice utterance is excluded from being responded to.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 2, 2025
From: KLEIN, TUDOR BUZASU; TARANOV, VIKTORIYA; GAVRYLENKO, SERGIY; ALLENDE, JOSE MIGUEL RADY; BURKETT, DAVID ERNESTO HEEKIN; CARLOMAGNO, DIEGO HERNAN; ESTES, CHAD STEVEN; HARRIS, JONATHAN REED; IRELAND, TALON EDWARD; KNAPP, JACLYN CARLEY; KRISHNAMURTHY, JAYANT; KUPPUSAMY, ASHOK; MCGOVERN, ANDREW PAUL; O'MAHONY, ALLISON ANNE; PALERMITI, MICHAEL FRANCIS, II; RUSAK, JESSE DANIEL ESKES; SYED, HARRIS; MAINS, RICHARD LEIGH
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 072810/0695 →
Continuity (4)
Continuation 17508762 · Oct 22, 2021
Continuation In Part 17364362 · Jun 30, 2021
Provisional Application 63165037 · Mar 23, 2021
Related Publication 20240241624A1 · Jul 18, 2024
References Cited (13)
US 5231691A · Yasuda · 1993 [cited by examiner]
US 8874447B2 · Da Palma · 2014 [cited by examiner]
US 20050171664A1 · Konig · 2005 [cited by examiner]
US 20110214162A1 · Brakensiek · 2011 [cited by examiner]
US 20140075330A1 · Kwon · 2014 [cited by examiner]
US 20180164957A1 · Schon · 2018 [cited by examiner]
US 20180173405A1 · Pereira · 2018 [cited by examiner]
US 20180335921A1 · Karunamuni · 2018 [cited by examiner]
US 20180335939A1 · Karunamuni · 2018 [cited by examiner]
US 20200110532A1 · Mani · 2020 [cited by examiner]
US 20210093794A1 · Volkar · 2021 [cited by examiner]
US 20220308718A1 · Klein · 2022 [cited by examiner]
Notice of Allowance mailed on Mar. 14, 2024, in U.S. Appl. No. 18/231,333, 10 pages. [cited by applicant]