IP Library Granted Patent US 12670900
Granted Patent B2
US 12670900 · App. 18/535,725 · Granted Jun 30, 2026

Intent evaluation for smart assistant computing system

Inventors: Devin Samuel Jacob Caplow-Munro (Seattle, WA); Amit Kaistha (Coppell, TX); Denys V Yaremenko (Bellevue, WA); Brett Andrew Tomky (Seattle, WA); Charbel Khawand (Sammamish, WA); Andrew James Hillenius (Woodinville, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/1815G06V10/764G06V40/10G06V40/20G10L15/22G10L15/25G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670900
App. No.
18/535,725
Granted
Jun 30, 2026
Kind
B2
Abstract

A method for user intent evaluation includes receiving recorded speech of a human user. One or more attention indicators are detected in an image of the human user. Using a trained command recognition model, a command confidence is estimated indicating a confidence that the recorded human speech includes a command for a smart assistant computing system. Based at least in part on detecting the one or more attention indicators, and the command confidence exceeding a command confidence threshold, the human user is classified as intending to interact with the smart assistant computing system.

Claims (39)

1 . A method for user intent evaluation, the method comprising:

receiving recorded human speech of a human user;

detecting one or more attention indicators in an image of the human user;

using a trained command recognition model, estimating a command confidence that the recorded human speech includes a command for a smart assistant computing system;

based at least in part on detecting the one or more attention indicators, and the command confidence exceeding a command confidence threshold, classifying the human user as intending to interact with the smart assistant computing system; and

upon classifying the human user as intending to interact with the smart assistant computing system, reducing the command confidence threshold for a subsequent time interval.

2 . The method of claim 1 , wherein the one or more attention indicators include a determination that a gaze vector of the human user is directed toward the smart assistant computing system.

3 . The method of claim 1 , wherein the one or more attention indicators include a determination that the human user is performing an interaction-initiating gesture.

4 . The method of claim 3 , wherein the interaction-initiating gesture is one of a plurality of predefined gestures, the plurality of predefined gestures including at least the interaction-initiating gesture and an interaction-terminating gesture, and wherein the method further comprises, upon detecting the interaction-terminating gesture in a subsequent image captured at a subsequent time, classifying the human user as no longer intending to interact with the smart assistant computing system at the subsequent time.

5 . The method of claim 1 , further comprising, upon classifying the human user as intending to interact with the smart assistant computing system, displaying an intent recognition notification at the smart assistant computing system.

6 . The method of claim 1 , further comprising, based at least in part on classifying the human user as intending to interact with the smart assistant computing system, outputting, via the smart assistant computing system, a response generated based at least in part on the recorded human speech.

7 . The method of claim 6 , further comprising, prior to outputting the response, prompting the human user to confirm whether they intend to interact with the smart assistant computing system, and outputting the response upon receiving an intent confirmation from the human user.

8 . The method of claim 6 , wherein the response is generated by a language model previously trained to receive a digital representation of human speech as an input, and generate natural language responses as an output.

9 . The method of claim 8 , wherein the language model is implemented by a server computing device, and wherein the method further comprises recording the recorded human speech at a microphone of the smart assistant computing system, transmitting the recorded human speech to the server computing device over a computer network, and receiving the response from the server computing device over the computer network.

10 . The method of claim 6 , wherein the response is generated prior to classification of the human user as intending to interact with the smart assistant computing system, and wherein the method further comprises, prior to outputting the response, displaying a response pending notification at the smart assistant computing system.

11 . A smart assistant computing system, comprising:

a logic subsystem; and

a storage subsystem holding instructions executable by the logic subsystem to:

receive recorded human speech of a human user;

detect one or more attention indicators in an image of the human user;

using a trained command recognition model, estimate a command confidence that the recorded human speech includes a command for the smart assistant computing system; and

based at least in part on detecting the one or more attention indicators, and the command confidence exceeding a command confidence threshold, classify the human user as intending to interact with the smart assistant computing system, wherein the command confidence threshold is dynamically changed based at least in part on a number of users in a surrounding environment of the smart assistant computing system.

12 . The smart assistant computing system of claim 11 , wherein the instructions are further executable to, upon classifying the human user as intending to interact with the smart assistant computing system, reduce the command confidence threshold for a subsequent time interval.

13 . The smart assistant computing system of claim 11 , wherein the one or more attention indicators include a determination that a gaze vector of the human user is directed toward the smart assistant computing system.

14 . The smart assistant computing system of claim 11 , wherein the one or more attention indicators include a determination that the human user is performing an interaction-initiating gesture.

15 . The smart assistant computing system of claim 11 , wherein the instructions are further executable to, based at least in part on classifying the human user as intending to interact with the smart assistant computing system, output, via the smart assistant computing system, a response generated based at least in part on the recorded human speech.

16 . The smart assistant computing system of claim 14 , wherein the response is generated prior to classification of the human user as intending to interact with the smart assistant computing system, and wherein the instructions are further executable to, prior to outputting the response, display a response pending notification at the smart assistant computing system.

17 . A method for user intent evaluation at a smart assistant computing system, the method comprising:

recording human speech of a human user via a microphone of the smart assistant computing system;

detecting one or more attention indicators in an image of the human user captured via a camera of the smart assistant computing system;

using a trained command recognition model, estimating a command confidence that the recorded human speech includes a command for the smart assistant computing system;

based at least in part on detecting the one or more attention indicators, and the command confidence exceeding a command confidence threshold, classifying the human user as intending to interact with the smart assistant computing system;

outputting a response generated based at least in part on the human speech;

detecting a subsequent one or more attention indicators in a subsequent image of the human user captured at a subsequent time; and

based at least in part on the subsequent one or more attention indicators, classifying the human user as still intending to interact with the smart assistant computing system at the subsequent time.

18 . The method of claim 17 , further comprising, upon classifying the human user as intending to interact with the smart assistant computing system, displaying an intent recognition notification at the smart assistant computing system.

19 . The method of claim 17 , further comprising, prior to outputting the response, prompting the human user to confirm whether they intend to interact with the smart assistant computing system, and outputting the response upon receiving an intent confirmation from the human user.

20 . The method of claim 17 , wherein the response is generated by a language model previously trained to receive a digital representation of human speech as an input, and generate natural language responses as an output;

wherein the language model is implemented by a server computing device, and wherein the method further comprises transmitting the recorded human speech to the server computing device over a computer network, and receiving the response from the server computing device over the computer network.