IP Library Granted Patent US 10,262,655
Granted Patent B2
US 10,262,655 · App. 14/827,154 · Granted Apr 16, 2019

Augmentation of key phrase user recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,262,655
App. No.
14/827,154
Granted
Apr 16, 2019
Kind
B2
Abstract

Examples for augmenting user recognition via speech are provided. One example method comprises, on a computing device, monitoring a use environment via one or more sensors including an acoustic sensor, detecting utterance of a key phrase via data from the acoustic sensor, and based upon the selected data from the acoustic sensor and also on other environmental sensor data collected at different times than the selected data from the acoustic sensor, determining a probability that the key phrase was spoken by an identified user. The method further includes, if the probability meets or exceeds a threshold probability, then performing an action on the computing device.

Claims (41)

1. On a computing device, a method comprising:

monitoring a use environment via one or more sensors including an acoustic sensor;

detecting via speech recognition an utterance of a key phrase followed by a command via selected data from the acoustic sensor;

based upon the selected data from the acoustic sensor and also on other environmental sensor data collected at different times than the selected data from the acoustic sensor, the other environmental sensor data comprising additional acoustic data, performing voice recognition to determine a probability that the key phrase was spoken by an identified user; and

if the probability meets or exceeds a threshold probability, then attributing the command to the identified user and performing an action specified by the command on the computing device,

wherein performing voice recognition comprises determining whether the identified user was also speaking before or after the key phrase was uttered based on analyzing the additional acoustic data, and

wherein determining the probability comprises determining a higher probability where the additional acoustic data indicates that the identified user was also speaking before or after the key phrase was uttered than where the additional acoustic data indicates that the identified user was not also speaking before or after the key phrase was uttered.

2. The method of claim 1 , wherein the other environmental sensor data further comprises image data.

3. The method of claim 2 , further comprising identifying one or more persons in the use environment based on the image data, and wherein determining the probability comprises determining the probability based at least in part upon a determined identity of the one or more persons in the use environment.

4. The method of claim 1 , wherein the other environmental sensor data further comprises location data.

5. The method of claim 4 , wherein the location data comprises proximity data from a proximity sensor.

6. The method of claim 4 , wherein the location data comprises calendar information for the identified user.

7. The method of claim 1 , further comprising detecting a user behavioral pattern, and wherein determining the probability comprises determining the probability based at least in part upon the user behavioral pattern.

8. The method of claim 7 , wherein the user behavioral pattern comprises information regarding how often the identified user speaks.

9. A computing system, comprising:

one or more sensors including at least an acoustic sensor;

a logic machine; and

a storage machine holding instructions executable by the logic machine to

monitor a use environment via the one or more sensors including the acoustic sensor;

detect via speech recognition an utterance of a key phrase followed by a command via selected data from the acoustic sensor;

based upon the selected data from the acoustic sensor and also on other environmental sensor data collected at different times than the selected data from the acoustic sensor, the other environmental sensor data comprising additional acoustic data, perform voice recognition to determine a probability that the key phrase was spoken by an identified user;

analyze the additional acoustic data;

determine that the identified user was also speaking before or after the key phrase was uttered based on analyzing the additional acoustic data;

adjust the probability in response to determining that the identified user was also speaking before or after the key phrase was uttered based on the other environmental sensor data; and

if the probability meets or exceeds a threshold probability, then attribute the command to the identified user and perform an action specified by the command on the computing system.

10. The computing system of claim 9 , wherein the other environmental sensor data further comprises image data, and wherein the instructions are further executable to identify one or more persons in the use environment based on the image data, and to determine the probability based at least in part upon a determined identity of the one or more persons in the use environment.

11. The computing system of claim 9 , wherein the other environmental sensor data further comprises location data, the location data comprising one or more of proximity data from a proximity sensor and calendar information for the identified user.

12. The computing system of claim 11 , wherein the instructions are further executable to determine whether the identified user is scheduled to be in the use environment during a time that the utterance of key phrase was detected based on the calendar information, and if the identified user is scheduled to be in the use environment, increase the probability that the key phrase was spoken by the identified user.

13. The computing system of claim 9 , wherein the instructions are further executable to detect a user behavioral pattern based upon prior user behaviors detected via environmental sensing, the user behavioral pattern including information regarding how frequently the identified user speaks, and to determine the probability based on the average frequency the identified user speaks.

14. The computing system of claim 9 , wherein the additional acoustic data is collected before and/or after the utterance of the key phrase.

15. A computing system, comprising:

one or more sensors including at least an acoustic sensor;

a logic machine; and

a storage machine holding instructions executable by the logic machine to

monitor a use environment via the one or more sensors including the acoustic sensor;

detect utterance of a key phrase followed by a command via selected data from the acoustic sensor;

based upon the selected data from the acoustic sensor and also on other environmental sensor data collected at different times than the selected data from the acoustic sensor, determine a probability that the key phrase was spoken by an identified user; and

if the probability meets or exceeds a threshold probability, then attribute the command to the identified user and perform an action specified by the command on the computing system,

wherein the other environmental sensor data collected at different times than the selected data from the acoustic sensor comprises additional acoustic data collected before and/or after the utterance of the key phrase, and wherein to determine the probability that the key phrase was spoken by the identified user, the instructions are further executable to analyze the additional acoustic data to determine if the identified user was also speaking before or after the key phrase was uttered, and

increase the probability that the key phrase was spoken by the identified user if the identified user was also speaking before or after the key phrase was uttered.

16. The computing system of claim 15 , wherein the instructions are further executable to decrease the probability that the key phrase was spoken by the identified user if the analysis indicates the identified user was not speaking before or after the utterance of the key phrase.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 14, 2015
From: LOVITT, ANDREW WILLIAM
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 036333/0275 →