SYSTEMS AND METHODS FOR DETERMINING WHETHER TO TRIGGER A VOICE CAPABLE DEVICE BASED ON SPEAKING CADENCE
Systems and methods are described for determining whether to activate a voice activated device based on a speaking cadence of the user. When the user speaks with a first cadence the system may determine that the user does not intend to activate the device and may accordingly not to trigger a voice activated device. When the user speaks with a second cadence the system may determine that the user does wish to trigger the device and may accordingly trigger the voice activated device.
1 . A method comprising:
monitoring for a plurality of voice inputs from a user to activate a voice capable device;
determining an average position of a trigger word based at least in part on the plurality of voice inputs;
determining a threshold maximum position of the trigger word based at least in part on the average position of the trigger word;
receiving a voice input from the user to activate the voice capable device, wherein the voice input comprises the trigger word;
determining a position of the trigger word in the received voice input; and
based at least in part on determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word, refraining from activating the voice capable device.
2 . The method of claim 1 , wherein the determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word comprises comparing the position of the trigger word in the received voice input to the threshold maximum position.
3 . The method of claim 1 , wherein the determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word comprises:
detecting a plurality of utterances comprising the trigger word from the plurality of voice inputs;
storing, in a profile of the user, a subset of utterances from the plurality of utterances comprising the trigger word, wherein the subset of utterances comprises the trigger word in a position that is within the threshold maximum position of the trigger word; and
comparing the received voice input to each utterance of the stored subset of utterances.
4 . The method of claim 1 , further comprising:
receiving the voice input from the user;
transmitting a region and a language associated with the user to a database; and
based at least in part on the transmitting and determining that the user is attempting to activate the voice capable device, receiving a value indicating the threshold maximum position in the received voice input where the trigger word appears.
5 . The method of claim 1 , wherein the determining that the position of the trigger word in the received voice input is greater than the threshold maximum position indicates that an inclusion of the trigger word in the received voice input is not intended to activate the voice capable device.
6 . The method of claim 1 , further comprising:
determining a threshold minimum position of the trigger word based at least in part on the average position of the trigger word; and
based at least in part on determining that the position of the trigger word in the received voice input is greater than the threshold minimum position of the trigger word, activating the voice capable device.
7 . The method of claim 6 , wherein the activating the voice capable device comprises:
determining whether a portion of the received voice input matches a word associated with a function of the voice capable device; and
in response to determining that the portion of the voice input matches the word associated with the function of the voice capable device, performing the function of the voice capable device.
8 . The method of claim 1 , wherein the average position of the trigger word is the threshold maximum position of the trigger word.
9 . The method of claim 1 , further comprising detecting the trigger word based at least in part on a fingerprint associated with the trigger word.
10 . The method of claim 1 , further comprising:
retrieving, from a profile of the user, demographic information corresponding to the user;
identifying, based at least in part on the demographic information corresponding to the user, a template speaking cadence;
storing, in the profile of the user, the template speaking cadence as the template speaking cadence of the user;
comparing a cadence associated with the received voice input to the template speaking cadence of the user; and
based at least in part on the comparing, determining whether the received voice input intends to activate the voice capable device.
11 . A system comprising:
a memory;
an input/output (I/O) circuitry; and
a control circuitry configured to:
monitor for a plurality of voice inputs from a user to activate a voice capable device;
determine an average position of a trigger word based at least in part on the plurality of voice inputs, wherein the average position of the trigger word is stored in the memory;
determine a threshold maximum position of the trigger word based at least in part on the average position of the trigger word;
wherein the I/O circuitry is configured to:
receive a voice input from the user to activate the voice capable device, wherein the voice input comprises the trigger word;
wherein the control circuitry is configured to:
determine a position of the trigger word in the received voice input; and
wherein the I/O circuitry is configured to:
based at least in part on determining that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word, refrain from activating the voice capable device.
12 . The system of claim 11 , wherein the control circuitry is configured to determine that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word by comparing the position of the trigger word in the received voice input to the threshold maximum position.
13 . The system of claim 11 , wherein the control circuitry is configured to determine that the position of the trigger word in the received voice input is greater than the threshold maximum position of the trigger word by:
detecting a plurality of utterances comprising the trigger word from the plurality of voice inputs;
storing, in a profile of the user, a subset of utterances from the plurality of utterances comprising the trigger word, wherein the subset of utterances comprises the trigger word in a position that is within the threshold maximum position of the trigger word; and
comparing the received voice input to each utterance of the stored subset of utterances.
14 . The system of claim 11 , wherein the I/O circuitry is further configured to:
receive the voice input from the user;
transmit a region and a language associated with the user to a database; and
based at least in part on the transmitting and determining that the user is attempting to activate the voice capable device, receive a value indicating the threshold maximum position in the received voice input where the trigger appears.
15 . The system of claim 11 , wherein the control circuitry is configured to determine that the position of the trigger word in the received voice input is greater than the threshold maximum position indicates that an inclusion of the trigger word in the received voice input is not intended to activate the voice capable device.
16 . The system of claim 11 , wherein the control circuitry is further configured to:
determine a threshold minimum position of the trigger word based at least in part on the average position of the trigger word; and
wherein the I/O circuitry is further configured to:
based at least in part on determining that the position of the trigger word in the received voice input is greater than the threshold minimum position of the trigger word, activate the voice capable device.
17 . The system of claim 16 , wherein the I/O circuitry is configured to activate the voice capable device by:
determining, via the control circuitry, whether a portion of the received voice input matches a word associated with a function of the voice capable device; and
in response to determining that the portion of the voice input matches the word associated with the function of the voice capable device, performing the function of the voice capable device.
18 . The system of claim 11 , wherein the average position of the trigger word is the threshold maximum position of the trigger word.
19 . The system of claim 11 , wherein the I/O circuitry is further configured to detect the trigger word based at least in part on a fingerprint associated with the trigger word.
20 . The system of claim 11 , wherein the I/O circuitry is further configured to:
retrieve, from a profile of the user, demographic information corresponding to the user; wherein the control circuitry is further configured to:
identify, based at least in part on the demographic information corresponding to the user, a template speaking cadence;
store, in the profile of the user, the template speaking cadence as the template speaking cadence of the user;
compare a cadence associated with the received voice input to the template speaking cadence of the user; and
based at least in part on the comparing, determine whether the received voice input intends to activate the voice capable device.