Voice control with contextual keywords
Systems and processes for operating an intelligent automated assistant are provided. In some embodiments, contextual data is obtained and used to select a set of keywords (e.g., words or phrases) for voice control of an electronic device. When a speech input is received by the electronic device, a determination is made whether the speech input includes any of the selected keywords. If the speech input does include a selected keyword, an action is performed in response.
1 . An electronic device, comprising:
one or more processors;
a memory; and
one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:
obtaining first contextual information;
selecting, based on the first contextual information, a first set of one or more keywords, wherein selecting the first set of one or more keywords includes:
determining, based on the first contextual information, a first action that is available for performance; and
in response to determining the first action, selecting a first keyword for the first set of one or more keywords, wherein the first keyword corresponds to the first action;
receiving a speech input;
determining whether the speech input includes at least one keyword of the first set of one or more keywords; and
in response to determining that the speech input includes at least one keyword of the first set of one or more keywords, performing a first an action that corresponds to the at least one keyword.
2 . The electronic device of claim 1 , wherein obtaining the first contextual information includes receiving at least a first portion of the first contextual information from one or more sensors.
3 . The electronic device of claim 1 , wherein the first contextual information includes current device context.
4 . The electronic device of claim 3 , wherein the current device context includes user interface information.
5 . The electronic device of claim 3 , wherein the current device context includes application information.
6 . The electronic device of claim 1 , wherein the first contextual information includes user data.
7 . The electronic device of claim 1 , wherein selecting the first set of one or more keywords includes:
selecting a respective keyword of the first set of one or more keywords from a plurality of keywords.
8 . The electronic device of claim 7 , the one or more programs including instructions for:
prior to selecting the first set of one or more keywords:
receiving a natural-language speech input directed to a digital assistant, wherein the natural-language speech input includes a request to perform a third action; and
in accordance with a determination that the natural-language speech input satisfies keyword criteria, adding a candidate keyword to the plurality of keywords, wherein the candidate keyword corresponds to at least a portion of the natural-language speech input.
9 . The electronic device of claim 7 , wherein at least one keyword of the plurality of keywords is obtained from an application.
10 . The electronic device of claim 1 , wherein determining, based on the first contextual information, the first action that is available for performance includes:
determining whether the first action satisfies attention criteria.
11 . The electronic device of claim 1 , the one or more programs including instructions for:
obtaining second contextual information;
selecting, based on the second contextual information, a second set of one or more keywords different from the first set of one or more keywords;
receiving a second speech input;
determining whether the second speech input includes at least one keyword of the second set of one or more keywords; and
in accordance with a determination that the speech input includes at least one keyword of the second set of one or more keywords, performing a second action.
12 . The electronic device of claim 11 , wherein selecting the second set of one or more keywords in accordance with a determination that the second contextual information differs from the first contextual information.
13 . The electronic device of claim 1 , the one or more programs including instructions for:
providing an output indicating at least a first portion of the first set of one or more keywords.
14 . The electronic device of claim 1 , the one or more programs including instructions for:
in accordance with a determination that the speech input includes a trigger word in addition to the at least one keyword of the first set of one or more keywords, providing an output indicating that performance of the action can be caused with the at least one keyword and without the trigger word.
15 . The electronic device of claim 1 , wherein determining whether the speech input includes at least one keyword of the first set of one or more keywords includes:
extracting a set of acoustic features from the speech input.
16 . The electronic device of claim 15 , wherein determining whether the speech input includes at least one keyword of the first set of one or more keywords includes:
determining, based on the set of acoustic features, a confidence score representing a likelihood that the speech input includes the at least one keyword; and
determining whether the confidence score satisfies confidence criteria.
17 . The electronic device of claim 1 , wherein determining whether the speech input includes at least one keyword of the first set of one or more keywords is performed using a first language model.
18 . The electronic device of claim 1 , the one or more programs including instructions for:
in accordance with a determination that the speech input does not include at least one keyword of the first set of one or more keywords, forgoing performance of the action.
19 . The electronic device of claim 1 , wherein performing the action includes:
determining, using a second language model, at least one word included in the speech input.
20 . The electronic device of claim 19 , wherein the at least one keyword included in the speech input corresponds to a first portion of the speech input; and
wherein determining, using the second language model, the at least one word included in the speech input is performed in accordance with a determination that the speech input includes a second portion not corresponding to at least one keyword of the first set of one or more keywords.
21 . A method, comprising:
at an electronic device with one or more processors and memory:
obtaining first contextual information;
selecting, based on the first contextual information, a first set of one or more keywords, wherein selecting the first set of one or more keywords includes:
determining, based on the first contextual information, a first action that is available for performance; and
in response to determining the first action, selecting a first keyword for the first set of one or more keywords, wherein the first keyword corresponds to the first action;
receiving a speech input;
determining whether the speech input includes at least one keyword of the first set of one or more keywords; and
in response to determining that the speech input includes at least one keyword of the first set of one or more keywords, performing a first an action that corresponds to the at least one keyword.
22 . The method of claim 21 , wherein obtaining the first contextual information includes receiving at least a first portion of the first contextual information from one or more sensors.
23 . The method of claim 21 , wherein the first contextual information includes current device context.
24 . The method of claim 23 , wherein the current device context includes user interface information.
25 . The method of claim 23 , wherein the current device context includes application information.
26 . The method of claim 21 , wherein the first contextual information includes user data.
27 . The method of claim 21 , wherein selecting the first set of one or more keywords includes:
selecting a respective keyword of the first set of one or more keywords from a plurality of keywords.
28 . The method of claim 27 , further comprising:
prior to selecting the first set of one or more keywords:
receiving a natural-language speech input directed to a digital assistant, wherein the natural-language speech input includes a request to perform a third action; and
in accordance with a determination that the natural-language speech input satisfies keyword criteria, adding a candidate keyword to the plurality of keywords, wherein the candidate keyword corresponds to at least a portion of the natural-language speech input.
29 . The method of claim 27 , wherein at least one keyword of the plurality of keywords is obtained from an application.
30 . The method of claim 21 , wherein determining, based on the first contextual information, the first action that is available for performance includes:
determining whether the first action satisfies attention criteria.
31 . The method of claim 21 , further comprising:
in accordance with a determination that the speech input does not include at least one keyword of the first set of one or more keywords, forgoing performance of the action.
32 . A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of a first electronic device, cause the first electronic device to:
obtain first contextual information;
select, based on the first contextual information, a first set of one or more keywords, wherein selecting the first set of one or more keywords includes:
determining, based on the first contextual information, a first action that is available for performance; and
in response to determining the first action, selecting a first keyword for the first set of one or more keywords, wherein the first keyword corresponds to the first action;
receive a speech input;
determine whether the speech input includes at least one keyword of the first set of one or more keywords; and
in response to determining that the speech input includes at least one keyword of the first set of one or more keywords, perform an action that corresponds to the at least one keyword.
33 . The non-transitory computer-readable storage medium of claim 32 , wherein obtaining the first contextual information includes receiving at least a first portion of the first contextual information from one or more sensors.
34 . The non-transitory computer-readable storage medium of claim 32 , wherein the first contextual information includes current device context.
35 . The non-transitory computer-readable storage medium of claim 34 , wherein the current device context includes user interface information.
36 . The non-transitory computer-readable storage medium of claim 34 , wherein the current device context includes application information.
37 . The non-transitory computer-readable storage medium of claim 32 , wherein the first contextual information includes user data.
38 . The non-transitory computer-readable storage medium of claim 32 , wherein selecting the first set of one or more keywords includes:
selecting a respective keyword of the first set of one or more keywords from a plurality of keywords.
39 . The non-transitory computer-readable storage medium of claim 38 , the one or more programs further including instructions for:
prior to selecting the first set of one or more keywords:
receiving a natural-language speech input directed to a digital assistant, wherein the natural-language speech input includes a request to perform a third action; and
in accordance with a determination that the natural-language speech input satisfies keyword criteria, adding a candidate keyword to the plurality of keywords, wherein the candidate keyword corresponds to at least a portion of the natural-language speech input.
40 . The non-transitory computer-readable storage medium of claim 38 , wherein at least one keyword of the plurality of keywords is obtained from an application.
41 . The non-transitory computer-readable storage medium of claim 32 , wherein determining, based on the first contextual information, the first action that is available for performance includes:
determining whether the first action satisfies attention criteria.
42 . The non-transitory computer-readable storage medium of claim 32 , the one or more programs further including instructions for:
in accordance with a determination that the speech input does not include at least one keyword of the first set of one or more keywords, forgoing performance of the action.