IP Library › Granted Patent US 10,803,869
Granted Patent B2
US 10,803,869 · App. 16/447,426 · Granted Oct 13, 2020

Voice enablement and disablement of speech processing functionality

Inventors: Shaman D'Souza (Seattle, WA); Ian Suttle (Seattle, WA); Srikanth Nori (Seattle, WA); Rajiv Reddy (Seattle, WA); Amol Kanitkar (Seattle, WA); Tina Orooji (Seattle, WA)
Assignee: AMAZON TECHNOLOGIES, INC.
G10L15/22G10L15/063G10L15/183G10L15/26G10L17/00G10L2015/0635G10L2015/223G10L2015/225H04M3/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,803,869
App. No.
16/447,426
Granted
Oct 13, 2020
Kind
B2
Abstract

Methods and devices for enabling and disabling applications using voice are described herein. In some embodiments, an individual speak an utterance to their electronic device, which may send audio data representing the utterance to a backend system. The backend system may generate text data representing the utterance, and may determine that an intent of the utterance was for an application to be enabled or disabled for their user account on the backend system. If, for instance, the intent was to enable the application, the backend system may receive one or more rules for performing functionalities of the application, as well as one or more sample templates of sample utterances and sample responses that future utterances may use when requesting the application. Furthermore, one or more invocation phrases that may be used within the future utterances to invoke the application may be received, along with slot values for the sample templates.

Claims (70)

1. A method performed by a computing system, comprising:

determining first audio data representing a first utterance;

generating first data representing the first audio data;

determining that the first data corresponds to an intent to enable a functionality of the computing system;

configuring, based at least in part on the first data corresponding to the intent, an updated speech processing component to cause a first action to be performed based at least in part on recognition of a first phrase;

detecting a first manual input;

determining, based at least in part on detecting the first manual input, second audio data representing a second utterance;

generating second data representing the second audio data;

processing the second data using the updated speech processing component to detect the first phrase; and

based at least in part on detection of the first phrase, causing the first action to be performed.

2. The method of claim 1 , further comprising invoking a first application associated with the first phrase.

3. The method of claim 1 , further comprising:

detecting a second manual input;

wherein determining the first audio data is further based at least in part on detecting the second manual input.

4. The method of claim 1 , further comprising:

determining the first phrase based at least in part on the first data.

5. The method of claim 1 , wherein processing the second data further comprises:

using the updated speech processing component to perform natural language understanding processing on the second data to determine an intent of the second utterance.

6. The method of claim 1 , further comprising:

receiving the first audio data from a first device;

determining that a first speech processing component is associated with the first device;

configuring the updated speech processing component at least in part by updating the first speech processing component;

receiving the second audio data from the first device; and

determining that the updated speech processing component is associated with the first device.

7. A computing system, comprising:

at least one processor; and

at least one computer-readable medium encoded with instructions which, when executed by the at least one processor, cause the computing system to:

determine first audio data representing a first utterance,

generate first data representing the first audio data,

determine that the first data corresponds to an intent to enable a functionality of the computing system,

configure, based at least in part on the first data corresponding to the intent, an updated speech processing component to cause a first action to be performed based at least in part on recognition of a first phrase,

detect a first manual input,

determine, based at least in part on the first manual input, second audio data representing a second utterance,

generate second data representing the second audio data,

process the second data using the updated speech processing component to detect the first phrase, and

based at least in part on detection of the first phrase, cause the first action to be performed.

8. The computing system of claim 7 , wherein the at least one computer-readable medium is encoded with additional instructions which, when executed by the at least one processor, further cause the computing system to:

invoke a first application associated with the first phrase.

9. The computing system of claim 7 , wherein the at least one computer-readable medium is encoded with additional instructions which, when executed by the at least one processor, further cause the computing system to:

detect a second manual input, and

determine the first audio data based at least in part on the second manual input.

10. The computing system of claim 7 , wherein the at least one computer-readable medium is encoded with additional instructions which, when executed by the at least one processor, further cause the computing system to:

determine the first phrase based at least in part on the first data.

11. The computing system of claim 7 , wherein the at least one computer-readable medium is encoded with additional instructions which, when executed by the at least one processor, further cause the computing system to process the second data at least in part by:

using the updated speech processing component to perform natural language understanding processing on the second data to determine an intent of the second utterance.

12. The computing system of claim 7 , wherein the at least one computer-readable medium is encoded with additional instructions which, when executed by the at least one processor, further cause the computing system to:

receive the first audio data from a first device;

determine that a first speech processing component is associated with the first device;

configure the updated speech processing component at least in part by updating the first speech processing component;

receive the second audio data from the first device; and

determine that the updated speech processing component is associated with the first device.

13. A method performed by a computing system, comprising:

determining first audio data representing a first utterance;

determining that the first audio data corresponds to an intent to enable a functionality of the computing system;

configuring, based at least in part on the first audio data corresponding to the intent, an updated speech processing component to cause a first action to be performed based at least in part on recognition of a first phrase;

detecting a first manual input;

determining, based at least in part on detecting the first manual input, second audio data representing to a second utterance;

processing the second audio data using the updated speech processing component to detect the first phrase; and

based at least in part on detection of the first phrase, causing the first action to be performed.

14. The method of claim 13 , further comprising:

receiving the first audio data from a first device;

determining that a first speech processing component is associated with the first device;

configuring the updated speech processing component at least in part by updating the first speech processing component;

receiving the second audio data from the first device; and

determining that the updated speech processing component is associated with the first device.

15. The method of claim 1 , wherein the first data comprises first text data and the second data comprises second text data.

16. The computing system of claim 7 , wherein the first data comprises first text data and the second data comprises second text data.

17. The method of claim 13 , further comprising:

detecting a second manual input;

wherein determining the first audio data is further based at least in part on detecting the second manual input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2019
From: D'SOUZA, SHAMAN; SUTTLE, IAN; NORI, SRIKANTH; REDDY, RAJIV; KANITKAR, AMOL; OROOJI, TINA
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 049541/0567 →
Continuity (2)
Continuation 15194453 · Jun 27, 2016
Related Publication 20190371329A1 · Dec 5, 2019
Cited By (1)
US 12,400,653