IP Library › Granted Patent US 11,455,987
Granted Patent B1
US 11,455,987 · App. 16/294,747 · Granted Sep 27, 2022

Multiple skills processing

Inventors: Rohin Dabas (Kirkland, WA); Troy Dean Schuring (Maple Valley, WA); Rashmi Tonge (Bellevue, WA); Michael James Montgomery (Seattle, WA); Kevindra Pal Singh (Seattle, WA); Adam Baran (Redmond, WA); David Thomas (Woodinville, WA); Nnenna Eleanya Okwara (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G10L15/22G10L15/1815G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,455,987
App. No.
16/294,747
Granted
Sep 27, 2022
Kind
B1
Abstract

Described herein is a system for enabling a user to perform complex goals using multiple skills/applications of an intelligent assistant device. Skills may register as consumers of an action or providers of an action, and the consumer skills may be configured to invoke provider skills to perform actions. The system receives a request to perform an action from a skill along with some action data. The system validates the action data, selects another skill to perform the action, and forwards the request to the selected skill to perform the action.

Claims (117)

1. A computer-implemented method comprising:

receiving, from a first skill capable of performing a first action, first data to enable the first skill to request performance of a second action that is performable by at least one other skill different than the first skill;

storing the first data to enable the first skill to request performance of the second action;

after storing the first data, receiving audio data representing an utterance;

associating the audio data with a session identifier for a dialog session;

processing the audio data to determine ASR data corresponding to the utterance;

determining, using the ASR data, an intent to perform the first action and entity data associated with execution of the first action;

determining the first skill, having one or more processing components is capable of performing the first action;

sending, to the first skill, a first request to perform the first action using the entity data, the first request including the session identifier;

receiving, from the first skill, a second request to perform the second action, the second action being determined by the first skill based on the first request, wherein the second request includes the session identifier;

determining, using the first data, that the first skill is authorized to request performance of the second action;

determining a second skill, having one or more processing components, is capable of performing the second action;

sending, to the second skill, a third request to perform the second action, the third request including at least the ASR data and the session identifier;

receiving, from the second skill, second data indicating an acknowledgment of the third request, the second data including the session identifier; and

sending third data to the first skill, the third data indicating the acknowledgement of the third request.

2. The computer-implemented method of claim 1 , further comprising:

querying a database to determine a list of candidate skills associated with the second action;

determining a filtered list of candidate skills based on the entity data;

determining a user profile associated with the audio data; and

determining a ranked list of candidate skills using preference data associated with the user profile; and

wherein determining the second skill includes selecting the second skill from the ranked list of candidate skills.

3. The computer-implemented method of claim 1 , further comprising:

determining fourth data indicating data parameters for the second skill to perform the second action; and

determining, using the fourth data, that the entity data is valid for performance of the second action.

4. The computer-implemented method of claim 1 , further comprising:

prior to receiving the audio data:

receiving fourth data from the second skill, the fourth data indicating action data required by the second skill to perform the second action; and

storing the fourth data in a database associating the second skill with the second action,

wherein determining the second skill is associated with the second action further comprises:

querying the database using the second action; and

determining the second skill is associated with the second action based on retrieving the fourth data.

5. A computer-implemented method comprising:

receiving, from a first skill capable of performing a first action, first data to enable the first skill to request performance of a second action that is performable by at least one other skill different than the first skill;

after receiving the first data, receiving input data associated with a session identifier;

determining, using the input data, that the first action is to be performed;

determining the first skill is capable of performing the first action;

sending, to the first skill, the session identifier and a first request to perform the first action;

receiving, from the first skill and after sending the first request, a second request to perform the second action, the second action being determined by the first skill;

determining, using the first data, that the first skill is authorized to request performance of the second action;

determining a second skill is capable of performing the second action; and

sending, to the second skill, the session identifier and a third request to perform the second action.

6. The computer-implemented method of claim 5 , further comprising:

querying a data source to determine a first list of skills capable of performing the second action;

determining a user profile associated with the input data; and

determining a second list of skills based on preference data associated with the user profile,

wherein determining the second skill includes selecting the second skill from the second list of skills.

7. The computer-implemented method of claim 5 , further comprising:

determining, using the input data, second data to be used to perform the second action;

receiving, from at least one data storage, fourth data indicating parameters needed by the second skill to perform the second action; and

determining, using the fourth data, that the second data is valid for the second action.

8. The computer-implemented method of claim 5 , further comprising:

prior to receiving the input data:

receiving, from the second skill, second data to enable the second skill to perform the second action; and

storing the second data in a data storage associating the second skill with the second action, the second data indicating action data required by the second skill to perform the second action,

wherein determining the second skill is associated with the second action further comprises:

querying the data storage using the second action; and

determining the second skill is associated with the second action based on retrieving the second data.

9. The computer-implemented method of claim 5 , further comprising:

determining the first skill having one or more components capable of performing the first action, the first skill configured to output natural language data.

10. The computer-implemented method of claim 5 , further comprising:

receiving audio data representing an utterance spoken by a user;

performing automatic speech recognition (ASR) using the audio data to determine the input data;

performing natural language understanding (NLU) on the input data to determine an intent; and

determining the first action based on the intent.

11. The computer-implemented method of claim 5 , further comprising:

generating audio output data requesting confirmation from a user to proceed with the second skill;

receiving input audio data;

performing ASR using the input audio data to determine text data representing an utterance in the input audio data;

performing NLU processing using the text data to determine affirmation from the user; and

sending the third request to the second skill in response to the affirmation from the user.

12. The computer-implemented method of claim 5 , further comprising:

determining that second data is required to perform the second action by the second skill; and

generating audio output data requesting the second data from a user.

13. A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

receive, from a first skill capable of performing a first action, first data to enable the first skill to request performance of a second action that is performable by at least one skill different than the first skill:

after receiving the first data, receive input data;

determining, using the input data, a first action is to be performed;

determine the first skill is capable of performing the first action;

send, to the first skill, a first request to perform the first action;

receive, from the first skill and after sending the first request, a second request to perform a second action, the second action being determined by the first skill;

determine, using the first data, that the first skill is authorized to request performance of the second action;

determine a second skill is capable of performing the second action; and

send, to the second skill, a third request to perform the second action.

14. The system of claim 13 , wherein the instructions, when executed by the at least one processor, further cause the system to:

query a data source to determine a first list of skills capable of performing the second action;

determine a user profile associated with the input data; and

determine a second list of skills based on preference data associated with the user profile,

wherein determine the second skill includes selecting the second skill from the second list of skills.

15. The system of claim 13 , wherein the instructions, when executed by the at least one processor, further cause the system to:

determine, using the input data, second data to be used to perform the second action;

receive, from at least one data storage, fourth data indicating parameters needed by the second skill to perform the second action; and

determine, using the fourth data, that the second data is valid for the second action.

16. The system of claim 13 , wherein the instructions, when executed by the at least one processor, further cause the system to:

prior to receiving the input data:

receive, from the second skill, second data to enable the second skill to perform the second action; and

store the second data in a data storage associating the second skill with the second action, the second data indicating action data required by the second skill to perform the second action,

wherein determine the second skill is associated with the second action further comprises:

query the data storage using the second action; and

determine the second skill is associated with the second action based on retrieving the second data.

17. The system of claim 13 , wherein the instructions, when executed by the at least one processor, further cause the system to:

determine the first skill having one or more components capable of performing the first action, the first skill configured to output natural language data.

18. The system of claim 13 , wherein the instructions, when executed by the at least one processor, further cause the system to:

receive audio data representing an utterance spoken by a user;

perform automatic speech recognition (ASR) using the audio data to determine the input data;

perform natural language understanding (NLU) on the input data to determine an intent; and

determine the first action based on the intent.

19. The system of claim 13 , wherein the instructions, when executed by the at least one processor, further cause the system to:

generate audio output data requesting confirmation from a user to proceed with the second action;

receive input audio data;

perform ASR using the input audio data to determine text data representing an utterance in the input audio data;

perform NLU processing using the text data to determine affirmation from the user; and

send the third request to the second skill in response to the affirmation from the user.

20. The system of claim 13 , wherein the instructions, when executed by the at least one processor, further cause the system to:

determine that second data is required to perform the second action by the second skill; and

generate audio output data requesting the second data from a user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2019
From: DABAS, ROHIN; SCHURING, TROY DEAN; TONGE, RASHMI; MONTGOMERY, MICHAEL JAMES; SINGH, KEVINDRA PAL; BARAN, ADAM; THOMAS, DAVID; OKWARA, NNENNA ELEANYA
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 048523/0672 →
Cited By (4)
US 12,230,266 US 12,271,407 US 12,456,454 US 12,597,425