IP Library Granted Patent US 10,504,513
Granted Patent B1
US 10,504,513 · App. 15/716,353 · Granted Dec 10, 2019

Natural language understanding with affiliated devices

Inventors: Timothy Thomas Gray (Seattle, WA); Michal Grzegorz Kurpanik (Bytow, PL); Jenny Toi Wah Lam (Bainbridge Island, WA); Sarveshwar Nigam (Yorba Linda, CA); Shirin Saleem (Belmont, MA); Jonhenry A. Righter (Mountlake Terrace, WA); Jeremy Richard Hill (Seattle, WA); Kavya Ravikumar (Mercer Island, WA); Joe Virgil Fernandez (Seattle, WA); Kynan Dylan Antos (Seattle, WA); Kelly James Vanee (Shoreline, WA)
Assignee: AMAZON TECHNOLOGIES, INC.
G10L15/22G10L15/02G10L15/1815G10L15/30G10L17/005G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,504,513
App. No.
15/716,353
Granted
Dec 10, 2019
Kind
B1
Abstract

A dock device connects participating devices such as a tablet device and an audio activated device, allowing them to operate as a single device. These participating devices may be associated with different accounts, each account being associated with particular “speechlets” or data processing functions. A natural language understanding (NLU) system uses NLU models to process text obtained from an automatic speech recognition (ASR) system to determine a set of possible intents. A second set of possible intents may then be generated that is limited to those possible intents that correspond to the speechlets associated with the docked device. The intents within the second set of possible intents are ranked, and the highest ranked intent may be deemed to be the intent of the user. Command data corresponding to the highest ranked intent may be generated and used to perform the action associated with that intent.

Claims (112)

1. A system comprising:

at least one memory storing computer-executable instructions; and

at least one processor in communication with the at least one memory, the at least one processor executing the computer-executable instructions to:

determine a first device is associated with a user account;

determine a first speechlet associated with the user account, wherein the first speechlet comprises a first set of data processing functions available to the first device;

determine a second device is associated with the user account;

determine a second speechlet associated with the user account, wherein the second speechlet comprises a second set of data processing functions available to the second device;

generate a speechlet set comprising the first set of data processing functions and the second set of data processing functions;

determine a third device that is associated with the first device;

determine an output capability that is indicative of a type of output that the third device is able to present;

receive first data from the first device;

process the first data using a first natural language understanding (NLU) model to determine a set of possible intents that are representative of intended actions as expressed in the first data that are available in the speechlet set;

based at least in part on the output capability, determine a first intent of the set of possible intents as a first ranked intent associated with performing an action;

generate command data corresponding to the first ranked intent; and

send the command data to another device.

2. The system of claim 1 , further comprising:

at least one memory storing second computer-executable instructions; and

at least one processor in communication with the at least one memory, the at least one processor executing the second computer-executable instructions to:

determine a dock device that is associated with the user account; and

wherein a rank of an intent in the set of possible intents is based at least in part on an output capability of the dock device.

3. A system comprising:

at least one memory storing computer-executable instructions; and

at least one processor in communication with the at least one memory, the at least one processor executing the computer-executable instructions to:

determine a first device is associated with a user account;

determine a second device is associated with the user account;

determine speechlet data indicative of one or more speechlets available to the first device and the second device to process one or more intents;

determine a third device that is associated with the first device;

determine an output capability that is indicative of a type of output that the third device is able to present;

receive first data from one or more of the first device or the second device;

process the first data using a first natural language understanding (NLU) model to determine a set of possible intents that are representative of intended actions as expressed in the first data that are available to the one or more speechlets indicated by the speechlet data;

rank the set of possible intents;

based at least in part on the output capability, select, from the set of possible intents, a first ranked intent associated with performing an intended action; and

generate command data corresponding to the first ranked intent.

4. The system of claim 3 , the at least one processor further executing the computer-executable instructions to:

determine the one or more speechlets comprise a first speechlet associated with the first device, wherein the first speechlet comprises a first set of data processing functions; and

determine the one or more speechlets comprise a second speechlet associated with the second device, wherein the second speechlet comprises a second set of data processing functions;

wherein the speechlet data comprises a third set of data processing functions from the first speechlet and from the second speechlet.

5. The system of claim 3 , the at least one processor further executing the computer-executable instructions to:

determine a second natural language understanding (NLU) model associated with the first device;

determine a third NLU model associated with the second device; and

wherein the first NLU model comprises the second NLU model and the third NLU model.

6. The system of claim 3 , the at least one processor executing the computer-executable instructions to determine the speechlet data by executing instructions to:

send at least one command to present a user interface to one or more of the first device or the second device;

receive from the one or more of the first device or the second device, selection data indicative of designation of the user account as obtained with the user interface; and

determine one or more data processing functions associated with the user account.

7. The system of claim 3 , the at least one processor further executing the computer-executable instructions to:

determine identity data indicative of an identity of a speaker as represented by audio data;

determine the user account is associated with the identity data;

determine one or more data processing functions associated with the user account; and

designate the one or more data processing functions associated with the user account as the speechlet data.

8. The system of claim 3 , the at least one processor further executing the computer-executable instructions to:

determine the one or more speechlets comprise a first speechlet that is associated with the user account;

determine the one or more speechlets comprise a second speechlet that is associated with the user account; and

wherein the speechlet data comprises the first speechlet and the second speechlet.

9. The system of claim 3 , the at least one processor executing the computer-executable instructions to:

determine a dock identifier that is associated with one or more of the first device or the second device; and

the at least one processor executing the computer-executable instructions to determine the speechlet data by executing instructions to:

determine one or more data processing functions that are associated with the dock identifier, wherein the speechlet data comprises information indicative of availability of the one or more data processing functions.

10. The system of claim 3 , the at least one processor further executing the computer-executable instructions to:

determine the one or more speechlets comprise a first speechlet available to the first device and a second speechlet available to the second device;

determine first data indicative of first content available to the user account by using the first speechlet;

determine second data indicative of second content available to the user account by using the second speechlet;

select the first content based at least in part on the first data and the second data;

determine data processing functions that are accessible to the user account; and

designate the data processing functions that are accessible to the user account as the speechlet data.

11. The system of claim 3 , the at least one processor further executing the computer-executable instructions to:

receive image data from the one or more of the first device or the second device;

determine a count of people represented in the image data;

determine the count exceeds a threshold value; and

responsive to the count exceeding the threshold value, the at least one processor further executing the computer-executable instructions to determine the speechlet data by executing instructions to:

determine the one or more speechlets comprise a first speechlet that is available to the first device;

determine the one or more speechlets comprise a second speechlet that is available to the first device; and

wherein the speechlet data comprises the first speechlet and the second speechlet.

12. The system of claim 3 , wherein each intent of the set of possible intents is associated with a confidence value; and further wherein the rank of the set of possible intents is based on the confidence value for each of the intents in the set of possible intents.

13. A method comprising:

determining a first device is associated with a user account;

determining a second device is associated with the user account;

determining a third device that is associated with the first device;

determining speechlet data that is indicative of one or more data processing functions available to the first device and the second device to process one or more intents;

determining an output capability that is indicative of a type of output that the third device is able to present;

receiving first data from one or more of the first device or the second device;

processing the first data using one or more natural language understanding (NLU) models to determine a set of possible intents that are representative of intended actions as expressed in the first data that are available to the one or more data processing functions indicated by the speechlet data;

based at least in part on the output capability and from the set of possible intents, determining a first ranked intent associated with performing an intended action; and

generating command data corresponding to the first ranked intent.

14. The method of claim 13 , further comprising:

receiving user input indicative of the user account; and

determining information indicative of one or more data processing functions that are associated with the user account.

15. The method of claim 13 , further comprising:

determining identity data indicative of an identity of a speaker as represented by the first data;

determining the user account is associated with the identity data; and

determining information indicative of one or more data processing functions that are associated with the user account.

16. The method of claim 13 , further comprising:

determining the one or more data processing functions comprise a first speechlet that is available to the first device; and

determining the one or more data processing functions comprise a second speechlet that is available to the second device.

17. The method of claim 13 , further comprising:

determining a dock identifier that is indicative of a dock device; and

determining one or more data processing functions that are associated with the dock identifier; and

wherein the speechlet data is indicative of the one or more data processing functions that are associated with the dock identifier.

18. The method of claim 13 , further comprising:

determining first content data indicative of first content available to the first device;

determining second content data indicative of second content available to the second device;

selecting the first content based at least in part on the first content data and the second content data; and

wherein the speechlet data is indicative of one or more data processing functions that are accessible to the user account.

19. The method of claim 13 , further comprising:

sending the command data to a computing device associated with the one or more data processing functions;

performing the one or more data processing functions; and

presenting output using the one or more of the first device or the second device.

20. The method of claim 13 , further comprising:

determining a second natural language understanding (NLU) model associated with the first device;

determining a third NLU model associated with the second device; and

wherein the NLU model comprises the second NLU model and the third NLU model.

21. The method of claim 13 , wherein each intent of the set of possible intents is associated with a confidence value; and further wherein a rank of the set of possible intents is based on the confidence value for each of the intents in the set of possible intents.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNOR NAME PREVIOUSLY RECORDED AT REEL: 050616 FRAME: 0320. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 11, 2019
From: GRAY, TIMOTHY THOMAS; KURPANIK, MICHAL GRZEGORZ; LAM, JENNY TOI WAH; NIGAM, SARVESHWAR; SALEEM, SHIRIN; RIGHTER, JONHENRY A.; HILL, JEREMY RICHARD; RAVIKUMAR, KAVYA; FERNANDEZ, JOE VIRGIL; ANTOS, KYNAN DYLAN; VANEE, KELLY JAMES
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 050711/0039 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2019
From: GRAY, TIMOTHY THOMAS; KURPANIK, MICHAL GRZEGORZ; WAH LAM, JENNY TOI; NIGAM, SARVESHWAR; SALEEM, SHIRIN; RIGHTER, JONHENRY A.; HILL, JEREMY RICHARD; RAVIKUMAR, KAVYA; FERNANDEZ, JOE VIRGIL; ANTOS, KYNAN DYLAN; VANEE, KELLY JAMES
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 050616/0320 →
Cited By (9)
US 12,198,413 US 12,211,502 US 12,333,404 US 12,335,550 US 12,374,097 US 12,386,434 US 12,406,316 US 12,475,698 US 12,477,470