IP Library › Granted Patent US 10,567,515
Granted Patent B1
US 10,567,515 · App. 15/794,315 · Granted Feb 18, 2020

Speech processing performed with respect to first and second user profiles in a dialog session

Inventor: Yu Bao (Issaquah, WA)
Assignee: Amazon Technologies, Inc.
H04L67/14G06F3/167G10L15/22G10L17/22H04L67/306G10L15/26G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,567,515
App. No.
15/794,315
Granted
Feb 18, 2020
Kind
B1
Abstract

Techniques for implementing a “volatile” user ID are described. A system receives first input audio data and determines first speech processing results therefrom. The system also determines a first user that spoke an utterance represented in the first input audio data. The system establishes a multi-turn dialog session with a first content source and receives first output data from the first content source based on the first speech processing results and the first user. The system causes a device to present first output content associated with the first output data. The system then receives second input audio data and determines second speech processing results therefrom. The system also determines the second input audio data corresponds to the same multi-turn dialog session. The system determines a second user that spoke an utterance represented in the second input audio data and receives second output data from the first content source based on the second speech processing results and the second user. The system causes the device to present second output content associated with the second output data.

Claims (117)

1. A computer-implemented method comprising:

during a first time period:

receiving, from a device, first input audio data corresponding to a first utterance;

determining first audio characteristics corresponding to the first utterance;

determining the first audio characteristics correspond to first stored audio characteristic data associated with first user profile data;

performing automatic speech recognition (ASR) on the first input audio data to generate first input text data;

performing natural language understanding (NLU) on the first input text data to generate first speech processing results data;

determining a first speechlet device associated with the speech processing results;

sending, to the first speechlet device, the first user profile data, the first speech processing results data, and a dialog session identifier (ID);

receiving first output data and the dialog session ID from the first speechlet device; and

sending, to the device, the dialog session ID and the first output data; and

during a second time period after the first time period:

receiving, from the device, the dialog session ID and second input audio data corresponding to a second utterance;

performing ASR on the second input audio data to generate second input text data;

performing NLU on the second input text data to generate second speech processing results data;

determining second audio characteristics corresponding to the second utterance;

determining the second audio characteristics correspond to second stored audio characteristic data associated with second user profile data;

sending, to the first speechlet device, the second user profile data, the second speech processing results data, and the dialog session ID;

receiving second output data and the dialog session ID from the first speechlet device; and

sending to the device, the dialog session ID and the second output data.

2. The computer-implemented method of claim 1 , further comprising:

determining the first audio characteristics correspond to at least one individual under at least eighteen years of age; and

performing NLU on the first input text data to with respect to a subset of functions corresponding to content selected for at least one individual user under at least eighteen years of age.

3. The computer-implemented method of claim 1 , further comprising:

determining the first user profile data indicates a first user age of less than eighteen years;

determining the second user profile data indicates a second user age of at least eighteen years; and

based on the first user age being less than eighteen years, sending an indication that a user is a child to the first speechlet device.

4. The computer-implemented method of claim 1 , further comprising:

determining the first audio characteristics correspond to first stored audio characteristic data associated with at least one individual under at least eighteen years of age;

determining the second audio characteristics corresponding to second stored audio characteristic data associated with at least one individual eighteen years of age or older; and

sending an indication that a user is a child to the first speechlet device.

5. A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

during a first time period:

receive, from a device, first input audio data corresponding to a first utterance;

determine first user profile data associated with the first input audio data;

determine first speech processing results data based on the first input audio data;

associate the first speech processing results data with a dialog session identifier (ID);

generate first output data based on the first speech processing results data and the first user profile data; and

send, to the device, the first output data and the dialog session ID; and

during a second time period after the first time period:

receive, from the device, the dialog session ID and second input audio data corresponding to a second utterance;

determine second speech processing results data based on the second input audio data;

associate the second speech processing results data with the dialog session ID

determine second user profile data associated with the second input audio data;

generate second output data based on the second speech processing results data and the second user profile data; and

send, to the device, the second output data and the dialog session ID.

6. The system of claim 5 , wherein the instructions, when executed by the at least one processor, further cause the system to:

determine first audio characteristics corresponding to the first utterance; and

determine the first audio characteristics correspond to stored audio characteristic data associated with at least one individual under at least eighteen years of age,

wherein the instructions further cause the system to determine the first speech processing results data with respect to a subset of functions corresponding to content selected for at least one individual user under at least eighteen years of age.

7. The system of claim 5 , wherein the instructions, when executed by the at least one processor, further cause the system to:

determine a device ID associated with the device; and

determine the device ID is associated with an indication of a user under at least eighteen years of age,

wherein the instructions further cause the system to determine the first speech processing results data with respect to a subset of functions corresponding to content selected for at least one individual user under at least eighteen years of age.

8. The system of claim 5 , wherein:

the first speech processing results data is associated with the dialog session ID based at least in part on the first utterance including a wakeword; and

the second speech processing results data is associated with the dialog session ID based at least in part on the dialog session ID being received with the second input audio data, the second utterance not including the wakeword.

9. The system of claim 5 , wherein the instructions, when executed by the at least one processor, further cause the system to:

determine the second speech processing results data correspond to a command inappropriate for at least one individual under at least eighteen years of age,

wherein the second output content corresponds to an indication that the command is inappropriate for at least one individual under at least eighteen years of age.

10. The system of claim 5 , wherein the instructions causing the system to generate the first output data further include instructions to:

send, to a first content source device, the first speech processing results data, the first user profile data, and the dialog session ID; and

receive, from the first content source, the first output data and the dialog session ID.

11. The system of claim 5 , wherein the instructions causing the system to generate the first output data further include instructions to:

determine the first user profile data indicates a first user age of less than eighteen years;

determine the second user profile data indicates a second user age of at least eighteen years; and

generate the second output data based at least in part on the first user age being less than eighteen years.

12. The system of claim 5 , wherein the instructions, when executed by the at least one processor, further cause the system to:

determine first audio characteristics corresponding to the first utterance;

determine the first audio characteristics correspond to first stored audio characteristic data associated with at least one individual under at least eighteen years of age;

determine second audio characteristics corresponding to the second utterance;

determine the second audio characteristics corresponding to second stored audio characteristic data associated with at least one individual eighteen years of age or older; and

generate the second output data based at least in part on the first audio characteristics corresponding to the first stored audio characteristic data.

13. A computer-implemented method comprising:

during a first time period:

receiving, from a device, first input audio data corresponding to a first utterance;

determining first user profile data associated with the first input audio data;

determining first speech processing results data based on the first input audio data;

associating the first speech processing results data with a dialog session identifier (ID);

generating first output data based on the first speech processing results data and the user profile data; and

sending, to the device, the first output data and the dialog session ID; and

during a second time period after the first time period:

receiving, from the device, the dialog session ID and second input audio data corresponding to a second utterance;

determining second speech processing results data based on the second input audio data;

associating the second speech processing results data with the dialog session ID;

determining second user profile data associated with the second input audio data;

generating second output data based on the second speech processing results data and the second user profile data; and

sending, to the device, the second output data and the dialog session ID.

14. The computer-implemented method of claim 13 , further comprising

determining first audio characteristics corresponding to the first utterance;

determining the first audio characteristics correspond to stored audio characteristic data associated with at least one individual under at least eighteen years of age; and

determining the first speech processing results data with respect to a subset of functions corresponding to content selected for at least one individual user under at least eighteen years of age.

15. The computer-implemented method of claim 13 , further comprising:

determining a device ID associated with the device;

determining the device ID is associated with an indication of a user under at least eighteen years of age; and

determining the first speech processing results data with respect to a subset of functions corresponding to content selected for at least one individual user under at least eighteen years of age.

16. The computer-implemented method of claim 13 , further comprising:

associating the first speech processing results data with the dialog session ID based at least in part on the first utterance including a wakeword; and

associating the second speech processing results data with the dialog session ID based at least in part on the dialog session ID being received with the second input audio data, the second utterance not including the wakeword.

17. The computer-implemented method of claim 13 , further comprising:

determining the second speech processing results data correspond to a command inappropriate for at least one individual under at least eighteen years of age,

wherein the second output content corresponds to an indication that the command is inappropriate for at least one individual under at least eighteen years of age.

18. The computer-implemented method of claim 13 , wherein generating the first output audio data further comprises:

sending, to a first content source device, the first speech processing results data, the first user profile data, and the dialog session ID; and

receiving, from the first content source, the first output data and the dialog session ID.

19. The computer-implemented method of claim 13 , further comprising:

determining the first user profile data indicates a first user age of less than eighteen years;

determining the second user profile data indicates a second user age of at least eighteen years; and

generating the second output data based at least in part on the first user age being less than eighteen years.

20. The computer-implemented method of claim 13 , further comprising:

determining first audio characteristics corresponding to the first utterance;

determining the first audio characteristics correspond to first stored audio characteristic data associated with at least one individual under at least eighteen years of age;

determining second audio characteristics corresponding to the second utterance;

determining the second audio characteristics corresponding to second stored audio characteristic data associated with at least one individual eighteen years of age or older; and

generating the second output data based at least in part on the first audio characteristics corresponding to the first stored audio characteristic data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2017
From: BAO, YU
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 043957/0811 →
Cited By (35)
US 12,192,713 US 12,210,801 US 12,211,490 US 12,217,748 US 12,217,765 US 12,223,228 US 12,230,291 US 12,231,859 US 12,236,932 US 12,277,368 US 12,288,558 US 12,322,390 US 12,327,189 US 12,340,802 US 12,360,734 US 12,374,334 US 12,375,052 US 12,381,880 US 12,387,716 US 12,405,717 US 12,424,220 US 12,438,977 US 12,462,802 US 12,505,832 US 12,513,466 US 12,513,479 US 12,518,755 US 12,518,756 US 12,578,779 US 12,579,978 US 12,626,717 US 12,699,543 US 12,711,962 US 12,732,547 US 12,744,035