IP Library Granted Patent US 11,289,082
Granted Patent B1
US 11,289,082 · App. 16/676,888 · Granted Mar 29, 2022

Speech processing output personalization

Inventors: Andrea Klein Lacy (Mountain View, CA); Timothy Whalin (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G10L15/22G10L13/02G10L15/02G10L2015/225G10L2015/227
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,289,082
App. No.
16/676,888
Granted
Mar 29, 2022
Kind
B1
Abstract

Described herein is a system for adapting an output to a user input over a period of time based on how often the user interacts with the system. The system may determine a user's level of familiarity of the system, and may determine to personalize the output to a user request based on his level of familiarity. The user's level of familiarity may be determined by analyzing historical interactions between the user and the system. The level of personalization applied to the output may be determined based on the user's level of familiarity. As user becomes more familiar with the system, the output may be more personalized.

Claims (96)

1. A computer-implemented method comprising:

receiving first user profile data corresponding to a first user profile;

determining one or more second user profiles similar to the first user profile;

determining historical interaction data, using second user profile data corresponding to the one or more second user profiles, the historical interaction data relating to a system within a time period;

determining, using the historical interaction data, familiarity data for the first user profile, wherein the familiarity data represents an application familiarity corresponding to the first user profile;

receiving, from a device, input audio data representing an utterance;

performing speech processing on the input audio data to determine that the utterance corresponds to an application and the first user profile; and

generating, using the familiarity data, first output data responsive to the utterance.

2. The computer-implemented method of claim 1 , further comprising:

receiving, from a system, first data representing a substantive response to the utterance; and

determining, using the familiarity data, second data representing a descriptive format for the first data,

wherein generating the first output data comprises generating the first output data using the first data and the second data.

3. The computer-implemented method of claim 2 , further comprising:

determining a location associated with the device; and

determining that the utterance represents a request for information corresponding to the location,

wherein determining the second data comprises determining that the descriptive format indicates removal of text from second output data, the text representing the location and the second output data corresponding to a past utterance.

4. The computer-implemented method of claim 1 , further comprising:

determining, using the first user profile data, a number of utterances received corresponding to the application;

determining, based on the number of utterances, second familiarity data corresponding to the first user profile;

receiving, from the device, second input audio data representing a second utterance;

performing speech processing on the second input audio data to determine that the second utterance corresponds to the application;

receiving, from a system, first data corresponding to the second utterance; and

generating, based on the second familiarity data, second output data responsive to the second utterance, the second output data including the first output data and the first data.

5. The computer-implemented method of claim 4 , further comprising:

receiving feedback data representing user feedback with respect to the second output data;

determining that the feedback data represents negative feedback;

receiving, from the device, third input audio data representing a third utterance, the third utterance representing a command represented by the second utterance; and

generating, based on the feedback data, representing negative feedback third output data responsive to the third utterance, the third output data corresponding to the first output data.

6. The computer-implemented method of claim 1 , further comprising:

determining, using the first user profile data, a total number of utterances received within the time period;

determining a second familiarity data based on the total number of utterances, the second familiarity data representing a system familiarity corresponding to the first user profile;

receiving, from the device, second input audio data corresponding to a second utterance;

performing speech processing on the second input audio data to determine that the second utterance corresponds to the application; and

generating, using the second familiarity data, second output data responsive to the second utterance.

7. The computer-implemented method of claim 1 , further comprising:

retrieving second user profile data associated with a second user profile;

determining, using the second user profile data, a number of utterances corresponding to the application received within the time period;

determining, based on the number of utterances, a second familiarity data representing an application familiarity corresponding to the second user profile;

receiving, from a second device, second input audio data corresponding to a second utterance, the second utterance representing a command represented by the utterance; and

generating, using the second familiarity data, second output data responsive to the command, the second output data being different than the first output data.

8. The computer-implemented method of claim 1 , further comprising:

determining that the familiarity data satisfies a condition;

receiving, from a first system associated with the application, first data responsive to the utterance;

determining, using the first user profile data, a second application; and

receiving, from a second system associated with the second application, second data,

wherein generating the first output data comprises generating the first output data including the first data and the second data.

9. The computer-implemented method of claim 1 , wherein the familiarity data corresponds to a first application, and

wherein the utterance corresponds to a second application.

10. A system comprising:

at least one processor; and

at least one memory including instructions that, when executed by the at least one processor, cause the system to:

receive first user profile data corresponding to a first user profile;

determine one or more second user profiles similar to the first user profile;

determine historical interaction data, using second user profile data corresponding to the one or more second user profiles, the historical interaction data relating to a system within a time period;

determine, using the historical interaction data, familiarity data for the first user profile, wherein the familiarity data represents an application familiarity corresponding to the first user profile;

receive, from a device, input audio data representing an utterance;

perform speech processing on the input audio data to determine that the utterance corresponds to an application and the first user profile; and

generate, using the familiarity data, first output data responsive to the utterance.

11. The system of claim 10 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

receive, from a system, first data representing a substantive response to the utterance; and

determine, using the familiarity data, second data representing a descriptive format for the first data,

wherein the instructions that cause the system to generate the first output data further cause the system to generate the first output data using the first data and the second data.

12. The system of claim 11 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

determine a location associated with the device; and

determine that the utterance represents a request for information corresponding to the location,

wherein the instructions that cause the system to determine the second data further cause the system to determine that the descriptive format indicates removal of text from second output data, the text representing the location and the second output data corresponding to a past utterance.

13. The system of claim 10 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

determine, using the first user profile data, a number of utterances received corresponding to the application;

determine, based on the number of utterances, second familiarity data corresponding to the first user profile;

receive, from the device, second input audio data representing a second utterance;

perform speech processing on the second input audio data to determine that the second utterance corresponds to the application;

receive, from a system, first data corresponding to the second utterance; and

generate, based on the second familiarity data, second output data responsive to the second utterance, the second output data including the first output data and the first data.

14. The system of claim 13 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

receive feedback data representing user feedback with respect to the second output data;

determine that the feedback data represents negative feedback;

receive, from the device, third input audio data representing a third utterance, the third utterance representing a command represented by the second utterance; and

generate, based on the feedback data representing negative feedback, third output data responsive to the third utterance, the third output data corresponding to the first output data.

15. The system of claim 10 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

determine, using the first user profile data, a total number of utterances received within the time period;

determine a second familiarity data based on the total number of utterances, the second familiarity data representing a system familiarity corresponding to the first user profile;

receive, from the device, second input audio data corresponding to a second utterance;

perform speech processing on the second input audio data to determine that the second utterance corresponds to the application; and

generate, using the second familiarity data, second output data responsive to the second utterance.

16. The system of claim 10 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

retrieve second user profile data associated with a second user profile;

determine, using the second user profile data, a number of utterances relating to the application received within the time period;

determine, based on the number of utterances, second familiarity data representing an application familiarity corresponding to the second user profile;

receive, from a second device, second input audio data corresponding to a second utterance, the second utterance representing a command represented by the utterance; and

generate, using the second familiarity data, second output data responsive to the command, the second output data being different than the first output data.

17. The system of claim 10 , wherein the instructions that, when executed by the at least one processor, further cause the system to:

determine that the familiarity data satisfies a condition;

receive, from a first system associated with the application, first data responsive to the utterance;

determine, using the first user profile data, a second application; and

receive, from a second system associated with the second application, second data,

wherein the instructions that cause the system to generate the first output data further causes the system to generate the first output data including the first data and the second data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2019
From: WHALIN, TIMOTHY; LACY, ANDREA KLEIN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 050947/0862 →
Cited By (20)
US 12,197,817 US 12,200,297 US 12,211,502 US 12,216,894 US 12,219,314 US 12,236,952 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,333,404 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,437,746 US 12,477,470 US 12,559,026 US 12,608,171 US 12,619,452