IP Library › Granted Patent US 11,004,454
Granted Patent B1
US 11,004,454 · App. 16/182,021 · Granted May 11, 2021

Voice profile updating

Inventors: Sundararajan Srinivasan (Sunnyvale, CA); Arindam Mandal (Redwood City, CA); Krishna Subramanian (Cupertino, CA); Spyridon Matsoukas (Hopkinton, MD); Aparna Khare (San Jose, CA); Rohit Prasad (Lexington, MA)
Assignee: Amazon Technologies, Inc.
G10L17/04G06F3/16G10L15/06G10L17/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,004,454
App. No.
16/182,021
Granted
May 11, 2021
Kind
B1
Abstract

Techniques for updating voice profiles used to perform user recognition are described. A system may use clustering techniques to update voice profiles. When the system receives audio data representing a spoken user input, the system may store the audio data. Periodically, the system may recall, from storage, audio data (representing previous user inputs). The system may identify clusters of the audio data, with each cluster including similar or identical speech characteristics. The system may determine a cluster is substantially similar to an existing voice profile. If this occurs, the system may create an updated voice profile using the original voice profile and the cluster of audio data.

Claims (94)

1. A method, comprising:

generating first voice profile data associated with a group profile identifier and a first user identifier;

identifying a first stored representation of a first user input received after the first voice profile data is generated, the first stored representation being associated with a group profile identifier and the first user identifier;

identifying a second stored representation of a second user input received after the first voice profile data is generated, the second stored representation being associated with the group profile identifier and the first user identifier;

determining the first stored representation is similar to the second stored representation;

generating updated first voice profile data using the first voice profile data, the first stored representation, and the second stored representation; and

storing an association between the updated first voice profile data and the first user identifier.

2. The method of claim 1 , further comprising:

determining a first number of stored representations used to generate the first voice profile data;

determining a second number corresponding to an amount of stored representations, the second number representing the first number, the first stored representation, and the second stored representation; and

based at least in part on the second number, generating the updated first voice profile data.

3. The method of claim 1 , further comprising:

determining the first stored representation is associated with a first intent indicator;

determining the second stored representation is associated with a second intent indicator;

determining the first intent indicator is different from the second intent indicator; and

based at least in part on determining the first intent indicator is different from the second intent indicator, generating the updated first voice profile data.

4. The method of claim 1 , further comprising:

determining a third stored representation of a third user input;

determining a fourth stored representation of a fourth user input;

determining the third stored representation is similar to the fourth stored representation;

determining the third stored representation and the fourth stored representation are dissimilar to the first voice profile data; and

based at least in part on determining the third stored representation and the fourth stored representation are dissimilar to the first voice profile data, generating second voice profile data using the third stored representation and the fourth stored representation.

5. The method of claim 1 , wherein the first stored representation corresponds to first audio data, the second stored representation corresponds to second audio data, and wherein the method further comprises:

generating a first user recognition feature vector representing the first audio data;

generating a second user recognition feature vector representing the second audio data; and

determining the first user recognition feature vector is similar to the second user recognition feature vector.

6. The method of claim 1 , wherein the first stored representation is a first user recognition feature vector, the second stored representation is a second user recognition feature vector, and the method further comprises:

determining a distance between the first user recognition feature vector and the second user recognition feature vector; and

based at least in part on the distance, determining the first user recognition feature vector is similar to the second user recognition feature vector.

7. A method, comprising:

generating first voice profile data associated with a first user identifier;

identifying a first stored representation of a first user input received after the first voice profile data is generated, the first stored representation being associated with the first user identifier;

identifying a second stored representation of a second user input received after the first voice profile data is generated, the second stored representation being associated with the first user identifier;

determining the first stored representation is similar to the second stored representation;

generating updated first voice profile data using the first voice profile data, the first stored representation, and the second stored representation; and

storing an association between the updated first voice profile data and the first user identifier.

8. The method of claim 7 , wherein the first stored representation corresponds to first audio data, the second stored representation corresponds to second audio data, and wherein the method further comprises:

generating a first user recognition feature vector representing the first audio data;

generating a second user recognition feature vector representing the second audio data; and

determining the first user recognition feature vector is similar to the second user recognition feature vector.

9. The method of claim 7 , further comprising:

determining a first number of stored representations used to generate the first voice profile data;

determining a second number corresponding to an amount of stored representations, the second number representing the first number, the first stored representation, and the second stored representation; and

based at least in part on the second number, generating the updated first voice profile data.

10. The method of claim 7 , further comprising:

determining the first stored representation is associated with an intent indicator; and

based at least in part on the first stored representation being associated with the intent indicator, generating the updated first voice profile data.

11. The method of claim 7 , further comprising:

identifying a third stored representation of a third user input;

determining the third stored representation is dissimilar to the first voice profile data; and

generating second voice profile data using the third stored representation.

12. The method of claim 7 , further comprising:

identifying the first stored representation based at least in part on the first stored representation being associated with a group profile identifier.

13. The method of claim 7 , wherein the first stored representation is a first user recognition feature vector, wherein the second stored representation is a second user recognition feature vector, and wherein the method further comprises:

determining a distance between the first user recognition feature vector and the second user recognition feature vector; and

based at least in part on the distance, determining the first user recognition feature vector is similar to the second user recognition feature vector.

14. The method of claim 7 , further comprising:

receiving audio data representing a third user input;

storing the audio data;

after storing the audio data, receiving an indicator representing the audio data is to be associated with a second user identifier;

identifying second voice profile data associated with the second user identifier; and

using the audio data, generating updated second voice profile data.

15. The method of claim 7 , further comprising:

receiving first audio data representing a third user input;

determining the first audio data corresponds to second voice profile data associated with a second user identifier;

determining user profile data, associated with the second user identifier, corresponds to a user name;

generating second audio data including the user name;

causing a first device to output audio corresponding to the second audio data;

receiving, from the first device, third audio data representing the user name is incorrect; and

based at least in part on the third audio data representing the user name is incorrect, using the first audio data as a negative utterance with respect to the second voice profile data.

16. A system, comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

generate first voice profile data associated with a first user identifier;

identify a first stored representation of a first user input received after the first voice profile data is generated, the first stored representation being associated with the first user identifier;

identify a second stored representation of a second user input received after the first voice profile data is generated, the second stored representation being associated with the first user identifier;

determine the first stored representation is similar to the second stored representation;

generate updated first voice profile data using the first voice profile data, the first stored representation, and the second stored representation; and

store an association between the updated first voice profile data and the first user identifier.

17. The system of claim 16 , wherein the first stored representation corresponds to first audio data, the second stored representation corresponds to second audio data, and wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

generate a first user recognition feature vector representing the first audio data;

generate a second user recognition feature vector representing the second audio data; and

determine the first user recognition feature vector is similar to the second user recognition feature vector.

18. The system of claim 16 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine a first number of stored representations used to generate the first voice profile data;

determine a second number corresponding to an amount of stored representations, the second number representing the first number, the first stored representation, and the second stored representation; and

based at least in part on the second number, generate the updated first voice profile data.

19. The system of claim 16 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine the first stored representation is associated with an intent indicator; and

based at least in part on the first stored representation being associated with the intent indicator, generate the updated first voice profile data.

20. The system of claim 16 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

identify a third stored representation of a third user input;

determine the third stored representation is dissimilar to the first voice profile data; and

generate second voice profile data using the second stored representation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2018
From: SRINIVASAN, SUNDARARAJAN; MANDAL, ARINDAM; SUBRAMANIAN, KRISHNA; MATSOUKAS, SPYRIDON; KHARE, APARNA; PRASAD, ROHIT
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 047424/0693 →
Cited By (3)
US 12,531,068 US 12,694,209 US 12,749,480