IP Library Granted Patent US 9,424,846
Granted Patent B2
US 9,424,846 · App. 14/447,115 · Granted Aug 23, 2016

Segment-based speaker verification using dynamically generated phrases

Inventors: Dominik Roblek (Mountain View, CA); Matthew Sharifi (Palo Alto, CA)
Assignee: Google Inc.
G10L17/24G10L17/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,424,846
App. No.
14/447,115
Granted
Aug 23, 2016
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for verifying an identity of a user. The methods, systems, and apparatus include actions of receiving a request for a verification phrase for verifying an identity of a user. Additional actions include, in response to receiving the request for the verification phrase for verifying the identity of the user, identifying subwords to be included in the verification phrase and in response to identifying the subwords to be included in the verification phrase, obtaining a candidate phrase that includes at least some of the identified subwords as the verification phrase. Further actions include providing the verification phrase as a response to the request for the verification phrase for verifying the identity of the user.

Claims (66)

1. A computer-implemented method, comprising:

providing a speaker identification verification phrase;

obtaining audio data representing a candidate user speaking the speaker identification verification phrase;

obtaining, for each of multiple subwords associated with the speaker identification phrase, sample acoustic features that are derived from the audio data representing the candidate user speaking the speaker identification verification phrase;

obtaining, for each of the multiple subwords associated with the speaker identification verification phrase, reference acoustic features that (i) are stored in a collection of acoustic features for a target user, and (ii) are derived from audio data of the target user speaking one or more words that include the subword;

determining, for each of the multiple subwords associated with the speaker identification verification phrase, that the sample acoustic features are associated with the reference acoustic features;

in response to determining that the sample acoustic features are associated with the reference acoustic features, identifying the candidate user as the target user;

obtaining, for each of one or more other subwords associated with the speaker identification verification phrase, other acoustic features that (i) are not already stored in the collection of acoustic features for the target user, and (ii) are derived from the audio data representing the candidate user speaking the speaker identification verification phrase;

in response to identifying the candidate user as the target user, storing one or more of the other acoustic features in the collection of acoustic features for the target user; and

using at least one of the one or more other acoustic features stored in the collection of acoustic features for the target user to verify an identity of a user speaking a subsequent utterance that includes one or more subwords that correspond to the at least one of the one or more of the other acoustic features stored in the collection of acoustic features for the target user.

2. The method of claim 1 , wherein in response to identifying the candidate user as the target user, storing one or more of the other acoustic features in the collection of acoustic features for the target user comprises:

in response to identifying the candidate user as the target user based on determining that the sample acoustic features that are derived from the audio data representing the candidate user speaking the speaker identification verification phrase are associated with the reference acoustic features, storing, in the collection of acoustic features for the target user, one or more of the other acoustic features that are derived from the audio data representing the candidate user speaking the speaker identification verification phrase.

3. The method of claim 1 , wherein the speaker identification verification phrase includes the multiple subwords and the one or more other subwords.

4. The method of claim 1 , comprising:

in response to identifying the candidate user as the target user, updating the collection of acoustic features for the target user with the one or more of the acoustic features.

5. The method of claim 4 , wherein updating the collection of acoustic features for the target user with the one or more of the acoustic features comprises:

storing the one or more of the acoustic features in the collection of acoustic features for the target user.

6. The method of claim 1 , wherein determining, for each of the multiple subwords associated with the speaker identification verification phrase, that the sample acoustic features are associated with the reference acoustic features, comprises:

determining, for each of the multiple subwords associated with the speaker identification verification phrase, a distance between the sample acoustic features and the reference acoustic features; and

based on at least the distance, determining, for each of the multiple subwords associated with the speaker identification verification phrase, that the sample acoustic features are associated with the reference acoustic features.

7. The method of claim 6 , wherein based on at least the distance, determining, for each of the multiple subwords associated with the speaker identification verification phrase, that the sample acoustic features are associated with the reference acoustic features, comprises:

determining sound discriminativeness of the subword; and

based on at least the distance and the sound discriminativeness, determining that the sample acoustic features are associated with the reference acoustic features.

8. The method of claim 1 , comprising:

determining, for each of the one or more other subwords associated with the speaker identification verification phrase, that the sample acoustic features are not associated with the reference acoustic features.

9. The method of claim 1 , comprising:

providing an indication whether the user speaking the subsequent utterance is verified as the target user in response to using at least one of the one or more other acoustic features stored in the collection of acoustic features for the target user to verify the identify of the user speaking the subsequent utterance that includes one or more subwords that correspond to the at least one of the one or more of the other acoustic features stored in the collection of acoustic features for the target user.

10. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

providing a speaker identification verification phrase;

obtaining audio data representing a candidate user speaking the speaker identification verification phrase;

obtaining, for each of multiple subwords associated with the speaker identification phrase, sample acoustic features that are derived from the audio data representing the candidate user speaking the speaker identification verification phrase;

obtaining, for each of the multiple subwords associated with the speaker identification verification phrase, reference acoustic features that (i) are stored in a collection of acoustic features for a target user, and (ii) are derived from audio data of the target user speaking one or more words that include the subword;

determining, for each of the multiple subwords associated with the speaker identification verification phrase, that the sample acoustic features are associated with the reference acoustic features;

in response to determining that the sample acoustic features are associated with the reference acoustic features, identifying the candidate user as the target user;

obtaining, for each of one or more other subwords associated with the speaker identification verification phrase, other acoustic features that (i) are not already stored in the collection of acoustic features for the target user, and (ii) are derived from the audio data representing the candidate user speaking the speaker identification verification phrase;

in response to identifying the candidate user as the target user, storing one or more of the other acoustic features in the collection of acoustic features for the target user; and

using at least one of the one or more other acoustic features stored in the collection of acoustic features for the target user to verify an identity of a user speaking a subsequent utterance that includes one or more subwords that correspond to the at least one of the one or more of the other acoustic features stored in the collection of acoustic features for the target user.

11. The system of claim 10 , the operations comprising:

in response to identifying the candidate user as the target user, updating the collection of acoustic features for the target user with the one or more of the acoustic features.

12. The system of claim 11 , wherein updating the collection of acoustic features for the target user with the one or more of the acoustic features comprises:

storing the one or more of the acoustic features in the collection of acoustic features for the target user.

13. The system of claim 10 , wherein determining, for each of the multiple subwords associated with the speaker identification verification phrase, that the sample acoustic features are associated with the reference acoustic features, comprises:

determining, for each of the multiple subwords associated with the speaker identification verification phrase, a distance between the sample acoustic features and the reference acoustic features; and

based on at least the distance, determining, for each of the multiple subwords associated with the speaker identification verification phrase, that the sample acoustic features are associated with the reference acoustic features.

14. The system of claim 13 , wherein based on at least the distance, determining, for each of the multiple subwords associated with the speaker identification verification phrase, that the sample acoustic features are associated with the reference acoustic features, comprises:

determining sound discriminativeness of the subword; and

based on at least the distance and the sound discriminativeness, determining that the sample acoustic features are associated with the reference acoustic features.

15. The system of claim 10 , the operations comprising:

determining, for each of the one or more other subwords associated with the speaker identification verification phrase, that the sample acoustic features are not associated with the reference acoustic features.

16. The system of claim 10 , wherein the speaker identification verification phrase includes the multiple subwords and the one or more other subwords.

17. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

providing a speaker identification verification phrase;

obtaining audio data representing a candidate user speaking the speaker identification verification phrase;

obtaining, for each of multiple subwords associated with the speaker identification phrase, sample acoustic features that are derived from the audio data representing the candidate user speaking the speaker identification verification phrase;

obtaining, for each of the multiple subwords associated with the speaker identification verification phrase, reference acoustic features that (i) are stored in a collection of acoustic features for a target user, and (ii) are derived from audio data of the target user speaking one or more words that include the subword;

determining, for each of the multiple subwords associated with the speaker identification verification phrase, that the sample acoustic features are associated with the reference acoustic features;

in response to determining that the sample acoustic features are associated with the reference acoustic features, identifying the candidate user as the target user;

obtaining, for each of one or more other subwords associated with the speaker identification verification phrase, other acoustic features that (i) are not already stored in the collection of acoustic features for the target user, and (ii) are derived from the audio data representing the candidate user speaking the speaker identification verification phrase;

in response to identifying the candidate user as the target user, storing one or more of the other acoustic features in the collection of acoustic features for the target user; and

using at least one of the one or more other acoustic features stored in the collection of acoustic features for the target user to verify an identity of a user speaking a subsequent utterance that includes one or more subwords that correspond to the at least one of the one or more of the other acoustic features stored in the collection of acoustic features for the target user.

18. The medium of claim 17 , the operations comprising:

in response to identifying the candidate user as the target user, updating the collection of acoustic features for the target user with the one or more of the acoustic features.

19. The medium of claim 18 , wherein updating the collection of acoustic features for the target user with the one or more of the acoustic features comprises:

storing the one or more of the acoustic features in the collection of acoustic features for the target user.

20. The medium of claim 17 , wherein the speaker identification verification phrase includes the multiple subwords and the one or more other subwords.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 22, 2014
From: ROBLEK, DOMINIK; SHARIFI, MATTHEW
To: GOOGLE INC.
Reel/Frame 033593/0814 →
Continuity (2)
Continuation 14242098 · Apr 1, 2014
Related Publication 20150279374A1 · Oct 1, 2015