IP Library Granted Patent US 10,037,760
Granted Patent B2
US 10,037,760 · App. 15/669,701 · Granted Jul 31, 2018

Segment-based speaker verification using dynamically generated phrases

Inventors: Dominik Roblek (Meilen, CH); Matthew Sharifi (Kilchberg, CH)
G10L17/24G10L15/02G10L17/04G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,037,760
App. No.
15/669,701
Granted
Jul 31, 2018
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for verifying an identity of a user. The methods, systems, and apparatus include actions of receiving a request for a verification phrase for verifying an identity of a user. Additional actions include, in response to receiving the request for the verification phrase for verifying the identity of the user, identifying subwords to be included in the verification phrase and in response to identifying the subwords to be included in the verification phrase, obtaining a candidate phrase that includes at least some of the identified subwords as the verification phrase. Further actions include providing the verification phrase as a response to the request for the verification phrase for verifying the identity of the user.

Claims (53)

1. A computer-implemented method comprising:

identifying candidate enrollment phrases to enroll a particular user for voice verification, each candidate enrollment phrase includes at least one subword and at least one candidate enrollment phrase includes at least one subword for which no stored enrollment audio data representing the user speaking the subword has been obtained;

prompting the particular user to speak candidate phrases including the at least one candidate enrollment phrase that contains at least one subword for which no stored enrollment audio data representing the particular user speaking the subword has been obtained;

obtaining and storing enrollment audio data representing the particular user speaking the candidate enrollment phrases until enrollment audio data has been obtained that meets a certain threshold;

dynamically generating a verification phrase based at least on one or more of the subwords included in the candidate enrollment phrases uttered by the particular user in the enrollment acoustic data;

prompting a user to speak the dynamically generated verification phrase;

obtaining verification audio data representing the user speaking the dynamically generated verification phrase;

comparing the obtained verification audio data with the enrollment verification audio data to determine whether the user speaking the dynamically generated verification phrase is the particular user who spoke the candidate enrollment phrases; and

in response to determining that the user speaking the dynamically generated verification phrase is the particular user who spoke the candidate enrollment phrases, verifying an identity of the user as the particular user.

2. The computer-implemented method of claim 1 , wherein the subwords comprise phonemes.

3. The computer-implemented method of claim 1 , wherein obtaining and storing enrollment audio data representing the particular user speaking the candidate enrollment phrases until enrollment audio data has been obtained that meets a certain threshold includes:

determining that the obtained enrollment audio data contains a minimum number of subwords spoken by the particular user a minimum number of times.

4. The computer-implemented method of claim 1 , wherein obtaining and storing enrollment audio data representing the particular user speaking the candidate enrollment phrases continues until enrollment audio data has been obtained that satisfies an utterance quality threshold.

5. The computer-implemented method of claim 1 , wherein dynamically generating a verification phrase based at least on one or more of the particular subwords included in the candidate enrollment phrases uttered by the particular user in the enrollment audio data comprises:

generating a verification phrase that includes at least one or more of the particular subwords.

6. The computer-implemented method of claim 1 , wherein dynamically generating a verification phrase based at least on one or more of the particular subwords included in the candidate enrollment phrases uttered by the particular user in the enrollment audio data comprises:

generating a verification phrase that includes (i) at least one or more of the particular subwords and (ii) one or more subwords that are not any of the one or more particular subwords.

7. A system comprising:

one or more data processing apparatuses; and

one or more storage devices storing instructions that are operable, when executed by one or more data processing apparatuses, to cause the one or more data processing apparatuses to perform operations comprising:

identifying candidate enrollment phrases to enroll a particular user for voice verification, each candidate enrollment phrase includes at least one subword and at least one candidate enrollment phrase includes at least one subword for which no stored enrollment audio data representing the user speaking the subword has been obtained;

prompting the particular user to speak candidate phrases including the at least one candidate enrollment phrase that contains at least one subword for which no stored enrollment audio data representing the particular user speaking the subword has been obtained;

obtaining and storing enrollment audio data representing the particular user speaking the candidate enrollment phrases until enrollment audio data has been obtained that meets a certain threshold;

dynamically generating a verification phrase based at least on one or more of the subwords included in the candidate enrollment phrases uttered by the particular user in the enrollment acoustic data;

prompting a user to speak the dynamically generated verification phrase;

obtaining verification audio data representing the user speaking the dynamically generated verification phrase;

comparing the obtained verification audio data with the enrollment verification audio data to determine whether the user speaking the dynamically generated verification phrase is the particular user who spoke the candidate enrollment phrases; and

in response to determining that the user speaking the dynamically generated verification phrase is the particular user who spoke the candidate enrollment phrases, verifying an identity of the user as the particular user.

8. The system of claim 7 , wherein subwords comprise phonemes.

9. The system of claim 7 , wherein obtaining and storing enrollment audio data representing the particular user speaking the candidate enrollment phrases until enrollment audio data has been obtained that meets a certain threshold includes:

determining that the obtained enrollment audio data contains a minimum number of subwords spoken by the particular user a minimum number of times.

10. The system of claim 7 , wherein obtaining and storing enrollment audio data representing the particular user speaking the candidate enrollment phrases continues until enrollment audio data has been obtained that satisfies an utterance quality threshold.

11. The system of claim 7 , wherein dynamically generating a verification phrase based at least on one or more of the particular subwords included in the candidate enrollment phrases uttered by the particular user in the enrollment audio data comprises:

generating a verification phrase that includes at least one or more of the particular subwords.

12. The system of claim 7 , wherein dynamically generating a verification phrase based at least on one or more of the particular subwords included in the candidate enrollment phrases uttered by the particular user in the enrollment audio data comprises:

generating a verification phrase that includes (i) at least one or more of the particular subwords and (ii) one or more subwords that are not any of the one or more particular subwords.

13. One or more non-transitory computer-readable storage mediums storing instructions thereon that are executable by a processing device and upon such execution cause the processing device to perform operations comprising:

identifying candidate enrollment phrases to enroll a particular user for voice verification, each candidate enrollment phrase includes at least one subword and at least one candidate enrollment phrase includes at least one subword for which no stored enrollment audio data representing the user speaking the subword has been obtained;

prompting the particular user to speak candidate phrases including the at least one candidate enrollment phrase that contains at least one subword for which no stored enrollment audio data representing the particular user speaking the subword has been obtained;

obtaining and storing enrollment audio data representing the particular user speaking the candidate enrollment phrases until enrollment audio data has been obtained that meets a certain threshold;

dynamically generating a verification phrase based at least on one or more of the subwords included in the candidate enrollment phrases uttered by the particular user in the enrollment acoustic data;

prompting a user to speak the dynamically generated verification phrase;

obtaining verification audio data representing the user speaking the dynamically generated verification phrase;

comparing the obtained verification audio data with the enrollment verification audio data to determine whether the user speaking the dynamically generated verification phrase is the particular user who spoke the candidate enrollment phrases; and

in response to determining that the user speaking the dynamically generated verification phrase is the particular user who spoke the candidate enrollment phrases, verifying an identity of the user as the particular user.

14. The non-transitory computer-readable medium of claim 13 , wherein the subwords comprise phonemes.

15. The non-transitory computer-readable storage medium of claim 13 , wherein obtaining and storing enrollment audio data representing the particular user speaking the candidate enrollment phrases until enrollment audio data has been obtained that meets a certain threshold includes:

determining that the obtained enrollment audio data contains a minimum number of subwords spoken by the particular user a minimum number of times.

16. The non-transitory computer-readable storage medium of claim 13 , wherein obtaining audio data representing the particular user speaking the candidate enrollment phrases continues until enrollment audio data has been obtained that satisfies an utterance quality threshold.

17. The non-transitory computer-readable storage medium of claim 13 , wherein dynamically generating a verification phrase based at least on one or more of the particular subwords included in the candidate enrollment phrases uttered by the particular user in the enrollment audio data comprises:

generating a verification phrase that includes at least one or more of the particular subwords.

18. The non-transitory computer-readable storage medium of claim 13 , wherein dynamically generating a verification phrase based at least on one or more of the particular subwords included in the candidate enrollment phrases uttered by the particular user in the enrollment audio data comprises:

generating a verification phrase that includes (i) at least one or more of the particular subwords and (ii) one or more subwords that are not any of the one or more particular subwords.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2017
From: ROBLEK, DOMINIK; SHARIFI, MATTHEW
To: GOOGLE INC.
Reel/Frame 043338/0847 →
Continuity (4)
Continuation 15191886 · Jun 24, 2016
Continuation 14447115 · Jul 30, 2014
Continuation 14242098 · Apr 1, 2014
Related Publication 20180025734A1 · Jan 25, 2018