IP Library Granted Patent US 11,568,879
Granted Patent B2
US 11,568,879 · App. 17/303,928 · Granted Jan 31, 2023

Segment-based speaker verification using dynamically generated phrases

Inventors: Dominik Roblek (Meilen, CH); Matthew Sharifi (Kilchberg, CH)
Assignee: Google LLC
G10L17/24G10L15/02G10L17/04G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,568,879
App. No.
17/303,928
Granted
Jan 31, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for verifying an identity of a user. The methods, systems, and apparatus include actions of receiving a request for a verification phrase for verifying an identity of a user. Additional actions include, in response to receiving the request for the verification phrase for verifying the identity of the user, identifying subwords to be included in the verification phrase and in response to identifying the subwords to be included in the verification phrase, obtaining a candidate phrase that includes at least some of the identified subwords as the verification phrase. Further actions include providing the verification phrase as a response to the request for the verification phrase for verifying the identity of the user.

Claims (40)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving a command input by a particular user;

in response to receiving the command, generating a verification phrase for the particular user, the verification phrase comprising at least one subword obtained from stored enrollment audio data representing the particular user speaking the subword;

providing, as output from a user interface, a message informing the particular user to speak the verification phrase to verify the identity of the particular user;

receiving verification audio data representing an unverified user speaking the verification phrase;

determining whether the unverified user speaking the verification phrase comprises the particular user based on the stored enrollment audio data and the verification audio data; and

in response to determining that the unverified user speaking the verification phrase comprises the particular user, verifying an identity of the unverified user as the particular user.

2. The computer-implemented method of claim 1 , wherein providing the message as output from the user interface comprises displaying the message as output on a display in communication with the data processing hardware.

3. The computer-implemented method of claim 1 , wherein providing the message as output from the user interface comprises outputting the message as synthesized speech.

4. The computer-implemented method of claim 1 , wherein the operations further comprise, after verifying the identity of the unverified user as the particular user, displaying a welcome interface on a display in communication with the data processing hardware.

5. The computer-implemented method of claim 1 , wherein the operations further comprise:

receiving the enrollment audio data representing the particular user speaking the at least one subword; and

storing the received enrollment audio data in memory hardware in communication with the data processing hardware.

6. The computer-implemented method of claim 5 , wherein the operations further comprises determining that the received enrollment audio data contains a minimum number of subwords spoken by the particular user a minimum number of times.

7. The computer-implemented method of claim 6 , wherein the subwords comprise phonemes.

8. The computer-implemented method of claim 5 , wherein receiving the enrollment audio data continues until enrollment audio data has been received that satisfies an utterance quality threshold.

9. The computer-implemented method of claim 1 , wherein determining whether the unverified user speaking the verification phrase comprises the particular user comprises:

determining whether the verification audio data matches the stored enrollment audio data; and

in response to determining that the verification audio data matches the stored enrollment audio data, classifying the unverified user as the particular user.

10. A system:

data processing hardware: and

memory hardware in communication with the data processing hardware and storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a command input by a particular user;

in response to receiving the command, generating a verification phrase for the particular user, the verification phrase comprising at least one subword obtained from stored enrollment audio data representing the particular user speaking the subword;

providing, as output from a user interface, a message informing the particular user to speak the verification phrase to verify the identity of the particular user;

receiving verification audio data representing an unverified user speaking the verification phrase;

determining whether the unverified user speaking the verification phrase comprises the particular user based on the stored enrollment audio data and the verification audio data; and

in response to determining that the unverified user speaking the verification phrase comprises the particular user, verifying an identity of the unverified user as the particular user.

11. The system of claim 10 , wherein determining whether the unverified user speaking the verification phrase comprises the particular user comprises:

determining whether the verification audio data matches the stored enrollment audio data; and

in response to determining that the verification audio data matches the stored enrollment audio data, classifying the unverified user as the particular user.

12. The system of claim 10 , wherein providing the message as output from the user interface comprises displaying the message as output on a display in communication with the data processing hardware.

13. The system of claim 10 , wherein providing the message as output from the user interface comprises outputting the message as synthesized speech.

14. The system of claim 10 , wherein the operations further comprise, after verifying the identity of the unverified user as the particular user, displaying a welcome interface on a display in communication with the data processing hardware.

15. The system of claim 10 , wherein the operations further comprise:

receiving the enrollment audio data representing the particular user speaking the at least one subword; and

storing the received enrollment audio data in memory hardware in communication with the data processing hardware.

16. The system of claim 15 , wherein the operations further comprises determining that the received enrollment audio data contains a minimum number of subwords spoken by the particular user a minimum number of times.

17. The system of claim 16 , wherein the subwords comprise phonemes.

18. The system of claim 16 , wherein receiving the enrollment audio data continues until enrollment audio data has been received that satisfies an utterance quality threshold.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 10, 2021
From: ROBLEK, DOMINIK; SHARIFI, MATTHEW
To: GOOGLE INC.
Reel/Frame 056502/0005 →
CHANGE OF NAME Recorded Jun 10, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 056540/0789 →
Continuity (7)
Continuation 16675420 · Nov 6, 2019
Continuation 16017690 · Jun 25, 2018
Continuation 15669701 · Aug 4, 2017
Continuation 15191886 · Jun 24, 2016
Continuation 14447115 · Jul 30, 2014
Continuation 14242098 · Apr 1, 2014
Related Publication 20210295850A1 · Sep 23, 2021