IP Library › Granted Patent US 10,255,922
Granted Patent B1
US 10,255,922 · App. 15/191,892 · Granted Apr 9, 2019

Speaker identification using a text-independent model and a text-dependent model

Inventors: Matthew Sharifi (Kilchberg, CH); Dominik Roblek (Meilen, CH)
Assignee: Google LLC
G10L17/24G10L15/02G10L15/22G10L17/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,255,922
App. No.
15/191,892
Granted
Apr 9, 2019
Kind
B1
Abstract

In some implementations, a single registration utterance that includes a hotword and an introduction declaration is received. A user is registered, including training a text-dependent speaker identification model using the hotword of the single registration utterance and training a text-independent speaker identification model using the introduction declaration of the single registration utterance. An authentication utterance by the user that includes the hotword and a voice command that is different from the introduction declaration is received. The user is authenticated, including processing the hotword of the authentication utterance using the text-dependent speaker identification model and processing the voice command using the text-independent speaker identification model. Access to an access-controlled personal resource of the user is provided without requiring the user to submit any further authentication information other than the single registration utterance by the user that includes the hotword and the introduction declaration to the speech-enabled home device.

Claims (55)

1. A computer-implemented method comprising:

receiving, during a speech registration session and by a speech-enabled home device that includes one or more microphones for detecting utterances spoken in a home environment, a single registration utterance by a user that includes a hotword and an introduction declaration;

registering the user by a server-based voice authentication device that includes an automated speech recognizer and that is associated with the speech-enabled home device, wherein registering includes training, by the server-based voice authentication device, a text-dependent speaker identification model using the hotword of the single registration utterance and training, by the server-based voice authentication device, a text-independent speaker identification model using the introduction declaration of the single registration utterance;

after the speech registration session is concluded, receiving, by the speech-enabled home device, an authentication utterance by the user that includes the hotword and a voice command that is different from the introduction declaration;

in response to receiving the authentication utterance by the user that includes the hotword and the voice command that is different from the introduction declaration, authenticating, by the server-based voice authentication device, the user, wherein authenticating includes processing the hotword of the authentication utterance by the server-based voice authentication device using the text-dependent speaker identification model and processing the voice command using the text-independent speaker identification model;

in response to authenticating the user by the server-based voice authentication device, providing, by the server-based voice authentication device, access to an access-controlled personal resource of the user without requiring the user to submit any further authentication information other than the single registration utterance by the user that includes the hotword and the introduction declaration to the speech-enabled home device; and

providing a personalized response to the voice command to the speech-enabled home device, for output.

2. The method of claim 1 , wherein the hotword comprises:

one or more terms that are used to both (i) trigger processing of an utterance and (ii) perform speaker identification.

3. The method of claim 1 , wherein the introduction declaration comprises:

one or more terms in the utterance that follow the hotword.

4. The method of claim 1 , wherein the voice command comprises:

one or more terms that indicate an action to be performed.

5. The method of claim 1 , further comprising providing access to a resource of the user that is not accessible until the user is authenticated.

6. The method of claim 1 , comprising:

during the speech registration session, requesting that the user speak the single registration utterance.

7. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, during a speech registration session and by a speech-enabled home device that includes one or more microphones for detecting utterances spoken in a home environment, a single registration utterance by a user that includes a hotword and an introduction declaration;

registering the user by a server-based voice authentication device that includes an automated speech recognizer and that is associated with the speech-enabled home device, wherein registering includes training, by the server-based voice authentication device, a text-dependent speaker identification model using the hotword of the single registration utterance and training, by the server-based voice authentication device, a text-independent speaker identification model using the introduction declaration of the single registration utterance;

after the speech registration session is concluded, receiving, by the speech-enabled home device, an authentication utterance by the user that includes the hotword and a voice command that is different from the introduction declaration;

in response to receiving the authentication utterance by the user that includes the hotword and the voice command that is different from the introduction declaration, authenticating, by the server-based voice authentication device, the user, wherein authenticating includes processing the hotword of the authentication utterance by the server-based voice authentication device using the text-dependent speaker identification model and processing the voice command using the text-independent speaker identification model;

in response to authenticating the user by the server-based voice authentication device, providing, by the server-based voice authentication device, access to an access-controlled personal resource of the user without requiring the user to submit any further authentication information other than the single registration utterance by the user that includes the hotword and the introduction declaration to the speech-enabled home device; and

providing a personalized response to the voice command to the speech-enabled home device, for output.

8. The system of claim 7 , wherein the hotword comprises:

one or more terms that are used to both (i) trigger processing of an utterance and (ii) perform speaker identification.

9. The system of claim 7 , wherein the introduction declaration comprises:

one or more terms in the utterance that follow the hotword.

10. The system of claim 7 , wherein the voice command comprises:

one or more terms that indicate an action to be performed.

11. The system of claim 7 , further comprising: providing access to a resource of the user that is not accessible until the user is authenticated.

12. The system of claim 7 , the operations comprising:

during the speech registration session, requesting that the user speak the single registration utterance.

13. One or more non-transitory computer-readable media storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, during a speech registration session and by a speech-enabled home device that includes one or more microphones for detecting utterances spoken in a home environment, a single registration utterance by a user that includes a hotword and an introduction declaration;

registering the user by a server-based voice authentication device that includes an automated speech recognizer and that is associated with the speech-enabled home device, wherein registering includes training, by the server-based voice authentication device, a text-dependent speaker identification model using the hotword of the single registration utterance and training, by the server-based voice authentication device, a text-independent speaker identification model using the introduction declaration of the single registration utterance;

after the speech registration session is concluded, receiving, by the speech-enabled home device, an authentication utterance by the user that includes the hotword and a voice command that is different from the introduction declaration;

in response to receiving the authentication utterance by the user that includes the hotword and the voice command that is different from the introduction declaration, authenticating, by the server-based voice authentication device, the user, wherein authenticating includes processing the hotword of the authentication utterance by the server-based voice authentication device using the text-dependent speaker identification model and processing the voice command using the text-independent speaker identification model;

in response to authenticating the user by the server-based voice authentication device, providing, by the server-based voice authentication device, access to an access-controlled personal resource of the user without requiring the user to submit any further authentication information other than the single registration utterance by the user that includes the hotword and the introduction declaration to the speech-enabled home device; and

providing a personalized response to the voice command to the speech-enabled home device, for output.

14. The media of claim 13 , wherein the hotword comprises:

one or more terms that are used to both (i) trigger processing of an utterance and (ii) perform speaker identification.

15. The media of claim 13 , wherein the introduction declaration comprises:

one or more terms in the utterance that follow the hotword.

16. The media of claim 13 , wherein the voice command comprises:

one or more terms that indicate an action to be performed.

17. The method of claim 1 , comprising determining a weighted combination of a confidence level associated with the text-dependent speaker identification model and a confidence level associated with the text-independent speaker identification model;

wherein authenticating the user comprises evaluating the weighted combination.

18. The method of claim 1 , wherein authenticating the user based on the authentication utterance comprises weighting a confidence level associated with the text-dependent speaker identification model and a confidence level associated with the text-independent speaker identification model using a first weighting;

wherein the method includes:

training the text-independent speaker identification model based on one or more subsequent utterances;

after training the text-independent speaker identification model based on one or more subsequent utterances, receiving a second authentication utterance; and

authenticating the user based on the second authentication utterance, comprising weighting a confidence level associated with the text-dependent speaker identification model and a confidence level associated with the text-independent speaker identification model using a second weighting that is different from the first weighting.

19. The method of claim 18 , wherein the first weighting more heavily weights the confidence level associated with the text-dependent model compared to the confidence level associated with the text-independent model.

20. The method of claim 18 , further comprising determining to use the second weighting for authentication based on the second authentication utterance based on determining that at least a predetermined number of utterances of the user have been processed by the server-based voice authentication device.

Assignments (2)
CHANGE OF NAME Recorded Oct 20, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044567/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2016
From: SHARIFI, MATTHEW; ROBLEK, DOMINIK
To: GOOGLE INC.
Reel/Frame 039111/0692 →
Continuity (1)
Continuation 13944975 · Jul 18, 2013
Cited By (28)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,502 US 12,216,894 US 12,219,314 US 12,223,282 US 12,236,952 US 12,248,551 US 12,254,887 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,333,404 US 12,361,943 US 12,367,879 US 12,380,895 US 12,386,434 US 12,386,491 US 12,431,128 US 12,451,140 US 12,477,470 US 12,556,890 US 12,608,171 US 12,613,730 US 12,619,452