IP Library Granted Patent US 11,017,784
Granted Patent B2
US 11,017,784 · App. 16/557,390 · Granted May 25, 2021

Speaker verification across locations, languages, and/or dialects

Inventors: Ignacio Lopez Moreno (New York, NY); Li Wan (Forest Hills, NY); Quan Wang (Jersey City, NJ)
Assignee: Google LLC
G10L17/24G10L17/02G10L17/08G10L17/14G10L17/18G10L17/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,017,784
App. No.
16/557,390
Granted
May 25, 2021
Kind
B2
Abstract

Methods, systems, apparatus, including computer programs encoded on computer storage medium, to facilitate language independent-speaker verification. In one aspect, a method includes actions of receiving, by a user device, audio data representing an utterance of a user. Other actions may include providing, to a neural network stored on the user device, input data derived from the audio data and a language identifier. The neural network may be trained using speech data representing speech in different languages or dialects. The method may include additional actions of generating, based on output of the neural network, a speaker representation and determining, based on the speaker representation and a second representation, that the utterance is an utterance of the user. The method may provide the user with access to the user device based on determining that the utterance is an utterance of the user.

Claims (44)

1. A method comprising:

receiving audio data representing an utterance, spoken by a user of a user device, of a predetermined word or phrase designated as a hotword for a language or location associated with the user, wherein the user device is configured to perform an action or change a state of the user device in response to detecting an utterance of the hotword;

providing, as input to a speaker recognition system comprising at least one neural network, a set of input data derived from the audio data and a language identifier or location identifier associated with the user device, the speaker recognition system being trained using speech data representing speech in different languages or different dialects;

determining an identity of the user based on output of the speaker recognition system and a reference speaker representation derived from a previous utterance of the predetermined word or phrase designated as a hotword for the language or location associated with the user; and

providing a personalized response to the user based on determining the identity of the user.

2. The method of claim 1 , wherein the reference speaker representation is a speaker representation that was generated using output of the neural network generated in response to receiving (i) a set of input data derived from audio data of the previous utterance and (ii) the language identifier or location identifier associated with the user device.

3. The method of claim 1 , wherein the reference speaker representation is stored by the user device prior to receiving the audio data representing the utterance.

4. The method of claim 1 , wherein parameters of the neural network have been trained using training examples including utterances of a particular word or phrase designated as the hotword for multiple different languages or locations, wherein the particular word or phrase has a different pronunciation in at least some of multiple different languages or locations.

5. The method of claim 1 , wherein parameters of the neural network have been trained using training examples including utterances of different words or phrases designated as hotwords for different languages or locations.

6. The method of claim 1 , comprising:

determining a language of the utterance based on the audio data representing the utterance; and

determining the language identifier or location identifier based on determining the language of the utterance based on the audio data representing the utterance.

7. The method of claim 1 , wherein the set of input data derived from the audio data and the language identifier or location identifier includes:

a first vector that is derived from the audio data, and

a second vector corresponding to a language identifier or location identifier.

8. The method of claim 7 , comprising generating an input vector by concatenating the first vector and the second vector into a single concatenated vector;

wherein providing the set of input data comprises providing, to the neural network, the generated input vector; and

wherein the output of the speaker recognition system comprises a speaker representation generated based on output of the neural network produced in response to receiving the input vector, the speaker representation being indicative of characteristics of a voice of the user.

9. The method of claim 7 , comprising generating an input vector based on a weighted sum of the first vector and the second vector;

wherein providing the set of input data comprises providing, to the neural network, the generated input vector; and

wherein the output of the speaker recognition system comprises a speaker representation generated based on output of the neural network produced in response to receiving the input vector, the speaker representation being indicative of characteristics of a voice of the user.

10. The method of claim 1 , wherein the output of the speaker recognition system comprises output of the neural network, produced in response to receiving the set of input data, includes data indicating a set of activations at a layer of the neural network that was used as a hidden layer during training of the neural network.

11. A system comprising:

one or more computers; and

one or more computer-readable media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving audio data representing an utterance, spoken by a user of a user device, of a predetermined word or phrase designated as a hotword for a language or location associated with the user, wherein the user device is configured to perform an action or change a state of the user device in response to detecting an utterance of the hotword;

providing, as input to a speaker recognition system comprising at least one neural network, a set of input data derived from the audio data and a language identifier or location identifier associated with the user device, the speaker recognition system being trained using speech data representing speech in different languages or different dialects;

determining an identity of the user based on output of the speaker recognition system and a reference speaker representation derived from a previous utterance of the predetermined word or phrase designated as a hotword for the language or location associated with the user; and

providing a personalized response to the user based on determining the identity of the user.

12. The system of claim 11 , wherein the reference speaker representation is a speaker representation that was generated using output of the neural network generated in response to receiving (i) a set of input data derived from audio data of the previous utterance and (ii) the language identifier or location identifier associated with the user device.

13. The system of claim 11 , wherein the reference speaker representation is stored by the user device prior to receiving the audio data representing the utterance.

14. The system of claim 11 , wherein parameters of the neural network have been trained using training examples including utterances of a particular word or phrase designated as the hotword for multiple different languages or locations, wherein the particular word or phrase has a different pronunciation in at least some of multiple different languages or locations.

15. The system of claim 11 , wherein parameters of the neural network have been trained using training examples including utterances of different words or phrases designated as hotwords for different languages or locations.

16. The system of claim 11 , wherein the operations comprise:

determining a language of the utterance based on the audio data representing the utterance; and

determining the language identifier or location identifier based on determining the language of the utterance based on the audio data representing the utterance.

17. One or more non-transitory computer-readable media storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

receiving audio data representing an utterance, spoken by a user of a user device, of a predetermined word or phrase designated as a hotword for a language or location associated with the user, wherein the user device is configured to perform an action or change a state of the user device in response to detecting an utterance of the hotword;

providing, as input to a speaker recognition system comprising at least one neural network, a set of input data derived from the audio data and a language identifier or location identifier associated with the user device, the speaker recognition system being trained using speech data representing speech in different languages or different dialects;

determining an identity of the user based on output of the speaker recognition system and a reference speaker representation derived from a previous utterance of the predetermined word or phrase designated as a hotword for the language or location associated with the user; and

providing a personalized response to the user based on determining the identity of the user.

18. The one or more non-transitory computer-readable media of claim 17 , wherein the reference speaker representation is a speaker representation that was generated using output of the neural network generated in response to receiving (i) a set of input data derived from audio data of the previous utterance and (ii) the language identifier or location identifier associated with the user device.

19. The one or more non-transitory computer-readable media of claim 17 , wherein the reference speaker representation is stored by the user device prior to receiving the audio data representing the utterance.

20. The one or more non-transitory computer-readable media of claim 17 , wherein parameters of the neural network have been trained using training examples including utterances of a particular word or phrase designated as the hotword for multiple different languages or locations, wherein the particular word or phrase has a different pronunciation in at least some of multiple different languages or locations.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2019
From: MORENO, IGNACIO LOPEZ; WAN, LI; WANG, QUAN
To: GOOGLE INC.
Reel/Frame 050370/0036 →
CHANGE OF NAME Recorded Sep 13, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 050376/0066 →
Continuity (4)
Continuation 15995480 · Jun 1, 2018
Continuation PCTUS2017040906 · Jul 6, 2017
Continuation 15211317 · Jul 15, 2016
Related Publication 20190385619A1 · Dec 19, 2019