IP Library Granted Patent US 10,986,498
Granted Patent B2
US 10,986,498 · App. 16/573,581 · Granted Apr 20, 2021

Speaker verification using co-location information

Inventors: Raziel Alvarez Guevara (Menlo Park, CA); Othar Hansson (Sunnyvale, CA)
Assignee: Google LLC
H04W12/06G06F21/32G10L15/08G10L15/18G10L17/00G10L17/20G10L17/22G10L17/24G10L19/00H04L63/0861G06F2221/2111G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,986,498
App. No.
16/573,581
Granted
Apr 20, 2021
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for identifying a user in a multi-user environment. One of the methods includes receiving, by a first user device, an audio signal encoding an utterance, obtaining, by the first user device, a first speaker model for a first user of the first user device, obtaining, by the first user device for a second user of a second user device that is co-located with the first user device, a second speaker model for the second user or a second score that indicates a respective likelihood that the utterance was spoken by the second user, and determining, by the first user device, that the utterance was spoken by the first user using (i) the first speaker model and the second speaker model or (ii) the first speaker model and the second score.

Claims (46)

1. A method comprising:

for each user of a plurality of different users of a user device:

receiving, at data processing hardware of a server in communication with the user device, one or more sample utterances from the corresponding user of the plurality of different users during an enrollment process; and

generating, by the data processing hardware, corresponding speaker verification data for the corresponding user of the plurality of different users of the user device based on the one or more sample utterances received from the corresponding user of the plurality of different users of the user device;

receiving, at the data processing hardware, audio data corresponding to an utterance captured by the user device having the plurality of different users, each user of the plurality of different users having different corresponding user permissions to access a plurality of applications on the user device;

determining, by the data processing hardware, a speaker of the utterance from one of the plurality of different users of the user device based on a comparison between the received audio data and the corresponding speaker verification data generated for each user of the plurality of different users of the user device; and

processing, by the data processing hardware, the audio data corresponding to the utterance using a speech recognition module to identify a particular action for the user device to execute, the particular action, when executed by the user device, launching a particular application of the plurality of applications on the user device based on the corresponding user permissions associated with the determined speaker to access the particular application.

2. The method of claim 1 , wherein the corresponding speaker verification data for each corresponding user of the plurality of different users of the user device is configured to discriminate utterances spoken by the corresponding user from the other users of the plurality of different users of the user device.

3. The method of claim 1 , wherein generating the corresponding speaker verification data comprises generating, for each user of the plurality of different users of the user device, a corresponding speaker verification model based on the one or more sample utterances received from the corresponding user of the plurality of different users during the enrollment process.

4. The method of claim 3 , wherein at least one of the corresponding speaker verification models comprises an i-vector speaker verification model.

5. The method of claim 3 , wherein at least one of the corresponding speaker verification models comprises a d-vector speaker verification model.

6. The method of claim 1 , wherein receiving the audio data corresponding to the utterance comprises receiving the audio data corresponding to the utterance of a particular, predefined hotword followed by a voice command.

7. The method of claim 6 , wherein the user device is configured to respond to voice commands while in a locked state upon receipt of the particular, predefined hotword.

8. The method of claim 1 , wherein determining the speaker of the utterance from one of the plurality of different users of the user device comprises:

for each user of the plurality of different users of the user device:

obtaining the corresponding speaker verification data from memory hardware in communication with the data processing hardware; and

generating a corresponding speaker verification score by comparing the corresponding speaker verification data and the audio data, the corresponding speaker verification score indicating a likelihood that the utterance was spoken by the corresponding user of the plurality of different users of the user device; and

identifying the speaker of the utterance as the user of the plurality of different users of the user device associated with a highest corresponding speaker verification score.

9. The method of claim 8 , further comprising, prior to identifying the speaker of the utterance, determining, by the data processing hardware, that the highest corresponding speaker verification score satisfies an acceptance threshold.

10. The method of claim 1 , further comprising, for each user of the plurality of different users of the user device:

after generating the corresponding speaker verification data for the corresponding user of the plurality of different users of the user device, storing, by the data processing hardware, the corresponding speaker verification data in memory hardware in communication with the data processing hardware; and

associating, by the data processing hardware, the corresponding speaker verification data stored in the memory hardware with a corresponding user identifier associated with the corresponding user of the plurality of different users of the user device.

11. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware and storing instructions, that when executed by the data processing hardware, cause the data processing hardware to perform operations comprising:

for each user of a plurality of different users of a user device:

receiving one or more sample utterances from the corresponding user of the plurality of different users during an enrollment process, each user of the plurality of different users having different corresponding user permissions to access a plurality of applications on the user device; and

generating corresponding speaker verification data for the corresponding user of the plurality of different users of the user device based on the one or more sample utterances received from the corresponding user of the plurality of different users of the user device;

receiving audio data corresponding to an utterance captured by the user device having the plurality of different users;

determining a speaker of the utterance from one of the plurality of different users of the user device based on a comparison between the received audio data and the corresponding speaker verification data generated for each user of the plurality of different users of the user device; and

processing the audio data corresponding to the utterance using a speech recognition module to identify a particular action for the user device to execute, the particular action, when executed by the user device, launching a particular application of the plurality of applications on the user device based on the corresponding user permissions associated with the determined speaker to access the particular application.

12. The system of claim 11 , wherein the corresponding speaker verification data for each corresponding user of the plurality of different users of the user device is configured to discriminate utterances spoken by the corresponding user from the other users of the plurality of different users of the user device.

13. The system of claim 11 , wherein generating the corresponding speaker verification data comprises generating, for each user of the plurality of different users of the user device, a corresponding speaker verification model based on the one or more sample utterances received from the corresponding user of the plurality of different users during the enrollment process.

14. The system of claim 13 , wherein at least one of the corresponding speaker verification models comprises an i-vector speaker verification model.

15. The system of claim 13 , wherein at least one of the corresponding speaker verification models comprises a d-vector speaker verification model.

16. The system of claim 11 , wherein receiving the audio data corresponding to the utterance comprises receiving the audio data corresponding to the utterance of a particular, predefined hotword followed by a voice command.

17. The system of claim 16 , wherein the user device is configured to respond to voice commands while in a locked state upon receipt of the particular, predefined hotword.

18. The system of claim 11 , wherein determining the speaker of the utterance from one of the plurality of different users of the user device comprises:

for each user of the plurality of different users of the user device:

obtaining the corresponding speaker verification data from the memory hardware; and

generating a corresponding speaker verification score by comparing the corresponding speaker verification data and the audio data, the corresponding speaker verification score indicating a likelihood that the utterance was spoken by the corresponding user of the plurality of different users of the user device; and

identifying the speaker of the utterance as the user of the plurality of different users of the user device associated with a highest corresponding speaker verification score.

19. The system of claim 18 , wherein the operations further comprise, prior to identifying the speaker of the utterance, determining that the highest corresponding speaker verification score satisfies an acceptance threshold.

20. The system of claim 11 , wherein the operations further comprise, for each user of the plurality of different users of the user device:

after generating the corresponding speaker verification data for the corresponding user of the plurality of different users of the user device, storing the corresponding speaker verification data in the memory hardware; and

associating the corresponding speaker verification data stored in the memory hardware with a corresponding user identifier associated with the corresponding user of the plurality of different users of the user device.

Continuity (6)
Continuation 16172221 · Oct 26, 2018
Continuation 15697052 · Sep 6, 2017
Continuation 15201972 · Jul 5, 2016
Continuation 14805687 · Jul 22, 2015
Continuation 14335380 · Jul 18, 2014
Related Publication 20200013412A1 · Jan 9, 2020