IP Library Granted Patent US 10,152,974
Granted Patent B2
US 10,152,974 · App. 15/457,911 · Granted Dec 11, 2018

Unobtrusive training for speaker verification

Inventors: Todd F. Mozer (Los Altos Hills, CA); Bryan Pellom (Erie, CO)
Assignee: Sensory, Incorporated
G10L17/04G10L17/02G10L17/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,152,974
App. No.
15/457,911
Granted
Dec 11, 2018
Kind
B2
Abstract

Techniques for implementing unobtrusive training for speaker verification are provided. In one embodiment, an electronic device can receive a plurality of voice samples uttered by one or more users as they interact with a voice command-and-control feature of the electronic device and, for each voice sample, assign the voice sample to one of a plurality of voice type categories. The electronic device can further group the voice samples assigned to each voice type category into one or more user sets, where each user set comprises voice samples likely to have been uttered by a unique user. The electronic device can then, for each user set: (1) generate a voice model, (2) issue, to the unique user, a request to provide an identity or name, and (3) label the voice model with the identity or name provided by the unique user.

Claims (65)

1. A method comprising:

receiving, by an electronic device, a plurality of voice samples uttered by one or more users as they interact with a voice command-and-control feature of the electronic device;

for each voice sample, assigning, by the electronic device, the voice sample to one of a plurality of voice type categories;

for each voice type category, grouping, by the electronic device, voice samples assigned to the voice type category into one or more user sets, each user set comprising one or more voice samples likely to have been uttered by a unique user in the one or more users; and

for each user set:

generating, by the electronic device, a voice model;

issuing, by the electronic device to the unique user, a request to provide an identity or name; and

labeling, by the electronic device, the voice model with the identity or name provided by the unique user.

2. The method of claim 1 wherein the assigning of the voice sample to one of the plurality of voice type categories is based on vocal characteristics of the voice sample.

3. The method of claim 1 wherein the assigning of the voice sample to one of the plurality of voice type categories is further based on one or more other biometric factors associated with the utterance of the voice sample.

4. The method of claim 1 wherein the grouping of voice samples assigned to the voice type category into one or more user sets is based on a quantity of the voice samples and similarity of vocal characteristics between the voice samples.

5. The method of claim 1 further comprising:

using the labeled voice models to verify the one or more users' identities with respect to future voice queries or actions.

6. The method of claim 5 wherein using the labeled voice models to verify the one or more users' identities with respect to future voice queries or actions comprises, for a particular user:

allowing or disallowing a voice query or an action initiated by the particular user based on the particular user's identity and one or more permissions defined on the electronic device.

7. The method of claim 5 further comprising:

upon verifying the identity of a particular user using the user's labeled voice model, creating or adapting a speech recognition model that is specific to the particular user based on voice samples received from the particular user.

8. The method of claim 1 further comprising, subsequently to the labeling:

receiving another voice sample that belongs to a particular user set; and

updating the voice model generated from the particular user set based on vocal characteristics of said another voice sample.

9. The method of claim 1 wherein the generated voice model is a text-dependent or text-independent voice model.

10. A non-transitory computer readable storage medium having stored thereon program code executable by a processor of an electronic device, the program code causing the processor to:

receive a plurality of voice samples uttered by one or more users as they interact with a voice command-and-control feature of the electronic device;

for each voice sample, assign the voice sample to one of a plurality of voice type categories;

for each voice type category, group voice samples assigned to the voice type category into one or more user sets, each user set comprising one or more voice samples likely to have been uttered by a unique user in the one or more users; and

for each user set:

generate a voice model;

issue, to the unique user, a request to provide an identity or name; and

label the voice model with the identity or name provided by the unique user.

11. The non-transitory computer readable storage medium of claim 10 wherein the assigning of the voice sample to one of the plurality of voice type categories is based on vocal characteristics of the voice sample.

12. The non-transitory computer readable storage medium of claim 10 wherein the assigning of the voice sample to one of the plurality of voice type categories is further based on one or more other biometric factors associated with the utterance of the voice sample.

13. The non-transitory computer readable storage medium of claim 10 wherein the grouping of voice samples assigned to the voice type category into one or more user sets is based on a quantity of the voice samples and similarity of vocal characteristics between the voice samples.

14. The non-transitory computer readable storage medium of claim 10 wherein the program code further causes the processor to:

use the labeled voice models to verify the one or more users' identities with respect to future voice queries or actions.

15. The non-transitory computer readable storage medium of claim 14 wherein using the labeled voice models to verify the one or more users' identities with respect to future voice queries or actions comprises, for a particular user:

allowing or disallowing a voice query or an action initiated by the particular user based on the particular user's identity and one or more permissions defined on the electronic device.

16. The non-transitory computer readable storage medium of claim 14 wherein the program code further causes the processor to:

upon verifying the identity of a particular user using the user's labeled voice model, create or adapt a speech recognition model that is specific to the particular user based on voice samples received from the particular user.

17. The non-transitory computer readable storage medium of claim 10 wherein the program code further causes the processor to, subsequent to the labeling:

receive another voice sample that belongs to a particular user set; and

update the voice model generated from the particular user set based on vocal characteristics of said another voice sample.

18. The non-transitory computer readable storage medium of claim 10 wherein the generated voice model is a text-dependent or text-independent voice model.

19. An electronic device comprising:

a processor; and

a non-transitory computer readable medium having stored thereon program code that, when executed by the processor, causes the processor to:

receive a plurality of voice samples uttered by one or more users as they interact with a voice command-and-control feature of the electronic device;

for each voice sample, assign the voice sample to one of a plurality of voice type categories;

for each voice type category, group voice samples assigned to the voice type category into one or more user sets, each user set comprising one or more voice samples likely to have been uttered by a unique user in the one or more users; and

for each user set:

generate a voice model;

issue, to the unique user, a request to provide an identity or name; and

label the voice model with the identity or name provided by the unique user.

20. The electronic device of claim 19 wherein the assigning of the voice sample to one of the plurality of voice type categories is based on vocal characteristics of the voice sample.

21. The electronic device of claim 19 wherein the assigning of the voice sample to one of the plurality of voice type categories is further based on one or more other biometric factors associated with the utterance of the voice sample.

22. The electronic device of claim 19 wherein the grouping of voice samples assigned to the voice type category into one or more user sets is based on a quantity of the voice samples and similarity of vocal characteristics between the voice samples.

23. The electronic device of claim 19 wherein the program code further causes the processor to:

use the labeled voice models to verify the one or more users' identities with respect to future voice queries or actions.

24. The electronic device of claim 23 wherein using the labeled voice models to verify the one or more users' identities with respect to future voice queries or actions comprises, for a particular user:

allowing or disallowing a voice query or an action initiated by the particular user based on the particular user's identity and one or more permissions defined on the electronic device.

25. The electronic device of claim 23 wherein the program code further causes the processor to:

upon verifying the identity of a particular user using the user's labeled voice model, create or adapt a speech recognition model that is specific to the particular user based on voice samples received from the particular user.

26. The electronic device of claim 19 wherein the program code further causes the processor to, subsequent to the labeling:

receive another voice sample that belongs to a particular user set; and

update the voice model generated from the particular user set based on vocal characteristics of said another voice sample.

27. The electronic device of claim 19 wherein the generated voice model is a text-dependent or text-independent voice model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2017
From: MOZER, TODD F.; PELLOM, BRYAN
To: SENSORY, INCORPORATED
Reel/Frame 041563/0547 →
Continuity (2)
Provisional Application 62323038 · Apr 15, 2016
Related Publication 20170301353A1 · Oct 19, 2017