IP Library › Granted Patent US 10,127,911
Granted Patent B2
US 10,127,911 · App. 14/835,169 · Granted Nov 13, 2018

Speaker identification and unsupervised speaker adaptation techniques

Inventors: Yoon Kim (Cupertino, CA); Sachin S. Kajarekar (Sunnyvale, CA)
Assignee: Apple Inc.
G10L17/26G10L15/26G10L17/04G10L17/06G10L15/1822
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,127,911
App. No.
14/835,169
Filed
Aug 25, 2015
Granted
Nov 13, 2018
Kind
B2
Art Unit
2658
USPC
704/235
Abstract

Systems and processes for generating a speaker profile for use in performing speaker identification for a virtual assistant are provided. One example process can include receiving an audio input including user speech and determining whether a speaker of the user speech is a predetermined user based on a speaker profile for the predetermined user. In response to determining that the speaker of the user speech is the predetermined user, the user speech can be added to the speaker profile and operation of the virtual assistant can be triggered. In response to determining that the speaker of the user speech is not the predetermined user, the user speech can be added to an alternate speaker profile and operation of the virtual assistant may not be triggered. In some examples, contextual information can be used to verify results produced by the speaker identification process.

Claims (170)

1. A method for operating a virtual assistant, the method comprising:

at an electronic device:

receiving, at the electronic device, an audio input comprising user speech, wherein the audio input is associated with a contextual data;

determining whether the user speech contains one or more predetermined words;

in response to determining that the user speech contains one or more predetermined words:

determining whether a speaker of the user speech is a predetermined user based at least in part on a speaker profile for the predetermined user; and

in accordance with a determination that the speaker of the user speech is the predetermined user, adding the audio input comprising user speech to the speaker profile for the predetermined user, wherein adding the audio input comprising user speech to the speaker profile includes annotating the audio input in the speaker profile with the contextual data;

receiving a second audio input comprising a second user speech;

determining whether a second contextual data associated with the second audio input matches the contextual data;

in accordance with a determination that the second contextual data associated with the second audio input matches the contextual data:

determining whether a speaker of the second user speech is the predetermined user based at least in part on the audio input added to the speaker profile; and

in accordance with a determination that the speaker of the second user speech is the predetermined user, activating the virtual assistant and processing a spoken command received subsequent to the second user speech.

2. The method of claim 1 , wherein the speaker profile for the predetermined user comprises a plurality of voice prints.

3. The method of claim 2 , wherein each of the plurality of voice prints of the speaker profile for the predetermined user was generated from previously received audio inputs comprising user speech.

4. The method of claim 2 , wherein determining whether the speaker of the user speech is the predetermined user based at least in part on the speaker profile for the predetermined user comprises:

determining whether the audio input comprising user speech matches at least a threshold number of the plurality of voice prints;

in accordance with a determination that the audio input comprising user speech matches at least the threshold number of the plurality of voice prints, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match at least the threshold number of the plurality of voice prints, determining that the speaker of the user speech is not the predetermined user.

5. The method of claim 2 , wherein determining whether the speaker of the user speech is the predetermined user based at least in part on the speaker profile for the predetermined user comprises:

determining whether the audio input comprising user speech matches at least a threshold number of the plurality of voice prints;

in accordance with a determination that the audio input comprising user speech matches at least the threshold number of the plurality of voice prints:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match at least the threshold number of the plurality of voice prints:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is the predetermined user.

6. The method of claim 1 , wherein adding the audio input comprising user speech to the speaker profile for the predetermined user comprises:

generating a voice print from the audio input comprising user speech; and

storing the voice print in association with the speaker profile for the predetermined user.

7. The method of claim 1 , wherein the method further comprises:

in accordance with a determination that the speaker of the user speech is not the predetermined user, adding the audio input comprising user speech to a speaker profile for an alternate user.

8. The method of claim 7 , wherein the speaker profile for the alternate user comprises a plurality of voice prints.

9. The method of claim 8 , wherein each of the plurality of voice prints of the speaker profile for the alternate user was generated from previously received audio inputs comprising user speech.

10. The method of claim 7 , wherein determining whether the speaker of the user speech is the predetermined user is further based at least in part on the speaker profile for the alternate user.

11. The method of claim 7 , wherein determining whether the speaker of the user speech is the predetermined user comprises:

determining whether the audio input comrising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user;

in accordance with a determination that the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user, determining that the speaker of the user speech is not the predetermined user.

12. The method of claim 7 , wherein determining whether the speaker of the user speech is the predetermined user comprises:

determining whether the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user;

in accordance with a determination that the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is the predetermined user.

13. The method of claim 1 , wherein the method further comprises:

in accordance with a determination that the speaker of the user speech is the predetermined user:

performing speech-to-text conversion on a third audio input comprising a third user speech, wherein the third audio input is received after receiving the audio input comprising user speech;

determining a user intent based on the third user speech;

determining a task to be performed based on the third user speech;

determining a parameter for the task to be performed based on the third user speech; and

performing the task to be performed in accordance with the determined parameter.

14. A system comprising:

one or more processors;

memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving an audio input comprising user speech, wherein the audio input is associated with a contextual data;

determining whether the user speech contains one or more predetermined words;

in response to determining that the user speech contains one or more predetermined words:

determining whether a speaker of the user speech is a predetermined user based at least in part on a speaker profile for the predetermined user; and

in accordance with a determination that the speaker of the user speech is the predetermined user, adding the audio input comprising user speech to the speaker profile for the predetermined user, wherein adding the audio input comprising user speech to the speaker profile includes annotating the audio input in the speaker profile with the contextual data;

receiving a second audio input comprising a second user speech;

determining whether a second contextual data associated with the second audio input matches the contextual data;

in accordance with a determination that the second contextual data associated with the second audio input matches the contextual data:

determining whether a speaker of the second user speech is the predetermined user based at least in part on the audio input added to the speaker profile; and

in accordance with a determination that the speaker of the second user speech is the predetermined user, activating the virtual assistant and processing a spoken command received subsequent to the second user speech.

15. The system of claim 14 , wherein adding the audio input comprising user speech to the speaker profile for the predetermined user comprises:

generating a voice print from the audio input comprising user speech; and

storing the voice print in association with the speaker profile for the predetermined user.

16. The system of claim 14 , wherein the one or more programs further includes instructions for:

in accordance with a determination that the speaker of the user speech is the predetermined user:

performing speech-to-text conversion on a third audio input comprising a third user speech, wherein the third audio input is received after receiving the audio input comprising user speech;

determining a user intent based on the third user speech;

determining a task to be performed based on the third user speech;

determining a parameter for the task to be performed based on the third user speech; and

performing the task to be performed in accordance with the determined parameter.

17. The system of claim 14 , wherein the speaker profile for the predetermined user comprises a plurality of voice prints.

18. The system of claim 17 , wherein each of the plurality of voice prints of the speaker profile for the predetermined user was generated from previously received audio inputs comprising user speech.

19. The system of claim 17 , wherein determining whether the speaker of the user speech is the predetermined user based at least in part on the speaker profile for the predetermined user comprises:

determining whether the audio input comprising user speech matches at least a threshold number of the plurality of voice prints;

in accordance with a determination that the audio input comprising user speech matches at least the threshold number of the plurality of voice prints, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match at least the threshold number of the plurality of voice prints, determining that the speaker of the user speech is not the predetermined user.

20. The system of claim 17 , wherein determining whether the speaker of the user speech is the predetermined user based at least in part on the speaker profile for the predetermined user comprises:

determining whether the audio input comprising user speech matches at least a threshold number of the plurality of voice prints;

in accordance with a determination that the audio input comprising user speech matches at least the threshold number of the plurality of voice prints:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match at least the threshold number of the plurality of voice prints:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is the predetermined user.

21. The system of claim 14 , wherein the one or more programs further include instructions for:

in accordance with a determination that the speaker of the user speech is not the predetermined user, adding the audio input comprising user speech to a speaker profile for an alternate user.

22. The system of claim 21 , wherein determining whether the speaker of the user speech is the predetermined user is further based at least in part on the speaker profile for the alternate user.

23. The system of claim 21 , wherein determining whether the speacker of the user speech is the predetermined user comprises:

determining whether the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user;

in accordance with a determination that the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user, determining that the speaker of the user speech is not the predetermined user.

24. The system of claim 21 , wherein determining whether the speaker of the user speech is the predetermined user comprises:

determining whether the audio input comprising user speech matches a great number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user;

in accordance with a determination that the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user:

deterining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is the predetermined user.

25. The system of claim 21 , wherein the speaker profile for the alternate user comprises a plurality of voice prints.

26. The system of claim 21 , wherein each of the plurality of voice prints of the speaker profile for the alternate user was generated from previously received audio inputs comprising user speech.

27. A non-transitory computer-readable storage medium comprising instructions for:

receiving an audio input comprising user speech, wherein the audio input is associated with a contextual data;

determining whether the user speech contains one or more predetermined words;

in response to determining that the user speech contains one or more predetermined words:

determining whether a speaker of the user speech is a predetermined user based at least in part on a speaker profile for the predetermined user; and

in accordance with a determination that the speaker of the user speech is the predetermined user, adding the audio input comprising user speech to the speaker profile for the predetermined user, wherein adding the audio input comprising user speech to the speaker profile includes annotating the audio input in the speaker profile with the contextual data;

receiving a second audio input comprising a second user speech;

determining whether a second contextual data associated with the second audio input matches the contextual data;

in accordance with a determination that the second contextual data associated with the second audio input matches the contextual data:

determining whether a speaker of the second user speech is the predetermined user based at least in part on the audio input added to the speaker profile; and

in accordance with a determination that the speaker of the second user speech is the predetermined user, activating the virtual assistant and processing a spoken command received subsequent to the second user speech.

28. The non-transitory computer-readable storage medium of claim 27 , wherein determining whether the speaker of the user speech is the predetermined user is further based at least in part on the speaker profile for the alternate user.

29. The non-transitory computer-readable storage medium of claim 27 , wherein determining whether the speaker of the user speech is the predetermined user comprises:

determining whether the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user;

in accordance with a determination that the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user, determining that the speaker of the user speech is not the predetermined user.

30. The non-transitory computer-readable storage medium of claim 27 , wherein the speaker profile for the predetermined user comprises a plurality of voice prints.

31. The non-transitory computer-readable storage medium of claim 30 , wherein adding the audio input comprising user speech to the speaker profile for the predetermined user comprises:

generating a voice print from the audio input comprising user speech; and

storing the voice print in association with the speaker profile for the predetermined user.

32. The non-transitory computer-readable storage medium of claim 30 , wherein instructions further comprise:

in accordance with a determination that the speaker of the user speech is the predetermined user:

performing speech-to-text conversion on a third audio input comprising a third user speech, wherein the third audio input is received after receiving the audio input comprising user speech;

determining a user intent based on the third user speech;

determining a task to be performed based on the third user speech;

determining a parameter for the task to be performed based on the third user speech; and

performing the task to be performed in accordance with the determined parameter.

33. The non-transitory computer-readable storage medium of claim 30 , wherein each of the plurality of voice prints of the speaker profile for the predetermined user was generated from previously received audio inputs comprising user speech.

34. The non-transitory computer-readable storage medium of claim 29 , wherein determining whether the speaker of the user speech is the predetermined user based at least in part on the speaker profile for the predetermined user comprises:

determining whether the audio input comprising user speech matches at least a threshold number of the plurality of voice prints;

in accordance with a determination that the audio input comprising user speech matches at least the threshold number of the plurality of voice prints, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match at least the threshold number of the plurality of voice prints, determining that the speaker of the user speech is not the predetermined user.

35. The non-transitory computer-readable storage medium of claim 33 , wherein determining whether the speaker of the user speech is the predetermined user based at least in part on the speaker profile for the predetermined user comprises:

determining whether the audio input comprising user speech matches at least a threshold number of the plurality of voice prints;

in accordance with a determination that the audio input comprising user speech matches at least the threshold number of the plurality of voice prints:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match at least the threshold number of the plurality of voice prints:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is the predetermined user.

36. The non-transitory computer-readable storage medium of claim 30 , wherein the instructions further comprise:

in accordance with a determination that the speaker of the user speech is not the predetermined user, adding the audio input comprising user speech to a speaker profile for an alternate user.

37. The non-transitory computer-readable storage medium of claim 36 , wherein the speaker profile for the alternate user comprises a plurality of voice prints.

38. The non-transitory computer-readable storage medium of claim 37 , wherein each of the plurality of voice prints of the speaker profile for the alternate user was generated from previously received audio inputs comprising user speech.

39. The non-transitory computer-readable storage medium of claim 27 , wherein determining whether the speaker of the user speech is the predetermined user comprises:

determining whether the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user;

in accordance with a determination that the audio input comprising user speech matches a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that the audio input comprising user speech does not match a greater number of voice prints of the speaker profile for the predetermined user than a number of voice prints of the speaker profile for the alternate user:

determining whether an erroneous speaker determination was made based on the contextual data;

in accordance with a determination that an erroneous speaker determination was not made based on the contextual data, determining that the speaker of the user speech is not the predetermined user; and

in accordance with a determination that an erroneous speaker determination was made based on the contextual data, determining that the speaker of the user speech is the predetermined user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 25, 2015
From: KIM, YOON; KAJAREKAR, SACHIN S.
To: APPLE INC.
Reel/Frame 036417/0623 →
Continuity (2)
Provisional Application 62057990 · Sep 30, 2014
Related Publication 20160093304A1 · Mar 31, 2016
Cited By (21)
US 12,211,490 US 12,217,748 US 12,230,291 US 12,236,932 US 12,283,269 US 12,283,277 US 12,327,549 US 12,327,556 US 12,360,734 US 12,387,716 US 12,401,744 US 12,424,220 US 12,451,143 US 12,505,832 US 12,513,479 US 12,518,756 US 12,579,978 US 12,699,543 US 12,711,962 US 12,732,547 US 12,748,566