IP Library › Granted Patent US 10,726,829
Granted Patent B2
US 10,726,829 · App. 15/907,630 · Granted Jul 28, 2020

Performing speaker change detection and speaker recognition on a trigger phrase

Inventor: John Paul Lesso (Edinburgh, GB)
Assignee: Cirrus Logic, Inc.
G10L15/08G10L17/06G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,726,829
App. No.
15/907,630
Granted
Jul 28, 2020
Kind
B2
Abstract

A method of speaker recognition comprises receiving an audio signal representing speech. A speaker change detection process is performed on the received audio signal. A trigger phrase detection process is also performed on the received audio signal. On detecting the trigger phrase in the received audio signal, a speaker recognition process is performed on the detected trigger phrase and on any speech preceding the detected trigger phrase and following an immediately preceding speaker change.

Claims (93)

1. A method of speaker recognition, the method comprising:

receiving an audio signal representing speech;

buffering the received audio signal;

attempting to detect a predetermined trigger phrase in the received audio signal; and

in response to detecting the predetermined trigger phrase in the received audio signal:

retrieving the buffered audio signal;

performing a speaker change detection process on the received audio signal including the retrieved buffered audio signal; and

performing a first speaker recognition process on the detected predetermined trigger phrase;

performing a second speaker recognition process on speech preceding the detected predetermined trigger phrase and following an immediately preceding speaker change; and

fusing results of the first and second speaker recognition processes.

2. A method according to claim 1 , wherein the first speaker recognition process is a text-dependent speaker recognition process and the second speaker recognition process is a text-independent speaker recognition process.

3. A method according to claim 1 , comprising buffering the received audio signal for a fixed time period on a first-in, first-out basis.

4. A method according to claim 1 , comprising performing the speaker change detection process on the received audio signal continuously, and generating a speaker change detection flag when a speaker change is recognised.

5. A method according to claim 1 , wherein the speaker change detection process is based on a detected angle of arrival of sound that gives rise to the audio signal.

6. A method according to claim 1 , wherein the speaker change detection process is based on a detected characteristic frequency of the speech.

7. A method according to claim 1 , wherein the speaker change detection process is based on extracting feature vectors for respective time windows of the received audio signal, and determining when a statistical difference between feature vectors of successive time windows exceeds a threshold.

8. A method according to claim 1 , further comprising:

on detecting the predetermined trigger phrase in the received audio signal, no longer attempting to detect the predetermined trigger phrase in the received audio signal.

9. A method according to claim 1 , wherein the speaker recognition process determines whether the received audio signal is derived from the speech of an enrolled user, the method further comprising, if it is determined that the received audio signal is derived from the speech of an enrolled user:

performing a speech recognition process on speech following the detected predetermined trigger phrase.

10. A method according to claim 9 , further comprising extracting a command from the speech following the detected predetermined trigger phrase.

11. A speaker recognition system, comprising:

an input, for receiving an audio signal representing speech;

a buffer for storing the received audio signal;

and wherein the system is configured for:

attempting to detect a predetermined trigger phrase in the received audio signal; and

in response to detecting the predetermined trigger phrase in the received audio signal:

retrieving the stored audio signal;

performing a speaker change detection process on the received audio signal including the retrieved stored audio signal; and

performing a first speaker recognition process on the detected predetermined trigger phrases;

performing a second speaker recognition process on speech preceding the detected predetermined trigger phrase and following an immediately preceding speaker change; and

fusing results of the first and second speaker recognition processes.

12. A device comprising a speaker recognition system according to claim 11 .

13. A non-transitory computer-readable medium, comprising code stored thereon, for causing a processor to perform a method comprising:

receiving an audio signal representing speech;

buffering the received audio signal;

attempting to detect a predetermined trigger phrase in the received audio signal; and

in response to detecting the predetermined trigger phrase in the received audio signal:

retrieving the buffered audio signal;

performing a speaker change detection process on the received audio signal including the retrieved buffered audio signal; and

performing a first speaker recognition process on the detected predetermined trigger phrase;

performing a second speaker recognition process on speech preceding the detected predetermined trigger phrase and following an immediately preceding speaker changes and

fusing results of the first and second speaker recognition processes.

14. A method of speaker recognition, the method comprising:

receiving an audio signal representing speech;

buffering the received audio signal;

attempting to detect a predetermined trigger phrase in the received audio signal; and

in response to detecting the predetermined trigger phrase in the received audio signal:

retrieving the buffered audio signal;

performing a speaker change detection process on the received audio signal including the retrieved buffered audio signal; and

performing a speaker recognition process on the detected predetermined trigger phrase, and on any speech preceding the detected predetermined trigger phrase and following an immediately preceding speaker change, and on any speech following the detected predetermined trigger phrase and preceding an immediately following speaker change.

15. A method according to claim 14 , comprising:

performing a first speaker recognition process on the detected predetermined trigger phrase;

performing a second speaker recognition process on speech preceding the detected predetermined trigger phrase and following an immediately preceding speaker change and on speech following the detected predetermined trigger phrase and preceding an immediately following speaker change; and

fusing results of the first and second speaker recognition processes.

16. A method according to claim 15 , wherein the first speaker recognition process is a text-dependent speaker recognition process and the second speaker recognition process is a text-independent speaker recognition process.

17. A method according to claim 14 , comprising:

performing a first speaker recognition process on the detected predetermined trigger phrase;

performing a second speaker recognition process on speech preceding the detected predetermined trigger phrase and following an immediately preceding speaker change;

performing a third speaker recognition process on speech following the detected predetermined trigger phrase and preceding an immediately following speaker change; and

fusing results of the first, second and third speaker recognition processes.

18. A method according to claim 17 , wherein the first speaker recognition process is a text-dependent speaker recognition process and the second and third speaker recognition processes are text-independent speaker recognition processes.

19. A method according to claim 14 , comprising buffering the received audio signal for a fixed time period on a first-in, first-out basis.

20. A method according to claim 14 , comprising performing the speaker change detection process on the received audio signal continuously, and generating a speaker change detection flag when a speaker change is recognised.

21. A method according to claim 14 , wherein the speaker change detection process is based on a detected angle of arrival of sound that gives rise to the audio signal.

22. A method according to claim 14 , wherein the speaker change detection process is based on a detected characteristic frequency of the speech.

23. A method according to claim 14 , wherein the speaker change detection process is based on extracting feature vectors for respective time windows of the received audio signal, and determining when a statistical difference between feature vectors of successive time windows exceeds a threshold.

24. A method according to claim 14 , further comprising:

on detecting the predetermined trigger phrase in the received audio signal, no longer attempting to detect the predetermined trigger phrase in the received audio signal.

25. A method according to claim 14 , wherein the speaker recognition process determines whether the received audio signal is derived from the speech of an enrolled user, the method further comprising,

if it is determined that the received audio signal is derived from the speech of an enrolled user:

performing a speech recognition process on speech following the detected predetermined trigger phrase.

26. A method according to claim 25 , further comprising extracting a command from the speech following the detected predetermined trigger phrase.

27. A speaker recognition system, comprising:

an input, for receiving an audio signal representing speech;

a buffer for storing the received audio signal;

and wherein the system is configured for:

receiving an audio signal representing speech;

buffering the received audio signal;

attempting to detect a predetermined trigger phrase in the received audio signal; and

in response to detecting the predetermined trigger phrase in the received audio signal:

retrieving the buffered audio signal;

performing a speaker change detection process on the received audio signal including the retrieved buffered audio signal; and

performing a speaker recognition process on the detected predetermined trigger phrase, and on any speech preceding the detected predetermined trigger phrase and following an immediately preceding speaker change, and on any speech following the detected predetermined trigger phrase and preceding an immediately following speaker change.

28. A device comprising a speaker recognition system according to claim 27 .

29. A non-transitory computer-readable medium, comprising code stored thereon, for causing a processor to perform a method comprising:

receiving an audio signal representing speech;

buffering the received audio signal;

attempting to detect a predetermined trigger phrase in the received audio signal; and

in response to detecting the predetermined trigger phrase in the received audio signal:

retrieving the buffered audio signal;

performing a speaker change detection process on the received audio signal including the retrieved buffered audio signal; and

performing a speaker recognition process on the detected predetermined trigger phrase, and on any speech preceding the detected predetermined trigger phrase and following an immediately preceding speaker change, and on any speech following the detected predetermined trigger phrase and preceding an immediately following speaker change.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 11, 2020
From: CIRRUS LOGIC INTERNATIONAL SEMICONDUCTOR LTD.
To: CIRRUS LOGIC, INC.
Reel/Frame 052907/0088 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2019
From: LESSO, JOHN PAUL
To: CIRRUS LOGIC INTERNATIONAL SEMICONDUCTOR LTD.
Reel/Frame 050734/0324 →
Continuity (1)
Related Publication 20190266996A1 · Aug 29, 2019