IP Library Granted Patent US 9,779,726
Granted Patent B2
US 9,779,726 · App. 15/105,882 · Granted Oct 3, 2017

Voice command triggered speech enhancement

Inventors: Robert James Hatfield (Edinburgh, GB); Michael Page (Oxfordshire, GB)
Assignee: Cirrus Logic International Semiconductor Ltd.
G10L15/08G10L15/063G10L15/20G10L15/22G10L15/285G10L21/0208G10L21/0216G10L25/84G10L2015/088G10L2021/02166
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,779,726
App. No.
15/105,882
Granted
Oct 3, 2017
Kind
B2
Abstract

Received data representing speech is stored, and a trigger detection block detects a presence of data representing a trigger phrase in the received data. In response, a first part of the stored data representing at least a part of the trigger phrase is supplied to an adaptive speech enhancement block, which is trained on the first part of the stored data to derive adapted parameters for the speech enhancement block. A second part of the stored data, overlapping with the first part of the stored data, is supplied to the adaptive speech enhancement block operating with said adapted parameters, to form enhanced stored data. A second trigger phrase detection block detects the presence of data representing the trigger phrase in the enhanced stored data. In response, enhanced speech data are output from the speech enhancement block for further processing, such as speech recognition.

Claims (67)

1. A method of processing received data representing speech, comprising:

storing the received data;

detecting a presence of data representing a trigger phrase in the received data;

in response to said detecting, supplying a first part of the stored data representing at least a part of the trigger phrase to an adaptive speech enhancement block;

training the speech enhancement block on the first part of the stored data to derive adapted parameters for the speech enhancement block;

supplying a second part of the stored data to the adaptive speech enhancement block operating with said adapted parameters, to form enhanced stored data, wherein the second part of the stored data overlaps with the first part of the stored data;

detecting the presence of data representing the trigger phrase in the enhanced stored data; and

outputting enhanced speech data from the speech enhancement block for further processing, in response to detecting the presence of data representing the trigger phrase in the enhanced stored data;

wherein the detecting the presence of data representing the trigger phrase in the received data is carried out by means of a first trigger phrase detection block; and

wherein the detecting the presence of data representing the trigger phrase in the enhanced stored data is carried out by means of a second trigger phrase detection block, and wherein the second trigger phrase detection block operates with different detection criteria from the first trigger phrase detection block.

2. A method as claimed in claim 1 , comprising, in response to failing to detect the presence of data representing the trigger phrase in the enhanced stored data, resetting the first trigger phrase detection block.

3. A method as claimed in claim 1 , wherein the second trigger phrase detection block operates with more rigorous detection criteria than the first trigger phrase detection block.

4. A method as claimed in claim 1 , comprising:

receiving and storing data from multiple microphones;

supplying data received from a subset of said microphones to the first trigger phrase detection block for detecting the presence of data representing the trigger phrase in the data received from the subset of said microphones;

in response to said detecting, supplying the first part of the stored data from said multiple microphones, representing at least a part of the trigger phrase, to the adaptive speech enhancement block;

training the speech enhancement block on the first part of the stored data from said multiple microphones to derive adapted parameters for the speech enhancement block; and

supplying the second part of the stored data from said multiple microphones to the adaptive speech enhancement block operating with said adapted parameters, to form said enhanced stored data.

5. A method as claimed in claim 4 , wherein the speech enhancement block is a beamformer.

6. A method as claimed in claim 1 , wherein the first part of the stored data is the data stored from a first defined starting point.

7. A method as claimed in claim 6 , wherein the second part of the stored data is the data stored from a second defined starting point, and the second defined starting point is later than the first defined starting point.

8. A method as claimed in claim 1 , comprising supplying the second part of the stored data to the speech enhancement block and outputting the enhanced speech data from the speech enhancement block at a higher rate than real time.

9. A method as claimed in claim 8 , comprising supplying the second part of the stored data to the speech enhancement block and outputting the enhanced speech data from the speech enhancement block at a higher rate than real time until the data being supplied is substantially time aligned with the data being stored.

10. A speech processor, comprising:

an input, for receiving data representing speech; and

a speech processing block, wherein the speech processing block is configured to perform a method of processing received data representing speech, comprising:

storing the received data;

detecting a presence of data representing a trigger phrase in the received data;

in response to said detecting, supplying a first part of the stored data representing at least a part of the trigger phrase to an adaptive speech enhancement block;

training the speech enhancement block on the first part of the stored data to derive adapted parameters for the speech enhancement block;

supplying a second part of the stored data to the adaptive speech enhancement block operating with said adapted parameters, to form enhanced stored data, wherein the second part of the stored data overlaps with the first part of the stored data;

detecting the presence of data representing the trigger phrase in the enhanced stored data; and

outputting enhanced speech data from the speech enhancement block for further processing, in response to detecting the presence of data representing the trigger phrase in the enhanced stored data;

wherein the detecting the presence of data representing the trigger phrase in the received data is carried out by means of a first trigger phrase detection block; and

wherein the detecting the presence of data representing the trigger phrase in the enhanced stored data is carried out by means of a second trigger phrase detection block, and wherein the second trigger phrase detection block operates with different detection criteria from the first trigger phrase detection block.

11. A speech processor, comprising:

an input, for receiving data representing speech; and

an output, for connection to a speech processing block, wherein the processing block is configured to perform a method of processing received data representing speech, comprising:

storing the received data;

detecting a presence of data representing a trigger phrase in the received data;

in response to said detecting, supplying a first part of the stored data representing at least a part of the trigger phrase to an adaptive speech enhancement block;

training the speech enhancement block on the first part of the stored data to derive adapted parameters for the speech enhancement block;

supplying a second part of the stored data to the adaptive speech enhancement block operating with said adapted parameters, to form enhanced stored data, wherein the second part of the stored data overlaps with the first part of the stored data;

detecting the presence of data representing the trigger phrase in the enhanced stored data; and

outputting enhanced speech data from the speech enhancement block for further processing, in response to detecting the presence of data representing the trigger phrase in the enhanced stored data;

wherein the detecting the presence of data representing the trigger phrase in the received data is carried out by means of a first trigger phrase detection block; and

wherein the detecting the presence of data representing the trigger phrase in the enhanced stored data is carried out by means of a second trigger phrase detection block, and wherein the second trigger phrase detection block operates with different detection criteria from the first trigger phrase detection block.

12. A mobile device, comprising a speech processor, wherein the speech processor is configured to perform a method of processing received data representing speech, comprising:

storing received data;

detecting a presence of data representing a trigger phrase in the received data;

in response to said detecting, supplying a first part of the stored data representing at least a part of the trigger phrase to an adaptive speech enhancement block;

training the speech enhancement block on the first part of the stored data to derive adapted parameters for the speech enhancement block;

supplying a second part of the stored data to the adaptive speech enhancement block operating with said adapted parameters, to form enhanced stored data, wherein the second part of the stored data overlaps with the first part of the stored data;

detecting the presence of data representing the trigger phrase in the enhanced stored data; and

outputting enhanced speech data from the speech enhancement block for further processing, in response to detecting the presence of data representing the trigger phrase in the enhanced stored data;

wherein the detecting the presence of data representing the trigger phrase in the received data is carried out by means of a first trigger phrase detection block; and

wherein the detecting the presence of data representing the trigger phrase in the enhanced stored data is carried out by means of a second trigger phrase detection block, and wherein the second trigger phrase detection block operates with different detection criteria from the first trigger phrase detection block.

13. A computer program product, comprising computer readable code embodied in non-transitory computer-readable media, for causing a processing device to perform a method of processing received data representing speech, comprising:

storing the received data;

detecting a presence of data representing a trigger phrase in the received data;

in response to said detecting, supplying a first part of the stored data representing at least a part of the trigger phrase to an adaptive speech enhancement block;

training the speech enhancement block on the first part of the stored data to derive adapted parameters for the speech enhancement block;

supplying a second part of the stored data to the adaptive speech enhancement block operating with said adapted parameters, to form enhanced stored data, wherein the second part of the stored data overlaps with the first part of the stored data;

detecting the presence of data representing the trigger phrase in the enhanced stored data; and

outputting enhanced speech data from the speech enhancement block for further processing, in response to detecting the presence of data representing the trigger phrase in the enhanced stored data;

wherein the detecting the presence of data representing the trigger phrase in the received data is carried out by means of a first trigger phrase detection block; and

wherein the detecting the presence of data representing the trigger phrase in the enhanced stored data is carried out by means of a second trigger phrase detection block, and wherein the second trigger phrase detection block operates with different detection criteria from the first trigger phrase detection block.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2017
From: CIRRUS LOGIC INTERNATIONAL SEMICONDUCTOR LTD.
To: CIRRUS LOGIC INC.
Reel/Frame 044002/0755 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2016
From: HATFIELD, ROBERT JAMES; PAGE, MICHAEL
To: CIRRUS LOGIC INTERNATIONAL SEMICONDUCTOR LTD.
Reel/Frame 039278/0013 →
Priority Claims (1)
GB 1322349.0 · Dec 18, 2013 · national
Continuity (1)
Related Publication 20160322045A1 · Nov 3, 2016