IP Library Granted Patent US 11,308,946
Granted Patent B2
US 11,308,946 · App. 16/521,641 · Granted Apr 19, 2022

Methods and apparatus for ASR with embedded noise reduction

Inventors: Jianzhong Teng (Shanghai, CN); Xiao-Lin Ren (Shanghai, CN); Xingui Zeng (Shanghai, CN); Yi Gao (Shanghai, CN)
Assignee: Cerence Operating Company
G10L15/20G10L15/02G10L15/08G10L15/22G10L21/0232G10L25/24G10L25/84G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,308,946
App. No.
16/521,641
Granted
Apr 19, 2022
Kind
B2
Abstract

Methods and an apparatus for performing feature extraction on speech in a microphone signal with embedded noise processing to reduce the amount of processing are provided. In embodiments, feature extraction and the noise estimate use an output of the same Fourier Transform, such that the noise filtering of the speech is embedded with the feature extraction of the speech.

Claims (46)

1. A method, comprising:

receiving a microphone signal;

determining whether a noise level of the microphone signal exceeds a noise threshold;

selecting between a first signal processing procedure and a second signal processing procedure, different from the first signal processing procedure, based on whether the noise level exceeds the threshold;

performing the first signal processing procedure when the noise level of the microphone signal exceeds the noise threshold, the first signal processing procedure including:

determining whether the microphone signal contains speech;

determining a noise estimate for the microphone signal when speech is found not to be present;

performing noise filtering using the noise estimate on the microphone signal when speech is found to be present; and

performing feature extraction on the microphone signal when speech is found to be present, such that the noise filtering of the speech is embedded with the feature extraction of the speech; and

performing the second signal processing procedure when the noise level of the microphone signal does not exceed the noise threshold, the second signal processing procedure including performing feature extraction on the microphone signal.

2. The method according to claim 1 , wherein the noise estimate and the noise filtering are not performed on a same frame of the microphone signal.

3. The method according to claim 1 , further including determining whether the microphone signal contains a wake-up phrase after the noise filtering.

4. The method according to claim 1 , wherein the feature extraction is performed while a device containing the microphone is in a sleep state.

5. The method according to claim 1 , wherein the feature extraction includes the use of mel-frequency cepstral coefficients (MFCCs).

6. The method according to claim 1 , further including using a main processor and a lower power processor to provide processing of a wakeup phrase for a device.

7. The method according to claim 6 , wherein the lower power processor performs the feature extraction and noise processing to identify the wakeup phrase.

8. A system comprising:

a processor and memory configured to:

determine whether a noise level of a microphone signal exceeds a noise threshold;

select between a first signal processing procedure and a second signal processing procedure, different from the first signal processing procedure, based on whether the noise level exceeds the threshold;

perform the first signal processing procedure when the noise level of the microphone signal exceeds the noise threshold, the first signal processing procedure including:

determining whether the microphone signal contains speech;

determining a noise estimate for the microphone signal when speech is found not to be present;

performing noise filtering using the noise estimate on the microphone signal when speech is found to be present; and

performing feature extraction on the microphone signal when speech is found to be present, such that the noise filtering of the speech is embedded with the feature extraction of the speech; and

perform the second signal processing procedure when the noise level of the microphone signal does not exceed the noise threshold, the second signal processing procedure including performing feature extraction on the microphone signal.

9. The system according to claim 8 , wherein the noise estimate and the noise filtering are not performed on a same frame of the microphone signal.

10. The system according to claim 8 , further including determining whether the microphone signal contains a wake-up phrase after the noise filtering.

11. The system according to claim 8 , wherein the feature extraction is performed while a device containing the microphone is in a sleep state.

12. The system according to claim 8 , wherein the feature extraction includes the use of mel-frequency cepstral coefficients (MFCCs).

13. The system according to claim 8 , further including using a main processor and a lower power processor to provide processing of a wakeup phrase for a device.

14. The system according to claim 13 , wherein the lower power processor performs the feature extraction and noise processing to identify the wakeup phrase.

15. An article, comprising: a non-transitory computer readable medium having stored instructions that enable a machine to:

determine whether a microphone signal has a noise level above a noise threshold;

select between a first signal processing procedure and a second signal processing procedure, different from the first signal processing procedure, based on whether the noise level is above the threshold;

perform the first signal processing procedure when the noise level of the microphone signal is above the noise threshold, the first signal processing procedure including:

determining whether the microphone signal contains speech;

determining a noise estimate for the microphone signal when speech is found not to be present; perform noise filtering using the noise estimate on the microphone signal when speech is found to be present; and

performing feature extraction on the microphone signal when speech is found to be present, such that the noise filtering of the speech is embedded with the feature extraction of the speech; and

perform a second signal processing procedure when the noise level of the microphone signal is not above the noise threshold, the second signal processing procedure including performing feature extraction on the microphone signal.

16. The article according to claim 15 , wherein the noise estimate and the noise filtering are not performed on a same frame of the microphone signal.

17. The article according to claim 15 , further including instructions for determining whether the microphone signal contains a wake-up phrase after the noise filtering.

18. The article according to claim 15 , wherein the feature extraction is performed while a device containing the microphone is in a sleep state.

19. The article according to claim 15 , wherein the feature extraction includes the use of mel-frequency cepstral coefficients (MFCCs).

20. The article according to claim 15 , further including using a main processor and a lower power processor to provide processing of a wakeup phrase for a device.

21. The method of claim 6 , wherein the lower power processor performs denoising and feature extraction for an automatic speech recognition operation performed by the main processor.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2019
From: TENG, JIANZHONG; REN, XIAO-LIN; ZENG, XINGUI; GAO, YI
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 049856/0391 →