IP Library Granted Patent US 9,741,341
Granted Patent B2
US 9,741,341 · App. 14/600,503 · Granted Aug 22, 2017

System and method for dynamic noise adaptation for robust automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,741,341
App. No.
14/600,503
Granted
Aug 22, 2017
Kind
B2
Abstract

A speech processing method and arrangement are described. A dynamic noise adaptation (DNA) model characterizes a speech input reflecting effects of background noise. A null noise DNA model characterizes the speech input based on reflecting a null noise mismatch condition. A DNA interaction model performs Bayesian model selection and re-weighting of the DNA model and the null noise DNA model to realize a modified DNA model characterizing the speech input for automatic speech recognition and compensating for noise to a varying degree depending on relative probabilities of the DNA model and the null noise DNA model.

Claims (55)

1. A method comprising:

characterizing, by a computing device, a speech input based on a dynamic noise adaptation (DNA) model reflecting effects of background noise;

characterizing the speech input based on a null noise DNA model reflecting a null noise mismatch condition;

performing Bayesian model averaging using a first weighting of the DNA model and a second weighting of the null noise DNA model;

re-weighting the DNA model and the null noise DNA model by adjusting the first weighting to increase a probability of the DNA model predicted by the Bayesian model averaging when the DNA model is more likely than the null noise DNA model to best characterize the speech input; and

performing recognition of the speech input using the re-weighted DNA model and the re-weighted null noise DNA model.

2. The method of claim 1 , comprising re-weighting the DNA model and the null noise DNA model by adjusting the first weighting to decrease the probability of the DNA model predicted by the Bayesian model averaging when the DNA model is less likely than the null noise DNA model to best characterize the speech input.

3. The method of claim 2 , comprising re-weighting the DNA model and the null noise DNA model by adjusting the first weighting to decrease the probability of the DNA model predicted by the Bayesian model averaging to be zero when the DNA model is less likely than the null noise DNA model to best characterize the speech input.

4. The method of claim 1 , comprising re-weighting the DNA model and the null noise DNA model by adjusting the second weighting to decrease the probability of the null noise DNA model to be zero when the DNA model is more likely than the null noise DNA model to best characterize the speech input.

5. The method of claim 1 , wherein the DNA model comprises a probability-based noise model reflecting transient components and evolving components of a current noise estimate of the background noise.

6. The method of claim 1 , comprising:

modeling a transient component of a noise process of the speech input at each frequency band as zero mean and Gaussian;

modeling a channel distortion of the speech input as a stochastically adapted parameter;

approximating a noise posterior at each given frame of the speech input as Gaussian; and

iteratively estimating a conditional posterior of a level of the background noise and a speech of the speech input for each speech Gaussian.

7. A system comprising:

at least one processor; and

one or more non-transitory computer-readable media storing executable instructions that, when executed by the at least one processor, cause the system to:

characterize a speech input based on a dynamic noise adaptation (DNA) model reflecting effects of background noise;

characterize the speech input based on a null noise DNA model reflecting a null noise mismatch condition;

perform Bayesian model averaging using a first weighting of the DNA model and a second weighting of the null noise DNA model;

re-weight the DNA model and the null noise DNA model by adjusting the first weighting to increase a probability of the DNA model predicted by the Bayesian model averaging when the DNA model is more likely than the null noise DNA model to best characterize the speech input; and

perform recognition of the speech input using the re-weighted DNA model and the re-weighted null noise DNA model.

8. The system of claim 7 , wherein the one or more non-transitory computer-readable media store executable instructions that, when executed by the at least one processor, cause the system to:

re-weight the DNA model and the null noise DNA model by adjusting the first weighting to decrease the probability of the DNA model predicted by the Bayesian model averaging when the DNA model is less likely than the null noise DNA model to best characterize the speech input.

9. The system of claim 8 , wherein the one or more non-transitory computer-readable media store executable instructions that, when executed by the at least one processor, cause the system to:

re-weight the DNA model and the null noise DNA model by adjusting the first weighting to decrease the probability of the DNA model predicted by the Bayesian model averaging to be zero when the DNA model is less likely than the null noise DNA model to best characterize the speech input.

10. The system of claim 7 , wherein the one or more non-transitory computer-readable media store executable instructions that, when executed by the at least one processor, cause the system to:

re-weight the DNA model and the null noise DNA model by adjusting the second weighting of the null noise DNA model to be zero when the DNA model is more likely than the null noise DNA model to best characterize the speech input.

11. The system of claim 7 , wherein the DNA model comprises a probability-based noise model reflecting transient components and evolving components of a current noise estimate of the background noise.

12. The system of claim 7 , wherein the DNA model includes a speech model, a noise model, a channel model, and an interaction model that describes how the speech model, the noise model, and the channel model combine to generate the speech input.

13. The system of claim 7 , wherein the one or more non-transitory computer-readable media store executable instructions that, when executed by the at least one processor, cause the system to:

model a transient component of a noise process of the speech input at each frequency band as zero mean and Gaussian;

model a channel distortion of the speech input as a stochastically adapted parameter;

approximate a noise posterior at each given frame of the speech input as Gaussian; and

iteratively estimate a conditional posterior of a level of the background noise and a speech of the speech input for each speech Gaussian.

14. One or more non-transitory computer-readable media storing executable instructions that, when executed by a processor, cause a device to:

characterize a speech input based on a dynamic noise adaptation (DNA) model reflecting effects of background noise;

characterize the speech input based on a null noise DNA model reflecting a null noise mismatch condition;

perform Bayesian model averaging using a first weighting of the DNA model and a second weighting of the null noise DNA model;

re-weight the DNA model and the null noise DNA model by adjusting the first weighting to increase a probability of the DNA model predicted by the Bayesian model averaging when the DNA model is more likely than the null noise DNA model to best characterize the speech input; and

perform recognition of the speech input using the re-weighted DNA model and the re-weighted null noise DNA model.

15. The one or more non-transitory computer-readable media of claim 14 , wherein the executable instructions, when executed by the processor, cause the device to:

re-weight the DNA model and the null noise DNA model by adjusting the first weighting to decrease the probability of the DNA model predicted by the Bayesian model averaging when the DNA model is less likely than the null noise DNA model to best characterize the speech input.

16. The one or more non-transitory computer-readable media of claim 15 , wherein the executable instructions, when executed by the processor, cause the device to:

re-weight the DNA model and the null noise DNA model by adjusting the first weighting to decrease the probability of the DNA model predicted by the Bayesian model averaging to be zero when the DNA model is less likely than the null noise DNA model to best characterize the speech input.

17. The one or more non-transitory computer-readable media of claim 14 , wherein the executable instructions, when executed by the processor, cause the device to:

re-weight the DNA model and the null noise DNA model by adjusting the second weighting to decrease the probability of the null noise DNA model to be zero when the DNA model is more likely than the null noise DNA model to best characterize the speech input.

18. The method of claim 1 , wherein the DNA model includes a speech model, a noise model, a channel model, and an interaction model that describes how the speech model, the noise model, and the channel model combine to generate the speech input.

19. The one or more non-transitory computer-readable media of claim 14 , wherein the DNA model includes a speech model, a noise model, a channel model, and an interaction model that describes how the speech model, the noise model, and the channel model combine to generate the speech input.

20. The one or more non-transitory computer-readable media of claim 14 , wherein the executable instructions, when executed by the processor, cause the device to:

model a transient component of a noise process of the speech input at each frequency band as zero mean and Gaussian;

model a channel distortion of the speech input as a stochastically adapted parameter;

approximate a noise posterior at each given frame of the speech input as Gaussian; and

iteratively estimate a conditional posterior of a level of the background noise and a speech of the speech input for each speech Gaussian.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2015
From: RENNIE, STEVEN J.; DOGNIN, PIERRE; FOUSEK, PETR
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 035350/0551 →