IP Library Granted Patent US 9,520,138
Granted Patent B2
US 9,520,138 · App. 14/210,036 · Granted Dec 13, 2016

Adaptive modulation filtering for spectral feature enhancement

Inventor: Bengt J. Borgstrom (Santa Monica, CA)
Assignee: Broadcom Corporation
G10L21/0208G10L17/00G10L25/06G10L17/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,520,138
App. No.
14/210,036
Granted
Dec 13, 2016
Kind
B2
Abstract

Techniques described herein are directed to the enhancement of spectral features of an audio signal via adaptive modulation filtering. The adaptive modulation filtering process is based on observed modulation envelope autocorrelation coefficients obtained from the audio signal. The modulation envelope autocorrelation coefficients are used to determine parameters of an adaptive filter configured to filter the spectral features of the audio signal to provide filtered spectral features. The parameters are updated based on the observed modulation envelope autocorrelation coefficients to adapt to changing acoustic conditions, such as signal-to-noise ratio (SNR) or reverberation time. Accordingly, such acoustic conditions are not required to be estimated explicitly. Techniques described herein also allow for the estimation of useful side information, e.g., signal-to-noise ratios, based on the observed spectral features of the audio signal and the filtered spectral features, which can be used to improve speaker identification algorithms and/or other audio processing algorithms.

Claims (54)

1. A method for identifying a target speaker, comprising:

obtaining spectral features of an audio signal;

obtaining a signal-to-noise ratio for the audio signal that is based at least on the spectral features;

adapting a speaker model based on the signal-to-noise ratio; and

determining a likelihood that the audio signal is associated with the target speaker based on the adapted speaker model.

2. The method of claim 1 , wherein the signal-to-noise ratio is based on spectral features of the audio signal and a filtered version of the spectral features, and wherein the filtered version of the spectral features is provided by a filter that is based on a model of a modulation spectrum for a clean audio signal and a model of a modulation spectrum of the audio signal.

3. The method of claim 1 , wherein adapting a speaker model based on the signal-to-noise estimate comprises:

adapting a covariance matrix associated with the speaker model based on the signal-to-noise ratio.

4. The method of claim 3 , wherein determining a likelihood that audio signal is associated with the target speaker based on the adapted speaker model comprises:

determining the likelihood based on the adapted covariance matrix and an inverse of the adapted covariance matrix.

5. The method of claim 1 , wherein said obtaining spectral features of the audio signal comprises:

determining first spectral features of one or more first frames in a series of frames that represent the audio signal;

obtaining one or more autocorrelation coefficients of a modulation envelope of the audio signal based on at least the first spectral features of the one or more first frames; and

modifying parameters of an adaptive modulation filter used to filter spectral features of the audio signal based on the one or more autocorrelation coefficients, the filtered spectral features being the obtained spectral features.

6. The method of claim 5 , wherein the adaptive modulation filter comprises one of an inverse filter and a Wiener filter.

7. The method of claim 1 , wherein the spectral features comprise:

linear spectral amplitudes;

warped-frequency spectral amplitudes;

log-spectral amplitudes; or

cepstra.

8. A system for identifying a target speaker, comprising:

filtering logic configured to obtain spectral features of an audio signal;

an SNR estimator configured to obtain a signal-to-noise ratio for the audio signal that is based at least on the spectral features; and

pattern matching logic configured to adapt a speaker model based on the signal-to-noise ratio,

wherein the pattern matching logic is further configured to determine a likelihood that the audio signal is associated with the target speaker based on the adapted speaker model.

9. The system of claim 8 , wherein the signal-to-noise ratio is based on spectral features of the audio signal and a filtered version of the spectral features, and wherein the filtered version of the spectral features is provided by a filter that is based on a model of a modulation spectrum for a clean audio signal and a model of a modulation spectrum of the audio signal.

10. The system of claim 8 , wherein the pattern matching logic is configured to adapt the speaker model based on the signal-to-noise estimate by:

adapting a covariance matrix associated with the speaker model based on the signal-to-noise ratio.

11. The system of claim 10 , wherein the pattern matching logic is configured to determine the likelihood based on the adapted covariance matrix and an inverse of the adapted covariance matrix.

12. The system of claim 8 , wherein the filtering logic is configured to obtain the spectral features of the audio signal by:

determining first spectral features of one or more first frames in a series of frames that represent the audio signal;

obtaining one or more autocorrelation coefficients of a modulation envelope of the audio signal based on at least the first spectral features of the one or more first frames; and

modifying parameters of an adaptive modulation filter used to filter spectral features of the audio signal based on the one or more autocorrelation coefficients, the filtered spectral features being the obtained spectral features.

13. The system of claim 12 , wherein the adaptive modulation filter comprises one of an inverse filter and a Wiener filter.

14. The system of claim 8 , wherein the spectral features comprise:

linear spectral amplitudes;

warped-frequency spectral amplitudes;

log-spectral amplitudes; or

cepstra.

15. A non-transitory computer-readable storage medium having program instructions recorded thereon that, when executed by a processing device, perform a method for identifying a target speaker, the method comprising:

obtaining spectral features of an audio signal;

obtaining a signal-to-noise ratio for the audio signal that is based at least on the spectral features;

adapting a speaker model based on the signal-to-noise ratio; and

determining a likelihood that the audio signal is associated with the target speaker based on the adapted speaker model.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the signal-to-noise ratio is based on spectral features of the audio signal and a filtered version of the spectral features, and wherein the filtered version of the spectral features is provided by a filter that is based on a model of a modulation spectrum for a clean audio signal and a model of a modulation spectrum of the audio signal.

17. The non-transitory computer-readable storage medium of claim 15 , wherein adapting a speaker model based on the signal-to-noise estimate comprises:

adapting a covariance matrix associated with the speaker model based on the signal-to-noise ratio.

18. The non-transitory computer-readable storage medium of claim 17 , wherein determining a likelihood that audio signal is associated with the target speaker based on the adapted speaker model comprises:

determining the likelihood based on the adapted covariance matrix and an inverse of the adapted covariance matrix.

19. The non-transitory computer-readable storage medium of claim 15 , wherein said obtaining spectral features of the audio signal comprises:

determining first spectral features of one or more first frames in a series of frames that represent the audio signal;

obtaining one or more autocorrelation coefficients of a modulation envelope of the audio signal based on at least the first spectral features of the one or more first frames; and

modifying parameters of an adaptive modulation filter used to filter spectral features of the audio signal based on the one or more autocorrelation coefficients, the filtered spectral features being the obtained spectral features.

20. The non-transitory computer-readable storage medium of claim 15 , wherein the adaptive modulation filter comprises one of an inverse filter and a Wiener filter.

Assignments (6)
CORRECTIVE ASSIGNMENT TO CORRECT THE EXECUTION DATE PREVIOUSLY RECORDED AT REEL: 047422 FRAME: 0464. ASSIGNOR(S) HEREBY CONFIRMS THE MERGER. Recorded Mar 6, 2019
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 048883/0702 →
MERGER Recorded Oct 5, 2018
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 047422/0464 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 3, 2017
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: BROADCOM CORPORATION
Reel/Frame 041712/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2017
From: BROADCOM CORPORATION
To: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
Reel/Frame 041706/0001 →
PATENT SECURITY AGREEMENT Recorded Feb 11, 2016
From: BROADCOM CORPORATION
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 037806/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2014
From: BORGSTROM, BENGT J.
To: BROADCOM CORPORATION
Reel/Frame 032750/0577 →
Continuity (3)
Provisional Application 61793324 · Mar 15, 2013
Provisional Application 61945544 · Feb 27, 2014
Related Publication 20140270226A1 · Sep 18, 2014