IP Library Granted Patent US 8,645,129
Granted Patent B2
US 8,645,129 · App. 12/464,624 · Granted Feb 4, 2014

Integrated speech intelligibility enhancement system and acoustic echo canceller

Inventors: Wilfrid LeBlanc (Vancouver, CA); Jes Thyssen (Laguna Niguel, CA); Juin-Hwey Chen (Irvine, CA)
Assignee: Broadcom Corporation
G10L21/0208
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,645,129
App. No.
12/464,624
Filed
May 12, 2009
Granted
Feb 4, 2014
Kind
B2
Art Unit
2658
USPC
704/233
Abstract

A system and method is described that improves the intelligibility of a far-end telephone speech signal to a user of a telephony device in the presence of near-end background noise. As described herein, the system and method improves the intelligibility of the far-end telephone speech signal in a manner that does not require user input and that minimizes the distortion of the far-end telephone speech signal. The system is integrated with an acoustic echo canceller and shares information therewith.

Claims (81)

1. A system, comprising:

estimation logic configured to calculate characteristics associated with at least a near-end speech signal to be transmitted by an audio device, the calculated characteristics including an estimated level of near-end background noise that is associated with the near-end speech signal, the estimated level including a plurality of estimations of the near-end background noise corresponding to a plurality of respective sub-band components of the near-end speech signal;

a processing module configured to receive the calculated characteristics and to modify a far-end speech signal, which is received for playback by the audio device, based on at least the calculated characteristics to increase the intelligibility thereof by boosting the far-end speech signal over the near-end background noise by performing spectral shaping on the far-end speech signal based on one or more of the estimations corresponding to one or more of the sub-band components; and

an acoustic echo canceller configured to receive the calculated characteristics and to suppress acoustic echo present in the near-end speech signal based on at least the calculated characteristics.

2. The system of claim 1 , wherein the estimated level of the near-end background noise that is associated with the near-end speech signal comprises a measure of loudness obtained by applying a weight to one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal.

3. The system of claim 1 , wherein the calculated characteristics comprise a determination of whether voice activity is present in the far-end speech signal; and

wherein the processing module is configured to control the operation of a level estimator based on the determination, the level estimator being configured to calculate an estimated signal level associated with the far-end speech signal, and to apply a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

4. The system of claim 3 , wherein the estimation logic is configured to determine whether voice activity is present in the far-end speech signal by analyzing one or more sub-band components of the far-end speech signal.

5. The system of claim 1 , wherein the calculated characteristics comprise a determination of whether voice activity is present in the near-end speech signal; and

wherein the processing module is configured to control the operation of a level estimator based on the determination, the level estimator being configured to calculate an estimated signal level associated with the far-end speech signal, and to apply a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

6. The system of claim 5 , wherein the estimation logic is configured to determine whether voice activity is present in the near-end speech signal by analyzing the one or more sub-band components of the near-end speech signal.

7. The system of claim 1 , further comprising:

a plurality of microphones; and

a beamformer connected to the plurality of microphones, the beamformer being configured to perform spatial filtering on signals received from the plurality of microphones to generate the near-end speech signal;

wherein the estimation logic is configured to calculate the estimated level of the near-end background noise that is associated with the near-end speech signal by calculating an estimated level of the near-end background noise one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal at one or more of the microphones in the plurality of microphones.

8. The system of claim 7 , wherein the estimation logic is configured to calculate the estimated level of the near-end background noise at one or more of the microphones in the plurality of microphones by modifying one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal to account for a noise changing effect produced by the beamformer.

9. A method, comprising:

calculating characteristics associated with at least a near-end speech signal to be transmitted by an audio device, said calculating comprising:

calculating an estimated level of near-end background noise that is associated with the near-end speech signal, the estimated level including a plurality of estimations of the near-end background noise corresponding to a plurality of respective sub-band components of the near-end speech signal;

modifying a far-end speech signal, which is received for playback by the audio device, based on at least the calculated characteristics to increase the intelligibility thereof by boosting the far-end speech signal over the near-end background noise, said modifying comprising:

performing spectral shaping on the far-end speech signal based on one or more of the estimations of the near-end background noise corresponding to one or more of the sub-band components; and

suppressing acoustic echo present in the near-end speech signal based on at least the calculated characteristics.

10. The method of claim 9 , wherein calculating the estimated level of the near-end background noise that is associated with the near-end speech signal comprises calculating a measure of loudness by applying a weight to one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal.

11. The method of claim 9 , wherein calculating characteristics associated with at least the near-end speech signal comprises determining whether voice activity is present in the far-end speech signal; and

wherein modifying the far-end speech signal based on at least the calculated characteristics comprises controlling the operation of a level estimator based on the determination, wherein the level estimator calculates an estimated signal level associated with the far-end speech signal, and applying a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

12. The method of claim 11 , wherein determining whether voice activity is present in the far-end speech signal comprises analyzing one or more sub-band components of the far-end speech signal.

13. The method of claim 9 , wherein calculating characteristics associated with at least the near-end speech signal comprises determining whether voice activity is present in the near-end speech signal; and

wherein modifying the far-end speech signal based on at least the calculated characteristics comprises controlling the operation of a level estimator based on the determination, wherein the level estimator calculates an estimated signal level associated with the far-end speech signal, and applying a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

14. The method of claim 13 , wherein determining whether voice activity is present in the near-end speech signal comprises analyzing the one or more sub-band components of the near-end speech signal.

15. The method of claim 9 , wherein calculating characteristics associated with at least the near-end speech signal comprises:

calculating the estimated level of the near-end background noise by calculating one or more estimations of the plurality of estimations of the near-end background noise at one or more microphones in a plurality of microphones associated with the audio device.

16. The method of claim 15 , wherein calculating the estimated level of the near-end background noise at one or more microphones in the plurality of microphones associated with the audio device comprises:

modifying one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal to account for a noise changing effect produced by a beamformer coupled to the plurality of microphones.

17. A system, comprising:

estimation logic configured to calculate characteristics associated with at least a far-end speech signal received for playback by an audio device, the calculated characteristics including an estimated level of near-end background noise that is associated with a near-end speech signal, the estimated level including a plurality of estimations of the near-end background noise corresponding to a plurality of respective sub-band components of the near-end speech signal;

a processing module configured to receive the calculated characteristics and to modify the far-end speech signal based on at least the calculated characteristics to increase the intelligibility thereof by boosting the far-end speech signal over the near-end background noise by performing spectral shaping on the far-end speech signal based on one or more of the estimations of the near-end background noise corresponding to one or more of the sub-band components; and

an acoustic echo canceller configured to receive the calculated characteristics and to suppress acoustic echo present in the near-end speech signal, which is to be transmitted by the audio device, based on at least the calculated characteristics.

18. The system of claim 17 , wherein the estimated level of the near-end background noise that is associated with the near-end speech signal comprises a measure of loudness obtained by applying a weight to one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal.

19. The system of claim 17 , wherein the calculated characteristics comprise a determination of whether voice activity is present in the near-end speech signal; and

wherein the processing module is configured to control the operation of a level estimator based on the determination, the level estimator being configured to calculate an estimated signal level associated with the far-end speech signal, and to apply a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

20. The system of claim 19 , wherein the estimation logic is configured to determine whether voice activity is present in the near-end speech signal by analyzing the one or more sub-band components of the near-end speech signal.

21. The system of claim 17 , further comprising:

a plurality of microphones; and

a beamformer connected to the plurality of microphones, the beamformer being configured to perform spatial filtering on signals received from the plurality of microphones to generate the near-end speech signal;

wherein the estimation logic is configured to calculate the estimated level of the near-end background noise that is associated with the near-end speech signal by calculating one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal at one or more of the microphones in the plurality of microphones.

22. The system of claim 21 , wherein the estimation logic is configured to calculate the estimated level of the near-end background noise at one or more of the microphones in the plurality of microphones by modifying one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal to account for a noise changing effect produced by the beamformer.

23. A method, comprising:

calculating characteristics associated with at least a far-end speech signal received for playback by an audio device, said calculating comprising:

calculating an estimated level of near-end background noise that is associated with a near-end speech signal, the estimated level including a plurality of estimations of the near-end background noise corresponding to a plurality of respective sub-band components of the near-end speech signal;

modifying the far-end speech signal based on at least the calculated characteristics to increase the intelligibility thereof by boosting the far-end speech signal over the near-end background noise, said modifying comprising:

performing spectral shaping on the far-end speech signal based on one or more of the estimations of the near-end background noise corresponding to one or more of the sub-band components; and

suppressing acoustic echo present in the near-end speech signal based on at least the calculated characteristics.

24. The method of claim 23 , wherein calculating characteristics associated with at least the far-end speech signal comprises determining whether voice activity is present in the far-end speech signal; and

wherein modifying the far-end speech signal based on at least the calculated characteristics comprises controlling the operation of a level estimator based on the determination, wherein the level estimator calculates an estimated signal level associated with the far-end speech signal, and applying a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

25. The method of claim 24 , wherein determining whether voice activity is present in the far-end speech signal comprises analyzing one or more sub-band components of the far-end speech signal.

26. The method of claim 23 , wherein calculating characteristics associated with at least the far-end speech signal comprises determining whether voice activity is present in the near-end speech signal; and

wherein modifying the far-end speech signal based on at least the calculated characteristics comprises controlling the operation of a level estimator based on the determination, wherein the level estimator calculates an estimated signal level associated with the far-end speech signal, and applying a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

27. The method of claim 26 , wherein determining whether voice activity is present in the near-end speech signal comprises analyzing the one or more sub-band components of the near-end speech signal.

28. The method of claim 23 , wherein calculating characteristics associated with at least the far-end speech signal comprises:

calculating the estimated level of the near-end background noise by calculating one or more estimations of the plurality of estimations of the near-end background noise at one or more microphones in a plurality of microphones associated with the audio device.

29. The method of claim 28 , wherein calculating the estimated level of the near-end background noise at one or more microphones in the plurality of microphones associated with the audio device comprises:

modifying one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal to account for a noise changing effect produced by a beamformer coupled to the plurality of microphones.

30. The method of claim 23 , wherein calculating the estimated level of the near-end background noise that is associated with the near-end speech signal comprises calculating a measure of loudness by applying a weight to one or more estimations of the plurality of estimations of the near-end background noise corresponding to one or more respective sub-band components of the plurality of sub-band components of the near-end speech signal.

31. The method of claim 23 , wherein calculating characteristics associated with at least the near-end speech signal comprises determining whether voice activity is present in the near-end speech signal; and

wherein modifying the far-end speech signal based on at least the calculated characteristics comprises controlling the operation of a level estimator based on the determination, wherein the level estimator calculates an estimated signal level associated with the far-end speech signal, and applying a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

32. A system, comprising:

estimation logic configured to calculate characteristics associated with at least one of a near-end speech signal to be transmitted by an audio device or a far-end speech signal received for playback by the audio device, the calculated characteristics including an estimated level of near-end background noise that is associated with the near-end speech signal, the estimated level including a plurality of estimations of the near-end background noise corresponding to a plurality of respective sub-band components of the near-end speech signal;

a processing module configured to receive the calculated characteristics and to modify the far-end speech signal based on at least the calculated characteristics to increase the intelligibility thereof by applying at least one of automatic volume boosting, amplitude compression, dispersion filtering or spectral shaping to the far-end speech signal based on one or more of the estimations corresponding to one or more of the sub-band components; and

an acoustic echo canceller configured to suppress acoustic echo present in the near-end speech signal based on at least the calculated characteristics.

33. The system of claim 32 , wherein the calculated characteristics include a measure of loudness of the near-end background noise that is associated with the near-end speech signal, the measure of loudness obtained by applying a weight to one or more estimated levels of the near-end background noise corresponding to one or more sub-band components of the near-end speech signal.

34. The system of claim 32 , wherein the calculated characteristics comprise a determination of whether voice activity is present in the far-end speech signal; and

wherein the processing module is configured to control the operation of a level estimator based on the determination, the level estimator being configured to calculate an estimated signal level associated with the far-end speech signal, and to apply a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

35. The system of claim 34 , wherein the estimation logic is configured to determine whether voice activity is present in the far-end speech signal by analyzing one or more sub-band components of the far-end speech signal.

36. The system of claim 32 , wherein the calculated characteristics comprise a determination of whether voice activity is present in the near-end speech signal; and

wherein the processing module is configured to control the operation of a level estimator based on the determination, the level estimator being configured to calculate an estimated signal level associated with the far-end speech signal, and to apply a gain to the far-end speech signal wherein the amount of gain applied is based on the estimated signal level.

37. The system of claim 36 , wherein the estimation logic is configured to determine whether voice activity is present in the near-end speech signal by analyzing one or more sub-band components of the near-end speech signal.

38. The system of claim 32 , further comprising:

a plurality of microphones; and

a beamformer connected to the plurality of microphones, the beamformer being configured to perform spatial filtering on signals received from the plurality of microphones to generate the near-end speech signal;

wherein the calculated characteristics include an estimated level of the near-end background noise that is associated with the near-end speech signal and wherein the estimation logic is configured to calculate the estimated level of the near-end background noise that is associated with the near-end speech signal by calculating an estimated level of the near-end background noise at one or more of the microphones in the plurality of microphones.

39. The system of claim 38 , wherein the estimation logic is configured to calculate the estimated level of the near-end background noise at one or more of the microphones in the plurality of microphones by modifying the estimated level of the near-end background noise that is associated with the near-end speech signal to account for a noise changing effect produced by the beamformer.

Assignments (7)
CORRECTIVE ASSIGNMENT TO CORRECT THE ERROR IN RECORDING THE MERGER IN THE INCORRECT US PATENT NO. 8,876,094 PREVIOUSLY RECORDED ON REEL 047351 FRAME 0384. ASSIGNOR(S) HEREBY CONFIRMS THE MERGER. Recorded Mar 8, 2019
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 049248/0558 →
CORRECTIVE ASSIGNMENT TO CORRECT THE EFFECTIVE DATE OF THE MERGER PREVIOUSLY RECORDED AT REEL: 047230 FRAME: 0910. ASSIGNOR(S) HEREBY CONFIRMS THE MERGER. Recorded Oct 29, 2018
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 047351/0384 →
MERGER Recorded Oct 4, 2018
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 047230/0910 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 3, 2017
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: BROADCOM CORPORATION
Reel/Frame 041712/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2017
From: BROADCOM CORPORATION
To: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
Reel/Frame 041706/0001 →
PATENT SECURITY AGREEMENT Recorded Feb 11, 2016
From: BROADCOM CORPORATION
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 037806/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2009
From: LEBLANC, WILFRID; THYSSEN, JES; CHEN, JUIN-HWEY
To: BROADCOM CORPORATION
Reel/Frame 022673/0422 →
Continuity (2)
Provisional Application 61052553 · May 12, 2008
Related Publication 20090281805A1 · Nov 12, 2009