IP Library Granted Patent US 7,945,006
Granted Patent B2
US 7,945,006 · App. 10/875,553 · Granted May 17, 2011

Data-driven method and apparatus for real-time mixing of multichannel signals in a media server

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,945,006
App. No.
10/875,553
Granted
May 17, 2011
Kind
B2
Abstract

An apparatus for mixing audio signals in a voice-over-IP teleconferencing environment comprises a preprocessor, a mixing controller, and a mixing processor. The preprocessor is divided into a media parameter estimator and a media preprocessor. The media parameter estimator estimates signal parameters such as signal-to-noise ratios, energy levels, and voice activity (i.e., the presence or absence of voice in the signal), which are used to control how different channels are mixed. The media preprocessor employs signal processing algorithms such as silence suppression, automatic gain control, and noise reduction, so that the quality of the incoming voice streams is optimized. Based on a function of the estimated signal parameters, the mixing controller specifies a particular mixing strategy and the mixing processor mixes the preprocessed voice streams according the strategy provided by the controller.

Claims (34)

1. A method for generating a mixed audio channel signal from a plurality of incoming audio channel signals, the method comprising the steps of:

determining a corresponding signal-to-noise ratio estimate for each of said plurality of incoming audio channel signals;

selecting a proper subset of said plurality of incoming audio channel signals based on said corresponding plurality of signal-to-noise ratio estimates; and

generating the mixed audio channel signal by combining only said selected proper subset of said incoming audio channel signals,

wherein said proper subset of said plurality of incoming audio channel signals is selected by choosing a plural number of said plurality of incoming audio channel signals having higher signal-to-noise ratio estimates than other ones of said plurality of incoming audio channel signals.

2. The method of claim 1 wherein said incoming audio channel signals comprise voice-over-IP packet-based voice signals.

3. The method of claim 1 wherein said signal-to-noise ratio estimates comprise both short-term signal-to-noise estimates and long-term signal-to-noise estimates.

4. The method of claim 1 further comprising the step of determining a corresponding power estimate for each of said plurality of incoming audio channel signals, and wherein said step of selecting said proper subset of said plurality of incoming audio channel signals is further based on said corresponding plurality of power estimates.

5. The method of claim 4 wherein said power estimates comprise both instantaneous energy estimates and long-term average energy estimates.

6. The method of claim 1 further comprising the step of applying voice activity detection to each of said plurality of incoming audio channel signals to determine whether each of said plurality of incoming audio channel signals comprises speech, and wherein said step of selecting said proper subset of said plurality of incoming audio channel signals is further based on said determination of whether each of said plurality of incoming audio channel signals comprises speech.

7. The method of claim 1 wherein said step of selecting the proper subset of said plurality of incoming audio channel signals is further based on a priori predetermined knowledge.

8. The method of claim 7 wherein said a priori predetermined knowledge comprises an indication that one or more of said plurality of incoming audio channel signals is to always be included in said selected proper subset of said plurality of incoming audio channel signals.

9. The method of claim 1 further comprising the step of preprocessing each of said plurality of incoming audio channel signals to produce corresponding preprocessed incoming audio channel signals, and wherein said step of generating said mixed audio channel signal comprises combining said preprocessed incoming audio channel signals corresponding to only said selected proper subset of said incoming audio channel signals.

10. The method of claim 9 wherein said step of preprocessing each of said plurality of incoming audio channel signals comprises performing automatic gain control on each of said plurality of incoming audio channel signals such that each of said preprocessed incoming audio channel signals has a similar volume level.

11. The method of claim 9 wherein said step of preprocessing each of said plurality of incoming audio channel signals comprises performing silence suppression on each of said plurality of incoming audio channel signals to attenuate noise levels in an absence of speech.

12. The method of claim 9 wherein said step of preprocessing each of said plurality of incoming audio channel signals comprises performing noise reduction on each of said plurality of incoming audio channel signals.

13. The method of claim 9 wherein said step of preprocessing each of said plurality of incoming audio channel signals comprises performing speech enhancement on each of said plurality of incoming audio channel signals.

14. An apparatus for generating a mixed audio channel signal from a plurality of incoming audio Channel signals, the apparatus comprising:

a plurality of signal-to-noise ratio estimators which determine a corresponding signal-to-noise ratio estimate for each of said plurality of incoming audio channel signals;

a mixing controller which selects a proper subset of said plurality of incoming audio channel signals based on said corresponding plurality of signal-to-noise ratio estimates;

a mixing processor which generates the mixed audio channel signal by combining only said selected proper subset of said incoming audio channel signals,

wherein said proper subset of said plurality of incoming audio channel signals is selected by choosing a plural number of said plurality of incoming audio channel signals having higher signal-to-noise ratio estimates than other ones of said plurality of incoming audio channel signals.

15. The apparatus of claim 14 wherein said incoming audio channel signals comprise voice-over-IP packet-based voice signals.

16. The apparatus of claim 14 wherein said signal-to-noise ratio estimates comprise both short-term signal-to-noise estimates and long-term signal-to-noise estimates.

17. The apparatus of claim 14 further comprising a plurality of power estimators which determine a corresponding power estimate for each of said plurality of incoming audio channel signals, and wherein said mixing controller selects said proper subset of said plurality of incoming audio channel signals further based oh said corresponding plurality of power estimates.

18. The apparatus of claim 17 wherein said power estimates comprise both instantaneous energy estimates and long-term average energy estimates.

19. The apparatus of claim 14 further comprising a plurality of voice activity detectors which are applied to each of said plurality of incoming audio channel signals to determine whether each of said plurality of incoming audio channel signals comprises speech, and wherein said mixing controller selects said proper subset of said plurality of incoming audio channel signals further based on said determination of whether each of said plurality of incoming audio channel signals comprises speech.

20. The apparatus of claim 14 further comprising a plurality of signal preprocessors applied to said plurality of incoming audio channel signals to produce corresponding preprocessed incoming audio channel signals, and wherein said mixing processor combines said preprocessed incoming audio channel signals corresponding to only said selected proper subset of said incoming audio channel signals.

21. The apparatus of claim 20 wherein said signal preprocessors comprise means for applying automatic gain control to said plurality of incoming audio channel signals such that each of said preprocessed incoming audio channel signals has a similar volume level.

22. The apparatus of claim 20 wherein said signal preprocessors comprise means for performing silence suppression on said plurality of incoming audio channel signals to attenuate noise levels in an absence of speech.

23. The apparatus of claim 20 wherein said signal preprocessors comprise means for applying noise reduction on said plurality of incoming audio channel signals.

24. The apparatus of claim 20 wherein said signal preprocessors comprise means for performing speech enhancement on said plurality of incoming audio channel signals.

25. The apparatus of claim 14 wherein said mixing controller selects the proper subset of said plurality of incoming audio channel signals further based on a priori predetermined knowledge.

26. The apparatus of claim 25 wherein said a priori predetermined knowledge comprises an indication that one or more of said plurality of incoming audio channel signals is to always be included in said selected proper subset of said plurality of incoming audio channel signals.

Assignments (7)
SECURITY INTEREST Recorded Jun 1, 2021
From: WSOU INVESTMENTS, LLC
To: OT WSOU TERRIER HOLDINGS, LLC
Reel/Frame 056990/0081 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 10, 2020
From: NOKIA OF AMERICA CORPORATION
To: WSOU INVESTMENTS, LLC
Reel/Frame 052372/0577 →
CHANGE OF NAME Recorded Nov 20, 2019
From: ALCATEL-LUCENT USA INC.
To: NOKIA OF AMERICA CORPORATION
Reel/Frame 051061/0753 →
RELEASE OF SECURITY INTEREST Recorded Oct 9, 2014
From: CREDIT SUISSE AG
To: ALCATEL-LUCENT USA INC.
Reel/Frame 033949/0531 →
SECURITY INTEREST Recorded Mar 7, 2013
From: ALCATEL-LUCENT USA INC.
To: CREDIT SUISSE AG
Reel/Frame 030510/0627 →
MERGER Recorded Mar 17, 2011
From: LUCENT TECHNOLOGIES INC.
To: ALCATEL-LUCENT USA INC.
Reel/Frame 025975/0172 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 2, 2004
From: CHEN, JINGDONG; HUANG, YITENG ARDEN; WOO, THOMAS Y.
To: LUCENT TECHNOLOGIES INC.
Reel/Frame 015752/0437 →