IP Library Granted Patent US 10,375,131
Granted Patent B2
US 10,375,131 · App. 15/599,677 · Granted Aug 6, 2019

Selectively transforming audio streams based on audio energy estimate

Inventor: Marcello Caramma (Binfield, GB)
Assignee: Cisco Technology, Inc.
H04L65/403G10L19/167H04L65/605H04W88/02G10L19/012G10L19/0212H04N21/4394
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,375,131
App. No.
15/599,677
Granted
Aug 6, 2019
Kind
B2
Abstract

A server receives, from each of a plurality of participant devices in a communication session, a respective one of a plurality of audio streams. The server estimates an audio energy of each of the plurality of audio streams and determines whether to perform a transform on at least one of the plurality of audio streams. If so, the server performs the transform on the at least one of the plurality of audio streams and transmits the at least one of the plurality of audio streams to at least one of the plurality of participant devices.

Claims (71)

1. A method comprising:

receiving, from each of a plurality of participant devices in a communication session, a respective one of a plurality of audio streams;

estimating an audio energy of each of the plurality of audio streams;

based on the estimated audio energy of each of the plurality of audio streams, determining whether to perform a transform on at least one of the plurality of audio streams, the determining including determining whether performing the transform would require a processing load that is less than a processing load threshold; and

if it is determined to perform the transform on the at least one of the plurality of audio streams:

performing the transform on the at least one of the plurality of audio streams; and

transmitting the at least one of the plurality of audio streams to at least one of the plurality of participant devices.

2. The method of claim 1 , wherein the audio streams are modified discrete cosine transformed frequency domain audio streams, the method further comprising:

partially decoding the plurality of audio streams so as to determine modified discrete cosine transform frequency coefficients corresponding to each of the plurality of audio streams, wherein:

estimating includes estimating the audio energy of each of the plurality of audio streams based on the coefficients;

determining includes determining whether to perform an inverse modified discrete cosine transform on the at least one of the plurality of audio streams; and

performing includes performing the inverse modified discrete cosine transform on the at least one of the plurality of audio streams.

3. The method of claim 2 , wherein determining whether to perform the inverse modified discrete cosine transform on the at least one of the plurality of audio streams further includes:

determining whether the at least one of the plurality of audio streams among the plurality of audio streams has a highest estimated audio energy.

4. The method of claim 2 , wherein determining whether to perform the inverse modified discrete cosine transform on the at least one of the plurality of audio streams further includes:

determining whether the at least one of the plurality of audio streams has an estimated audio energy that exceeds an audio energy threshold.

5. The method of claim 1 , wherein:

estimating includes estimating the audio energy of each of the plurality of audio streams in the time domain;

determining includes determining whether to perform a modified discrete cosine transform on the at least one of the plurality of audio streams; and

performing includes performing the modified discrete cosine transform on the at least one of the plurality of audio streams.

6. The method of claim 5 , further comprising:

if it is determined not to perform the transform on an audio stream of the plurality of audio streams, replacing the audio stream with comfort noise.

7. The method of claim 1 , wherein the audio streams are modified discrete cosine transformed frequency domain audio streams, and wherein:

performing includes performing the inverse modified discrete cosine transform on the at least one of the plurality of audio streams.

8. An apparatus comprising:

a network interface configured to enable communications over a network in order to receive, from each of a plurality of participant devices in a communication session, a respective one of a plurality of audio streams; and

one or more processors coupled to the network interface, wherein the one or more processors are configured to:

estimate an audio energy of each of the plurality of audio streams;

based on the estimated audio energy of each of the plurality of audio streams, determine whether to perform a transform on at least one of the plurality of audio streams, including determining whether performing the transform would require a processing load that is less than a processing load threshold; and

if it is determined to perform the transform on the at least one of the plurality of audio streams:

perform the transform on the at least one of the plurality of audio streams; and

transmit the at least one of the plurality of audio streams to at least one of the plurality of participant devices.

9. The apparatus of claim 8 , wherein the audio streams are modified discrete cosine transformed frequency domain audio streams, the one or more processors further configured to:

partially decode the plurality of audio streams so as to determine modified discrete cosine transform frequency coefficients corresponding to each of the plurality of audio streams, wherein the one or more processors are configured to:

estimate by estimating the audio energy of each of the plurality of audio streams based on the coefficients;

determine by determining whether to perform an inverse modified discrete cosine transform on the at least one of the plurality of audio streams; and

perform by performing the inverse modified discrete cosine transform on the at least one of the plurality of audio streams.

10. The apparatus of claim 9 , wherein the one or more processors are further configured to determine whether to perform the inverse modified discrete cosine transform on the at least one of the plurality of audio streams by:

determining whether the at least one of the plurality of audio streams among the plurality of audio streams has a highest estimated audio energy.

11. The apparatus of claim 9 , wherein the one or more processors are further configured to determine whether to perform the inverse modified discrete cosine transform on the at least one of the plurality of audio streams by:

determining whether the at least one of the plurality of audio streams has an estimated audio energy that exceeds an audio energy threshold.

12. The apparatus of claim 8 , wherein the one or more processors are configured to:

estimate by estimating the audio energy of each of the plurality of audio streams in the time domain;

determine by determining whether to perform a modified discrete cosine transform on the at least one of the plurality of audio streams; and

perform by performing the modified discrete cosine transform on the at least one of the plurality of audio streams.

13. The apparatus of claim 12 , wherein the one or more processors are further configured to:

if it is determined not to perform the transform on an audio stream of the plurality of audio streams, replace the audio stream with comfort noise.

14. The apparatus of claim 8 , wherein the audio streams are modified discrete cosine transformed frequency domain audio streams, and wherein the one or more processors are configured to:

perform by performing the inverse modified discrete cosine transform on the at least one of the plurality of audio streams.

15. One or more non-transitory computer readable storage media encoded with instructions that, when executed by a processor, cause the processor to:

obtain, from each of a plurality of participant devices in a communication session, a respective one of a plurality of audio streams;

estimate an audio energy of each of the plurality of audio streams;

based on the estimated audio energy of each of the plurality of audio streams, determine whether to perform a transform on at least one of the plurality of audio streams, including determining whether performing the transform would require a processing load that is less than a processing load threshold; and

if it is determined to perform the transform on the at least one of the plurality of audio streams:

perform the transform on the at least one of the plurality of audio streams; and

transmit the at least one of the plurality of audio streams to at least one of the plurality of participant devices.

16. The non-transitory computer readable storage media of claim 15 , wherein the audio streams are modified discrete cosine transformed frequency domain audio streams, the instructions further causing the processor to:

partially decode the plurality of audio streams so as to determine modified discrete cosine transform frequency coefficients corresponding to each of the plurality of audio streams, wherein the instructions cause the processor to:

estimate by estimating the audio energy of each of the plurality of audio streams based on the coefficients;

determine by determining whether to perform an inverse modified discrete cosine transform on the at least one of the plurality of audio streams; and

perform by performing the inverse modified discrete cosine transform on the at least one of the plurality of audio streams.

17. The non-transitory computer readable storage media of claim 16 , wherein the instructions that cause the processor to determine whether to perform the inverse modified discrete cosine transform on the at least one of the plurality of audio streams cause the processor to:

determine whether the at least one of the plurality of audio streams among the plurality of audio streams has a highest estimated audio energy.

18. The non-transitory computer readable storage media of claim 16 , wherein the instructions that cause the processor to determine whether to perform the inverse modified discrete cosine transform on the at least one of the plurality of audio streams cause the processor to:

determine whether the at least one of the plurality of audio streams has an estimated audio energy that exceeds an audio energy threshold.

19. The non-transitory computer readable storage media of claim 15 , the instructions further causing the processor to:

estimate by estimating the audio energy of each of the plurality of audio streams in the time domain;

determine by determining whether to perform a modified discrete cosine transform on the at least one of the plurality of audio streams; and

perform by performing the modified discrete cosine transform on the at least one of the plurality of audio streams.

20. The non-transitory computer readable storage media of claim 15 , wherein the audio streams are modified discrete cosine transformed frequency domain audio streams, and wherein the instructions cause the processor to:

perform by performing the inverse modified discrete cosine transform on the at least one of the plurality of audio streams.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2017
From: CARAMMA, MARCELLO
To: CISCO TECHNOLOGY, INC.
Reel/Frame 042504/0495 →
Continuity (1)
Related Publication 20180337964A1 · Nov 22, 2018