IP Library Granted Patent US 12706111
Granted Patent B2
US 12706111 · App. 18/425,122 · Granted Aug 11, 2026

Receive-side audio processing for calls in a web conferencing client

Inventors: Scott Plude (San Jose, CA); Mathew Shaji Kavalekalam (Warsaw, PL); Nathan Allan Rickey (San Marcos, CA); Kamil Krzysztof Wojcicki (Kangaroo Point, AU); Mansur Yesilbursa (Krakow, AR); Rafal Pilarczyk (Plock, PL)
Assignee: Cisco Technology, Inc.
G10L21/0232G10L15/063G10L25/18G10L25/21G10L25/51H04M3/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12706111
App. No.
18/425,122
Granted
Aug 11, 2026
Kind
B2
Abstract

According to one or more embodiments of the disclosure, receive-side audio processing for calls in a web conferencing client is provided by a method that includes receiving, by a device, an audio signal from a sending device and detecting, by the device, a particular telephony audio cue present in the audio signal, wherein the particular telephony audio cue has similar characteristics of noise that is to be removed by a noise removal operation. The method further includes preserving, by the device, the particular telephony audio cue in the audio signal throughout the noise removal operation on the audio signal and causing, by the device, an enhanced audio signal to be produced in conjunction with the noise removal operation, wherein the enhanced audio signal includes the particular telephony audio cue to be delivered to a receiver device.

Claims (50)

1 . A method comprising:

receiving, by a device, an audio signal from a sending device;

detecting, by the device, a particular telephony audio cue present in the audio signal, wherein the particular telephony audio cue has similar characteristics of noise that is to be removed by a noise removal operation;

preserving, by the device, the particular telephony audio cue in the audio signal throughout the noise removal operation on the audio signal by automatically bypassing, responsive to detecting the particular telephony audio cue, the noise removal operation during one or more time intervals of the audio signal that correspond to the detected particular telephony audio cue, while continuing to apply the noise removal operation during other time intervals of the audio signal that do not contain the particular telephony audio cue; and

causing, by the device, an enhanced audio signal to be produced in conjunction with the noise removal operation, wherein the enhanced audio signal includes the particular telephony audio cue to be delivered to a receiver device.

2 . The method as in claim 1 , further comprising:

applying a warped discrete Fourier transform to the audio signal to enhance frequency resolution for lower frequency bins.

3 . The method as in claim 2 , wherein applying the warped discrete Fourier transform is preceded by applying an all-pass filter to the audio signal to adjust a phase response of the audio signal prior to detecting the particular telephony audio cue.

4 . The method as in claim 1 , wherein detecting comprises:

differentiating the particular telephony audio cue from noise based on identifying frequency bins within the audio signal that have power levels exceeding a predefined threshold.

5 . The method as in claim 1 , further comprising:

utilizing a machine learning model trained to differentiate between music-on-hold and background music.

6 . The method as in claim 1 , wherein detecting comprises:

training a machine learning model to detect the particular telephony audio cue.

7 . The method as in claim 1 , further comprising:

determining whether the audio signal is a narrow band signal; and

extending bandwidth of the audio signal to full band for narrow band signals.

8 . The method as in claim 1 , further comprising:

tracking a power level of each frequency bin in the audio signal over consecutive short-time Fourier transform (STFT) frames to identify constant amplitude signals indicative of telephony audio cues.

9 . The method as in claim 1 , wherein the particular telephony audio cue is a tone indicating a status of a communication associated with the audio signal.

10 . The method as in claim 1 , wherein the noise removal operation on the audio signal is a receiver-side audio processing operation for a call within a web conferencing client.

11 . The method as in claim 1 , wherein the particular telephony audio cue is at least one selected from a group of music on hold, answering service tones, touch tones, and combinations thereof, and wherein the device determines the one or more time intervals by segmenting the audio signal into analysis frames and identifying the one or more time intervals as one or more contagious analysis frames in which the particular telephony audio cue is detected.

12 . An apparatus, comprising:

one or more network interfaces;

a processor coupled to the one or more network interfaces and configured to execute one or more processes; and

a memory configured to store a process that is executable by the processor, the process when executed configured to:

receive an audio signal from a sending device;

detect a particular telephony audio cue present in the audio signal, wherein the particular telephony audio cue has similar characteristics of noise that is to be removed by a noise removal operation;

preserve the particular telephony audio cue in the audio signal throughout the noise removal operation on the audio signal by automatically bypassing, responsive to detecting the particular telephony audio cue, the noise removal operation during one or more time intervals of the audio signal that correspond to the detected particular telephony audio cue, while continuing to apply the noise removal operation during other time intervals of the audio signal that do not contain the particular telephony audio cue; and

cause an enhanced audio signal to be produced in conjunction with the noise removal operation, wherein the enhanced audio signal includes the particular telephony audio cue to be delivered to a receiver device.

13 . The apparatus as in claim 12 , wherein the process when executed is further configured to:

apply a warped discrete Fourier transform to the audio signal to enhance frequency resolution for lower frequency bins.

14 . The apparatus as in claim 13 , wherein the process when executed is further configured to:

apply an all-pass filter to the audio signal to adjust a phase response of the audio signal prior to detecting the particular telephony audio cue before application of the warped discrete Fourier transform.

15 . The apparatus as in claim 12 , wherein to detect the particular telephony audio cue further comprises to:

differentiate the particular telephony audio cue from noise based on identifying frequency bins within the audio signal that have power levels exceeding a predefined threshold.

16 . The apparatus as in claim 12 , wherein the process when executed is further configured to:

train a machine learning model to detect the particular telephony audio cue.

17 . The apparatus as in claim 12 , wherein to detect the particular telephony audio cue further comprises to:

utilize a machine learning model trained to differentiate between music-on-hold and background music.

18 . The apparatus as in claim 12 , wherein the process when executed is further configured to:

determine whether the audio signal is a narrow band signal; and

extend bandwidth of the audio signal to full band for narrow band signals.

19 . The apparatus as in claim 12 , wherein the process when executed is further configured to:

track a power level of each frequency bin in the audio signal over consecutive short-time Fourier transform (STFT) frames to identify constant amplitude signals indicative of telephony audio cues.

20 . A tangible, non-transitory, computer-readable medium storing program instructions that cause a device to execute a process comprising:

receiving an audio signal from a sending device;

detecting a particular telephony audio cue present in the audio signal, wherein the particular telephony audio cue has similar characteristics of noise that is to be removed by a noise removal operation;

preserving the particular telephony audio cue in the audio signal throughout the noise removal operation on the audio signal by automatically bypassing, responsive to detecting the particular telephony audio cue, the noise removal operation during one or more time intervals of the audio signal that correspond to the detected particular telephony audio cue, while continuing to apply the noise removal operation during other time intervals of the audio signal that do not contain the particular telephony audio cue; and

causing an enhanced audio signal to be produced in conjunction with the noise removal operation, wherein the enhanced audio signal includes the particular telephony audio cue to be delivered to a receiver device.