IP Library Granted Patent US 12670921
Granted Patent B1
US 12670921 · App. 18/469,691 · Granted Jun 30, 2026

Structured extension of speech bandwidth with denoising

Inventors: Kamil K. Wojcicki (Kangaroo Point, AU); Pramod Bhaskar Bachhav (Cracow, PL); Mansur Yesilbursa (Cracow, PL); Samer Lutfi Hijazi (San Jose, CA)
Assignee: CISCO TECHNOLOGY, INC.
G10L21/0232G10L25/18G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670921
App. No.
18/469,691
Granted
Jun 30, 2026
Kind
B1
Abstract

Techniques for speech bandwidth extension and denoising. The techniques integrate data-driven artificial intelligence (AI) models specifically trained to be robust to myriad of distortions. The system is capable of producing high-fidelity wideband speech from real-life narrowband inputs. The output is consistently preferred by listeners over the narrowband input, as well as over denoising alone.

Claims (43)

1 . A method comprising:

obtaining a narrowband signal containing speech audio and noise from a communication channel;

generating a spectra of the narrowband signal;

applying the narrowband signal to at least one neural network-based classifier that derives a high frequency replacement spectra that is a prediction of high frequency content for replacement in the spectra of the narrowband signal, wherein the at least one neural network-based classifier generates shape probabilities and level probabilities that are each aggregated to produce an aggregated shape component and an aggregated level component that are combined to generate the high frequency replacement spectra;

replacing content above a cut-off frequency of the spectra of the narrowband signal with the high frequency replacement spectra to produce a bandwidth extended spectra; and

processing the bandwidth extended spectra with a deep neural network spectral mask to generate a bandwidth extended enhanced spectra from which noise is suppressed at lower frequencies and enhancement is made at higher frequencies.

2 . The method of claim 1 , wherein the at least one neural network-based classifier performs a first classification to generate high frequency shape probabilities that are weighted to produce the aggregated shape component and a second classification to generate level probabilities that are weighted to produce the aggregated level component.

3 . The method of claim 2 , wherein the first classification of the at least one neural network-based classifier generates the high frequency shape probabilities which are a prediction of high frequency shapes, among a plurality of stored high frequency shapes, that are present in the narrowband signal, and the second classification of the at least one neural network-based classifier generates the level probabilities which are a prediction of levels, among a plurality of stored levels, of high frequency shapes predicted to be present in the narrowband signal.

4 . The method of claim 3 , wherein the aggregated shape component is a vector and the aggregated level component is a scalar, and when combined, produce the high frequency replacement spectra that is in a log-magnitude domain.

5 . The method of claim 3 , wherein generating a spectra of the narrowband signal comprises applying a Short-Time Fourier Transform operation on the narrowband signal to produce a complex spectra.

6 . The method of claim 5 , further comprising:

generating from the complex spectra a magnitude spectra and a phase spectra; and

performing a logarithm operation on the magnitude spectra to produce log-magnitude spectra,

wherein replacing comprises replacing content in the log-magnitude spectra above a cut-off frequency with the high frequency replacement spectra to produce bandwidth extended log-magnitude spectra.

7 . The method of claim 6 , further comprising:

performing a low-to-high frequency translation or high frequency randomization on the phase spectra to produce high frequency phase spectra;

converting the high frequency phase spectra to phase spectra;

converting the bandwidth extended log-magnitude spectra to bandwidth extended magnitude spectra; and

multiplying the phase spectra with the bandwidth extended magnitude spectra to produce bandwidth extended complex spectra.

8 . The method of claim 7 , wherein processing comprises processing the bandwidth extended complex spectra with the deep neural network spectral mask to produce bandwidth extended enhanced complex spectra.

9 . The method of claim 8 , wherein the deep neural network spectral mask is predicted using a generative adversarial network (GAN)-trained neural network.

10 . The method of claim 9 , wherein the GAN-trained neural network is trained based on exposure to one or more of: different audio coder/decoder processes, different bitrates, different cut-off frequencies, different spectral shapes, different noises, different reverb and different levels to achieve noise reduction/speech enhancement training.

11 . The method of claim 8 , further comprising:

transforming the bandwidth extended enhanced complex spectra to a wideband enhanced speech audio signal in the time domain.

12 . The method of claim 11 , wherein transforming the bandwidth extended enhanced complex spectra comprises performing an inverse Short-Time Fourier Transform on the bandwidth extended enhanced complex spectra to produce the wideband enhanced speech audio signal in the time domain.

13 . An apparatus comprising:

a communication interface configured to receive signals over a communication channel, the signals including a narrowband signal containing speech audio and noise from the communication channel; and

a processor coupled to the communication interface, the processor configured to perform operations on the narrowband signal including:

generating a spectra of the narrowband signal;

applying the narrowband signal to at least one neural network-based classifier that derives a high frequency replacement spectra that is a prediction of high frequency content for replacement in the spectra of the narrowband signal, wherein the at least one neural network-based classifier generates shape probabilities and level probabilities that are each aggregated to produce an aggregated shape component and an aggregated level component that are combined to generate the high frequency replacement spectra;

replacing content above a cut-off frequency of the spectra of the narrowband signal with the high frequency replacement spectra to produce a bandwidth extended spectra; and

processing the bandwidth extended spectra with a deep neural network spectral mask to generate a bandwidth extended enhanced spectra from which noise is suppressed at lower frequencies and enhancement is made at higher frequencies.

14 . The apparatus of claim 13 , wherein the at least one neural network-based classifier performs a first classification to generate high frequency shape probabilities that are weighted to produce the aggregated shape component and a second classification to generate level probabilities that are weighted to produce the aggregated level component, wherein the aggregated shape component and the aggregated level component are combined to produce the high frequency replacement spectra.

15 . The apparatus of claim 14 , wherein the first classification of the at least one neural network-based classifier generates the high frequency shape probabilities which are a prediction of high frequency shapes, among a plurality of stored high frequency shapes, that are present in the narrowband signal, and the second classification of the at least one neural network-based classifier generates the level probabilities which are a prediction of levels, among a plurality of stored levels, of high frequency shapes predicted to be present in the narrowband signal.

16 . The apparatus of claim 15 , wherein the aggregated shape component is a vector and the aggregated level component is a scalar, and when combined, produce the high frequency replacement spectra that is in a log-magnitude domain.

17 . One or more non-transitory computer readable storage media encoded with instructions that, when executed by a processor, cause the processor to perform operations including:

generating a spectra of a narrowband signal containing speech audio and noise from a communication channel;

applying the narrowband signal to at least one neural network-based classifier that derives a high frequency replacement spectra that is a prediction of high frequency content for replacement in the spectra of the narrowband signal, wherein the at least one neural network-based classifier generates shape probabilities and level probabilities that are each aggregated to produce an aggregated shape component and an aggregated level component that are combined to generate the high frequency replacement spectra;

replacing content above a cut-off frequency of the spectra of the narrowband signal with the high frequency replacement spectra to produce a bandwidth extended spectra; and

processing the bandwidth extended spectra with a deep neural network spectral mask to generate a bandwidth extended enhanced spectra from which noise is suppressed at lower frequencies and enhancement is made at higher frequencies.

18 . The one or more non-transitory computer readable storage media of claim 17 , wherein the at least one neural network-based classifier performs a first classification to generate high frequency shape probabilities that are weighted to produce the aggregated shape component and a second classification to generate level probabilities that are weighted to produce the aggregated level component, wherein the aggregated shape component and the aggregated level component are combined to produce the high frequency replacement spectra.

19 . The one or more non-transitory computer readable storage media of claim 18 , wherein the first classification of the at least one neural network-based classifier generates the high frequency shape probabilities which are a prediction of high frequency shapes, among a plurality of stored high frequency shapes, that are present in the narrowband signal, and the second classification of the at least one neural network-based classifier generates the level probabilities which are a prediction of levels, among a plurality of stored levels, of high frequency shapes predicted to be present in the narrowband signal.

20 . The one or more non-transitory computer readable storage media of claim 19 , wherein the aggregated shape component is a vector and the aggregated level component is a scalar, and when combined, produce the high frequency replacement spectra that is in a log-magnitude domain.