IP Library Granted Patent US 11,621,009
Granted Patent B2
US 11,621,009 · App. 16/719,857 · Granted Apr 4, 2023

Audio processing for voice encoding and decoding using spectral shaper model

Inventors: Lars Villemoes (Jarfalla, SE); Janusz Klejsa (Solna, SE); Per Hedelin (Gothenburg, SE)
Assignee: Dolby International AB
G10L19/032G10L19/02G10L19/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,621,009
App. No.
16/719,857
Granted
Apr 4, 2023
Kind
B2
Abstract

The present disclosure relates to an audio encoding and decoding (codec) system for voice encoding/decoding using a spectral shaper model. In an embodiment, a method of audio signal decoding comprises: receiving a bit stream associated with an audio signal, the bit stream including encoded transform coefficients, spectral envelope data and one or more parameters of a spectral shaper model, the spectral shaper model indicative of a fundamental frequency of a multi-sinusoidal signal model, where the fundamental frequency corresponds to a time domain delay; decoding the encoded transform coefficients; adjusting the decoded transform coefficients using the spectral envelope data and the spectral shaper model; reconstructing transform coefficients of the audio signal using the adjusted, decoded transform coefficients; and transforming the reconstructed transform coefficients into a time domain audio signal.

Claims (32)

1. A method of audio signal decoding, comprising:

receiving a bit stream associated with an audio signal, the bit stream including encoded transform coefficients, spectral envelope data and one or more parameters of a spectral shaper, the spectral shaper indicative of a fundamental frequency of a multi-sinusoidal signal model, where the fundamental frequency corresponds to a time domain delay;

decoding the encoded transform coefficients;

adjusting, with a subband predictor, the decoded transform coefficients using the spectral envelope data and the spectral shaper;

reconstructing transform coefficients of the audio signal using the adjusted, decoded transform coefficients; and

transforming the reconstructed transform coefficients into a time domain audio signal.

2. The method of claim 1 , further comprising:

dequantizing the transform coefficients with a quantizer selected from a plurality of quantizers.

3. The method of claim 1 , wherein adjusting the transform coefficients includes unflattening the transform coefficients.

4. The method of claim 1 , wherein the one or more parameters include a gain parameter and a lag parameter in stride units.

5. A system comprising:

one or more processors; and

a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, causes the one or more processors to perform operations comprising:

receiving a bit stream associated with an audio signal, the bit stream including encoded transform coefficients, spectral envelope data and one or more parameters of a spectral shaper, the spectral shaper indicative of a fundamental frequency of a multi-sinusoidal signal model, where the fundamental frequency corresponds to a time domain delay;

decoding the encoded transform coefficients;

adjusting, with a subband predictor, the decoded transform coefficients using the spectral envelope data and the spectral shaper;

reconstructing transform coefficients of the audio signal using the adjusted, decoded transform coefficients; and

transforming the reconstructed transform coefficients into a time domain audio signal.

6. The system of claim 5 , the operations further comprising:

dequantizing the transform coefficients with a quantizer selected from a plurality of quantizers.

7. The system of claim 5 , wherein adjusting the transform coefficients includes unflattening the transform coefficients.

8. The system of claim 5 , wherein the one or more parameters include a gain parameter and a lag parameter in stride units.

9. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, causes the one or more processors to perform operations comprising:

receiving a bit stream associated with an audio signal, the bit stream including encoded transform coefficients, spectral envelope data and one or more parameters of a spectral shaper, the spectral shaper indicative of a fundamental frequency of a multi-sinusoidal signal model, where the fundamental frequency corresponds to a time domain delay;

decoding the encoded transform coefficients;

adjusting, with a subband predictor, the decoded transform coefficients using the spectral envelope data and the spectral shaper;

reconstructing transform coefficients of the audio signal using the adjusted, decoded transform coefficients; and

transforming the reconstructed transform coefficients into a time domain audio signal.

10. The non-transitory computer-readable medium of claim 9 , the operations further comprising:

dequantizing the transform coefficients with a quantizer selected from a plurality of quantizers.

11. The non-transitory computer-readable medium of claim 9 , wherein adjusting the transform coefficients includes unflattening the transform coefficients.

12. The non-transitory computer-readable medium of claim 9 , wherein the one or more parameters include a gain parameter and a lag parameter in stride units.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2020
From: VILLEMOES, LARS; KLEJSA, JANUSZ; HEDELIN, PER
To: DOLBY INTERNATIONAL AB
Reel/Frame 051609/0549 →
Continuity (5)
Continuation 16032921 · Jul 11, 2018
Continuation 14781219
Provisional Application 61808675 · Apr 5, 2013
Provisional Application 61875553 · Sep 9, 2013
Related Publication 20200126574A1 · Apr 23, 2020
Cited By (3)
US 12,380,901 US 12,444,426 US 12,718,825