IP Library Granted Patent US 12,205,602
Granted Patent B2
US 12,205,602 · App. 18/601,546 · Granted Jan 21, 2025

Backward-compatible integration of high frequency reconstruction techniques for audio signals

Inventors: Kristofer Kjoerling (Solna, SE); Lars Villemoes (Järfälla, SE); Heiko Purnhagen (Sundbyberg, SE); Per Ekstrand (Saltsjöbaden, SE)
Assignee: Dolby International AB
G10L19/02G10L19/167G10L19/26H03M7/6005H03M7/6011
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,205,602
App. No.
18/601,546
Granted
Jan 21, 2025
Kind
B2
Abstract

A method for decoding an encoded audio bitstream is disclosed. The method includes receiving the encoded audio bitstream and decoding the audio data to generate a decoded lowband audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded lowband audio signal with an analysis filterbank to generate a filtered lowband audio signal. The method also includes extracting a flag indicating whether either spectral translation or harmonic transposition is to be performed on the audio data and regenerating a highband portion of the audio signal using the filtered lowband audio signal and the high frequency reconstruction metadata in accordance with the flag.

Claims (112)

1. A method for performing high frequency reconstruction of an audio signal, the method comprising:

receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a lowband portion of the audio signal and high frequency reconstruction metadata;

decoding the audio data to generate a decoded lowband audio signal;

extracting from the encoded audio bitstream the high frequency reconstruction metadata, the high frequency reconstruction metadata including operating parameters for a high frequency reconstruction process, the operating parameters including a patching mode parameter located in a backward-compatible extension container of the encoded audio bitstream, wherein a first value of the patching mode parameter indicates spectral translation and a second value of the patching mode parameter indicates harmonic transposition by phase-vocoder frequency spreading, wherein the encoded audio bitstream further includes a fill element with an identifier indicating a start of the fill element and fill data after the identifier, wherein the fill data includes the backward-compatible extension container, and wherein the identifier is a three bit unsigned integer transmitted most significant bit first and having a value of 0x6;

filtering the decoded lowband audio signal to generate a filtered lowband audio signal, wherein the filtering is performed by an analysis filterbank that includes analysis filters, h k (n), that are modulated versions of a prototype filter, p 0 (n), according to:

h

k

(

n

)

=

p

0

(

n

)

exp

{

i

π

M

(

k

+

1

2

)

(

n

-

N

2

)

}

,

0

n

N

;

0

k

<

M

wherein p 0 (n) is a real-valued symmetric or asymmetric prototype filter, M is a number of channels in the analysis filterbank and N is an order of the prototype filter, wherein the prototype filter, p 0 (n), is derived from coefficients of Table 4 herein by one or more mathematical operations selected from the group consisting of rounding, subsampling, interpolation, or decimation; and

regenerating a highband portion of the audio signal using the filtered lowband audio signal and the high frequency reconstruction metadata, wherein the regenerating includes spectral translation if the patching mode parameter is the first value and the regenerating includes harmonic transposition by phase-vocoder frequency spreading if the patching mode parameter is the second value; and

combining the filtered lowband audio signal with the regenerated highband portion to form a wideband audio signal.

2. The method of claim 1 wherein the backward-compatible extension container includes inverse filtering control data to be used when the patching mode parameter equals the second value.

3. The method of claim 1 wherein the backward-compatible extension container further includes missing harmonic control data to be used when the patching mode parameter equals the second value.

4. The method of claim 1 wherein a phase shift is added to the filtered lowband audio signal after the filtering and compensated for before the combining to reduce a complexity of the method.

5. A non-transitory computer readable medium containing instructions that when executed by a processor perform the method of claim 1 .

6. An audio processing unit for performing high frequency reconstruction of an audio signal, the audio processing unit comprising:

an input interface for receiving an encoded audio bitstream, the encoded audio bitstream including audio data representing a lowband portion of the audio signal and high frequency reconstruction metadata;

a core audio decoder for decoding the audio data to generate a decoded lowband audio signal;

a deformatter for extracting from the encoded audio bitstream the high frequency reconstruction metadata, the high frequency reconstruction metadata including operating parameters for a high frequency reconstruction process, the operating parameters including a fill element with an identifier indicating a start of the fill element and fill data after the identifier, wherein the fill data includes a backward-compatible extension container including a patching mode parameter, wherein a first value of the patching mode parameter indicates spectral translation and a second value of the patching mode parameter indicates harmonic transposition by phase-vocoder frequency spreading, and wherein the identifier is a three bit unsigned integer transmitted most significant bit first and having a value of 0x6;

an analysis filterbank for filtering the decoded lowband audio signal to generate a filtered lowband audio signal, wherein the filtering is performed by an analysis filterbank that includes analysis filters, h k (n), that are modulated versions of a prototype filter, p 0 (n), according to:

h

k

(

n

)

=

p

0

(

n

)

exp

{

i

π

M

(

k

+

1

2

)

(

n

-

N

2

)

}

,

0

n

N

;

0

k

<

M

wherein p 0 (n) is a real-valued symmetric or asymmetric prototype filter, M is a number of channels in the analysis filterbank and N is an order of the prototype filter, wherein the prototype filter, p 0 (n), is derived from coefficients of Table 4 herein by one or more mathematical operations selected from the group consisting of rounding, subsampling, interpolation, or decimation; and

a high frequency regenerator for reconstructing a highband portion of the audio signal using the filtered lowband audio signal and the high frequency reconstruction metadata, wherein the reconstructing includes a spectral translation if the patching mode parameter is the first value and the reconstructing includes harmonic transposition by phase-vocoder frequency spreading if the patching mode parameter is the second value; and

a synthesis filterbank for combining the filtered lowband audio signal with the regenerated highband portion to form a wideband audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2024
From: KJOERLING, KRISTOFER; VILLEMOES, LARS; PURNHAGEN, HEIKO; EKSTRAND, PER
To: DOLBY INTERNATIONAL AB
Reel/Frame 067597/0931 →
Continuity (5)
Continuation 18357679 · Jul 24, 2023
Continuation 17677608 · Feb 22, 2022
Continuation 16965291
Provisional Application 62622205 · Jan 26, 2018
Related Publication 20240212695A1 · Jun 27, 2024
References Cited (68)
US 6449596B1 · Ejima · 2002 [cited by applicant]
US 8532983B2 · Gao · 2013 [cited by applicant]
US 10818306B2 · Villemoes · 2020 [cited by applicant]
US 11341977B2 · Ishikawa · 2022 [cited by applicant]
US 11626121B2 · Kjoerling · 2023 [cited by applicant]
US 20020087304A1 · Kjörling · 2002 [cited by applicant]
US 20030016772A1 · Ekstrand · 2003 [cited by applicant]
US 20040008615A1 · Oh · 2004 [cited by applicant]
US 20040145478A1 · Frederick · 2004 [cited by applicant]
US 20060031075A1 · Oh · 2006 [cited by applicant]
US 20080260048A1 · Oomen · 2008 [cited by applicant]
US 20090006103A1 · Koishida · 2009 [cited by applicant]
US 20090063159A1 · Crockett · 2009 [cited by applicant]
US 20090319259A1 · Liljeryd · 2009 [cited by applicant]
US 20100063802A1 · Gao · 2010 [cited by applicant]
US 20110178810A1 · Villemoes · 2011 [cited by applicant]
US 20110216918A1 · Nagel · 2011 [cited by applicant]
US 20120328124A1 · Kjoerling · 2012 [cited by applicant]
US 20130028426A1 · Purnhagen · 2013 [cited by applicant]
US 20140019146A1 · Neuendorf · 2014 [cited by applicant]
US 20140039890A1 · Mundt · 2014 [cited by applicant]
US 20140365231A1 · Hoerich · 2014 [cited by applicant]
US 20150248894A1 · Ishikawa · 2015 [cited by applicant]
US 20150317986A1 · Kristofer · 2015 [cited by applicant]
US 20160021476A1 · Robinson · 2016 [cited by applicant]
US 20160142845A1 · Dick · 2016 [cited by applicant]
US 20160203826A1 · Kaniewska · 2016 [cited by applicant]
US 20160254007A1 · Guo · 2016 [cited by applicant]
US 20170110135A1 · Disch · 2017 [cited by applicant]
US 20170178651A1 · Davis · 2017 [cited by applicant]
US 20170178655A1 · Kjoerling · 2017 [cited by applicant]
US 20170206909A1 · Truman · 2017 [cited by applicant]
US 20180025737A1 · Villemoes · 2018 [cited by applicant]
US 20190156845A1 · Nagel · 2019 [cited by applicant]
US 20190385624A1 · Kjoerling · 2019 [cited by applicant]
US 20230049358A1 · Kjoerling · 2023 [cited by applicant]
CN 1555046 · 2004 [cited by applicant]
CN 1766993 · 2006 [cited by applicant]
CN 101140759 · 2008 [cited by applicant]
CN 103155033 · 2013 [cited by applicant]
CN 104217730 · 2014 [cited by applicant]
JP 2004533155 · 2004 [cited by applicant]
JP 2011529578 · 2011 [cited by applicant]
JP 2013521538 · 2013 [cited by applicant]
JP 2015534112A · 2015 [cited by applicant]
JP 2017526004 · 2017 [cited by applicant]
KR 20170115101A · 2017 [cited by applicant]
RU 2244386 · 2005 [cited by applicant]
RU 2547220 · 2015 [cited by applicant]
RU 2601188C2 · 2016 [cited by applicant]
TW 201322743 · 2013 [cited by applicant]
TW 201523590 · 2015 [cited by applicant]
TW 201640357 · 2016 [cited by applicant]
TW 201643864 · 2016 [cited by applicant]
TW 201724085A · 2017 [cited by applicant]
WO 2014161993 · 2014 [cited by applicant]
WO 2016146492A1 · 2016 [cited by applicant]
WO 2018175347A1 · 2018 [cited by applicant]
WO 2019148112A1 · 2019 [cited by applicant]
Bleidt, R. L., Sen, D., Niedermeier, A., Czelhan, B., Fug, S., Disch, S., . . . & Kim, M. Y. (2017). Development of the MPEG-H TV audio system for ATSC 3.0. IEEE Transactions on broadcasting, 63(1), 202-236. [cited by examiner]
Anonymous: “ISO/IEC 14496-3:200x, Fourth Edition, part 4”, 82. MPEG MEETING;Oct. 22, 2007-Oct. 26, 2007; Shenzhen; (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11 ),, May 15, 2009 (May 15, 2009), XP030017007. [cited by applicant]
Gampp, P., Uhle, C., Herre, J., Disch, S., Hellmuth, 0., Prokein, P., . . . & Karampourniotis, A (Aug. 2017). Methods for low bitrate coding enhancement part I: spectral restoration. In Audio Engineering Society Confere… [cited by applicant]
Haishan, Z et al “QMF Based Harmonic Spectral Band Replication” AES presented at the 131st Convention, Oct. 20-23, 2011, New York, NY, USA, pp. 1-9. [cited by applicant]
Information Technology—MPEG Audio technologies—Part 3: Unified speech and Audio Coding, ISO/IEC DIS 23003-3, Jul. 2011. [cited by applicant]
ISO/IEC 23003-3, Information Technology—MPEG Audio Technologies, Part 3, Unified Speech and Audio Coding, 2012. [cited by applicant]
ISO/IEC JTC 1/SC 29/WG 11 N9500, ISO/IEC 14496-3, MPEG-4 Audio Fourth Edition, Subpart 4: General Audio Coding (GA)—AAC, TwinVQ, BSAC, May 15, 2019. [cited by applicant]
MPEG-4 Advanced Audio Coding (AAC) format, described in the MPEG standard ISO/IEC 14496-3:2009. [cited by applicant]
Ryu, Sang-Uk, et al (Oct. 2005). “Effective high frequency regeneration based on sinusoidal modeling for MPEG-4 HE-AAC” In IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, 2005. (pp. 211-214). … [cited by applicant]