IP Library › Granted Patent US 12,573,409
Granted Patent B2
US 12,573,409 · App. 18/582,428 · Granted Mar 10, 2026

Audio encoder, method for providing an encoded representation of an audio information, computer program and encoded audio representation using immediate playout frames

Inventors: Max Neuendorf (Erlangen, DE); Nikolaus Rettelbach (Erlangen, DE); Christina Mittag (Erlangen, DE); Daniel Richter (Erlangen, DE); Agathe Deniau (Erlangen, DE); Wahaj Aslam (Erlangen, DE); Ingo Hofmann (Erlangen, DE); Bernd Herrmann (Erlangen, DE)
Assignee: Fraunhofer-Gesellschaft zur Förderung der angewandten Forschung e.V.
G10L19/002G10L19/032G10L19/12G10L2019/0001
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,409
App. No.
18/582,428
Granted
Mar 10, 2026
Kind
B2
Abstract

An audio encoder is disclosed for providing an encoded representation of an audio information encodes a sequence of audio frames. The audio encoder provides one or more immediate playout frames including a representation of a current audio frame, preceding the current audio frame. The audio encoder provides the representations of the current frame and of the one or more audio frames preceding the current audio frame, such that these representations are decodable using a same decoder configuration. The audio encoder provides the representations of the one or more audio frames preceding the current audio frame, which are included into the immediate playout frame, using a modified encoding functionality, which encodes an audio frame using a smaller number of bits than a normal encoding functionality, which is used for the encoding of the current audio frame.

Claims (51)

1 . An audio encoder for providing an encoded representation of an audio information on the basis of an input audio information,

wherein the audio encoder is configured to encode a sequence of audio frames,

wherein the audio encoder is configured to provide one or more immediate playout frames comprising a representation of a current audio frame and encoded representations of one or more audio frames preceding the current audio frame,

wherein the audio encoder is configured to provide the representation of the current frame and the representations of the one or more audio frames preceding the current audio frame such that the representation of the current frame and the representations of the one or more audio frames preceding the current audio frame are decodable using a same decoder configuration, and

wherein the audio encoder is configured to provide the representations of the one or more audio frames preceding the current audio frame, which are included into the immediate playout frame, using a modified encoding functionality which is adapted to encode an audio frame using a smaller number of bits than a normal encoding functionality which is used for the encoding of the current audio frame.

2 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, in which a bitrate setting or a bitrate limit is reduced when compared to the normal encoding functionality, for providing the representations of the one or more audio frames preceding the current audio frame.

3 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use the bitrate setting or bitrate limit for deciding how many bits are allocated to an encoding of different spectral values.

4 . The audio encoder according to claim 1 , wherein the audio encoder is configured to leave encoding parameters, a change of which would result in a change of a decoder configuration unchanged between the encoding of the current frame and the encoding of the one or more audio frames preceding the current audio frame.

5 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, in which a number of bits available for a quantization or for an encoding of one or more parameters is reduced or limited when compared to normal encoding functionality, for providing the representations of the one or more audio frames preceding the current audio frame.

6 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, in which a coarser quantization of a MDCT spectrum is used when compared to the normal encoding functionality, for providing the representations of the one or more audio frames preceding the current audio frame.

7 . The audio encoder according to claim 1 , wherein:

the audio encoder is configured to change a global gain parameter, in order to acquire a coarser quantization, when using the modified encoding functionality; and/or

the audio encoder is configured to use a modified encoding functionality, in which a masking threshold acquired using a psychoacoustic model is changed to acquire a coarser quantization, for providing the representations of the one or more audio frames preceding the current audio frame.

8 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, and a bandwidth extension bit load is reduced, for providing the representations of the one or more audio frames preceding the current audio frame.

9 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, and wherein:

a spectral band replication bit load is reduced, for providing the representations of the one or more audio frames preceding the current audio frame; and/or

a plurality of spectral band replication parameters are set to a predetermined value which allows for a reduction or for a minimization of a number of bits required for an encoding of the spectral band replication parameters, for providing the representations of the one or more audio frames preceding the current audio frame.

10 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, and wherein:

a number of spectral band replication bands or a number of spectral band replication envelopes is reduced, for providing the representations of the one or more audio frames preceding the current audio frame; and/or

a frequency resolution of spectral band replication data is reduced, for providing the representations of the one or more audio frames preceding the current audio frame.

11 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, and wherein:

a bit load in a UsacSbrData( ) syntax element is reduced, for providing the representations of the one or more audio frames preceding the current audio frame, while keeping spectral band replication parameters which are part of an usacConfig( ) syntax element and/or of a SbrConfig( ) syntax element unchanged; and/or

a multi-channel encoding bit load is reduced, for providing the representations of the one or more audio frames preceding the current audio frame.

12 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, and wherein:

a transform-coded excitation linear-prediction domain encoding is used instead of an ACELP linear predication domain encoding, for providing the representations of the one or more audio frames preceding the current audio frame; and/or

a transform-coded excitation linear-prediction domain encoding with a coarser quantization is used instead of a transform-coded excitation linear-prediction domain encoding with a finer quantization, for providing the representations of the one or more audio frames preceding the current audio frame.

13 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, and wherein:

a time domain resolution is reduced, for providing the representations of the one or more audio frames preceding the current audio frame; and/or

a usage of multiple TCX windows within a single audio frame is avoided, for providing the representations of the one or more audio frames preceding the current audio frame.

14 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, and wherein:

a single long TCX window is used instead of 2 medium sized TCX windows, and/or in which a single long TCX window is used instead of 4 short TCX windows, or in which a single long TCX window is used instead of a plurality of shorted TCX windows, for providing the representations of the one or more audio frames preceding the current audio frame; and/or

a usage of a plurality of short MDCT transform windows within a single audio frame is avoided, for providing the representations of the one or more audio frames preceding the current audio frame.

15 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, and wherein

a single long MDCT transform window is used instead a plurality of shorter MDCT transform windows, for providing the representations of the one or more audio frames preceding the current audio frame; and/or;

a “START_STOP” MDCT transform window is used instead of an “EIGHT_SHORT” MDCT transform window, for providing the representations of the one or more audio frames preceding the current audio frame.

16 . The audio encoder according to claim 1 , wherein the audio encoder is configured to use a modified encoding functionality, and wherein:

a reduced ACELP excitation codebook size is used, for providing the representations of the one or more audio frames preceding the current audio frame; and/or

a reduced number of bits is used for an encoding of an innovation codebook index representing an ACELP excitation, for providing the representations of the one or more audio frames preceding the current audio frame.

17 . The audio encoder according to claim 1 , wherein the audio encoder is configured to also encode the one or more audio frames preceding the current audio frame in the normal encoding mode, in order to acquire one or more non-immediate playout frames preceding the immediate playout frame.

18 . The audio encoder according to claim 1 , wherein the audio encoder is configured to re-use intermediate encoding results of an encoding of the one or more frames preceding the current frame using the normal encoding functionality, in order to determine the bitrate reduced encoded representation of the one or more frames preceding the current frame which is the result of the modified encoding functionality.

19 . The audio encoder according to claim 1 , wherein the audio encoder is configured to implement the normal encoding functionality using a first core coder instance, and to implement the modified encoding functionality using a second core coder instance.

20 . A method for providing an encoded representation of an audio information on the basis of an input audio information, the method comprising:

encoding a sequence of audio frames;

providing one or more immediate playout frames comprising a representation of a current audio frame and encoded representations of one or more audio frames preceding the current audio frame;

providing the representation of the current frame and the representations of the one or more audio frames preceding the current audio frame such that the representation of the current frame and the representations of the one or more audio frames preceding the current audio frame are decodable using a same decoder configuration, and

providing the representations of the one or more audio frames preceding the current audio frame, which are included into the immediate playout frame, using a modified encoding functionality which is adapted to encode an audio frame using a smaller number of bits than a normal encoding functionality which is used for the encoding of the current audio frame.

21 . A non-transitory digital storage medium having a computer program stored that when executed by a computer process provides an encoded representation of an audio information on the basis of an input audio information by:

encoding a sequence of audio frames;

providing one or more immediate playout frames comprising a representation of a current audio frame and encoded representations of one or more audio frames preceding the current audio frame;

providing the representation of the current frame and the representations of the one or more audio frames preceding the current audio frame such that the representation of the current frame and the representations of the one or more audio frames preceding the current audio frame are decodable using a same decoder configuration, and

providing the representations of the one or more audio frames preceding the current audio frame, which are included into the immediate playout frame, using a modified encoding functionality which is adapted to encode an audio frame using a smaller number of bits than a normal encoding functionality which is used for the encoding of the current audio frame.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2025
From: NEUENDORF, MAX; RETTELBACH, NIKOLAUS; MITTAG, CHRISTINA; RICHTER, DANIEL; DENIAU, AGATHE; ASLAM, WAHAJ; HOFMANN, INGO; HERRMANN, BERND
To: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 072732/0558 →
Priority Claims (1)
EP 21192257 · Aug 19, 2021 · regional
Continuity (2)
Continuation PCTEP2022073073 · Aug 18, 2022
Related Publication 20240194207A1 · Jun 13, 2024
References Cited (10)
US 10614824B2 · Fischer · 2020 [cited by examiner]
US 11972769B2 · Fersch · 2024 [cited by examiner]
US 20050261900A1 · Ojala · 2005 [cited by examiner]
US 20070223660A1 · Dei · 2007 [cited by examiner]
EP 2863386A1 · 2015 [cited by applicant]
WO 2018130577A1 · 2018 [cited by applicant]
WO 2020038938A1 · 2020 [cited by applicant]
International Search Report and Written Opinion in PCT/EP2022/073073, mailed Oct. 25, 2022, 18 pages. [cited by applicant]
“Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio”, International Standard, ISO/IEC 23008-3:2022(E), Third edition, Aug. 2022, pp. 1-882. [cited by applicant]
“Information technology—MPEG audio technologies Part 3: Unified speech and audio coding”, International Standard, ISO/IEC 23003-3:2020(E), Second edition, Jun. 2020, pp. 1-348. [cited by applicant]