IP Library Granted Patent US 12707222
Granted Patent B2
US 12707222 · App. 18/407,598 · Granted Aug 11, 2026

Ambience audio representation and associated rendering

Inventor: Lasse Laaksonen (Tampere, FI)
Assignee: Nokia Technologies Oy
H04S7/303G10L25/21H04R1/406H04R3/005H04R5/027H04S3/008G10L19/008H04S2400/03H04S2400/11H04S2400/15H04S2420/03H04S2420/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12707222
App. No.
18/407,598
Granted
Aug 11, 2026
Kind
B2
Abstract

An apparatus configured to: generate at least one ambience component representation, wherein the ambience component representation comprises: at least one respective diffuse background audio signal, and at least one parameter, wherein the at least one parameter is associated with the at least one respective diffuse background audio signal; and output the at least one ambience component representation, wherein the at least one ambience component representation is configured to be used in rendering an ambience audio signal, based on the at least one respective diffuse background audio signal, the at least one parameter of the at least one ambience component representation, and at least one of: a rendering position, or a rendering direction, relative to a defined position within an audio field.

Claims (98)

1 . An apparatus comprising:

at least one processor; and

at least one memory storing instructions that, when executed with the at least one processor, cause the apparatus at least to:

generate at least one ambience component representation, wherein the ambience component representation comprises:

at least one respective diffuse background audio signal, and

at least one parameter, wherein the at least one parameter is associated with the at least one respective diffuse background audio signal; and

output the at least one ambience component representation, wherein the at least one output ambience component representation is configured to be used in rendering at least one ambience audio signal based on processing of:

the at least one respective diffuse background audio signal,

the at least one parameter of the at least one ambience component representation, and

at least one of: a rendering position, or a rendering direction, relative to a defined position within an audio field.

2 . The apparatus of claim 1 , wherein the at least one parameter is associated with:

at least one frequency range or at least one part of the at least one frequency range,

at least one time period or at least one part of the at least one time period, and

a directional range for the defined position, wherein the directional range defines a range of angles.

3 . The apparatus of claim 1 , wherein the at least one ambience component representation further comprises at least one of:

a minimum distance threshold, over which the at least one ambience component representation is configured to be used in rendering the at least one ambience audio signal,

a maximum distance threshold, under which the at least one ambience component representation is configured to be used in rendering the at least one ambience audio signal, or

a distance weighting function, to be used in rendering the at least one ambience audio signal with a 6-degrees-of-freedom or an enhanced 3-degrees-of-freedom renderer, based on the at least one parameter of the at least one ambience component representation, the at least one of: the rendering position or the rendering direction, and the at least one respective diffuse background audio signal.

4 . The apparatus of claim 1 , wherein generating the at least one ambience component representation comprises the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:

obtain at least two audio signals captured with a microphone array;

analyze the at least two audio signals to determine at least one energy parameter;

obtain at least one close audio signal associated with an audio source; and

remove directional audio components associated with the at least one close audio signal from the at least one energy parameter to generate the at least one parameter.

5 . The apparatus of claim 4 , wherein generating the at least one ambience component representation comprises the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:

generate the at least one respective diffuse background audio signal of the at least one ambience component representation based on the at least two audio signals captured with the microphone array and the at least one close audio signal.

6 . The apparatus of claim 5 , wherein generating the at least one respective diffuse background audio signal comprises the at least one memory stores instructions that, when executed cause the apparatus to at least one of:

downmix the at least two audio signals captured with the microphone array;

select at least one audio signal from the at least two audio signals captured with the microphone array; or

beamform the at least two audio signals captured with the microphone array.

7 . A method comprising:

generating at least representation, wherein representation comprises:

one ambience component the ambience component

at least one respective diffuse background audio signal, and at least one parameter, wherein the at least one parameter is associated with the at least one respective diffuse background audio signal; and

outputting the at least one ambience component representation, wherein the at least one output ambience component representation is configured to be used in rendering at least one ambience audio signal based on processing of:

the at least one respective diffuse background audio signal,

the at least one parameter of the at least one ambience component representation, and

at least one of: a rendering position, or a rendering direction, relative to a defined position within an audio field.

8 . The method of claim 7 , wherein the at least one parameter is associated with:

at least one frequency range or at least one part of the at least one frequency range,

at least one time period or at least one part of the at least one time period, and

a directional range for the defined position, wherein the directional range defines a range of angles.

9 . The method of claim 7 , wherein the at least one ambience component representation further comprises at least one of:

a minimum distance threshold, over which the at least one ambience component representation is configured to be used in rendering the at least one ambience audio signal,

a maximum distance threshold, under which the at least one ambience component representation is configured to be used in rendering the at least one ambience audio signal, or

a distance weighting function, to be used in rendering the at least one ambience audio signal with a 6-degrees-of-freedom or an enhanced 3-degrees-of-freedom renderer, based on the at least one parameter of the at least one ambience component representation, the at least one of: the rendering position or the rendering direction, and the at least one respective diffuse background audio signal.

10 . The method of claim 7 , wherein the generating of the at least one ambience component representation comprises:

obtaining at least two audio signals captured with a microphone array;

analyzing the at least two audio signals to determine at least one energy parameter;

obtaining at least one close audio signal associated with an audio source; and

removing directional audio components associated with the at least one close audio signal from the at least one energy parameter to generate the at least one parameter.

11 . The method of claim 10 , wherein the generating of the at least one ambience component representation comprises:

generating the at least one respective diffuse background audio signal of the at least one ambience component representation based on the at least two audio signals captured with the microphone array and the at least one close audio signal.

12 . The method of claim 11 , wherein the generating of the at least one respective diffuse background audio signal comprises at least one of:

downmixing the at least two audio signals captured with the microphone array;

selecting at least one audio signal from the at least two audio signals captured with the microphone array; or

beamforming the at least two audio signals captured with the microphone array.

13 . An apparatus comprising:

at least one processor; and

at least one memory storing instructions that, when executed with the at least one processor, cause the apparatus at least to:

obtain at least one ambience component representation, wherein the ambience component representation comprises:

at least one respective diffuse background audio signal, and

at least one parameter, wherein the at least one parameter is associated with the at least one respective diffuse background audio signal;

obtain at least one of: a rendering position, or a rendering direction, relative to a defined position within an audio field; and

render at least one ambience audio signal, comprising processing the at least one respective diffuse background audio signal based on:

the at least one parameter, and

the at least one of: the rendering position, or the rendering direction, relative to the defined position within the audio field.

14 . The apparatus of claim 13 , wherein the at least one parameter is associated with:

at least one frequency range or at least one part of the at least one frequency range,

at least one time period or at least one part of the at least one time period, and

a directional range for the defined position, wherein the directional range defines a range of angles.

15 . The apparatus of claim 14 , wherein the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:

determine the at least one of: the rendering position, or the rendering direction, within the audio field;

wherein rendering the at least one ambience audio signal comprises the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:

render the at least one ambience audio signal based on the at least one of: the rendering position, or the rendering direction, being within a directional range.

16 . The apparatus of claim 13 , wherein obtaining the at least one of: the rendering position, or the rendering direction, comprises the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:

obtain the at least one of: the rendering position, or the rendering direction, within a 6-degrees-of-freedom or an enhanced 3-degrees-of-freedom audio field;

wherein rendering the at least one ambience audio signal comprises the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:

render the at least one ambience audio signal based on the at least one parameter and the at least one of: the rendering position, or the rendering direction, within the 6-degrees-of-freedom or the enhanced 3-degrees-of-freedom audio field.

17 . The apparatus of claim 16 , wherein rendering the at least one ambience audio signal comprises the at least one memory stores instructions that, when executed with the at least one processor, cause the apparatus to:

render the at least one ambience audio signal based on a distance defined with the at least one of: the rendering position, or the rendering direction, within the audio field being over a minimum distance threshold;

render the at least one ambience audio signal based on the distance defined with the at least one of: the rendering position, or the rendering direction, within the audio field being under a maximum distance threshold; and

render the at least one ambience audio signal based on a distance weighting function applied to the distance defined with the at least one of: the rendering position, or the rendering direction, within the audio field.

18 . A method comprising:

obtaining at least one ambience component representation, wherein the ambience component representation comprises:

at least one respective diffuse background audio signal, and

at least one parameter, wherein the at least one parameter is associated with the at least one respective diffuse background audio signal;

obtaining at least one of: a rendering position, or a rendering direction, relative to a defined position within an audio field; and

rendering at least one ambience audio signal, comprising processing the at least one respective diffuse background audio signal based on:

the at least one parameter, and

the at least one of: the rendering position, or the rendering direction, relative to the defined position within the audio field.

19 . The method of claim 18 , wherein the at least one parameter is associated with:

at least one frequency range or at least one part of the at least one frequency range,

at least one time period or at least one part of the at least one time period, and

a directional range for the defined position, wherein the directional range defines a range of angles.

20 . The method of claim 18 , wherein the obtaining of the at least one of: the rendering position, or the rendering direction, comprises:

obtaining the at least one of: the rendering position, or the rendering direction, within a 6-degrees-of-freedom or an enhanced 3-degrees-of-freedom audio field;

wherein the rendering of the at least one ambience audio signal comprises:

rendering the at least one ambience audio signal based on the at least one parameter and the at least one of: the rendering position, or the rendering direction, within the 6-degrees-of-freedom or the enhanced 3-degrees-of-freedom audio field.