AUDIO SCENE PROCESSING APPARATUS
An apparatus comprising: an audio source selector configured to select a set of audio signals from received audio signals; an audio source classifier configured to classify each of the set of audio signals dependent on at least one audio characteristic; and a classification selector configured to select from the set of audio signals at least one audio signal dependent on the audio characteristic.
1 - 81 . (canceled)
82 . Apparatus comprising at least one processor and at least one memory including computer code, the at least one memory and the computer code configured to with the at least one processor cause the apparatus to at least:
select a set of audio signals from received audio signals;
determine at least one audio characteristic value associated with each of the set of the audio signals compared to an associated reference signal;
classify each of the set of audio signals dependent on the at least one audio characteristic value associated with each audio signal;
determine for each received audio signal a location estimation; and
select the set of audio signals from the received audio signals dependent on the location estimation associated with the received audio signal.
83 . The apparatus as claimed in claim 82 , wherein select the set of audio signals from the received audio signals dependent on the location estimation associated with the received audio signal causes the apparatus to:
select the set of audio signals from the received audio signals dependent on the location estimation being within a determined audio scene area.
84 . The apparatus as claimed in claim 82 , wherein classify each of the set of audio signals dependent on the at least one audio characteristic value associated with each audio signal causes the apparatus to map the at least one audio characteristic value associated with each audio signal to one of a defined number of audio characteristic levels.
85 . The apparatus as claimed in claim 84 , wherein map the at least one audio characteristic value associated with each audio signal to one of the defined number of audio characteristic levels causes the apparatus to:
map a first audio characteristic value associated with each audio signal to one of a first defined number of levels associated with the first classification;
map a second audio characteristic value associated with each audio signal to one of a second number of levels associated with the second classification; and
combine the first classification mapping level and the second classification mapping level.
86 . The apparatus as claimed in claim 85 , wherein combine the first characteristic value mapping level and the second characteristic value mapping level causes the apparatus to average the first characteristic value mapping level and the second characteristic value mapping level.
87 . The apparatus as claimed in claim 82 , wherein determine at least one audio characteristic value associated with each of the set of the audio signals compared to a reference signal causes the apparatus to at least one of:
determine a spectral distance associated with each of the set of audio signals compared to the associated reference signal; and
determine a frequency response distance with each of the set of audio signals compared to the associated reference signal.
88 . The apparatus as claimed in claim 87 , wherein determine a spectral distance associated with each of the set of audio signals compared to the associated reference signal causes the apparatus to determine the spectral distance, Xdist, for each audio signal, x m , according to the following equations:
X
diff
m
[
k
,
r
]
=
∑
tl
=
l
·
L
(
l
+
1
)
·
L
-
1
(
X
ref
[
k
,
r
]
-
X
tile
m
[
k
,
tl
]
)
,
Xdist
m
[
sb
,
r
]
=
X
diff
m
[
binIdx
,
r
]
2
,
sbOffset
[
sb
]
≤
binIdx
<
sbOffset
[
sb
+
1
]
,
where sbOffset describes frequency band boundaries,
X
ref
[
k
,
r
]
=
∑
tl
=
l
·
L
(
l
+
1
)
·
L
-
1
∑
m
=
90
N
-
1
X
tile
m
[
k
.
tl
]
N
,
X
tile
m
[
k
,
r
]
=
X
m
[
k
,
tl
]
,
l
·
L
≤
tl
<
(
l
+
1
)
·
L
,
where
r
=
1
,
2
,
3
,
…
for
every
t
1
=
L
,
2
L
,
3
L
,
X
m
[
k
,
l
]
=
TF
(
x
m
,
l
,
T
)
,
where m is the signal index, k is a frequency bin index, l is a time frame index, T is a hop size between successive segments and TF is a time to frequency operator.
89 . The apparatus as claimed in claim 87 , wherein determine a frequency response distance with each of the set of audio signals compared to the associated reference signal causes the apparatus to determine the difference signal, Xdist, for each audio signal, x m , according to the following equations:
X
diff
m
[
sb
,
r
]
=
∑
binIdx
=
sbOffset
[
sb
]
sbOffset
[
sb
+
1
]
-
1
∑
tl
=
l
·
L
(
l
+
1
)
·
L
-
1
X
tile
m
[
binIdx
,
tl
]
2
,
Xdist
m
[
sb
,
r
]
=
X
ref
[
sb
,
r
]
-
X
diff
m
[
sb
,
r
]
,
where sbOffset describes frequency band boundaries,
X
ref
[
sb
,
r
]
=
∑
binIdx
=
sbOffset
[
sb
]
sbOffset
[
sb
+
1
]
-
1
∑
tl
=
l
·
L
(
l
+
1
)
·
L
-
1
∑
m
=
0
N
-
1
X
tile
m
[
binIdx
,
tl
]
2
,
X
tile
m
[
k
,
r
]
=
X
m
[
k
,
tl
]
,
l
·
L
≤
tl
<
(
l
+
1
)
·
L
,
where r=1, 2, 3, . . . for every tl=L, 2L, 3L, X in [k,l]=TF(x m,l,T ), where m is the signal index, k is a frequency bin index, l is a time frame index, T is a hop size between successive segments and TF is a time to frequency operator.
90 . The apparatus as claimed in claim 82 , wherein classify each of the set of audio signals dependent on the at least one classification value associated with each audio signal causes the apparatus to:
further classify the each of the set of audio signals dependent on an orientation of the audio signal, and wherein select from the set of audio signals at least one audio signal dependent on the audio characteristic causes the apparatus to select from the set of audio signals at least one audio signal dependent on the characteristic value mapping level
91 . The apparatus as claimed in claim 82 , further caused to receive at least one audio scene parameter, wherein the audio scene parameter comprises at least one of:
an audio scene location;
an audio scene area;
an audio scene radius;
an audio scene direction; and
an audio scene perceptual relevance.
92 . A method comprising:
selecting a set of audio signals from received audio signals;
determining at least one audio characteristic value associated with each of the set of the audio signals compared to an associated reference signal;
classifying each of the set of audio signals dependent on the at least one audio characteristic value associated with each audio signal;
determining for each received audio signal a location estimation; and
selecting the set of audio signals from the received audio signals dependent on the location estimation associated with the received audio signal.
93 . The method as claimed in claim 92 , wherein selecting the set of audio signals from the received audio signals dependent on the location estimation associated with the received audio signal comprises:
selecting the set of audio signals from the received audio signals dependent on the location estimation being within a determined audio scene area.
94 . The method as claimed in claim 92 , wherein classifying each of the set of audio signals dependent on the at least one audio characteristic value associated with each audio signal comprises mapping the at least one audio characteristic value associated with each audio signal to one of a defined number of audio characteristic levels.
95 . The method as claimed in claim 94 , wherein mapping the at least one audio characteristic value associated with each audio signal to one of the defined number of audio characteristic levels comprises:
mapping a first audio characteristic value associated with each audio signal to one of a first defined number of levels associated with the first classification;
mapping a second audio characteristic value associated with each audio signal to one of a second number of levels associated with the second classification; and
combining the first classification mapping level and the second classification mapping level.
96 . The method as claimed in claim 95 , wherein combining the first characteristic value mapping level and the second characteristic value mapping level comprises averaging the first characteristic value mapping level and the second characteristic value mapping level.
97 . The method as claimed in claim 82 , wherein determining at least one audio characteristic value associated with each of the set of the audio signals compared to a reference signal comprises at least one of:
determining a spectral distance associated with each of the set of audio signals compared to the associated reference signal; and
determining a frequency response distance with each of the set of audio signals compared to the associated reference signal.
98 . The method as claimed in claim 97 , wherein determining a spectral distance associated with each of the set of audio signals compared to the associated reference signal comprises determining the spectral distance, Xdist, for each audio signal, x m , according to the following equations:
X
diff
m
[
k
,
r
]
=
∑
tl
=
l
·
L
(
l
+
1
)
·
L
-
1
(
X
ref
[
k
,
r
]
-
X
tile
m
[
k
,
tl
]
)
,
Xdist
m
[
sb
,
r
]
=
X
diff
m
[
binIdx
,
r
]
2
,
sbOffset
[
sb
]
≤
binIdx
<
sbOffset
[
sb
+
1
]
,
where sbOffset describes frequency band boundaries,
X
ref
[
k
,
r
]
=
∑
tl
=
l
·
L
(
l
+
1
)
·
L
-
1
∑
m
=
0
N
-
1
X
tile
m
[
k
.
tl
]
N
,
X
tile
m
[
k
,
r
]
=
X
m
[
k
,
tl
]
,
l
·
L
≤
tl
<
(
l
+
1
)
·
L
,
where r=1, 2, 3, . . . for every tl=L, 2L, 3L, X m [k,l]=TF(x m,l,T ), where m is the signal index, k is a frequency bin index, l is a time frame index, T is a hop size between successive segments and TF is a time to frequency operator.
99 . The method as claimed in claim 97 , wherein determining a frequency response distance with each of the set of audio signals compared to the associated reference signal comprises determining the difference signal, Xdist, for each audio signal, x m , according to the following equations:
X
diff
m
[
sb
,
r
]
=
∑
binIdx
=
sbOffset
[
sb
]
sbOffset
[
sb
+
1
]
-
1
∑
tl
=
l
·
L
(
l
+
1
)
·
L
-
1
X
tile
m
[
binIdx
,
tl
]
2
,
Xdist
m
[
sb
,
r
]
=
X
ref
[
sb
,
r
]
-
X
diff
m
[
sb
,
r
]
,
where sbOffset describes frequency band boundaries,
X
ref
[
sb
,
r
]
=
∑
binIdx
=
sbOffset
[
sb
]
sbOffset
[
sb
+
1
]
-
1
∑
tl
=
l
·
L
(
l
+
1
)
·
L
-
1
∑
m
=
0
N
-
1
X
tile
m
[
binIdx
,
tl
]
2
,
X
tile
m
[
k
,
r
]
=
X
m
[
k
,
tl
]
,
l
·
L
≤
tl
<
(
l
+
1
)
·
L
,
where r=1, 2, 3, . . . for every tl=L, 2L, 3L, X m [k,l]=TF(x m,l,T ), where m is the signal index, k is a frequency bin index, l is a time frame index, T is a hop size between successive segments and TF is a time to frequency operator.
100 . The method as claimed in claim 92 , wherein classifying each of the set of audio signals dependent on the at least one classification value associated with each audio signal comprises:
further classifying the each of the set of audio signals dependent on an orientation of the audio signal, wherein selecting from the set of audio signals at least one audio signal dependent on the audio characteristic comprises selecting from the set of audio signals at least one audio signal dependent on the characteristic value mapping level
101 . The method as claimed in claim 92 , further comprising receiving at least one audio scene parameter, wherein the audio scene parameter comprises at least one of:
an audio scene location;
an audio scene area;
an audio scene radius;
an audio scene direction; and
an audio scene perceptual relevance.