IP Library Patent Application 13979791
Patent Application
App. No. 13/979,791

AUDIO SCENE PROCESSING APPARATUS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/979,791
Abstract

An apparatus comprising: an audio source selector configured to select a set of audio signals from received audio signals; an audio source classifier configured to classify each of the set of audio signals dependent on at least one audio characteristic; and a classification selector configured to select from the set of audio signals at least one audio signal dependent on the audio characteristic.

Claims (829)

1 - 81 . (canceled)

82 . Apparatus comprising at least one processor and at least one memory including computer code, the at least one memory and the computer code configured to with the at least one processor cause the apparatus to at least:

select a set of audio signals from received audio signals;

determine at least one audio characteristic value associated with each of the set of the audio signals compared to an associated reference signal;

classify each of the set of audio signals dependent on the at least one audio characteristic value associated with each audio signal;

determine for each received audio signal a location estimation; and

select the set of audio signals from the received audio signals dependent on the location estimation associated with the received audio signal.

83 . The apparatus as claimed in claim 82 , wherein select the set of audio signals from the received audio signals dependent on the location estimation associated with the received audio signal causes the apparatus to:

select the set of audio signals from the received audio signals dependent on the location estimation being within a determined audio scene area.

84 . The apparatus as claimed in claim 82 , wherein classify each of the set of audio signals dependent on the at least one audio characteristic value associated with each audio signal causes the apparatus to map the at least one audio characteristic value associated with each audio signal to one of a defined number of audio characteristic levels.

85 . The apparatus as claimed in claim 84 , wherein map the at least one audio characteristic value associated with each audio signal to one of the defined number of audio characteristic levels causes the apparatus to:

map a first audio characteristic value associated with each audio signal to one of a first defined number of levels associated with the first classification;

map a second audio characteristic value associated with each audio signal to one of a second number of levels associated with the second classification; and

combine the first classification mapping level and the second classification mapping level.

86 . The apparatus as claimed in claim 85 , wherein combine the first characteristic value mapping level and the second characteristic value mapping level causes the apparatus to average the first characteristic value mapping level and the second characteristic value mapping level.

87 . The apparatus as claimed in claim 82 , wherein determine at least one audio characteristic value associated with each of the set of the audio signals compared to a reference signal causes the apparatus to at least one of:

determine a spectral distance associated with each of the set of audio signals compared to the associated reference signal; and

determine a frequency response distance with each of the set of audio signals compared to the associated reference signal.

88 . The apparatus as claimed in claim 87 , wherein determine a spectral distance associated with each of the set of audio signals compared to the associated reference signal causes the apparatus to determine the spectral distance, Xdist, for each audio signal, x m , according to the following equations:

X

diff

m

[

k

,

r

]

=

tl

=

l

·

L

(

l

+

1

)

·

L

-

1

(

X

ref

[

k

,

r

]

-

X

tile

m

[

k

,

tl

]

)

,

Xdist

m

[

sb

,

r

]

=

X

diff

m

[

binIdx

,

r

]

2

,

sbOffset

[

sb

]

binIdx

<

sbOffset

[

sb

+

1

]

,

where sbOffset describes frequency band boundaries,

X

ref

[

k

,

r

]

=

tl

=

l

·

L

(

l

+

1

)

·

L

-

1

m

=

90

N

-

1

X

tile

m

[

k

.

tl

]

N

,

X

tile

m

[

k

,

r

]

=

X

m

[

k

,

tl

]

,

l

·

L

tl

<

(

l

+

1

)

·

L

,

where

r

=

1

,

2

,

3

,

for

every

t

1

=

L

,

2

L

,

3

L

,

X

m

[

k

,

l

]

=

TF

(

x

m

,

l

,

T

)

,

where m is the signal index, k is a frequency bin index, l is a time frame index, T is a hop size between successive segments and TF is a time to frequency operator.

89 . The apparatus as claimed in claim 87 , wherein determine a frequency response distance with each of the set of audio signals compared to the associated reference signal causes the apparatus to determine the difference signal, Xdist, for each audio signal, x m , according to the following equations:

X

diff

m

[

sb

,

r

]

=

binIdx

=

sbOffset

[

sb

]

sbOffset

[

sb

+

1

]

-

1

tl

=

l

·

L

(

l

+

1

)

·

L

-

1

X

tile

m

[

binIdx

,

tl

]

2

,

Xdist

m

[

sb

,

r

]

=

X

ref

[

sb

,

r

]

-

X

diff

m

[

sb

,

r

]

,

where sbOffset describes frequency band boundaries,

X

ref

[

sb

,

r

]

=

binIdx

=

sbOffset

[

sb

]

sbOffset

[

sb

+

1

]

-

1

tl

=

l

·

L

(

l

+

1

)

·

L

-

1

m

=

0

N

-

1

X

tile

m

[

binIdx

,

tl

]

2

,

X

tile

m

[

k

,

r

]

=

X

m

[

k

,

tl

]

,

l

·

L

tl

<

(

l

+

1

)

·

L

,

where r=1, 2, 3, . . . for every tl=L, 2L, 3L, X in [k,l]=TF(x m,l,T ), where m is the signal index, k is a frequency bin index, l is a time frame index, T is a hop size between successive segments and TF is a time to frequency operator.

90 . The apparatus as claimed in claim 82 , wherein classify each of the set of audio signals dependent on the at least one classification value associated with each audio signal causes the apparatus to:

further classify the each of the set of audio signals dependent on an orientation of the audio signal, and wherein select from the set of audio signals at least one audio signal dependent on the audio characteristic causes the apparatus to select from the set of audio signals at least one audio signal dependent on the characteristic value mapping level

91 . The apparatus as claimed in claim 82 , further caused to receive at least one audio scene parameter, wherein the audio scene parameter comprises at least one of:

an audio scene location;

an audio scene area;

an audio scene radius;

an audio scene direction; and

an audio scene perceptual relevance.

92 . A method comprising:

selecting a set of audio signals from received audio signals;

determining at least one audio characteristic value associated with each of the set of the audio signals compared to an associated reference signal;

classifying each of the set of audio signals dependent on the at least one audio characteristic value associated with each audio signal;

determining for each received audio signal a location estimation; and

selecting the set of audio signals from the received audio signals dependent on the location estimation associated with the received audio signal.

93 . The method as claimed in claim 92 , wherein selecting the set of audio signals from the received audio signals dependent on the location estimation associated with the received audio signal comprises:

selecting the set of audio signals from the received audio signals dependent on the location estimation being within a determined audio scene area.

94 . The method as claimed in claim 92 , wherein classifying each of the set of audio signals dependent on the at least one audio characteristic value associated with each audio signal comprises mapping the at least one audio characteristic value associated with each audio signal to one of a defined number of audio characteristic levels.

95 . The method as claimed in claim 94 , wherein mapping the at least one audio characteristic value associated with each audio signal to one of the defined number of audio characteristic levels comprises:

mapping a first audio characteristic value associated with each audio signal to one of a first defined number of levels associated with the first classification;

mapping a second audio characteristic value associated with each audio signal to one of a second number of levels associated with the second classification; and

combining the first classification mapping level and the second classification mapping level.

96 . The method as claimed in claim 95 , wherein combining the first characteristic value mapping level and the second characteristic value mapping level comprises averaging the first characteristic value mapping level and the second characteristic value mapping level.

97 . The method as claimed in claim 82 , wherein determining at least one audio characteristic value associated with each of the set of the audio signals compared to a reference signal comprises at least one of:

determining a spectral distance associated with each of the set of audio signals compared to the associated reference signal; and

determining a frequency response distance with each of the set of audio signals compared to the associated reference signal.

98 . The method as claimed in claim 97 , wherein determining a spectral distance associated with each of the set of audio signals compared to the associated reference signal comprises determining the spectral distance, Xdist, for each audio signal, x m , according to the following equations:

X

diff

m

[

k

,

r

]

=

tl

=

l

·

L

(

l

+

1

)

·

L

-

1

(

X

ref

[

k

,

r

]

-

X

tile

m

[

k

,

tl

]

)

,

Xdist

m

[

sb

,

r

]

=

X

diff

m

[

binIdx

,

r

]

2

,

sbOffset

[

sb

]

binIdx

<

sbOffset

[

sb

+

1

]

,

where sbOffset describes frequency band boundaries,

X

ref

[

k

,

r

]

=

tl

=

l

·

L

(

l

+

1

)

·

L

-

1

m

=

0

N

-

1

X

tile

m

[

k

.

tl

]

N

,

X

tile

m

[

k

,

r

]

=

X

m

[

k

,

tl

]

,

l

·

L

tl

<

(

l

+

1

)

·

L

,

where r=1, 2, 3, . . . for every tl=L, 2L, 3L, X m [k,l]=TF(x m,l,T ), where m is the signal index, k is a frequency bin index, l is a time frame index, T is a hop size between successive segments and TF is a time to frequency operator.

99 . The method as claimed in claim 97 , wherein determining a frequency response distance with each of the set of audio signals compared to the associated reference signal comprises determining the difference signal, Xdist, for each audio signal, x m , according to the following equations:

X

diff

m

[

sb

,

r

]

=

binIdx

=

sbOffset

[

sb

]

sbOffset

[

sb

+

1

]

-

1

tl

=

l

·

L

(

l

+

1

)

·

L

-

1

X

tile

m

[

binIdx

,

tl

]

2

,

Xdist

m

[

sb

,

r

]

=

X

ref

[

sb

,

r

]

-

X

diff

m

[

sb

,

r

]

,

where sbOffset describes frequency band boundaries,

X

ref

[

sb

,

r

]

=

binIdx

=

sbOffset

[

sb

]

sbOffset

[

sb

+

1

]

-

1

tl

=

l

·

L

(

l

+

1

)

·

L

-

1

m

=

0

N

-

1

X

tile

m

[

binIdx

,

tl

]

2

,

X

tile

m

[

k

,

r

]

=

X

m

[

k

,

tl

]

,

l

·

L

tl

<

(

l

+

1

)

·

L

,

where r=1, 2, 3, . . . for every tl=L, 2L, 3L, X m [k,l]=TF(x m,l,T ), where m is the signal index, k is a frequency bin index, l is a time frame index, T is a hop size between successive segments and TF is a time to frequency operator.

100 . The method as claimed in claim 92 , wherein classifying each of the set of audio signals dependent on the at least one classification value associated with each audio signal comprises:

further classifying the each of the set of audio signals dependent on an orientation of the audio signal, wherein selecting from the set of audio signals at least one audio signal dependent on the audio characteristic comprises selecting from the set of audio signals at least one audio signal dependent on the characteristic value mapping level

101 . The method as claimed in claim 92 , further comprising receiving at least one audio scene parameter, wherein the audio scene parameter comprises at least one of:

an audio scene location;

an audio scene area;

an audio scene radius;

an audio scene direction; and

an audio scene perceptual relevance.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2015
From: NOKIA CORPORATION
To: NOKIA TECHNOLOGIES OY
Reel/Frame 035457/0916 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2013
From: OJANPERA, JUHA
To: NOKIA CORPORATION
Reel/Frame 030798/0914 →