IP Library Granted Patent US 9,311,925
Granted Patent B2
US 9,311,925 · App. 13/500,871 · Granted Apr 12, 2016

Method, apparatus and computer program for processing multi-channel signals

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,311,925
App. No.
13/500,871
Granted
Apr 12, 2016
Kind
B2
Abstract

The invention relates to a method and an apparatus in which samples of at least a part of an audio signal of a first channel and a part of an audio signal of a second channel are used to produce a sparse representation of the audio signals to increase the encoding efficiency. In an example embodiment one or more audio signals are input and relevant auditory cues are determined in a time-frequency plane. The relevant auditory cues are combined to form an auditory neurons map. Said one or more audio signals are transformed into a transform domain and the auditory neurons map is used to form a sparse representation of said one or more audio signal.

Claims (55)

1. A method comprising:

inputting one or more audio signals for an audio scene;

determining relevant auditory cues that preserve detailed information about sound features over time, said determining comprising:

windowing said one or more audio signals, wherein said windowing comprises first and second windowings of different bandwidths to produce a first windowed audio signal and a second windowed audio signal respectively;

transforming the first and second windowed audio signals into a transform domain; and

calculating said auditory cues based on said first and second windowed audio signals;

forming an auditory neurons map comprising paths in the transform domain of the relevant auditory cues;

transforming said one or more audio signals into a transform domain;

using the auditory neurons map to form a sparse representation of said one or more transformed audio signals; and

outputting said sparse representation of said one or more transformed audio signals for at least one of encoding by an encoder and storing in a storage device.

2. The method according to claim 1 , wherein said first windowing comprises using two or more windows of a first type having different bandwidths, and wherein said second windowing comprises using two or more analysis windows of a second type having different bandwidths.

3. The method according to claim 2 , said determining further comprising, for each of said one or more audio signals:

combining transformed windowed audio signals resulting from the first windowing; and

combining transformed windowed audio signals resulting from the second windowing.

4. The method according to claim 1 , said determining further comprising combining the respective auditory cues determined for each of said one or more audio signals.

5. The method according to claim 1 , said using comprising determining auditory cue threshold values based on the auditory neurons map.

6. The method according to claim 5 , wherein said determining auditory cue threshold values further comprises adjusting threshold values in response to a transient signal segment.

7. The method according to claim 5 , wherein said sparse representation is determined based at least partly on said auditory cue threshold values.

8. The method according to claim 1 wherein said one or more audio signals comprises a multi-channel audio signal.

9. An apparatus comprising

at least one processor; and

at least one non-transitory memory comprising computer program code, the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to:

input one or more audio signals for an audio scene;

determine relevant auditory cues that preserve detailed information about sound features over time, said determining comprising:

windowing said one or more audio signals, wherein said windowing comprises first and second windowings of different bandwidths to produce a first windowed audio signal and a second windowed audio signal respectively;

transforming the first and second windowed audio signals into a transform domain; and

calculating said auditory cues based on said first and second windowed audio signals;

form an auditory neurons map comprising paths in the transform domain of the relevant auditory cues;

transform said one or more audio signals into a transform domain;

use the auditory neurons map to form a sparse representation of said one or more audio signals; and

output said sparse representation of said one or more transformed audio signals for at least one of encoding by an encoder and storing in a storage device.

10. The apparatus according to claim 9 , wherein said first windowing comprises using two or more windows of a first type having different bandwidths, and wherein said second windowing comprises using two or more analysis windows of a second type having different bandwidths.

11. The apparatus according to claim 10 , wherein said determining further comprises the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to, for each of said one or more audio signals:

combine transformed windowed audio signals resulting from the first windowing; and

combine transformed windowed audio signals resulting from the second windowing.

12. The apparatus according to claim 9 , wherein said determining further comprises the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to combine the respective auditory cues determined for each of said one or more audio signals.

13. The apparatus according to claim 9 , wherein said forming comprises the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to determine maxima of the respective relevant auditory cues.

14. The apparatus according to claim 9 , wherein said using comprises the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to determine auditory cue threshold values based on the auditory neurons map.

15. The apparatus according to claim 14 , wherein said determining auditory cue threshold values comprises the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to determine threshold values based on median of respective values of one or more auditory neurons maps.

16. The apparatus according to claim 14 , wherein said determining auditory cue threshold values further comprises the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to adjust threshold values in response to a transient signal segment.

17. The apparatus according to claim 9 , wherein said one or more audio signals comprises a multi-channel audio signal.

18. A non-transitory computer program product comprising a computer program code configured to, with at least one processor, cause an apparatus to:

input one or more audio signals for an audio scene;

determine relevant auditory cues that preserve detailed information about sound features over time, said determining comprising:

windowing said one or more audio signals, wherein said windowing comprises first and second windowings of different bandwidths to produce a first windowed audio signal and a second windowed audio signal respectively;

transforming the first and second windowed audio signals into a transform domain; and

calculating said auditory cues based on said first and second windowed audio signals;

form an auditory neurons map comprising paths in the transform domain of the relevant auditory cues;

transform said one or more audio signals into a transform domain; and

use the auditory neurons map to form a sparse representation of said one or more transformed audio signals; and

output said sparse representation of said one or more transformed audio signals for at least one of encoding by an encoder and storing in a storage device.

19. A method according to claim 1 , wherein:

said forming the auditory neurons map comprises determining paths of auditory cues in a time-frequency plane.

20. An apparatus according to claim 9 , wherein

said forming the auditory neurons map comprises determining paths of auditory cues in a time-frequency plane.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2015
From: NOKIA CORPORATION
To: NOKIA TECHNOLOGIES OY
Reel/Frame 035512/0432 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2012
From: OJANPERA, JUHA
To: NOKIA CORPORATION
Reel/Frame 028236/0601 →