IP Library Granted Patent US 10,419,867
Granted Patent B2
US 10,419,867 · App. 16/034,373 · Granted Sep 17, 2019

Device and method for processing audio signal

Inventors: Jeonghun Seo (Seoul, KR); Taegyu Lee (Seoul, KR); Hyun Oh Oh (Seongnam-si, KR)
Assignee: GAUDIO LAB, INC.
H04S7/303H04R5/02H04R5/033H04S3/00H04S3/008H04S2400/01H04S2400/11H04S2420/01H04S2420/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,419,867
App. No.
16/034,373
Granted
Sep 17, 2019
Kind
B2
Abstract

The present invention relates to an apparatus and a method for processing an audio signal, and more particularly, to an apparatus and a method for efficiently rendering a higher order ambisonics signal. To this end, provided are an audio signal processing apparatus, including: a pre-processor configured to separate an input audio signal into a first component corresponding to at least one object signal and a second component corresponding to a residual signal and extract position vector information corresponding to the first component from the input audio signal; a first rendering unit configured to perform an object-based first rendering on the first component using the position vector information; and a second rendering unit configured to perform a channel-based second rendering on the second component and an audio signal processing method using the same.

Claims (63)

1. An audio signal processing apparatus, the apparatus comprising:

a pre-processor configured to separate an input audio signal into a first component corresponding to at least one object signal and a second component corresponding to a residual signal and extract position vector information corresponding to the first component from the input audio signal, wherein the input audio signal comprises higher order ambisonics (HOA) coefficients, and wherein the position vector information is obtained by decomposing the HOA coefficients into a first matrix representing a plurality of audio signals and a second matrix representing position vector information of each of the plurality of audio signals;

a first rendering unit configured to perform an object-based first rendering on the first component using position vector information of the second matrix corresponding to the first component; and

a second rendering unit configured to perform a channel-based second rendering on the second component,

wherein the first component is extracted from audio signals having a level equal to or higher than a threshold value among the plurality of audio signals represented by the first matrix.

2. The apparatus of claim 1 , wherein the pre-processor performs a matrix decomposition of the HOA coefficients using singular value decomposition (SVD).

3. The apparatus of claim 1 ,

wherein the first rendering is an object-based binaural rendering, and

wherein the first rendering unit performs the first rendering using a head related transfer function (HRTF) based on the position vector information corresponding to the first component.

4. The apparatus of claim 1 ,

wherein the second rendering is a channel-based binaural rendering, and

wherein the second rendering unit maps the second component to at least one virtual channel and performs the second rendering using an HRTF based on the mapped virtual channel.

5. The apparatus of claim 1 , wherein the first rendering unit performs the first rendering by referring to spatial information of at least one object obtained from a video signal corresponding to the input audio signal.

6. The apparatus of claim 5 , wherein the first rendering unit modifies at least one parameter related to the first component based on the spatial information obtained from the video signal, and performs an object-based rendering on the first component using the modified parameter.

7. An audio signal processing method, the method comprising:

separating an input audio signal into a first component corresponding to at least one object signal and a second component corresponding to a residual signal, wherein the input audio signal comprises higher order ambisonics (HOA) coefficients;

extracting position vector information corresponding to the first component from the input audio signal, wherein the position vector information is obtained by decomposing the HOA coefficients into a first matrix representing a plurality of audio signals and a second matrix representing position vector information of each of the plurality of audio signals;

performing an object-based first rendering on the first component using position vector information of the second matrix corresponding to the first component; and

performing a channel-based second rendering on the second component,

wherein the first component is extracted from audio signals having a level equal to or higher than a threshold value among the plurality of audio signals represented by the first matrix.

8. The method of claim 7 , further comprising

performing a matrix decomposition of the HOA coefficients using singular value decomposition (SVD).

9. The method of claim 7 ,

wherein the first rendering is an object-based binaural rendering, and

wherein the first rendering is performed using a head related transfer function (HRTF) based on the position vector information corresponding to the first component.

10. The method of claim 7 ,

wherein the second rendering is a channel-based binaural rendering, and

wherein the second rendering is performed by mapping the second component to at least one virtual channel and using an HRTF based on the mapped virtual channel.

11. The method of claim 7 , wherein the first rendering is performed by referring to spatial information of at least one object obtained from a video signal corresponding to the input audio signal.

12. The method of claim 11 , wherein performing the first rendering further comprises:

modifying at least one parameter related to the first component based on the spatial information obtained from the video signal; and

performing an object-based rendering on the first component using the modified parameter.

13. An audio signal processing apparatus, the apparatus comprising:

a pre-processor configured to separate an input audio signal into a first component corresponding to at least one object signal and a second component corresponding to a residual signal and extract position vector information corresponding to the first component from the input audio signal, wherein the input audio signal comprises higher order ambisonics (HOA) coefficients, and wherein the position vector information is obtained by decomposing the HOA coefficients into a first matrix representing a plurality of audio signals and a second matrix representing position vector information of each of the plurality of audio signals;

a first rendering unit configured to perform an object-based first rendering on the first component using position vector information of the second matrix corresponding to the first component; and

a second rendering unit configured to perform a channel-based second rendering on the second component,

wherein the first component is extracted from coefficients of a predetermined low order among the HOA coefficients.

14. The apparatus of claim 13 , wherein the pre-processor performs a matrix decomposition of the HOA coefficients using singular value decomposition (SVD).

15. The apparatus of claim 13 ,

wherein the first rendering is an object-based binaural rendering, and

wherein the first rendering unit performs the first rendering using a head related transfer function (HRTF) based on the position vector information corresponding to the first component.

16. The apparatus of claim 13 ,

wherein the second rendering is a channel-based binaural rendering, and

wherein the second rendering unit maps the second component to at least one virtual channel and performs the second rendering using an HRTF based on the mapped virtual channel.

17. The apparatus of claim 13 , wherein the first rendering unit performs the first rendering by referring to spatial information of at least one object obtained from a video signal corresponding to the input audio signal.

18. The apparatus of claim 17 , wherein the first rendering unit modifies at least one parameter related to the first component based on the spatial information obtained from the video signal, and performs an object-based rendering on the first component using the modified parameter.

19. An audio signal processing method, the method comprising:

separating an input audio signal into a first component corresponding to at least one object signal and a second component corresponding to a residual signal, wherein the input audio signal comprises higher order ambisonics (HOA) coefficients;

extracting position vector information corresponding to the first component from the input audio signal, wherein the position vector information is obtained by decomposing the HOA coefficients into a first matrix representing a plurality of audio signals and a second matrix representing position vector information of each of the plurality of audio signals;

performing an object-based first rendering on the first component using position vector information of the second matrix corresponding to the first component; and

performing a channel-based second rendering on the second component,

wherein the first component is extracted from coefficients of a predetermined low order among the HOA coefficients.

20. The method of claim 19 , further comprising performing a matrix decomposition of the HOA coefficients using singular value decomposition (SVD).

21. The method of claim 19 ,

wherein the first rendering is an object-based binaural rendering, and

wherein the first rendering is performed using a head related transfer function (HRTF) based on the position vector information corresponding to the first component.

22. The method of claim 19 ,

wherein the second rendering is a channel-based binaural rendering, and

wherein the second rendering is performed by mapping the second component to at least one virtual channel and using an HRTF based on the mapped virtual channel.

23. The method of claim 19 , wherein the first rendering is performed by referring to spatial information of at least one object obtained from a video signal corresponding to the input audio signal.

24. The method of claim 23 , wherein performing the first rendering further comprises:

modifying at least one parameter related to the first component based on the spatial information obtained from the video signal; and

performing an object-based rendering on the first component using the modified parameter.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2019
From: GAUDIO LAB, INC.
To: GAUDIO LAB, INC.
Reel/Frame 051155/0142 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2018
From: SEO, JEONGHUN; LEE, TAEGYU; OH, HYUN OH
To: GAUDIO LAB, INC.
Reel/Frame 046340/0110 →
Priority Claims (1)
KR 10-2016-0006650 · Jan 19, 2016 · national
Continuity (2)
Continuation PCTKR2017000633 · Jan 19, 2017
Related Publication 20180324542A1 · Nov 8, 2018
Cited By (4)
US 12,413,929 US 12,425,792 US 12,574,696 US 12,647,742