IP Library Granted Patent US 12,621,620
Granted Patent B2
US 12,621,620 · App. 18/685,762 · Granted May 5, 2026

Sound signal downmix method, sound signal coding method, sound signal downmix apparatus, sound signal coding apparatus, program

Inventors: Takehiro Moriya (Tokyo, JP); Yutaka Kamamoto (Tokyo, JP); Ryosuke Sugiura (Tokyo, JP)
Assignee: NTT, Inc.
H04S1/007G10L19/008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,621,620
App. No.
18/685,762
Granted
May 5, 2026
Kind
B2
Abstract

A sound signal downmixing method includes a step of obtaining, for each of two channels, a signal obtained by adding an input sound signal of one channel to a signal obtained by delaying an input sound signal of the other channel and multiplying the delayed input sound signal by a weight value as a delayed crosstalk-added signal of the one channel, a step of obtaining preceding channel information and a left-right correlation value, and step of obtaining a downmix signal by performing weighted addition on the input sound signals of the two channels based on the left-right correlation value and the preceding channel information such that more of a signal derived from an input sound signal of a preceding channel among the signals derived from the input sound signals of the two channels is included as the left-right correlation value becomes larger.

Claims (305)

1 . A sound signal downmixing method for obtaining a downmix signal that is a monaural sound signal from input sound signals of two channels, the method comprising:

a delayed crosstalk addition step of obtaining, for each of the two channels, a signal obtained by adding an input sound signal of one channel to a signal obtained by delaying an input sound signal of the other channel and multiplying the delayed input sound signal by a weight value that is a predetermined value having an absolute value smaller than 1, as a delayed crosstalk-added signal of the one channel;

a left-right relationship information acquisition step of obtaining preceding channel information that is information indicating which of the delayed crosstalk-added signals of the two channels is preceding and a left-right correlation value that is a value indicating a magnitude of correlation between the delayed crosstalk-added signals of the two channels; and

a downmixing step of obtaining the downmix signal by performing weighted addition on the input sound signals of the two channels based on the left-right correlation value and the preceding channel information such that more of a signal derived from an input sound signal of a preceding channel among the signals derived from the input sound signals of the two channels is included as the left-right correlation value becomes larger.

2 . The sound signal downmixing method according to claim 1 , wherein, in the delayed crosstalk addition step,

when the input sound signals of the two channels are respectively a left channel input sound signal and a right channel input sound signal, the delayed crosstalk-added signals of the two channels are respectively a left channel delayed crosstalk-added signal and a right channel delayed crosstalk-added signal, a sample number is t, each sample of the left channel input sound signal is x L (t), each sample of the right channel input sound signal is x R (t), each sample of the left channel delayed crosstalk-added signal is y L (t), each sample of the right channel delayed crosstalk-added signal is y R (t), predetermined positive values are a 1 and a2, and predetermined values having an absolute value smaller than 1 are w 1 and w2,

each sample y L (t) of the left channel delayed crosstalk-added signal is obtained by the following expression, and

[

Math

.

17

]

y

L

(

t

)

=

x

L

(

t

)

+

w

1

×

x

R

(

t

-

a

1

)

each sample y R (t) of the right channel delayed crosstalk-added signal is obtained by the following expression,

[

Math

.

18

]

y

R

(

t

)

=

x

R

(

t

)

+

w

2

×

x

L

(

t

-

a

2

)

.

3 . The sound signal downmixing method according to claim 1 , wherein, in the delayed crosstalk addition step,

when the input sound signals of the two channels are respectively a left channel input sound signal and a right channel input sound signal, the delayed crosstalk-added signals of the two channels are respectively a left channel delayed crosstalk-added signal and a right channel delayed crosstalk-added signal, a frequency number is k, each frequency spectrum sample of a frequency spectrum obtained by performing Fourier transform on the left channel input sound signal for each frame is X L (k), each frequency spectrum sample of a frequency spectrum obtained by performing Fourier transform on the right channel input sound signal for each frame is X R (k), each frequency spectrum sample of the left channel delayed crosstalk-added signal in a frequency domain for each frame is Y L (k), each frequency spectrum sample of the right channel delayed crosstalk-added signal in the frequency domain for each frame is Y R (k), predetermined positive values are a 1 and a 2 , and predetermined values having an absolute value smaller than 1 are w 1 and w 2 ,

each frequency spectrum sample Y L (k) of the left channel delayed crosstalk-added signal in the frequency domain for each frame is obtained by the following expression, and

[

Math

.

19

]

Y

L

(

k

)

=

X

L

(

k

)

+

w

1

×

X

R

(

k

)

×

e

-

j

2

a

1

π

T

k

each frequency spectrum sample Y R (k) of the right channel delayed crosstalk-added signal in the frequency domain for each frame is obtained by the following expression,

[

Math

.

20

]

Y

R

(

k

)

=

X

R

(

k

)

+

w

2

×

X

L

(

k

)

×

e

-

j

2

a

2

π

T

k

.

4 . A sound signal encoding method comprising the sound signal downmixing method according to claim 1 as a sound signal downmixing step,

wherein the sound signal encoding method further comprises:

a monaural encoding step of encoding the downmix signal obtained in the downmixing step to obtain a monaural code; and

a stereo encoding step of encoding the input sound signals of the two channels to obtain a stereo code.

5 . A non-transitory computer readable medium that stores a program for causing a computer to execute processing of each step of the sound signal encoding method according to claim 4 .

6 . A non-transitory computer readable medium that stores a program for causing a computer to execute processing of each step of the sound signal downmixing method according to claim 1 .

7 . A sound signal downmixing apparatus for obtaining a downmix signal that is a monaural sound signal from input sound signals of two channels, the sound signal downmixing apparatus comprising processing circuitry configured to:

obtain, for each of the two channels, a signal obtained by adding an input sound signal of one channel to a signal obtained by delaying an input sound signal of the other channel and multiplying the delayed input sound signal by a weight value that is a predetermined value having an absolute value smaller than 1, as a delayed crosstalk-added signal of the one channel;

obtain preceding channel information that is information indicating which of the delayed crosstalk-added signals of the two channels is preceding and a left-right correlation value that is a value indicating a magnitude of correlation between the delayed crosstalk-added signals of the two channels; and

obtain the downmix signal by performing weighted addition on the input sound signals of the two channels based on the left-right correlation value and the preceding channel information such that more of a signal derived from an input sound signal of a preceding channel among the signals derived from the input sound signals of the two channels is included as the left-right correlation value becomes larger.

8 . The sound signal downmixing apparatus according to claim 7 , wherein, in the processing circuitry,

when the input sound signals of the two channels are respectively a left channel input sound signal and a right channel input sound signal, the delayed crosstalk-added signals of the two channels are respectively a left channel delayed crosstalk-added signal and a right channel delayed crosstalk-added signal, a sample number is t, each sample of the left channel input sound signal is x L (t), each sample of the right channel input sound signal is x R (t), each sample of the left channel delayed crosstalk-added signal is y L (t), each sample of the right channel delayed crosstalk-added signal is y R (t), predetermined positive values are a 1 and a 2 , and predetermined values having an absolute value smaller than 1 are w 1 and w 2 , each sample y L (t) of the left channel delayed crosstalk-added signal is obtained by the following expression, and

[

Math

.

21

]

y

L

(

t

)

=

x

L

(

t

)

+

w

1

×

x

R

(

t

-

a

1

)

each sample y R (t) of the right channel delayed crosstalk-added signal is obtained by the following expression,

[

Math

.

22

]

y

R

(

t

)

=

x

R

(

t

)

+

w

2

×

x

L

(

t

-

a

2

)

.

9 . The sound signal downmixing apparatus according to claim 7 , wherein, in the processing circuitry,

when the input sound signals of the two channels are respectively a left channel input sound signal and a right channel input sound signal, the delayed crosstalk-added signals of the two channels are respectively a left channel delayed crosstalk-added signal and a right channel delayed crosstalk-added signal, a frequency number is k, each frequency spectrum sample of a frequency spectrum obtained by performing Fourier transform on the left channel input sound signal for each frame is X L (k), each frequency spectrum sample of a frequency spectrum obtained by performing Fourier transform on the right channel input sound signal for each frame is X R (k), each frequency spectrum sample of the left channel delayed crosstalk-added signal in a frequency domain for each frame is Y L (k), each frequency spectrum sample of the right channel delayed crosstalk-added signal in the frequency domain for each frame is Y R (k), predetermined positive values are a 1 and a 2 , and predetermined values having an absolute value smaller than 1 are w 1 and w 2 ,

each frequency spectrum sample Y L (k) of the left channel delayed crosstalk-added signal in the frequency domain for each frame is obtained by the following expression, and

[

Math

.

23

]

Y

L

(

k

)

=

X

L

(

k

)

+

w

1

×

X

R

(

k

)

×

e

-

j

2

a

1

π

T

k

each frequency spectrum sample Y R (k) of the right channel delayed crosstalk-added signal in the frequency domain for each frame is obtained by the following expression,

[

Math

.

24

]

Y

R

(

k

)

=

X

R

(

k

)

+

w

2

×

X

L

(

k

)

×

e

-

j

2

a

2

π

T

k

.

10 . A sound signal encoding apparatus comprising the sound signal downmixing apparatus according to claim 7 ,

wherein the sound signal encoding apparatus further comprises processing circuitry configured to:

encode the downmix signal obtained by the downmixing unit to obtain a monaural code; and

encode the input sound signals of the two channels to obtain a stereo code.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0623 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2024
From: MORIYA, TAKEHIRO; KAMAMOTO, YUTAKA; SUGIURA, RYOSUKE
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 066982/0584 →
Continuity (1)
Related Publication 20250126424A1 · Apr 17, 2025
References Cited (7)
US 3943293A · Bailey · 1976 [cited by examiner]
US 20080010072A1 · Yoshida · 2008 [cited by examiner]
US 20120072207A1 · Morii · 2012 [cited by examiner]
US 20140010375A1 · Usher · 2014 [cited by examiner]
EP 4120249B1 · 2023 [cited by applicant]
EP 4120250A1 · 2023 [cited by applicant]
WO 2006070751A1 · 2006 [cited by applicant]