IP Library Granted Patent US 12,482,473
Granted Patent B2
US 12,482,473 · App. 18/011,132 · Granted Nov 25, 2025

Sound signal encoding method, sound signal encoder, program, and recording medium

Inventors: Takehiro Moriya (Tokyo, JP); Ryosuke Sugiura (Tokyo, JP); Yutaka Kamamoto (Tokyo, JP)
Assignee: NTT, Inc.
G10L19/008G10L19/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,473
App. No.
18/011,132
Granted
Nov 25, 2025
Kind
B2
Abstract

There is provided such embedded encoding that the algorithmic delay of stereo coding/decoding is not larger than that of monaural coding/decoding. An encoding device ( 100 ) encodes a sound signal having a plurality of channels. A stereo encoding unit ( 110 ) obtains and outputs a stereo code representing a characteristic of difference between channels of the sound signal. A downmix unit ( 150 ) obtains a signal by mixing the sound signal as a downmix signal. A monaural encoding unit ( 120 ) encodes the downmix signal by an encoding scheme that includes processing of applying a window having overlap between frames to obtain and output a monaural code. An additional encoding unit ( 130 ) encodes a part of the downmix signal for a section corresponding to the overlap between a current frame and an immediately following frame to obtain and output an additional code.

Claims (46)

1 . A sound signal encoding method for encoding an inputted sound signal of C channels (C is an integer of 2 or larger) for each frame, the sound signal encoding method comprising:

as processing for a current frame,

a stereo encoding step of obtaining and outputting a stereo code representing a characteristic parameter which is a parameter representing a characteristic of difference between channels of the sound signal having the C channels;

a downmix step of obtaining a signal obtained by mixing the sound signal having the C channels as a downmix signal; and

a monaural encoding step of encoding the downmix signal to obtain and output a monaural code, wherein

the monaural encoding step encodes the downmix signal by an encoding scheme that includes processing of applying a window having overlap between frames to obtain the monaural code, and

the sound signal encoding method further comprises an additional encoding step of encoding a part of the downmix signal for a section corresponding to the overlap between the current frame and an immediately following frame (hereinafter referred to as “a section X”) to obtain and output an additional code.

2 . The sound signal encoding method according to claim 1 , wherein

the monaural encoding step also obtains a monaural locally decoded signal corresponding to the monaural code; and

the additional encoding step encodes a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between a part of the downmix signal for a section except the section X (hereinafter referred to as “a section Y”) and the monaural locally decoded signal for the section Y, and the downmix signal for the section X to obtain the additional code.

3 . The sound signal encoding method according to claim 2 , wherein

the downmix step obtains a signal obtained by mixing the sound signal having the C channels in a frequency domain as the downmix signal;

the sound signal encoding method further comprises a monaural encoding target signal generation step of obtaining a signal by mixing the sound signal having the C channels in a time domain as a monaural encoding target signal; and

the monaural encoding step encodes the monaural encoding target signal by the encoding scheme that includes the processing of applying a window having overlap between frames to obtain the monaural code.

4 . The sound signal encoding method according to claim 2 , wherein

the additional encoding step

encodes the section X of the downmix signal to obtain a first additional code and a locally decoded signal for the section X (hereinafter referred to as “a first additional locally decoded signal”) corresponding to the first additional code,

encodes a signal obtained by concatenating a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the monaural locally decoded signal for the section Y and a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the first additional locally decoded signal for the section X, by an encoding scheme of collectively encoding a sample sequence, to obtain a second additional code, and

obtain a combination of the first additional code and the second additional code as the additional code.

5 . The sound signal encoding method according to claim 2 , wherein

the additional encoding step

obtains a predicted signal for the section X of the monaural locally decoded signal, from the monaural locally decoded signal for the section Y or the monaural locally decoded signal for the section Y and the section X, and

encodes the signal configured by subtraction or weighted subtraction between the sample values of the corresponding samples between the downmix signal and the monaural locally decoded signal for the section Y and a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the predicted signal for the section X to obtain the additional code.

6 . A non-transitory computer-readable recording medium in which a program for causing a computer to execute each step of the sound signal encoding method according to claim 1 is recorded.

7 . A sound signal encoding device for encoding an inputted sound signal having C channels (C is an integer of 2 or larger) for each frame, the sound signal encoding device comprising:

a stereo encoding unit configured to obtain and output a stereo code representing a characteristic parameter which is a parameter representing a characteristic of difference between channels of the sound signal having the C channels, as processing for a current frame;

a downmix unit configured to obtain a signal by mixing the sound signal having the C channels as a downmix signal, as processing for the current frame; and

a monaural encoding unit configured to encode the downmix signal to obtain and output a monaural code, as processing for the current frame, wherein

the monaural encoding unit encodes the downmix signal by an encoding scheme that includes processing of applying a window having overlap between frames to obtain the monaural code, and

the sound signal encoding device further comprises an additional encoding unit configured to encode a part of the downmix signal for a section corresponding to the overlap between the current frame and an immediately following frame (hereinafter referred to as “a section X”) to obtain and output an additional code.

8 . The sound signal encoding device according to claim 7 , wherein

the monaural encoding unit also obtains a monaural locally decoded signal corresponding to the monaural code; and

the additional encoding unit encodes a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between a part of the downmix signal for a section except the section X (hereinafter referred to as “a section Y”) and the monaural locally decoded signal for the section Y, and the downmix signal for the section X to obtain the additional code.

9 . The sound signal encoding device according to claim 8 , wherein

the downmix unit obtains a signal obtained by mixing the sound signal having the C channels in a frequency domain as the downmix signal;

the sound signal encoding device further comprises a monaural encoding target signal generation unit configured to obtain a signal by mixing the sound signal having the C channels in a time domain as a monaural encoding target signal; and

the monaural encoding unit encodes the monaural encoding target signal by the encoding scheme that includes the processing of applying a window having overlap between frames to obtain the monaural code.

10 . The sound signal encoding device according to claim 8 , wherein

the additional encoding unit

encodes the section X of the downmix signal to obtain a first additional code and a locally decoded signal for the section X (hereinafter referred to as “a first additional locally decoded signal”), corresponding to the first additional code,

encodes a signal obtained by concatenating a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the monaural locally decoded signal for the section Y and a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the first additional locally decoded signal for the section X, by an encoding scheme of collectively encoding a sample sequence, to obtain a second additional code, and

obtain a combination of the first additional code and the second additional code as the additional code.

11 . The sound signal encoding device according to claim 8 , wherein

the additional encoding unit

obtains a predicted signal for the section X of the monaural locally decoded signal, from the monaural locally decoded signal for the section Y or the monaural locally decoded signal for the section Y and the section X, and

encodes the signal configured by subtraction or weighted subtraction between the sample values of the corresponding samples between the downmix signal and the monaural locally decoded signal for the section Y and a signal configured by subtraction or weighted subtraction between sample values of corresponding samples between the downmix signal and the predicted signal for the section X to obtain the additional code.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0623 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2022
From: MORIYA, TAKEHIRO; SUGIURA, RYOSUKE; KAMAMOTO, YUTAKA
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 062132/0123 →
Continuity (1)
Related Publication 20230178086A1 · Jun 8, 2023
References Cited (5)
US 20070121953A1 · Wei · 2007 [cited by examiner]
US 20100169102A1 · Samsudin · 2010 [cited by examiner]
US 20110182433A1 · Takada · 2011 [cited by examiner]
Breebaart et al., “Parametric Coding of Stereo Audio”, EURASIP Journal on Applied Signal Processing, pp. 1305-1322. 2005:9. [cited by applicant]
3GPP, “Codec for Enhanced Voice Services (EVS); Detailed algorithmic description”, TS 26.445. pp. 1-22, 270-282, 570-579. [cited by applicant]