IP Library Granted Patent US 12,586,591
Granted Patent B2
US 12,586,591 · App. 18/011,128 · Granted Mar 24, 2026

Sound signal decoding method, sound signal decoder, program, and recording medium

Inventors: Takehiro Moriya (Tokyo, JP); Ryosuke Sugiura (Tokyo, JP); Yutaka Kamamoto (Tokyo, JP)
Assignee: NTT, Inc.
G10L19/008
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,591
App. No.
18/011,128
Granted
Mar 24, 2026
Kind
B2
Abstract

The stereo decoding unit 220 performs steps S 222 - 1 and S 222 - 2 below (step S 222 ). The stereo decoding unit 220 obtains a signal concatenating a sum signal (a signal configured by addition of sample values of corresponding samples) of the monaural decoded sound signal for the section Y and the additional decoded signal for the section Y and the additional decoded signal for the section X, as a decoded downmix signal for the section Y+X (step S 222 - 1 ) instead of step S 221 - 1 performed by the stereo decoding unit 220 of the first embodiment and obtains and outputs the decoded sound signals of the two channels from the decoded downmix signal obtained at step S 222 - 1 by the upmix processing using the characteristic parameter obtained from the stereo code CS, using the decoded downmix signal obtained at step S 222 - 1 instead of the decoded downmix signal obtained at step S 221 - 1 (step S 222 - 2 ).

Claims (48)

1 . A sound signal decoding method for decoding an inputted code representing an encoded sound signal for each time frame to obtain a decoded sound signal having C channels (C is an integer of 2 or larger), the sound signal decoding method comprising:

as processing for a current frame,

decoding a monaural code included in the inputted code by a decoding scheme that includes processing of applying a window having overlap between frames to obtain a monaural decoded sound signal;

decoding an additional code included in the inputted code to obtain an additional decoded signal, wherein the additional decoded signal represents a monaural decoded signal of a section corresponding to the overlap between the current frame and an immediately following frame (hereinafter referred to as “a section X”);

obtaining a decoded downmix signal, wherein the decoded downmix signal represents a concatenation of a part of the monaural decoded sound signal of a section except the section X (hereinafter referred to as “a section Y”) and the additional decoded signal for the section X; and

obtaining and outputting the decoded sound signal having the C channels from a windowed signal of the decoded downmix signal by upmix processing using a characteristic parameter obtained from a stereo code included in the inputted code,

thereby a first time delay caused by the stereo decoding of the frame is not longer than a second time delay caused by the monaural decoding.

2 . The sound signal decoding method according to claim 1 , wherein

the decoding the additional code further comprises decoding the additional code included in the inputted code to obtain the additional decoded signal, and the additional decoded signal represents a monaural decoded signal for each of the section Y and the section X,

the decoded downmix signal represents a concatenation of a signal configured by addition or weighted addition of sample values of corresponding samples between the monaural decoded sound signal and the additional decoded signal for the section Y, and the additional decoded signal for the section X, and

the obtaining and outputting the decoded sound signal further comprises obtaining and outputting the decoded sound signal having the C channels from the windowed signal of the decoded downmix signal by the upmix processing using the characteristic parameter obtained from the stereo code included in the inputted code.

3 . The sound signal decoding method according to claim 2 , wherein

the decoding the additional code further comprises:

decoding a first additional code included in the additional code to obtain a first additional decoded signal for the section X,

decoding a second additional code included in the additional code by a decoding scheme of obtaining a collective sample sequence from a code to obtain a second additional decoded signal for each of the section Y and the section X,

obtaining the second additional decoded signal for the section Y as the additional decoded signal for the section Y, and

obtaining a signal configured by addition or weighted addition of sample values of corresponding samples between the first additional decoded signal and the second additional decoded signal for the section X, as the additional decoded signal for the section X.

4 . The sound signal decoding method according to claim 2 , wherein

the obtaining the decoded downmix signal further comprises:

obtaining a predicted signal for the section X from the monaural decoded sound signal for the section Y or the monaural decoded sound signals for the section Y and the section X,

concatenating the signal configured by addition or weighted addition of the sample values of the corresponding samples between the monaural decoded sound signal and the additional decoded signal for the section Y, and a signal configured by addition or weighted addition of sample values of corresponding samples between the predicted signal and the additional decoded signal for the section X to obtain the decoded downmix signal, and

the obtaining and outputting the decoded sound signal further comprises:

obtaining and outputting the decoded sound signal having the C channels from the windowed signal of the decoded downmix signal by the upmix processing using the characteristic parameter obtained from the stereo code included in the inputted code.

5 . A sound signal decoding device for decoding an inputted code representing an encoded sound signal for each time frame to obtain a decoded sound signal having C channels (C is an integer of 2 or larger), the sound signal decoding device comprising a processor configured to execute operations comprising, as processing for a current frame:

decoding a monaural code included in the inputted code by a decoding scheme that includes processing of applying a window having overlap between frames to obtain a monaural decoded sound signal, as processing for a current frame;

decoding an additional code included in the inputted code to obtain an additional decoded signal, wherein the additional decoded signal represents a monaural decoded signal of a section corresponding to the overlap between the current frame and an immediately following frame (hereinafter referred to as “a section X”), as processing for the current frame;

obtaining a decoded downmix signal, wherein the decoded downmix signal represents a concatenation of a part of the monaural decoded sound signal of a section except the section X (hereinafter referred to as “a section Y”) and the additional decoded signal for the section X; and

obtaining and output the decoded sound signal having the C channels from a windowed signal of the decoded downmix signal by upmix processing using a characteristic parameter obtained from a stereo code included in the inputted code,

thereby a first time delay caused by the stereo decoding of the frame is not longer than a second time delay caused by the monaural decoding.

6 . The sound signal decoding device according to claim 5 , wherein

the decoding the additional code further comprises decoding the additional code included in the inputted code to obtain the additional decoded signal, the additional decoded signal represents a monaural decoded signal for each of the section Y and the section X,

the decoded downmix signal represents a concatenation of a signal configured by addition or weighted addition of sample values of corresponding samples between the monaural decoded sound signal and the additional decoded signal for the section Y, and the additional decoded signal for the section X, and

the obtaining and outputting the decoded sound signal further comprises obtaining and outputting the decoded sound signal having the C channels from the windowed signal of the decoded downmix signal by the upmix processing using the characteristic parameter obtained from the stereo code included in the inputted code.

7 . The sound signal decoding device according to claim 6 ,

wherein

the decoding the additional code further comprises:

decoding a first additional code included in the additional code to obtain a first additional decoded signal for the section X,

decoding a second additional code included in the additional code by a decoding scheme of obtaining a collective sample sequence from a code to obtain a second additional decoded signal for each of the section Y and the section X,

obtaining the second additional decoded signal for the section Y as the additional decoded signal for the section Y, and

obtaining a signal configured by addition or weighted addition of sample values of corresponding samples between the first additional decoded signal and the second additional decoded signal for the section X, as the additional decoded signal for the section X.

8 . The sound signal decoding device according to claim 6 ,

wherein

the obtaining the decoded downmix signal further comprises:

obtaining a predicted signal for the section X from the monaural decoded sound signal for the section Y or the monaural decoded sound signals for the section Y and the section X,

obtaining the decoded downmix signal, which is a concatenation of the signal configured by addition or weighted addition of the sample values of the corresponding samples between the monaural decoded sound signal and the additional decoded signal for the section Y, and a signal configured by addition or weighted addition of sample values of corresponding samples between the predicted signal and the additional decoded signal for the section X, and

the obtaining and outputting the decoded sound signal further comprises:

obtaining and outputting the decoded sound signal having the C channels from the windowed signal of the decoded downmix signal by the upmix processing using the characteristic parameter obtained from the stereo code included in the inputted code.

9 . A non-transitory computer-readable recording medium in which a program for causing a computer to execute each step of the sound signal decoding method according to claim 1 is recorded.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0623 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2022
From: MORIYA, TAKEHIRO; SUGIURA, RYOSUKE; KAMAMOTO, YUTAKA
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 062132/0043 →
Continuity (1)
Related Publication 20230298598A1 · Sep 21, 2023
References Cited (12)
US 8332229B2 · Samsudin · 2012 [cited by examiner]
US 11657826B2 · Dick · 2023 [cited by examiner]
US 11832078B2 · Laitinen · 2023 [cited by examiner]
US 11962990B2 · Sen · 2024 [cited by examiner]
US 20100169102A1 · Samsudin · 2010 [cited by examiner]
US 20150380001A1 · Purnhagen et al. · 2015 [cited by applicant]
CN 108140393B · 2023 [cited by examiner]
S. Miyabe, T. Mihashi, T. Takatani, H. Saruwatari, K. Shikano and T. Nomura, “Compressive Coding of Stereo Audio Signals Extracting Sparseness among Sound Sources with Independent Component Analysis,” 2007 IEEE Workshop… [cited by examiner]
Breebaart et al., “Parametric Coding of Stereo Audio”, EURASIP Journal on Applied Signal Processing, pp. 1305-1322, 2005:9. [cited by applicant]
3GPP, “Codec for Enhanced Voice Services (EVS); Detailed algorithmic description”, TS 26.445. pp. 1-22, 270-282, 570-579. [cited by applicant]
International Standard, ISO/IEC 14496-3 (2009) “Information technology—Coding of audio-visual objects” [on-line] website: https://csclub.uwaterloo.ca/˜ehashman/ISO14496-3.2009.pdf. [cited by applicant]
Helmrich et al. (2015) “Low-complexity semi-parametric joint-stereo audio transform coding” 23rd European Signal Processing Conference, DOI: 10.1109/EUSIPCO.2015.7362492, pp. 799-803. [cited by applicant]