IP Library › Granted Patent US 12,400,670
Granted Patent B2
US 12,400,670 · App. 18/614,837 · Granted Aug 26, 2025

Linear prediction analysis device, method, program, and storage medium

Inventors: Yutaka Kamamoto (Atsugi, JP); Takehiro Moriya (Atsugi, JP); Noboru Harada (Atsugi, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L19/06G10L19/0212G10L19/032G10L21/04G10L25/06G10L25/12G10L25/18G10L25/27
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,400,670
App. No.
18/614,837
Granted
Aug 26, 2025
Kind
B2
Abstract

An autocorrelation calculation unit 21 calculates an autocorrelation R O (i) from an input signal. A prediction coefficient calculation unit 23 performs linear prediction analysis by using a modified autocorrelation R′ O (i) obtained by multiplying a coefficient w O (i) by the autocorrelation R O (i). It is assumed here, for each order i of some orders i at least, that the coefficient w O (i) corresponding to the order i is in a monotonically increasing relationship with an increase in a value that is negatively correlated with a fundamental frequency of the input signal of the current frame or a past frame.

Claims (56)

1. A linear prediction analysis method of obtaining, in each frame, which is a predetermined time interval, coefficients to be transformed to linear prediction coefficients corresponding to an input time-series signal, the linear prediction analysis method comprising:

a step of receiving the input time-series signal, the time-series signal being a speech signal or an acoustic signal;

an autocorrelation calculation step of calculating an autocorrelation R O (i) between an input time-series signal X O (n) of a current frame and an input time-series signal X O (n−i) i samples before the input time-series signal X O (n) or an input time-series signal X O (n+i) i samples after the input time-series signal X O (n), for each i of i=0, 1, . . . , P max at least, where the current frame includes parts of adjacent frame; and

a prediction coefficient calculation step of calculating coefficients to be transformed to first-order to P max -order linear prediction coefficients, by using a modified autocorrelation R′ O (i) obtained by multiplying a coefficient w O (i) by the autocorrelation R O (i) for each i,

wherein a coefficient table t0 stores a coefficient w t0 (i) and a coefficient table t1 stores a coefficient w t1 (i), w t0 (i)<w t1 (i) being satisfied for at least part of i other than i=0, w t0 (i)≤w t1 (i) being satisfied for the remaining each i other than i=0,

the linear prediction analysis method further comprises a coefficient determination step of, by using a period, a quantized value of the period, an estimated value of the period or a value that is negatively correlated with a fundamental frequency based on the input time-series signal of the current frame or a past frame,

(1) obtaining the coefficient w t0 (i) as the coefficient w O (i) from the coefficient table t0 when the period, the quantized value of the period, the estimated value of the period or the value that is negatively correlated with the fundamental frequency is less than or equal to a predetermined threshold or less than the predetermined threshold, and

(2) obtaining the coefficient w t1 (i) as the coefficient w O (i) from the coefficient table t1 when the period, the quantized value of the period, the estimated value of the period or the value that is negatively correlated with the fundamental frequency is more than the predetermined threshold or more than or equal to the predetermined threshold, and

the linear prediction analysis method further includes encoding or analyzing the speech signal or the acoustic signal using the calculated coefficients to be transformed to first order to Pmax-order linear prediction coefficients.

2. A linear prediction analysis method of obtaining, in each frame, which is a predetermined time interval, coefficients to be transformed to linear prediction coefficients corresponding to an input time-series signal, the linear prediction analysis method comprising:

a step of receiving the input time-series signal, the time-series signal being a speech signal or an acoustic signal;

an autocorrelation calculation step of calculating an autocorrelation R O (i) between an input time-series signal X O (n) of a current frame and an input time-series signal X O (n−i) i samples before the input time-series signal X O (n) or an input time-series signal X O (n+i) i samples after the input time-series signal X O (n), for each i of i=0, 1, . . . , Pmax at least, where the current frame includes parts of adjacent frame; and

a prediction coefficient calculation step of calculating coefficients to be transformed to first-order to P max -order linear prediction coefficients, by using a modified autocorrelation R′ O (i) obtained by multiplying a coefficient w O (i) by the autocorrelation R O (i) for each i;

wherein a coefficient table t0 stores a coefficient w t0 (i) and a coefficient table t1 stores a coefficient w t1 (i), w t0 (i)<w t1 (i) being satisfied for at least part of i other than i=0, w t0 (i)≤w t1 (i) being satisfied for the remaining each i other than i=0,

the linear prediction analysis method further comprises a coefficient determination step of, by using a fundamental frequency, a quantized value of the fundamental frequency, an estimated value of the fundamental frequency or a value that is positively correlated with a fundamental frequency based on the input time-series signal of the current frame or a past frame,

(1) obtaining the coefficient w t0 (i) as the coefficient w O (i) from the coefficient table t0 when the fundamental frequency, the quantized value of the fundamental frequency, the estimated value of the fundamental frequency or the value that is positively correlated with the fundamental frequency is more than or equal to a predetermined threshold or more than the predetermined threshold, and

(2) obtaining the coefficient w t1 (i) as the coefficient w O (i) from the coefficient table t1 when the fundamental frequency, the quantized value of the fundamental frequency, the estimated value of the fundamental frequency or the value that is positively correlated with the fundamental frequency is less than the predetermined threshold or less than or equal to the predetermined threshold, and

the linear prediction analysis method further includes encoding or analyzing the speech signal or the acoustic signal using the calculated coefficients to be transformed to first order to P max -order linear prediction coefficients.

3. A linear prediction analysis device that obtains, in each frame, which is a predetermined time interval, coefficients to be transformed to linear prediction coefficients corresponding to an input time-series signal, the linear prediction analysis device comprising:

processing circuitry configured to

receive the input time-series signal, the time-series signal being a speech signal or an acoustic signal;

calculate an autocorrelation R O (i) between an input time-series signal X O (n) of a current frame and an input time-series signal X O (n−i) i samples before the input time-series signal X O (n) or an input time-series signal X O (n+i) i samples after the input time-series signal X O (n), for each i of i=0, 1 . . . , P max at least, where the current frame includes parts of adjacent frame; and

calculate coefficients to be transformed to first-order to P max -order linear prediction coefficients, by using a modified autocorrelation R′ O (i) obtained by multiplying a coefficient w O (i) by the autocorrelation R O (i) for each i;

wherein a coefficient table t0 stores a coefficient w t0 (i) and a coefficient table t1 stores a coefficient w t1 (i), w t0 (i)<w t1 (i) being satisfied for at least part of i other than i=0, w t0 (i)≤w t1 (i) being satisfied for the remaining each i other than i=0,

the processing circuitry is further configured to, by using a period, a quantized value of the period, an estimated value of the period or a value that is negatively correlated with a fundamental frequency based on the input time-series signal of the current frame or a past frame,

(1) obtain the coefficient w t0 (i) as the coefficient w O (i) from the coefficient table t0 when the period, the quantized value of the period, the estimated value of the period or the value that is negatively correlated with the fundamental frequency is less than or equal to a predetermined threshold or less than the predetermined threshold, and

(2) obtain the coefficient w t1 (i) as the coefficient w O (i) from the coefficient table t1 when the period, the quantized value of the period, the estimated value of the period or the value that is negatively correlated with the fundamental frequency is more than the predetermined threshold or more than or equal to the predetermined threshold, and

the processing circuitry is configured to encode or analyze the speech signal or the acoustic signal using the calculated coefficients to be transformed to first order to Pmax-order linear prediction coefficients.

4. A linear prediction analysis device that obtains, in each frame, which is a predetermined time interval, coefficients to be transformed to linear prediction coefficients corresponding to an input time-series signal, the linear prediction analysis device comprising:

processing circuitry configured to

receive the input time-series signal, the time-series signal being a speech signal or an acoustic signal;

calculate an autocorrelation R O (i) between an input time-series signal X O (n) of a current frame and an input time-series signal X O (n−i) i samples before the input time-series signal X O (n) or an input time-series signal X O (n+i) i samples after the input time-series signal X O (n), for each i of i=0, 1, . . . , Pmax at least, where the current frame includes parts of adjacent frame; and

calculate coefficients to be transformed to first-order to P max -order linear prediction coefficients, by using a modified autocorrelation R′O(i) obtained by multiplying a coefficient w O (i) by the autocorrelation RO(i) for each i;

wherein a coefficient table t0 stores a coefficient w t0 (i) and a coefficient table t1 stores a coefficient w t1 (i), wt0(i)<w t1 (i) being satisfied for at least part of i other than i=0, w t0 (i)≤w t1 (i) being satisfied for the remaining each i other than i=0,

the processing circuitry is further configured to, by using a fundamental frequency, a quantized value of the fundamental frequency, an estimated value of the fundamental frequency or a value that is positively correlated with a fundamental frequency based on the input time-series signal of the current frame or a past frame,

(1) obtain the coefficient w t0 (i) as the coefficient w O (i) from the coefficient table t0 when the fundamental frequency, the quantized value of the fundamental frequency, the estimated value of the fundamental frequency or the value that is positively correlated with the fundamental frequency is more than or equal to a predetermined threshold or more than the predetermined threshold, and

(2) obtain the coefficient w t0 (i) as the coefficient w O (i) from the coefficient table t1 when the fundamental frequency, the quantized value of the fundamental frequency, the estimated value of the fundamental frequency or the value that is positively correlated with the fundamental frequency is less than the predetermined threshold or less than or equal to the predetermined threshold, and

the processing circuitry is configured to encode or analyze the speech signal or the acoustic signal using the calculated coefficients to be transformed to first order to P max -order linear prediction coefficients.

5. A non-transitory computer readable medium that stores a program for causing a computer to execute a linear prediction analysis method of obtaining, in each frame, which is a predetermined time interval, coefficients to be transformed to linear prediction coefficients corresponding to an input time-series signal, the linear prediction analysis method comprising:

a step of receiving the input time-series signal, the time-series signal being a speech signal or an acoustic signal;

an autocorrelation calculation step of calculating an autocorrelation R O (i) between an input time-series signal X O (n) of a current frame and an input time-series signal X O (n−i) i samples before the input time-series signal X O (n) or an input time-series signal X O (n+i) i samples after the input time-series signal X O (n), for each i of i=0, 1, . . . , P max at least, where the current frame includes parts of adjacent frame; and

a prediction coefficient calculation step of calculating coefficients to be transformed to first-order to P max -order linear prediction coefficients, by using a modified autocorrelation R′ O (i) obtained by multiplying a coefficient w O (i) by the autocorrelation R O (i) for each i,

wherein a coefficient table t0 stores a coefficient w t0 (i) and a coefficient table t1 stores a coefficient w t1 (i), w t0 (i)<w t1 (i) being satisfied for at least part of i other than i=0, w t0 (i)≤w t1 (i) being satisfied for the remaining each i other than i=0,

the linear prediction analysis method further comprises a coefficient determination step of, by using a period, a quantized value of the period, an estimated value of the period or a value that is negatively correlated with a fundamental frequency based on the input time-series signal of the current frame or a past frame,

(1) obtaining the coefficient w t0 (i) as the coefficient w O (i) from the coefficient table t0 when the period, the quantized value of the period, the estimated value of the period or the value that is negatively correlated with the fundamental frequency is less than or equal to a predetermined threshold or less than the predetermined threshold, and

(2) obtaining the coefficient w t1 (i) as the coefficient w O (i) from the coefficient table t1 when the period, the quantized value of the period, the estimated value of the period or the value that is negatively correlated with the fundamental frequency is more than the predetermined threshold or more than or equal to the predetermined threshold, and

the linear prediction analysis method further includes encoding or analyzing the speech signal or the acoustic signal using the calculated coefficients to be transformed to first order to Pmax-order linear prediction coefficients.

6. A non-transitory computer readable medium that stores a program for causing a computer to execute a linear prediction analysis method of obtaining, in each frame, which is a predetermined time interval, coefficients to be transformed to linear prediction coefficients corresponding to an input time-series signal, the linear prediction analysis method comprising:

a step of receiving the input time-series signal, the time-series signal being a speech signal or an acoustic signal;

an autocorrelation calculation step of calculating an autocorrelation R O (i) between an input time-series signal X O (n) of a current frame and an input time-series signal X O (n−i) i samples before the input time-series signal X O (n) or an input time-series signal X O (n+i) i samples after the input time-series signal X O (n), for each i of i=0, 1, . . . , Pmax at least, where the current frame includes parts of adjacent frame; and

a prediction coefficient calculation step of calculating coefficients to be transformed to first-order to P max -order linear prediction coefficients, by using a modified autocorrelation R′O(i) obtained by multiplying a coefficient w O (i) by the autocorrelation RO(i) for each i;

wherein a coefficient table t0 stores a coefficient w t0 (i) and a coefficient table t1 stores a coefficient w t1 (i), wt0(i)<w t1 (i) being satisfied for at least part of i other than i=0, w t0 (i)≤w t1 (i) being satisfied for the remaining each i other than i=0,

the linear prediction analysis method further comprises a coefficient determination step of, by using a fundamental frequency, a quantized value of the fundamental frequency, an estimated value of the fundamental frequency or a value that is positively correlated with a fundamental frequency based on the input time-series signal of the current frame or a past frame,

(1) obtaining the coefficient w t0 (i) as the coefficient w O (i) from the coefficient table t0 when the fundamental frequency, the quantized value of the fundamental frequency, the estimated value of the fundamental frequency or the value that is positively correlated with the fundamental frequency is more than or equal to a predetermined threshold or more than the predetermined threshold, and

(2) obtaining the coefficient w t1 (i) as the coefficient w O (i) from the coefficient table t1 when the fundamental frequency, the quantized value of the fundamental frequency, the estimated value of the fundamental frequency or the value that is positively correlated with the fundamental frequency is less than the predetermined threshold or less than or equal to the predetermined threshold, and

the linear prediction analysis method further includes encoding or analyzing the speech signal or the acoustic signal using the calculated coefficients to be transformed to first order to P max -order linear prediction coefficients.

Assignments (1)
CHANGE OF NAME Recorded Oct 10, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 073080/0163 →
Priority Claims (1)
JP 2013-149160 · Jul 18, 2013 · national
Continuity (4)
Continuation 17970879 · Oct 21, 2022
Continuation 17120462 · Dec 14, 2020
Continuation 14905158
Related Publication 20240233739A1 · Jul 11, 2024
References Cited (30)
US 5774846A · Morii · 1998 [cited by examiner]
US 5819209A · Inoue · 1998 [cited by examiner]
US 6167373A · Morii · 2000 [cited by examiner]
US 6199035B1 · Lakaniemi · 2001 [cited by examiner]
US 6959274B1 · Gao · 2005 [cited by examiner]
US 9928850B2 · Kamamoto · 2018 [cited by examiner]
US 9966083B2 · Kamamoto · 2018 [cited by examiner]
US 10115413B2 · Kamamoto · 2018 [cited by examiner]
US 10134419B2 · Kamamoto · 2018 [cited by examiner]
US 10134420B2 · Kamamoto · 2018 [cited by examiner]
US 10163450B2 · Kamamoto · 2018 [cited by examiner]
US 10170130B2 · Kamamoto · 2019 [cited by examiner]
US 10909996B2 · Kamamoto · 2021 [cited by examiner]
US 20010027391A1 · Yasunaga · 2001 [cited by examiner]
US 20040002856A1 · Bhaskar · 2004 [cited by examiner]
US 20090089051A1 · Ishii · 2009 [cited by examiner]
US 20110022924A1 · Malenovsky · 2011 [cited by examiner]
US 20160140975A1 · Kamamoto · 2016 [cited by examiner]
US 20160336019A1 · Kamamoto · 2016 [cited by examiner]
US 20160343387A1 · Kamamoto · 2016 [cited by examiner]
US 20240177720A1 · Markovic · 2024 [cited by examiner]
US 20240194208A1 · Markovic · 2024 [cited by examiner]
ITU-T Recommendation G.718, “Series G: Transmission Systems and Media, Digital Systems and Networks; Digital terminal equipments—Coding of voice and audio signals,” International Telecommunication Union, Jun. 2008 (255 … [cited by applicant]
ITU-T Recommendation G.729, “General Aspects of Digital Transmission Systems,” International Telecommunication Union, Mar. 1996 (38 pages). [cited by applicant]
Yoh'ichi TOHKURA, et al., “Spectral Smoothing Technique in PARCOR Speech Analysis-Synthesis,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. ASSP-26, No. 6, Dec. 1978 (pp. 587-596). [cited by applicant]
International Search Report issued Sep. 16, 2014 in PCT/JP2014/068895 filed Jul. 16, 2014. [cited by applicant]
Office Action issued on Nov. 21, 2016 in Korean Patent Application No. 10-2016-7001218 (w/English translation). [cited by applicant]
Kawahara (Kawahara, Hideki, Ikuyo Masuda-Katsuse, and Alain De Cheveigne. “Restructuring speech representations using a pitch-adaptive time-frequency smoothing and an instantaneous-frequency-based FO extraction: Possibl… [cited by applicant]
Saeidi, Rahim, et al. “Temporally weighted linear prediction features for tackling additive noise in speaker verification.” IEEE Signal Processing Letters 17.6 (2010): 599-602. (Year: 2010). [cited by applicant]
Magi, Carlo, et al. “Stabilised weighted linear prediction.” Speech Communication 51.5 (2009): 401-411. (Year: 2009). [cited by applicant]