System for speech encoding having an adaptive encoding arrangement
In accordance with one aspect of the invention, a selector supports the selection of a first encoding scheme or the second encoding scheme based upon the detection or absence of the triggering characteristic in the interval of the input speech signal. The first encoding scheme has a pitch pre-processing procedure for processing the input speech signal to form a revised speech signal biased toward an ideal voiced and stationary characteristic. The pre-processing procedure allows the encoder to fully capture the benefits of a bandwidth-efficient, long-term predictive procedure for a greater amount of speech components of an input speech signal than would otherwise be possible. In accordance with another aspect of the invention, the second encoding scheme entails a long-term prediction mode for encoding the pitch on a sub-frame by sub-frame basis. The long-term prediction mode is tailored to where the generally periodic component of the speech is generally not stationary or less than completely periodic and requires greater frequency of updates from the adaptive codebook to achieve a desired perceptual quality of the reproduced speech under a long-term predictive procedure.
1 . A speech encoding system comprising:
a detector for detecting whether an input speech signal generally has a triggering characteristic during an interval;
an encoder supporting at least one of a first encoding scheme and a first encoding scheme applicable to the speech signal for a frame associated with the interval, the first encoding scheme having a pre-processing procedure for processing the inputted speech signal to form a revised speech signal biased toward a generally ideal voiced and stationary characteristic; and
a selector for selecting one of the first encoding scheme and the second encoding scheme based upon the detection or absence of the triggering characteristic in the interval of the input speech signal.
2 . The speech encoding system according to claim 1 where the triggering characteristic comprises a generally voiced and generally stationary speech component of the speech signal.
3 . The speech encoding system according to claim 1 where the selector selects the first encoding scheme if the detector determines that the speech signal is generally stationary and generally periodic during the frame.
4 . The speech encoding system according to claim 1 where the selector selects the second encoding scheme if the detector determines that the speech signal is generally nonstationary during the frame.
5 . The speech encoding system according to claim 1 further comprising:
a perceptual weighting filter for filtering the input speech signal;
a pitch-preprocessing module having an input coupled to an output of the perceptual weighting filter, the pitch pre-processing module determining a target signal for time warping the weighted speech signal.
6 . The speech encoding system according to claim 1 further comprising a pitch pre-processing module for determining an input pitch track based on multiple frames of the speech signal and altering variations in the pitch lag associated with samples to track the input pitch track.
7 . The speech encoding system according to claim 1 where the first encoding scheme has a first allocation of storage units per frame between a fixed codebook index and an adaptive codebook index, the second scheme having a second allocation of storage units per the frame between the fixed codebook index and the adaptive codebook index, where the first allocation differs from the second allocation.
8 . The speech encoding system according to claim 7 where the second allocation of storage units per frame allocates a greater number of storage units to the adaptive codebook index than the first allocation of storage units to facilitate long-term predictive coding on a subframe-by-subframe basis.
9 . The speech encoding system according to claim 7 where the first allocation of storage units per frame allocates a greater number of storage units for the fixed codebook index than the second allocation does to reduce a quantization error associated with the fixed codebook index.
10 . The speech encoding system according to claim 7 where the second encoding scheme has a higher allocation ratio than the first encoding scheme, the allocation ratio defined by a number of storage units allocated to the adaptive codebook index divided by the number of storage units allocated to the adaptive codebook index plus the fixed codebook index.
11 . The speech encoding system according to claim 7 where, for full-rate coding, the first encoding scheme supports a first frame type and the second encoding scheme supports a second frame type different from the first frame type.
12 . The speech encoding system according to claim 7 where, for higher-rate coding, the first encoding scheme supports a first frame type and the second encoding scheme supports a second frame type, and for lower-rate coding the encoder supports a third frame type and a fourth frame type.
13 . A speech encoding system comprising:
a detector for detecting whether an input speech signal generally has a generally voiced and generally stationary characteristic during an interval;
an encoder supporting at least one of a first encoding scheme and a second encoding scheme applicable to the speech signal for a frame associated with the interval, the second encoding scheme having long-term prediction procedure for processing the inputted speech signal on a sub-frame-by-subframe basis;
a selector for selecting one of the first encoding scheme and the second encoding scheme based upon said detection or absence of the generally voiced and generally stationary characteristic in the interval of the input speech signal.
14 . The speech encoding system according to claim 13 where the selector selects the second encoding scheme if the detector determines that the speech signal is not generally periodic during the frame.
15 . The speech encoding system according to claim 13 where the selector selects the second encoding scheme if the detector determines that the speech signal is generally nonstationary during the frame.
16 . The speech encoding system according to claim 13 where the second encoding scheme has a pitch track with a greater number of bits per frame than the first encoding scheme to represent the pitch track.
17 . A speech encoding method comprising the steps of:
detecting whether an input speech signal has a triggering characteristic during an interval;
selecting one of a first encoding scheme and a second encoding scheme, for application to the input speech signal for a frame associated with the interval, based upon said detection of the triggering characteristic; and
processing the inputted speech signal in accordance with the first encoding scheme to form a revised speech signal biased toward a generally ideal voiced and stationary characteristic if the triggering characteristic is detected in the input speech signal.
18 . The method according to claim 17 where the detecting step comprises detecting whether the input speech signal generally has a generally voiced and generally stationary component as the triggering characteristic during an interval.
19 . The method according to claim 17 further comprising the step of supporting the first encoding scheme having a first allocation of storage units per the frame between a fixed codebook index and an adaptive codebook index, the second encoding scheme having a second allocation of storage units per the frame between the fixed codebook index and the adaptive codebook index, where the second allocation differs from the first allocation
20 . The method according to claim 17 further comprising the step of processing the inputted speech signal on a sub-frame-by-subframe basis in accordance with a long-term prediction procedure of the second encoding scheme if the triggering characteristic is not detected during the interval.