Channel decoding for wireless telephones with multiple microphones and multiple description transmission
An embodiment of the present invention provides a wireless telephone including a receiver module, a channel decoder, a speech decoder, and a speaker. The receiver module receives a plurality of versions of a voice signal, wherein each version of the voice signal comprises a plurality of speech frames. The channel decoder is configured to decode a speech parameter associated with a first speech frame from a first version of the voice signal, wherein decoding the speech parameter includes selecting an optimal bit sequence from a plurality of candidate bit sequences and wherein the selection of the optimal bit sequence is based in part on a second speech frame from a second version of the voice signal. The speech decoder decodes at least one of the first or second versions of the voice signal based on the speech parameter to generate an output signal. The speaker receives the output signal and produces a sound pressure wave corresponding thereto.
1 . A wireless telephone, comprising:
a receiver module that receives a plurality of versions of a voice signal, each version of the voice signal comprising a plurality of speech frames;
a channel decoder configured to decode a speech parameter associated with a speech frame from one of the plurality of versions of the voice signal, wherein decoding the speech parameter includes selecting an optimal bit sequence from a plurality of candidate bit sequences and wherein the selection of the optimal bit sequence is based in part on a corresponding speech frame from another version of the plurality of versions of the voice signal;
a speech decoder that decodes at least one of the plurality of versions of the voice signal based on the speech parameter to generate an output signal; and
a speaker that receives the output signal and produces a sound pressure wave corresponding thereto.
2 . The wireless telephone of claim 1 , wherein the channel decoder selects the optimal bit sequence based (i) in part on the corresponding speech frame and (ii) in part on a previous speech frame from at least one of the plurality of versions of the voice signal.
3 . The wireless telephone of claim 1 , wherein the speech decoder decodes at least two versions of the voice signal and combines the at least two decoded versions of the voice signal to produce the output signal.
4 . The wireless telephone of claim 1 , wherein the speech decoder is configured to estimate channel impairments and selectively decode a version of the voice signal from the plurality of versions with the least channel impairments, and wherein the decoded version is used as the output signal.
5 . The wireless telephone of claim 1 , wherein the speech decoder is configured to dynamically discard each version of the voice signal having channel impairments worse than a threshold of channel impairments; and
the speech decoder is still further configured to decode at least two non-discarded versions of the voice signal and combine the at least two decoded versions to produce the output signal.
6 . A multiple-description transmission system, comprising:
(a) a first wireless telephone comprising
(i) a microphone array, each microphone in the array configured to receive voice input from a user and to produce a voice signal corresponding thereto,
(ii) an encoder coupled to the microphone array and configured to encode each voice signal, and
(iii) a transmitter coupled to the encoder and configured to transmit each encoded voice signal; and
(b) a second wireless telephone comprising
(i) a receiver module that receives each transmitted voice signal, each transmitted voice signal comprising a plurality of speech frames;
(ii) a channel decoder configured to decode a speech parameter associated with a speech frame from one of the transmitted voice signals, wherein decoding the speech parameter includes selecting an optimal bit sequence from a plurality of candidate bit sequences and wherein the selection of the optimal bit sequence is based in part on a corresponding speech frame from another transmitted voice signal;
(iii) a speech decoder that decodes at least one of the transmitted voice signals based on the speech parameter to generate an output signal; and
(iv) a speaker that receives the output signal and produces a sound pressure wave corresponding thereto.
7 . The system of claim 6 , wherein the channel decoder selects the optimal bit sequence based (i) in part on the corresponding speech frame and (ii) in part on a previous speech frame from at least one of the transmitted voice signals.
8 . The system of claim 6 , wherein the speech decoder decodes at least two transmitted voice signals and combines the at least two decoded voice signals to produce the output signal.
9 . The system of claim 6 , wherein:
the speech decoder decodes each transmitted voice signal; and
the speech decoder is further configured to (i) detect a direction of arrival (DOA) of a sound wave emanating from the mouth of a user of the first wireless telephone based on the decoded voice signals and (ii) adaptively combine the decoded voice signals based on the DOA to produce the output signal.
10 . The system of claim 9 , wherein the speech decoder is still further configured to adaptively combine the decoded voice signals based on the DOA to effectively steer a maximum sensitivity angle of the microphone array so that the user's mouth is within the maximum sensitivity angle, wherein the maximum sensitivity angle is defined as an angle within which a sensitivity of the microphone array is above a threshold.
11 . The system of claim 6 ,
wherein the speech decoder is further configured to estimate channel impairments and decode a transmitted voice signal with the least channel impairments, and wherein the decoded version is used as the output signal.
12 . The system of claim 6 , wherein:
the speech decoder is configured to dynamically discard each transmitted voice signal having channel impairments worse than a threshold of channel impairments; and
the speech decoder is further configured to decode the non-discarded voice signals and combine the decoded signals to produce the output signal.
13 . The system of claim 6 , wherein:
the speech decoder is configured to dynamically discard each transmitted voice signal having channel impairments worse than a threshold of channel impairments; and
the speech decoder is further configured to detect a direction of arrival (DOA) of a sound wave emanating from the mouth of a user of the first wireless telephone based on the non-discarded signals and to adaptively combine the non-discarded signals based on the DOA to produce the output signal.
14 . The system of claim 13 , wherein the decoder is still further configured to adaptively combine the non-discarded signals based on the DOA to effectively steer a maximum sensitivity angle of the microphone array so the user's mouth is within the maximum sensitivity angle, wherein the maximum sensitivity angle is defined as an angle within which a sensitivity of the microphone array is above a threshold.
15 . The system of claim 6 , wherein the encoder is configured to encode the voice signals at different bit rates.
16 . The system of claim 15 , wherein:
the encoder is configured to encode one of the voice signals at a first bit rate for transmission over a main channel and each other voice signal at a second bit rate for transmission over a corresponding auxiliary channel; and
the speech decoder is configured to estimate channel impairments, and if (i) the main channel is corrupted by channel impairments and (ii) at least one auxiliary channel is not corrupted by channel impairments, to decode an auxiliary channel.
17 . The system of claim 15 , wherein:
the speech encoder is configured to encode one of the voice signals at a first bit rate for transmission over a main channel and each other voice signal at a second bit rate for transmission over a corresponding auxiliary channel; and
the speech decoder is further configured to estimate channel impairments, and if (i) side information corresponding to the main channel is corrupted by channel impairments and (ii) side information corresponding to at least one auxiliary channel is not corrupted by channel impairments, to process both the main channel and the at least one of the auxiliary channels to improve performance of a frame erasure concealment algorithm in the production of the output signal.
18 . A method in a wireless telephone, comprising:
(a) receiving a plurality of versions of a voice signal, each version of the voice signal comprising a plurality of speech frames;
(b) decoding a speech parameter associated with a speech frame from one of the plurality of versions of the voice signal, wherein decoding the speech parameter includes selecting an optimal bit sequence from a plurality of candidate bit sequences and wherein the selection of the optimal bit sequence is based in part on a corresponding speech frame from another version of the plurality of versions of the voice signal;
(c) decoding at least one of the plurality of versions of the voice signal based on the speech parameter to generate an output signal; and
(d) producing a sound pressure wave corresponding to the output signal.
19 . The method of claim 18 , wherein step (b) comprises:
selecting the optimal bit sequence based (i) in part on the corresponding speech frame and (ii) in part on a previous speech frame from at least one of the plurality of versions of the voice signal.
20 . The method of claim 18 , wherein step (c) comprises:
(c1) decoding at least two versions of the voice signal; and
(c2) combining the at least two decoded versions of the voice signal to produce the output signal.
21 . The method of claim 18 , wherein step (c) comprises:
(c1) estimating channel impairments and decoding a version of the voice signal from the plurality of versions with the least channel impairments, wherein the decoded version is used as the output signal.
22 . The method of claim 18 , wherein step (c) further comprises:
(c1) dynamically discarding each version of the voice signal having channel impairments worse than a threshold of channel impairments; and
(c2) decoding at least two non-discarded versions of the voice signal and combining the at least two decoded versions to produce the output signal.