IP Library Granted Patent US 12,380,902
Granted Patent B2
US 12,380,902 · App. 18/540,060 · Granted Aug 5, 2025

Vector quantizer correction for audio codec system

Inventors: Marcin Ciolek (Gdansk, PL); Michal Sulewski (Gdańsk, PL); Raul A. Casas (Doylestown, PA); Samer Lutfi Hijazi (San Jose, CA); Mihailo Kolundzija (Lausanne, CH)
Assignee: CISCO TECHNOLOGY, INC.
G10L19/038G10L19/00G10L19/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,902
App. No.
18/540,060
Granted
Aug 5, 2025
Kind
B2
Abstract

A method comprises: vector quantizing input vectors representative of audio into an original sequence including indices of codewords of a codebook; generating candidate sequences including the indices of the codewords of the codebook by evaluating, for each candidate sequence, transition costs for transitions between the indices based on (i) transition probabilities of the transitions, and (ii) distances between the codewords represented by the indices and the input vectors that corresponds to the indices; determining a preferred candidate sequence of the candidate sequences to replace the original sequence based on the transition costs for each candidate sequence; and transmitting the preferred candidate sequence in place of the original sequence.

Claims (51)

1. A method comprising:

vector quantizing input vectors representative of audio into an original sequence including original indices of codewords of a codebook;

generating candidate sequences including indices of the codewords and that have respective starting transitions from an initial index of the original sequence to respective ones of all possible indices, wherein generating includes evaluating, for each candidate sequence, transition costs for transitions between the indices based on linear combinations of (i) transition probabilities of the transitions, and (ii) distances between the codewords represented by the indices and the input vectors that correspond to the indices;

summing the transition costs for each candidate sequence into a total transition cost, to produce total transition costs for corresponding ones of the candidate sequences;

selecting a preferred candidate sequence of the candidate sequences that has a lowest total transition cost; and

transmitting or storing the preferred candidate sequence in place of the original sequence.

2. The method of claim 1 , wherein:

each linear combination includes subtracting a transition probability from a corresponding distance.

3. The method of claim 1 , wherein evaluating includes evaluating, for each candidate sequence, the transition costs for the transitions between each index and each next index based on the transition probabilities of the transitions, and the distances between the codewords represented by each next index and an input vector that corresponds to each next index.

4. The method of claim 3 , wherein generating further comprises, for each candidate sequence, determining each next index for each index by:

computing next index transition costs for test transitions from each index to possible next indices for the codewords available in the codebook; and

selecting as each next index a possible next index associated with a lowest next index transition cost.

5. The method of claim 4 , wherein:

computing includes computing the next index transition costs based on test transition probabilities of the test transitions that lead from each index to each of the possible next indices and the distances between the codewords represented by the possible next indices and corresponding input vector.

6. The method of claim 1 , wherein:

accessing the transition probabilities from a datastore of predetermined transition probabilities of the transitions from each index of the codewords in the codebook to the indices of all other codewords of the codebook.

7. The method of claim 1 , where evaluating each transition cost includes computing a difference between a first function of a transition probability of a transition and a second function of a distance between a codeword and a corresponding input vector.

8. The method of claim 1 , wherein:

the original sequence includes N indices;

the codebook includes M codewords; and

generating includes generating M candidate sequences of N indices.

9. The method of claim 1 , wherein evaluating includes evaluating the original sequence as one of the candidate sequences.

10. The method of claim 1 , wherein generating includes generating the candidate sequences incrementally index position-by-index position and performing evaluating at each index position.

11. The method of claim 1 , wherein generating by evaluating includes constructing, incrementally over time, a trellis structure of the indices and the transitions between the indices such that paths through the trellis structure represent the candidate sequences.

12. An apparatus comprising:

a network input/output interface to communicate with a network; and

a processor coupled to the network input/output interface and configured to perform:

vector quantizing input vectors representative of audio into an original sequence including original indices of codewords of a codebook;

generating candidate sequences including indices of the codewords and that have respective starting transitions from an initial index of the original sequence to respective ones of all possible indices, wherein generating includes evaluating, for each candidate sequence, transition costs for transitions between the indices based on a linear combination of (i) transition probabilities of the transitions, and (ii) distances between the codewords represented by the indices and the input vectors that correspond to the indices;

summing the transition costs for each candidate sequence into a total transition cost, to produce total transition costs for corresponding ones of the candidate sequences;

selecting a preferred candidate sequence of the candidate sequences that has a lowest total transition cost; and

transmitting or storing the preferred candidate sequence in place of the original sequence.

13. The apparatus of claim 12 , wherein:

each linear combination includes subtracting a transition probability from a corresponding distance.

14. The apparatus of claim 12 , wherein the processor is further configured to perform evaluating by evaluating, for each candidate sequence, the transition costs for the transitions between each index and each next index based on the transition probabilities of the transitions, and the distances between the codewords represented by each next index and an input vector that corresponds to each next index.

15. The apparatus of claim 14 , wherein the processor is further configured to perform generating by, for each candidate sequence, determining each next index for each index by:

computing next index transition costs for test transitions from each index to possible next indices for the codewords available in the codebook; and

selecting as each next index a possible next index associated with a lowest next index transition cost.

16. The apparatus of claim 15 , wherein:

the processor is further configured to perform computing by computing the next index transition costs based on test transition probabilities of the test transitions that lead from each index to each of the possible next indices and the distances between the codewords represented by the possible next indices and corresponding input vector.

17. The apparatus of claim 12 , wherein the processor is further configured to perform:

accessing the transition probabilities from a datastore of predetermined transition probabilities of the transitions from each index of the codewords in the codebook to the indices of all other codewords of the codebook.

18. A non-transitory computer medium encoded with instructions that, when executed by a processor, cause the processor to perform operations including:

vector quantizing input vectors representative of audio into an original sequence including original indices of codewords of a codebook;

generating candidate sequences including indices of the codewords and that have respective starting transitions from an initial index of the original sequence to respective ones of all possible indices, wherein generating includes evaluating, for each candidate sequence, transition costs for transitions between the indices based on a linear combination of (i) transition probabilities of the transitions, and (ii) distances between the codewords represented by the indices and the input vectors that correspond to the indices;

summing the transition costs for each candidate sequence into a total transition cost, to produce total transition costs for corresponding ones of the candidate sequences;

selecting a preferred candidate sequence of the candidate sequences that has a lowest total transition cost; and

transmitting or storing the preferred candidate sequence in place of the original sequence.

19. The non-transitory computer medium of claim 18 , wherein:

each linear combination includes subtracting a transition probability from a corresponding distance.

20. The non-transitory computer medium of claim 18 , wherein the instructions to cause the processor to perform evaluating include instructions to cause the processor to perform evaluating, for each candidate sequence, the transition costs for the transitions between each index and each next index based on the transition probabilities of the transitions, and the distances between the codewords represented by each next index and an input vector that corresponds to each next index.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2023
From: CIOLEK, MARCIN; SULEWSKI, MICHAL; CASAS, RAUL A.; HIJAZI, SAMER LUTFI; KOLUNDZIJA, MIHAILO
To: CISCO TECHNOLOGY, INC.
Reel/Frame 065873/0677 →
Continuity (2)
Provisional Application 63591249 · Oct 18, 2023
Related Publication 20250131931A1 · Apr 24, 2025
References Cited (78)
US 10991379B2 · Hijazi et al. · 2021 [cited by applicant]
US 11682400B1 · Liu et al. · 2023 [cited by applicant]
US 11901976B2 · Mielczarek · 2024 [cited by examiner]
US 20040102972A1 · Droppo et al. · 2004 [cited by applicant]
US 20070279261A1 · Todorov et al. · 2007 [cited by applicant]
US 20110029317A1 · Chen et al. · 2011 [cited by applicant]
US 20140191888A1 · Klejsa et al. · 2014 [cited by applicant]
US 20150095035A1 · Shechtman · 2015 [cited by applicant]
US 20150213809A1 · Peters · 2015 [cited by examiner]
US 20160148618A1 · Huang et al. · 2016 [cited by applicant]
US 20170078010A1 · Suh · 2017 [cited by examiner]
US 20180308497A1 · Chau et al. · 2018 [cited by applicant]
US 20200186164A1 · Marpe et al. · 2020 [cited by applicant]
US 20200219519A1 · Wang · 2020 [cited by applicant]
US 20200234720A1 · Sung et al. · 2020 [cited by applicant]
US 20210125622A1 · Zhao · 2021 [cited by applicant]
US 20210256961A1 · Garman et al. · 2021 [cited by applicant]
US 20230019128A1 · Zeghidour et al. · 2023 [cited by applicant]
US 20230148275A1 · Kim et al. · 2023 [cited by applicant]
US 20230186927A1 · Zeghidour · 2023 [cited by examiner]
US 20230238011A1 · Mllemoes et al. · 2023 [cited by applicant]
US 20230260524A1 · Multrus et al. · 2023 [cited by applicant]
WO 2010040522A2 · 2010 [cited by applicant]
WO WO2019213021A1 · 2019 [cited by examiner]
WO 2023177803A1 · 2023 [cited by applicant]
Buhmann, et al., “Vector Quantization with Complexity Costs,” IEEE Trans. Information Theory, Jul. 1993. (Year: 1993). [cited by examiner]
Buhmann, et al., “Vector Quantization with Complexity Costs,” IEEE Trans. Information Theory, 1993 (see attached reference in the previous Office action). (Year: 1993). [cited by examiner]
Valin, J.M., et al., “LPCNET: Improving Neural Speech Synthesis Through Linear Prediction,” https://arxiv.org/abs/1810.11846, Feb. 19, 2019, 5 pages. [cited by applicant]
Jiang, X., et al., “Latent-Domain Predictive Neural Speech Coding,” https://arxiv.org/pdf/2207.08363.pdf, May 22, 2023, 12 pages. [cited by applicant]
Zeghidour, N., et al., “SoundStream: An End-to-End Neural Audio Codec,” https://arxiv.org/pdf/2107.03312.pdf, IEEE/ACM Transactions on Audio, Speech, and Language Processing Jul. 7, 2021, 12 pages. [cited by applicant]
Défossez, A., et al., “High Fidelity Neural Audio Compression,” https://export.arxiv.org/abs/2210.13438, Oct. 24, 2022, 19 pages. [cited by applicant]
Dietz, M., et al., “Overview of the EVS codec architecture,” https://ieeexplore.ieee.org/document/7179063/, 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2015, 5 pages. [cited by applicant]
Valin, J.M., et al., “Definition of the Opus Audio Codec,” RFC 6716, https://datatracker.ietf.org/doc/html/rfc6716, Sep. 2012, 326 pages. [cited by applicant]
Mermelstein, P., “G.722: a new CCITT coding standard for digital transmission of wideband audio signals,” IEEE Communications Magazine, vol. 26, No. 1, https://ieeexplore.ieee.org/document/417, Jan. 1988, pp. 8-15. [cited by applicant]
Oord, A. v.d., et al., “WaveNet: A Generative Model for Raw Audio,” https://arxiv.org/abs/1609.03499, Sep. 19, 2016, 15 pages. [cited by applicant]
Kleijn, W.B., et al. “Wavenet Based Low Rate Speech Coding,” 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), https://arxiv.org/abs/1712.01120, Dec. 2017, 5 pages. [cited by applicant]
Gray, R., “Vector quantization,” https://ieeexplore.ieee.org/document/1162229, Abstract, Apr. 1984, 4 pages. [cited by applicant]
Valin, J.M., et al., “Low-Bitrate Redundancy Coding of Speech Using a Rate-Distortion-Optimized Variational Autoencoder,” ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP… [cited by applicant]
Bengio, Y., et al., “Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation,” https://arxiv.org/abs/1308.3432, Aug. 15, 2013, 12 pages. [cited by applicant]
Oord, A. v.d., et al., “Neural Discrete Representation Learning,” Advances in neural information processing systems, vol. 30, https://arxiv.org/abs/1711.00937, May 30, 2018, 11 pages. [cited by applicant]
Jang, E., et al., “Categorical Reparameterization with Gumbel-Softmax,” https://arxiv.org/abs/1611.01144, Aug. 5, 2017, 13 pages. [cited by applicant]
Liu, D., et al., “Discrete-Valued Neural Communication,” https://arxiv.org/abs/2107.02367, Advances in Neural Information Processing Systems, vol. 34, Jul. 10, 2021, pp. 2109-2121. [cited by applicant]
Vaswani, A., et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), https://arxiv.org/abs/1706.03762, Dec. 2017, 15 pages. [cited by applicant]
Vasuki, A., et al., “A review of vector quantization techniques,” https://ieeexplore-dev.ieee.org/document/1664069, Abstract, Jul. 31, 2006, 4 pages. [cited by applicant]
Choi, H.-S., et al., “Phase-aware Speech Enhancement with Deep Complex U-Net,” International Conference on Learning Representations, https://arxiv.org/abs/1903.03107, Mar. 7, 2019, 20 pages. [cited by applicant]
Salimans, T., et al., “Improved Techniques for Training GANs,” Advances in neural information processing systems, vol. 29, https://arxiv.org/abs/1606.03498, Jun. 10, 2016, 10 pages. [cited by applicant]
Kumar, K., et al., “MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,” https://arxiv.org/abs/1910.06711, Advances in neural information processing systems, vol. 32, Dec. 9, 2019, 14 pages. [cited by applicant]
Salimans, T., et al., “Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks,” Advances in neural information processing systems, vol. 29, https://arxiv.org/abs/1602.07868, Jun… [cited by applicant]
Habets, E.A.P., “Room Impulse Response Generator,” https://www.researchgate.net/publication/259991276_Room_Impulse_Response_Generator, Sep. 20, 2010, 21 pages. [cited by applicant]
Grondin, F., et al., “BIRD: Big Impulse Response Dataset,” https://arxiv.org/abs/2010.09930, Oct. 19, 2020, 5 pages. [cited by applicant]
Liu, L., et al., “On the Variance of the Adaptive Learning Rate and Beyond,” https://arxiv.org/abs/1908.03265, Oct. 26, 2021, 14 pages. [cited by applicant]
The ITU Radiocommunication Assembly, “Method for the subjective assessment of intermediate quality level of coding systems,” retrieved from https://www.itu.int/rec/R-REC-BS.1534-1-200301-S/en, Dec. 12, 2023, 18 pages. [cited by applicant]
International Telecommunication Union, P. 863, “Perceptual Objective Listening Quality Assessment,” P.863 : Perceptual objective listening quality assessment (itu.int), Jan. 2011, 76 pages. [cited by applicant]
Hines, A., et al., “ViSQOL: an objective speech quality model,” EURASIP Journal on Audio, Speech, and Music Processing, https://asmp-eurasipjournals.springeropen.com/articles/10.1186/s13636-015-0054-9, May 17, 2015, 18 … [cited by applicant]
Lin, J., et al., “A Time-Domain Convolutional Recurrent Network for Packet Loss Concealment,” Abstract, https://ieeexplore.ieee.org/document/9413595, Jun. 2021, 4 pages. [cited by applicant]
Pascual, S., et al., “Adversarial Auto-Encoding for Packet Loss Concealment,” 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), https://www.semanticscholar.org/reader/935c9e7a6d974… [cited by applicant]
Mohamed, M., et al., “ConcealNet: An End-to-end Neural Network for Packet Loss Concealment in Deep Speech Emotion Recognition,” https://arxiv.org/abs/2005.07777, May 15, 2020, 5 pages. [cited by applicant]
Lecomte, J., et al., “Packet-loss concealment technology advances in EVS,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), https://www.semanticscholar.org/paper/Packet-loss-concea… [cited by applicant]
Kazakos, E., et al., “Slow-Fast Auditory Streams for Audio Recognition,” https://arxiv.org/pdf/2103.03516.pdf, Mar. 5, 2021, 7 pages. [cited by applicant]
Li, B., et al., “TranSFormer: Slow-Fast Transformer for Machine Translation,” https://arxiv.org/abs/2305.16982, May 26, 2023, 12 pages. [cited by applicant]
Mujika, A., et al., “Fast-Slow Recurrent Neural Networks,” https://arxiv.org/pdf/1705.08639.pdf, Jun. 9, 2017, 10 pages. [cited by applicant]
Ba, J., et al., “Using Fast Weights to Attend to the Recent Past,” https://arxiv.org/abs/1610.06258, Dec. 5, 2016, 10 pages. [cited by applicant]
Tjandra, A., et al., “Unsupervised Learning of Disentangled Speech Content and Style Representation,” https://arxiv.org/pdf/2010.12973.pdf, Jun. 20, 2021, 5 pages. [cited by applicant]
An, X., et al., “Disentangling Style and Speaker Attributes for TTS Style Transfer,” https://arxiv.org/pdf/2201.09472.pdf, Jan. 24, 2022, 13 pages. [cited by applicant]
Yuan, S., et al., “Improving Zero-Shot Voice Style Transfer via Disentangled Representation Learning,” https://openreview.net/pdf?id=TgSVWXw22FQ, Jan. 2021, 12 pages. [cited by applicant]
Perkins, C., et al., “RTP Payload for Redundant Audio Data,” RFC 2198, https://dl.acm.org/doi/10.17487/RFC2198, Sep. 1997, 11 pages. [cited by applicant]
Valin, J.M., et al., “A Real-Time Wideband Neural Vocoder at 1.6 kb/s Using LPCNet,” Proc. Interspeech 2019, https://arxiv.org/abs/1903.12087, Jun. 2019, pp. 3406-3410. [cited by applicant]
Jiang, X., et al., “End-to-End Neural Speech Coding for Real-Time Communications,” 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), https://ieeexplore.ieee.org/document/9746296, Fe… [cited by applicant]
Schulzrinne, H., et al., “RTP: A Transport Protocol for Real-Time Applications,” RFC 3550, https://www.rfc-editor.org/rfc/rfc3550, Jul. 2003, 104 pages. [cited by applicant]
Panayotov, V., et al., “Librispeech: An ASR corpus based on public domain audio books,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), https://ieeexplore.ieee.org/document/717896… [cited by applicant]
Diener, L., et al., “INTERSPEECH 2022 Audio Deep Packet Loss Concealment Challenge,” https://arxiv.org/abs/2204.05222, Apr. 11, 2022, 5 pages. [cited by applicant]
Diener, L., et al., “PLCMOS—a data-driven non-intrusive metric for the evaluation of packet loss concealment algorithms,” https://arxiv.org/abs/2305.15127, May 24, 2023, 5 pages. [cited by applicant]
Gray, R., “Vector quantization,” https://ieeexplore.ieee.org/document/1162229, IEEE ASSP Magazine, Apr. 1984, 26 pages. [cited by applicant]
Vasuki, A., et al., “A review of vector quantization techniques,” https://ieeexplore-dev.ieee.org/document/1664069, IEEE Potentials, Jul. 2006, 9 pages. [cited by applicant]
Lin, J., et al., “A Time-Domain Convolutional Recurrent Network for Packet Loss Concealment,” https://ieeexplore.ieee.org/document/9413595, ICASSP, Jun. 2021, 5 pages. [cited by applicant]
Zeghidour N., et al., “SoundStream: An End-to-End Neural Audio Codec”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, Published on Nov. 23, 2021, pp. 495-507. [cited by applicant]
Lincoln B., “An Experimental High Fidelity Perceptual Audio Coder”, Music 420 Project, Stanford University., Mar. 1998, Retrieved from https://ccrma.stanford.edu/˜jos/bosse/bosse.pdf, pp. 1-18. [cited by applicant]
Mentzer F., et al., “Conditional Probability Models for Deep Image Compression”, Computer Vision and Pattern Recognition, Jun. 4, 2018, Retrieved from https://openaccess.thecvf.com/content_cvpr_2018/papers/Mentzer_Condi… [cited by applicant]