IP Library Granted Patent US 11,574,642
Granted Patent B2
US 11,574,642 · App. 16/915,160 · Granted Feb 7, 2023

System and method to correct for packet loss in ASR systems

Inventors: Srinath Cheluvaraja (Carmel, IN); Ananth Nagaraja Iyer (Carmel, IN); Aravind Ganapathiraju (Hyderabad, IN); Felix Immanuel Wyss (Zionsville, IN)
G10L19/005G10L15/08G10L15/14G10L15/142G10L15/20G10L15/02G10L25/18G10L25/21G10L2015/025G10L2019/0012
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,574,642
App. No.
16/915,160
Granted
Feb 7, 2023
Kind
B2
Abstract

A system and method are presented for the correction of packet loss in audio in automatic speech recognition (ASR) systems. Packet loss correction, as presented herein, occurs at the recognition stage without modifying any of the acoustic models generated during training. The behavior of the ASR engine in the absence of packet loss is thus not altered. To accomplish this, the actual input signal may be rectified, the recognition scores may be normalized to account for signal errors, and a best-estimate method using information from previous frames and acoustic models may be used to replace the noisy signal.

Claims (30)

1. A method to correct for packet loss in an audio signal in automatic speech recognition systems comprising rescoring of phoneme probability calculations for a section of the audio signal, comprising the steps of:

accumulating probabilities for a given word in a lexicon using the highest scoring tokens for every frame of the audio signal;

determining when the probabilities for the given word are statistically significant;

reporting matches when the probabilities are statistically significant;

normalizing values of probabilities that are not statistically significant by replacing values for section of the audio signal comprising packet loss; and

rescoring the probabilities for the given word.

2. The method of claim 1 wherein the normalizing comprises deleting values for sections of the audio signal comprising packet loss.

3. The method of claim 1 wherein replacing values for sections of the audio signal comprising packet loss includes replacing a noisy signal caused by packet loss by a best estimate selected from historical values obtained offline or selected from previous frames and acoustic models.

4. The method of claim 1 , wherein the reporting comprises omitting low confidence hits.

5. The method of claim 1 , wherein the probabilities are accumulated from historical values previously obtained for each phoneme from a corpus without packet loss.

6. The method of claim 1 , wherein the rescoring comprises tagging the matches for the given word whose confidence values through their scores have been affected by packet loss using tagged packet information.

7. The method of claim 6 , wherein the tagged packet information is determined by extracting mel frequency cepstral coefficient features, wherein the frames of audio are decomposed into overlapping frames.

8. The method of claim 7 , wherein the overlapping frames are 20 ms with an overlap factor.

9. The method of claim 8 , wherein the overlap factor is 50%.

10. An automatic speech recognition system that corrects for packet loss in an audio signal comprising rescoring of phoneme probability calculations for a section of the audio signal, comprising:

a processor; and

a memory in communication with the processor, the memory storing instructions that, when executed by the processor, causes the processor to rescore phoneme probability calculations for a section of the audio signal by:

accumulating probabilities for a given word in a lexicon using the highest scoring tokens for every frame of the audio signal;

determining when the probabilities for the given word are statistically significant;

reporting matches when the probabilities are statistically significant;

normalizing values of probabilities that are not statistically significant by replacing values for sections of the audio signal comprising packet loss; and

rescoring the probabilities for the given word.

11. The system of claim 10 wherein the normalizing comprises deleting values for sections of the audio signal comprising packet loss.

12. The system of claim 10 wherein replacing values for sections of the audio signal comprising packet loss includes replacing a noisy signal caused by packet loss by a best estimate selected from historical values obtained offline or selected from previous frames and acoustic models.

13. The system of claim 10 , wherein the reporting comprises omitting low confidence hits.

14. The system of claim 10 , wherein the probabilities are accumulated from historical values previously obtained for each phoneme from a corpus without packet loss.

15. The system of claim 10 , wherein the rescoring comprises tagging the matches for the given word whose confidence values through their scores have been affected by packet loss using tagged packet information.

16. The system of claim 15 , wherein the tagged packet information is determined by extracting mel frequency cepstral coefficient features, wherein the frames of audio are decomposed into overlapping frames.

17. The system of claim 16 , wherein the overlapping frames are 20 ms with an overlap factor.

18. The system of claim 17 , wherein the overlap factor is 50%.

Assignments (4)
NOTICE OF SUCCESSION OF SECURITY INTERESTS AT REEL/FRAME 053598/0521 Recorded Feb 3, 2025
From: BANK OF AMERICA, N.A., AS RESIGNING AGENT
To: GOLDMAN SACHS BANK USA, AS SUCCESSOR AGENT
Reel/Frame 070097/0036 →
CHANGE OF NAME Recorded May 13, 2024
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: GENESYS CLOUD SERVICES, INC.
Reel/Frame 067391/0073 →
SECURITY AGREEMENT Recorded Aug 25, 2020
From: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
To: BANK OF AMERICA, N.A.
Reel/Frame 053598/0521 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2020
From: CHELUVARAJA, SRINATH; IYER, ANANTH NAGARAJA; GANAPATHIRAJU, ARAVIND; WYSS, FELIX IMMANUEL
To: GENESYS TELECOMMUNICATIONS LABORATORIES, INC.
Reel/Frame 053082/0464 →
Continuity (2)
Division 16186851 · Nov 12, 2018
Related Publication 20200335110A1 · Oct 22, 2020