IP Library Granted Patent US 8,589,166
Granted Patent B2
US 8,589,166 · App. 12/887,353 · Granted Nov 19, 2013

Speech content based packet loss concealment

Inventor: Robert W. Zopf (Rancho Santa Margarita, CA)
Assignee: Broadcom Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,589,166
App. No.
12/887,353
Granted
Nov 19, 2013
Kind
B2
Abstract

Systems and methods are described for performing packet loss concealment (PLC) to mitigate the effect of one or more lost frames within a series of frames that represent a speech signal. In accordance with the exemplary systems and methods, PLC is performed by searching a codebook of speech-related parameter profiles to identify content that is being spoken and by selecting a profile associated with the identified content for use in predicting or estimating speech-related parameter information associated with one or more lost frames of a speech signal. The predicted/estimated speech-related parameter information is then used to synthesize one or more frames to replace the lost frame(s) of the speech signal.

Claims (68)

1. A method for concealing the effects of one or more lost frames within a series of frames that comprise a speech signal, comprising:

composing an input vector that includes a computed value of a speech-related parameter for each of a number of frames that precede the lost frame(s);

comparing the input vector to at least one portion of each vector in a codebook, each vector in the codebook representing a different model of how the speech-related parameter varies over time;

selecting one of the vectors in the codebook based on the comparison;

determining a value of the speech-related parameter for each of the lost frame(s) based on the selected vector in the codebook; and

synthesizing one or more frames to replace the lost frame(s) based on the determined value(s) of the speech-related parameter.

2. The method of claim 1 , wherein composing the input vector comprises composing an input vector that comprises a computed value of one of the following speech-related parameters for each of the number of frames that precede the lost frame(s):

a fundamental frequency;

a frame gain;

a voicing measure;

a spectral envelope; and

a pitch.

3. The method of claim 1 , wherein the speech-related parameter comprises a scalar parameter.

4. The method of claim 1 , wherein the speech-related parameter comprises a vector parameter.

5. The method of claim 1 , wherein the speech-related parameter comprises a parameter that has been normalized for speaker independence.

6. The method of claim 1 , wherein composing the input vector comprises:

composing an input vector that includes the computed value of the speech-related parameter for each of the number of frames that precede the lost frame(s) and a computed value of the speech-related parameter for each of a number of frames that follow the lost frame(s).

7. The method of claim 1 , wherein comparing the input vector to at least one portion of each vector in the codebook comprises calculating a distortion measure based on the input vector and at least one portion of each vector in the codebook; and

wherein selecting one of the vectors in the codebook based on the comparison comprises selecting a vector in the codebook having a portion that generates the smallest distortion measure.

8. The method of claim 1 , wherein comparing the input vector to at least one portion of each vector in the codebook comprises calculating a similarity measure based on the input vector and at least one portion of each vector in the codebook; and

wherein selecting one of the vectors in the codebook based on the comparison comprises selecting a vector in the codebook having a portion that generates the greatest similarity measure.

9. The method of claim 1 , wherein comparing the input vector to at least one portion of each vector in the codebook comprises comparing the input vector to each of a series of overlapping portions of each vector in the codebook within an analysis window.

10. A method for concealing the effects of one or more lost frames within a series of frames that comprise a speech signal, comprising:

composing an input vector that includes a set of computed values of a plurality of speech-related parameters for each of a number of frames that precede the lost frame(s);

comparing the input vector to at least one portion of each vector in a codebook, each vector in the codebook jointly representing a plurality of models of how the plurality of speech-related parameters vary over time;

selecting one of the vectors in the codebook based on the comparison;

determining a value of each of the plurality of speech-related parameters for each of the lost frame(s) based on the selected vector in the codebook; and

synthesizing one or more frames to replace the lost frame(s) based on the determined value(s) of each of the plurality of speech-related parameters.

11. The method of claim 10 , wherein composing the input vector comprises composing an input vector that comprises, for each of the number of frames that precedes the lost frame(s), a set of computed values corresponding to two or more of the following speech-related parameters:

a fundamental frequency;

a frame gain;

a voicing measure;

a spectral envelope; and

a pitch.

12. The method of claim 10 , wherein at least one speech-related parameter in the plurality of speech-related parameters comprises a scalar parameter.

13. The method of claim 10 , wherein at least one speech-related parameter in the plurality of speech-related parameters comprises a vector parameter.

14. The method of claim 10 , wherein at least one speech-related parameter in the plurality of speech-related parameters comprises a parameter that has been normalized for speaker independence.

15. The method of claim 10 , wherein composing the input vector comprises:

composing an input vector that includes the set of computed values of the plurality of speech-related parameters for each of the number of frames that precede the lost frame(s) and a set of computed values of the plurality of speech-related parameters for each of a number of frames that follow the lost frame(s).

16. The method of claim 10 , wherein comparing the input vector to at least one portion of each vector in the codebook comprises calculating a distortion measure based on the input vector and at least one portion of each vector in the codebook; and

wherein selecting one of the vectors in the codebook based on the comparison comprises selecting a vector in the codebook having a portion that generates the smallest distortion measure.

17. The method of claim 16 , wherein calculating the distortion measure based on the input vector and at least one portion of each vector in the codebook comprises calculating a distortion measure that is a function of one or more characteristics of the speech signal or a bit stream that represents an encoded version of the speech signal.

18. The method of claim 16 , wherein calculating the distortion measure based on the input vector and at least one portion of a vector in the codebook comprises:

calculating a parameter-specific distortion measure for each speech-related parameter in the plurality of speech-related parameters based on the input vector and a portion of a vector in the codebook; and

combining the parameter-specific distortion measures.

19. The method of claim 18 , wherein combining the parameter-specific distortion measures comprises:

applying a weight to each of the parameter-specific distortion measures, wherein the weight applied to each of the parameter-specific distortion measures is selected based on one or more characteristics of the speech signal or a bit stream that represents an encoded version of the speech signal.

20. The method of claim 10 , wherein comparing the input vector to at least one portion of each vector in the codebook comprises calculating a similarity measure based on the input vector and at least a portion of each vector in the codebook; and

wherein selecting one of the vectors in the codebook based on the comparison comprises selecting a vector in the codebook having a portion that generates the greatest similarity measure.

21. The method of claim 20 , wherein calculating the similarity measure based on the input vector and at least one portion of each vector in the codebook comprises calculating a similarity measure that is a function of one or more characteristics of the speech signal or a bit stream that represents an encoded version of the speech signal.

22. The method of claim 20 , wherein calculating the similarity measure based on the input vector and at least one portion of a vector in the codebook comprises:

calculating a parameter-specific similarity measure for each speech-related parameter in the plurality of speech-related parameters based on the input vector and a portion of a vector in the codebook; and

combining the parameter-specific similarity measures.

23. The method of claim 22 , wherein combining the parameter-specific distortion measures comprises:

applying a weight to each of the parameter-specific distortion measures, wherein the weight applied to each of the parameter-specific distortion measures is selected based on one or more characteristics of the speech signal or a bit stream that represents an encoded version of the speech signal.

24. The method of claim 10 , wherein comparing the input vector to at least one portion of each vector in the codebook comprises comparing the input vector to each of a series of overlapping portions of each vector in the codebook within an analysis window.

25. A system for concealing the effects of one or more lost frames within a series of frames that comprise a speech signal, comprising:

at least one processor; and

at least one memory that stores software that is executed by the at least one processor, the software comprising:

a vector generation module that composes an input vector that includes a computed value of a speech-related parameter for each of a number of frames that precede the lost frame(s);

a codebook search module that compares the input vector to at least one portion of each vector in a codebook, each vector in the codebook representing a different model of how the speech-related parameter varies over time, selects one of the vectors in the codebook based on the comparison, and determines a value of the speech-related parameter for each of the lost frame(s) based on the selected vector in the codebook; and

a synthesis module that synthesizes one or more frames to replace the lost frame(s) based on the determined value(s) of the speech-related parameter.

26. A system for concealing the effects of one or more lost frames within a series of frames that comprise a speech signal, comprising:

at least one processor; and

at least one memory that stores software that is executed by the at least one processor, the software comprising:

an input vector generation module that composes an input vector that includes a set of computed values of a plurality of speech-related parameters for each of a number of frames that precede the lost frame(s);

a codebook search module that compares the input vector to at least one portion of each vector in a codebook, each vector in the codebook jointly representing a plurality of models of how the plurality of speech-related parameters vary over time, selects one of the vectors in the codebook based on the comparison, and determines a value of each of the plurality of speech-related parameters for each of the lost frame(s) based on the selected vector in the codebook; and

a synthesis module that synthesizes one or more frames to replace the lost frame(s) based on the determined value(s) of each of the plurality of speech-related parameters.

Assignments (7)
CORRECTIVE ASSIGNMENT TO CORRECT THE ERROR IN RECORDING THE MERGER IN THE INCORRECT US PATENT NO. 8,876,094 PREVIOUSLY RECORDED ON REEL 047351 FRAME 0384. ASSIGNOR(S) HEREBY CONFIRMS THE MERGER. Recorded Mar 8, 2019
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 049248/0558 →
CORRECTIVE ASSIGNMENT TO CORRECT THE EFFECTIVE DATE OF THE MERGER PREVIOUSLY RECORDED AT REEL: 047230 FRAME: 0910. ASSIGNOR(S) HEREBY CONFIRMS THE MERGER. Recorded Oct 29, 2018
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 047351/0384 →
MERGER Recorded Oct 4, 2018
From: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 047230/0910 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS Recorded Feb 3, 2017
From: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
To: BROADCOM CORPORATION
Reel/Frame 041712/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2017
From: BROADCOM CORPORATION
To: AVAGO TECHNOLOGIES GENERAL IP (SINGAPORE) PTE. LTD.
Reel/Frame 041706/0001 →
PATENT SECURITY AGREEMENT Recorded Feb 11, 2016
From: BROADCOM CORPORATION
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 037806/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2010
From: ZOPF, ROBERT W.
To: BROADCOM CORPORATION
Reel/Frame 025241/0031 →
Continuity (2)
Provisional Application 61253950 · Oct 22, 2009
Related Publication 20110099014A1 · Apr 28, 2011