IP Library › Granted Patent US 12,315,518
Granted Patent B2
US 12,315,518 · App. 17/810,132 · Granted May 27, 2025

Apparatus and method for improved concealment of the adaptive codebook in a CELP-like concealment employing improved pitch lag estimation

Inventors: Jeremie Lecomte (Fuerth, DE); Michael Schnabel (Geroldsgruen, DE); Goran Markovic (Nuremberg, DE); Martin Dietz (Nuremberg, DE); Bernhard Neugebauer (Erlangen, DE)
Assignee: Fraunhofer-Gesellschaft zur Foerderung der angewandten Forschung e.V.
G10L19/005G10L19/107G10L19/125G10L25/90G10L2019/0002G10L2019/0003G10L2019/0008G10L19/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,518
App. No.
17/810,132
Granted
May 27, 2025
Kind
B2
Abstract

An apparatus for determining an estimated pitch lag is provided. The apparatus includes an input interface for receiving a plurality of original pitch lag values, and a pitch lag estimator for estimating the estimated pitch lag. The pitch lag estimator is configured to estimate the estimated pitch lag depending on a plurality of original pitch lag values and depending on a plurality of information values, wherein for each original pitch lag value of the plurality of original pitch lag values, an information value of the plurality of information values is assigned to the original pitch lag value.

Claims (205)

1. An apparatus for generating a speech signal, comprising:

a pitch lag estimator for estimating an estimated pitch lag, wherein the apparatus is generating the speech signal using the estimated pitch lag,

wherein the pitch lag estimator is configured to estimate the estimated pitch lag depending on a plurality of original pitch lag values and depending on a plurality of information values,

wherein, for each original pitch lag value of the plurality of original pitch lag values, an information value of the plurality of information values is assigned to said original pitch lag value,

wherein the pitch lag estimator is configured to estimate the pitch lag by determining two parameters a and b, by minimizing an error function:

e

⁢

r

⁢

r

=

∑

i

=

s

k

g

p

(

i

)

·

(

(

a

+

b

·

i

)

-

P

⁡

(

i

)

)

2

wherein a is a real number, wherein b is a real number, wherein s is a first integer, wherein k is a second integer, and wherein P(i) is the i-th original pitch lag 1value, wherein g˜(i) is the i-th pitch gain value being assigned to the i-th pitch lag value P(i).

2. An apparatus according to claim 1 , wherein the pitch lag estimator is configured to estimate the estimated pitch lag depending on the plurality of original pitch lag values and depending on a plurality of pitch gain values as the plurality of information values, wherein for each original pitch lag value of the plurality of original pitch lag values, a pitch gain value of the plurality of pitch gain values is assigned to said original pitch lag value.

3. An apparatus according to claim 2 , wherein each of the plurality of pitch gain values is an adaptive codebook gain.

4. An apparatus according to claim 1 , wherein the pitch lag estimator is configured to estimate the estimated pitch lag by determining two parameters a, b, by minimizing the error function

err

=

∑

i

=

0

4

g

p

(

i

)

·

(

(

a

+

b

·

i

)

-

P

⁡

(

i

)

)

2

,

wherein a is a real number,

wherein b is a real number,

wherein P(i) is the i-th original pitch lag value,

wherein g p (i) is the i-th pitch gain value being assigned to the i-th pitch lag value P(i).

5. An apparatus according to claim 1 , wherein the pitch lag estimator is configured to estimate the estimated pitch lag depending on the plurality of original pitch lag values and depending on a plurality of time values as the plurality of information values, wherein for each original pitch lag value of the plurality of original pitch lag values, a time value of the plurality of time values is assigned to said original pitch lag value.

6. An apparatus according to claim 5 , wherein the pitch lag estimator is configured to estimate the estimated pitch lag by minimizing the error function.

7. An apparatus according to claim 6 , wherein the pitch lag estimator is configured to estimate the estimated pitch lag by determining two parameters a, b, by minimizing the error function

e

⁢

r

⁢

r

=

∑

i

=

0

k

time

passed

(

i

)

·

(

(

a

+

b

·

i

)

-

P

⁡

(

i

)

)

2

,

wherein a is a real number,

wherein b is a real number,

wherein k is an integer with k ≥2, and

wherein time passed (i) is representing an inverse of an amount of time that has passed after correctly receiving a pitch lag,

wherein P (i) is the i-th original pitch lag value,

corresponding to the pich lag.

8. An apparatus according to claim 6 , wherein the pitch lag estimator is configured to estimate the estimated pitch lag by determining two parameters a, b, by minimizing the error function

err

=

∑

i

=

0

4

time

passed

(

i

)

·

(

(

a

+

b

·

i

)

-

P

⁡

(

i

)

)

2

,

wherein a is a real number,

wherein b is a real number,

wherein time passed (i) is representing an inverse of an amount of time that has passed after correctly receiving a pitch lag,

wherein P (i) is the i-th original pitch lag value;

corresponding to the pich lag.

9. An apparatus according to claim 7 , wherein the pitch lag estimator is configured to determine the estimated pitch lag p according to

p=a·i+b.

10. A system for reconstructing a frame comprising a speech signal, wherein the system comprises:

an apparatus according to claim 1 for determining an estimated pitch lag, and

an apparatus for reconstructing the frame, wherein the apparatus for reconstructing the frame is configured to reconstruct the frame depending on the estimated pitch lag,

wherein the estimated pitch lag is a pitch lag of the speech signal.

11. A system for reconstructing a frame according to claim 10 ,

wherein the reconstructed frame is associated with one or more available frames, said one or more available frames being at least one of one or more preceding frames of the reconstructed frame and one or more succeeding frames of the reconstructed frame, wherein the one or more available frames comprise one or more pitch cycles as one or more available pitch cycles, and

wherein the apparatus for reconstructing the frame comprises

a determination unit for determining a sample number difference indicating a difference between a number of samples of one of the one or more available pitch cycles and a number of samples of a first pitch cycle to be reconstructed, and

a frame reconstructor for reconstructing the reconstructed frame by reconstructing, depending on the sample number difference and depending on the samples of said one of the one or more available pitch cycles, the first pitch cycle to be reconstructed as a first reconstructed pitch cycle,

wherein the frame reconstructor is configured to reconstruct the reconstructed frame, such that the reconstructed frame completely or partially comprises the first reconstructed pitch cycle, such that the reconstructed frame completely or partially comprises a second reconstructed pitch cycle, and such that the number of samples of the first reconstructed pitch cycle differs from a number of samples of the second reconstructed pitch cycle,

wherein the determination unit is configured to determine the sample number difference depending on the estimated pitch lag.

12. A method for generating a speech signal, comprising:

estimating an estimated pitch lag; and

generating the speech signal using the estimated pitch lag,

wherein estimating the estimated pitch lag is conducted depending on a plurality of original pitch lag values and depending on a plurality of information values,

wherein, for each original pitch lag value of the plurality of original pitch lag values, an information value of the plurality of information values is assigned to said original pitch lag value,

wherein estimating the estimated pitch lag is conducted by determining two parameters a and b, by minimizing an error function:

err

=

∑

k

i

=

s

g

p

(

i

)

·

(

(

a

+

b

·

i

)

-

P

⁡

(

i

)

)

2

wherein a is a real number, wherein b is a real number, wherein s is a first integer, wherein k is a second integer, and wherein P(i) is the i-th original pitch lag value, wherein g˜(i) is the i-th pitch gain value being assigned to the i-th pitch lag value P(i).

13. A non-transitory computer-readable medium comprising a computer program for implementing the method of claim 12 when being executed on a computer or signal processor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2022
From: LECOMTE, JÉRÉMIE; SCHNABEL, MICHAEL; MARKOVIC, GORAN; DIETZ, MARTIN; NEUGEBAUER, BERNHARD
To: FRAUNHOFER-GESELLSCHAFT ZUR FÖRDERUNG DER ANGEWANDTEN FORSCHUNG E.V.
Reel/Frame 061509/0157 →
Priority Claims (2)
EP 13173157 · Jun 21, 2013 · regional
EP 14166990 · May 5, 2014 · regional
Continuity (4)
Continuation 16445052 · Jun 18, 2019
Continuation 14977224 · Dec 21, 2015
Continuation PCTEP2014062589 · Jun 16, 2014
Related Publication 20220343924A1 · Oct 27, 2022
References Cited (174)
US 5179594A · Yip et al. · 1993 [cited by applicant]
US 5187745A · Yip et al. · 1993 [cited by applicant]
US 5621853A · Gardner · 1997 [cited by applicant]
US 5657419A · Yoo et al. · 1997 [cited by applicant]
US 5657422A · Janiszewski et al. · 1997 [cited by applicant]
US 5699485A · Shoham · 1997 [cited by applicant]
US 5781880A · Su · 1998 [cited by applicant]
US 5792072A · Keefe · 1998 [cited by applicant]
US 5946650A · Wei · 1999 [cited by examiner]
US 6035271A · Chen · 2000 [cited by examiner]
US 6456964B2 · Manjunath et al. · 2002 [cited by applicant]
US 6507814B1 · Gao · 2003 [cited by applicant]
US 6556966B1 · Gao · 2003 [cited by applicant]
US 6584438B1 · Manjunath et al. · 2003 [cited by applicant]
US 6781880B2 · Roohparvar et al. · 2004 [cited by applicant]
US 6782360B1 · Gao et al. · 2004 [cited by applicant]
US 7346110B2 · Minde · 2008 [cited by examiner]
US 7426466B2 · Ananthapadmanabhan et al. · 2008 [cited by applicant]
US 7590525B2 · Chen · 2009 [cited by applicant]
US 7873064B1 · Li et al. · 2011 [cited by applicant]
US 8255207B2 · Vaillancourt et al. · 2012 [cited by applicant]
US 8532984B2 · Rajendran · 2013 [cited by examiner]
US 8560329B2 · Qi et al. · 2013 [cited by applicant]
US 8725501B2 · Ehara · 2014 [cited by examiner]
US 8781880B2 · Kocsor et al. · 2014 [cited by applicant]
US 9043214B2 · Vos · 2015 [cited by examiner]
US 9280982B1 · Kushner · 2016 [cited by applicant]
US 10013988B2 · Lecomte et al. · 2018 [cited by applicant]
US 10381011B2 · Lecomte · 2019 [cited by examiner]
US 11410663B2 · Lecomte · 2022 [cited by examiner]
US 20020147583A1 · Gao · 2002 [cited by applicant]
US 20040002855A1 · Jabri et al. · 2004 [cited by applicant]
US 20040017811A1 · Lam · 2004 [cited by applicant]
US 20040044524A1 · Minde et al. · 2004 [cited by applicant]
US 20050137864A1 · Valve et al. · 2005 [cited by applicant]
US 20050216262A1 · Fejzo · 2005 [cited by applicant]
US 20060074641A1 · Goudar et al. · 2006 [cited by applicant]
US 20060089833A1 · Su et al. · 2006 [cited by applicant]
US 20060259296A1 · Lin · 2006 [cited by applicant]
US 20060271356A1 · Vos · 2006 [cited by applicant]
US 20060271357A1 · Wang et al. · 2006 [cited by applicant]
US 20060271373A1 · Khalil et al. · 2006 [cited by applicant]
US 20070206645A1 · Sundqvist et al. · 2007 [cited by applicant]
US 20070219788A1 · Gao · 2007 [cited by applicant]
US 20070239462A1 · Makinen et al. · 2007 [cited by applicant]
US 20070282603A1 · Bessette · 2007 [cited by applicant]
US 20080027715A1 · Rajendran et al. · 2008 [cited by applicant]
US 20080049795A1 · Lakaniemi · 2008 [cited by applicant]
US 20080071530A1 · Ehara · 2008 [cited by applicant]
US 20080147414A1 · Son et al. · 2008 [cited by applicant]
US 20080189101A1 · Jabri et al. · 2008 [cited by applicant]
US 20080294429A1 · Su et al. · 2008 [cited by applicant]
US 20090232228A1 · Thyssen · 2009 [cited by applicant]
US 20090234644A1 · Reznik et al. · 2009 [cited by applicant]
US 20090240491A1 · Reznik · 2009 [cited by applicant]
US 20090326930A1 · Kawashima et al. · 2009 [cited by applicant]
US 20100049511A1 · Ma et al. · 2010 [cited by applicant]
US 20100280823A1 · Shlomot et al. · 2010 [cited by applicant]
US 20110022924A1 · Malenovsky et al. · 2011 [cited by applicant]
US 20110125505A1 · Vaillancourt · 2011 [cited by examiner]
US 20110196673A1 · Sharma et al. · 2011 [cited by applicant]
US 20120072209A1 · Krishnan et al. · 2012 [cited by applicant]
US 20120209604A1 · Sehlstedt · 2012 [cited by applicant]
US 20120239389A1 · Jeon et al. · 2012 [cited by applicant]
US 20130041657A1 · Bradley · 2013 [cited by examiner]
US 20130124215A1 · Lecomte et al. · 2013 [cited by applicant]
US 20130144632A1 · Sung · 2013 [cited by applicant]
US 20140188465A1 · Choo et al. · 2014 [cited by applicant]
US 20150255079A1 · Huang et al. · 2015 [cited by applicant]
US 20160111094A1 · Lecomte et al. · 2016 [cited by applicant]
US 20220343924A1 · Lecomte · 2022 [cited by examiner]
CA 2483791A1 · 2003 [cited by applicant]
CN 1331825A · 2002 [cited by applicant]
CN 1432175A · 2003 [cited by applicant]
CN 1432176A · 2003 [cited by applicant]
CN 1455917A · 2003 [cited by applicant]
CN 1468427A · 2004 [cited by applicant]
CN 1659625A · 2005 [cited by applicant]
CN 1653521A · 2005 [cited by applicant]
CN 1983909A · 2007 [cited by applicant]
CN 1989548A · 2007 [cited by applicant]
CN 101046964A · 2007 [cited by applicant]
CN 101167125A · 2008 [cited by applicant]
CN 101199003A · 2008 [cited by applicant]
CN 101261833A · 2008 [cited by applicant]
CN 101364854B · 2009 [cited by applicant]
CN 101379551A · 2009 [cited by applicant]
CN 101627423A · 2010 [cited by applicant]
CN 102057424A · 2011 [cited by applicant]
CN 102203855A · 2011 [cited by applicant]
CN 102324236A · 2012 [cited by applicant]
CN 102449690A · 2012 [cited by applicant]
CN 102576540B · 2012 [cited by applicant]
CN 102834863A · 2012 [cited by applicant]
CN 103109318A · 2013 [cited by applicant]
CN 103109321A · 2013 [cited by applicant]
CN 103117062A · 2013 [cited by applicant]
EP 0424016A2 · 1991 [cited by applicant]
EP 0985328B1 · 2006 [cited by applicant]
EP 1850327A1 · 2007 [cited by applicant]
EP 1088302B1 · 2008 [cited by applicant]
EP 2107556A1 · 2009 [cited by applicant]
EP 2002427B1 · 2011 [cited by applicant]
FO 2830970A1 · 2003 [cited by applicant]
JP 2009003387A · 2009 [cited by applicant]
RU 2389085C2 · 2010 [cited by applicant]
RU 2418324C2 · 2011 [cited by applicant]
RU 2437172C1 · 2011 [cited by applicant]
RU 2459282C2 · 2012 [cited by applicant]
RU 2461898C2 · 2012 [cited by applicant]
WO 9847313A2 · 1998 [cited by applicant]
WO WO0011653A1 · 2000 [cited by applicant]
WO WO2004034376A2 · 2004 [cited by applicant]
WO WO2008007699A1 · 2008 [cited by applicant]
WO WO2008049221A1 · 2008 [cited by applicant]
WO WO2009059333A1 · 2009 [cited by applicant]
WO 2011042464A1 · 2011 [cited by applicant]
WO 2011048094A1 · 2011 [cited by applicant]
WO 2012110415A1 · 2012 [cited by applicant]
WO 2012110448A1 · 2012 [cited by applicant]
WO WO2012158159A1 · 2012 [cited by applicant]
WO 2014096279A1 · 2014 [cited by applicant]
Office Action dated Dec. 28, 2022 issued in the parallel Chinese patent application No. 201910627552.8 (20 pages). [cited by applicant]
3GPP; “3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Audio codec processing functions; Extended Adaptive Multi-Rage - Wideband (AMR-WB+) codec; Transcoding functions (Rel… [cited by applicant]
3GPP; “3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Mandatory Speech Codec speech processing functions; Adaptive Multi-Rate (AMR) speech codec; Error concealment of lost… [cited by applicant]
3GPP; “3rd Generation Partnership Project; Technical Specification Group Services and System Aspects; Speech Codec speech processing functions; Adaptive Multi-Rate—Wideband (AMR-WB) speech codec; Error concealment of er… [cited by applicant]
ITU-T; “G.719—Low-complexity, full-band audio coding for high-quality, conversational applications,” Series G: Transmission Systems and Media, Digital Systems and Networks / Digital terminal equipments—Coding of analogu… [cited by applicant]
ITU-T; “G.722—7 kHz audio-coding within 64 kbit/s—Appendix III: A high-quality packet loss concealment algorithm for G.722,” Series G: Transmission Systems and Media, Digital Systems and Networks / Digital terminal equi… [cited by applicant]
ITU-T; “G.722—7 kHz audio-coding within 64 kbit/s—Appendix IV: A low-complexity algorithm for packet-loss concealment with ITU-T G.722,” Series G: Transmission Systems and Media, Digital Systems and Networks / Digital t… [cited by applicant]
ITU-T; “G.722.2—Wideband coding of speech at around 16 kbit/s using Adaptive Multi-Rate Wideband (AMR-WB),” Series G: Transmission Systems and Media, Digital Systems and Networks / Digital terminal equipments ˜˜ Coding … [cited by applicant]
ITU-T: “G.729—Coding of speech at 8 kbit/s using conjugate-structure algebraic-code-excited linear prediction (CS-ACELP),” Series G: Transmission Systems and Media, Digital Systems and Networks / Digital terminal equipm… [cited by applicant]
Marques et al.; “Improved Pitch Prediction With Fractional Delays In CELP Coding,” 1990 International Conference on Acoustics, Speech, and Signal Processing, 1990; vol. 2; pp. 665-668. [cited by applicant]
Chibani et al.; “Fast Recovery for a CELP-Like Speech Codec After a Frame Erasure,” IEEE Transactions on Audio, Speech, and Language Processing, Nov. 2007; 15(8):2485-2495. [cited by applicant]
International Search Report in related PCT Application No. PCT/EP2014/062589 dated Oct. 8, 2014 (8 pages). [cited by applicant]
ITU-T; “G.729-based embedded variable bit-rate coder: An 8-32 kbit/s scalable wideband coder bitstream interoperable with G.729,” Series G: Transmission Systems and Media, Digital Systems and Networks, Digital terminal … [cited by applicant]
ITU-T; “Frame error robust narrow-band and wideband embedded variable bit-rate coding of speech and audio from 8-32 kbit/s,” Series G: Transmission Systems and Media, Digital Systems and Networks, Digital terminal equip… [cited by applicant]
Mu et al.; “A Frame Erasure Concealment Method Based on Pitch and Gain Linear Prediction for AMR-WB Codec,” 2011 IEEE International Conference on Consumer Electronics (ICCE), Jan. 9, 2011; pp. 815-816. [cited by applicant]
Anderson, Kyle and Gournay, Philippe; Pitch Resynchronization While Recovering From A Late Frame In A Predictive Speech Decoder (Interspeech Sep. 17-21, 2006)—ICSLP; http://www.gel.usherbrooke.ca/gournay/documents/publi… [cited by applicant]
Office Action issued in parallel Japanese patent application No. 2016-520421 dated May 2, 2017 (8 pages). [cited by applicant]
Office Action issued in co-pending U.S. Appl. No. 14/977,195 dated May 26, 2017 (39 pages). [cited by applicant]
Notice of Allowance dated Feb. 20, 2018 issued in co-pending U.S. Appl. No. 14/977,195 (28 pages). [cited by applicant]
Corrected Notice of Allowability dated Mar. 16, 2018 issued in co-pending U.S. Appl. No. 14/977,195 (13 pages). [cited by applicant]
Office Action dated Sep. 3, 2018 in the parallel Chinese patent application No. 201480035427.3 (31 pages with English translation). [cited by applicant]
Office Action with Search Report dated Sep. 18, 2018 issued in the parallel Chinese patent application No. 201480035474.8 (21 pages). [cited by applicant]
Examination Report dated Mar. 4, 2019 issued in parallel Indian patent application No. 3984/KOLNP/2015 (6 pages). [cited by applicant]
Office Action dated Feb. 11, 2019 issued in the parallel TW patent application No. 106123342 (13 pages). [cited by applicant]
Decision to Grant dated Apr. 29, 2019 issued in the parallel Chinese patent application No. 201480035474.8. [cited by applicant]
Fraunhofer IIS: Tdoc S4-130345, Qualification Deliverables for the Fraunhofer IIS Candidate for EVS (inclusing Technical Description and Report on Compliance to Design Constraints), TSG SA4#72bis meeting, Mar. 11-15, 20… [cited by applicant]
3GPP TS 26.290 V2.0.0 (Sep. 2004). [cited by applicant]
3GPP TS 26.290 V10.0.0 (Mar. 2011). [cited by applicant]
3GPP TS 26.290 V6.1.0 (Dec. 2004). [cited by applicant]
3GPP TS 26.403 V6.0.0 (Sep. 2004). [cited by applicant]
3GPP TS 26.442 V14.0.0 (Mar. 2017). [cited by applicant]
3GPP TS 26.443 14.0.0 (Mar. 2017). [cited by applicant]
3GPP TS 26.445 V12.0.0 (Sep. 2014). [cited by applicant]
3GPP TS 26.445 V14.0.0 (Mar. 2017). [cited by applicant]
3GPP TS 26.445 V14.2.0 (Dec. 2017). [cited by applicant]
3GPP TS 26.445 V16.2.0 (Dec. 2021). [cited by applicant]
3GPP TS 26.445 V17.0.0 (Apr. 2022). [cited by applicant]
3GPP TS 26.447 V14.0.0 (Mar. 2017). [cited by applicant]
3GPP TS 26.447 V14.2.0 (Jun. 2020). [cited by applicant]
3GPP TS 26.447 V16.0.0 (Mar. 2019). [cited by applicant]
3GPP TS 26.952 v17.0.0 (Apr. 2022). [cited by applicant]
Fuchs et al, MDCT-Based Coder for Highly Adaptive Speech and Audio Coding, 17th European Signal Processing Conference (EUSIPCO 2009), Glasgow, Scotland, Aug. 24-28, 2009. [cited by applicant]
Schnell et al., Low Delay Filter banks for Enhanced Low Delay Audio Coding, 2007 IEEE Workshop on Applications of Signal Processing to Audio and Accoustics Oct. 21, 2007. [cited by applicant]
Marina Bosi and Richard E. Goldberg, Introduction to Digital Audio Coding and Standards, Springer 2003. [cited by applicant]
Convolution theorem—Wikipedia. [cited by applicant]
ITU-T G.722 (Jul. 2003). [cited by applicant]
ITU-T G.7222 (Jan. 2002) of the Telecommunication Standardization Sector of the International Telecommunication Union (“G.722.2”), Annex A. [cited by applicant]
Ravishankar, C., Hughes Network Systems, Germantown, MD. Speech coding. United States, https://doi.org/10.2172/325392. [cited by applicant]
Huan Hou and Weibei Dou, Real-time Audio Error Concealment Method Based on Sinusoidal Model, International Conference on audio Language and Image Processing, IEEE, Jul. 2008, Shanghai, P.R. China, DOI:10.1109/ICALIP.200… [cited by applicant]
ITU-T G.718 (Jun. 2008), Series G: Transmission Systems and Media, Digital Systems and Networks, Frame error robust narrow-brand and wideband embedded variable bit-rate coding of speech and audio from 8-32 kbit/s Digita… [cited by applicant]
Virette, D., Low Delay Transform for High Quality Low Delay Audio Coding, Signal and Image Processing, (Université de Rennes 1, 2012), 40-41. [cited by applicant]
Ostergaard, J., et al., Real-time perceptual moving-horizon multiple-description audio coding, IEEE Transactions on Signal Processing, 4286 (2011). [cited by applicant]