IP Library › Granted Patent US 10,923,117
Granted Patent B2
US 10,923,117 · App. 16/279,491 · Granted Feb 16, 2021

Best path change rate for unsupervised language model weight selection

Inventors: Peidong Wang (Columbus, OH); Jia Cui (Bellevue, WA); Chao Weng (Fremont, CA); Dong Yu (Bothell, WA)
Assignee: TENCENT AMERICA LLC
G10L15/183G10L15/063G10L25/03G10L25/27G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,923,117
App. No.
16/279,491
Granted
Feb 16, 2021
Kind
B2
Abstract

A method for selecting an optimal language model weight (LMW) used to perform automatic speech recognition, including decoding test audio into a lattice using a language model; analyzing the lattice using a first LMW of a plurality of LMWs to determine a first plurality of best paths; analyzing the lattice using a second LMW of the plurality of LMWs to determine a second plurality of best paths; determining a first best path change rate (BCPR) corresponding to the first LMW based on a number of best path changes between the first plurality of best paths and the second plurality of best paths; and determining the first LMW to be the optimal LMW based on the first BCPR being a lowest BCPR from among a plurality of BCPRs corresponding to the plurality of LMWs.

Claims (187)

1. A method for selecting an optimal language model weight (LMW) used to perform automatic speech recognition, the method comprising:

decoding test audio into a lattice using a language model;

analyzing the lattice using a first LMW of a plurality of LMWs to determine a first plurality of best paths;

analyzing the lattice using a second LMW of the plurality of LMWs to determine a second plurality of best paths;

determining a first best path change rate (BCPR) corresponding to the first LMW based on a number of best path changes between the first plurality of best paths and the second plurality of best paths; and

determining the first LMW to be the optimal LMW based on the first BCPR being a lowest BCPR from among a plurality of BCPRs corresponding to the plurality of LMWs.

2. The method of claim 1 , wherein the plurality of LMWs comprise a plurality of scaling factors used to combine the language model with an acoustic model.

3. The method of claim 1 , wherein each LMW of the plurality of LMWs is separated from neighboring LMWs of the plurality of LMWs by a fixed step size.

4. The method of claim 1 , wherein each score of a plurality of scores corresponding to the first plurality of best paths comprises a combination of an acoustic score and a language model score balanced by the first LMW, and wherein each best path of the first plurality of best paths comprises a path in the lattice with a largest score.

5. The method of claim 1 , wherein the number of the best path changes is represented as follows:

c

⁡

(

w

)

=

∑

s

∈

S

⁢

δ

⁡

(

p

s

⁡

(

w

)

≠

p

s

⁡

(

w

+

ϵ

)

)

,

wherein c(w) denotes the number of the best path changes, s is an index of an utterance, S denotes utterances in a data set corresponding to the test audio, δ( ) denotes a Kronecker delta function, p s ( ) denotes a best path, w denotes the first LMW, ε denotes a fixed step size, and w+ε denotes the second LMW.

6. The method of claim 5 , wherein the first BCPR is represented as follows:

r

⁡

(

w

)

=

c

⁡

(

w

)

N

,

wherein r(w) denotes the first BCPR and N denotes a total number of the utterances in the data set.

7. The method of claim 6 , wherein a continuous function corresponding to the BCPR is represented as follows:

R (ω c )= I r (ω c ) +D r (ω c ),

wherein R(w c ) denotes the continuous function corresponding to the BCPR, wherein I r (w c ) denotes a normalized number of best paths which increase a word error rate after a LMW change, and D r (w c ) denotes a normalized number of best paths which decrease the word error rate after the LMW change.

8. The method of claim 7 , wherein I r (w c ) and D r (w c ) are represented as follows:

I r (ω c )| ω c ∈[ω opt −ε ] =−k*ω c +b I , and

D r (ω c )| ω c ∈[ω opt −ε,ω opt +ε ] =k*ω c +b D ,

wherein w opt denotes the optimal LMW, k denotes (k≥0) a slope, b I denotes an intercept for I r (w c ), and b D denotes an intercept for D r (w c ).

9. The method of claim 8 , wherein R(w c ) is represented as follows:

R (ω c )| ω c ∈[ω opt −ε,ω opt ε ] =I r (ω c )+ D r (ω c )= b I +b D , and

wherein R′(w opt ) is represented as follows:

R

′

⁡

(

w

opt

)

=

dR

⁡

(

w

c

)

dw

c

⁢

|

w

opt

=

0.

10. The method of claim 9 , wherein the first LMW is determined to be the optimal LMW based on the first BCPR being closest to a minimum value of R′(w opt ) from among the plurality of BCPRs.

11. A device for selecting an optimal language model weight (LMW) used to perform automatic speech recognition, the device comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code including:

decoding code configured to cause the at least one processor to decode test audio into a lattice using a language model,

first analyzing code configured to cause the at least one processor to analyze the lattice using a first LMW of a plurality of LMWs to determine a first plurality of best paths,

second analyzing code configured to cause the at least one processor to analyze the lattice using a second LMW of the plurality of LMWs to determine a second plurality of best paths,

first determining code configured to cause the at least one processor to determine a first best path change rate (BCPR) corresponding to the first LMW based on a number of best path changes between the first plurality of best paths and the second plurality of best paths, and

second determining code configured to cause the at least one processor to determine the first LMW to be the optimal LMW based on the first BCPR being a lowest BCPR from among a plurality of BCPRs corresponding to the plurality of LMWs.

12. The device of claim 11 , wherein the plurality of LMWs comprise a plurality of scaling factors used to combine the language model with an acoustic model.

13. The device of claim 11 , wherein each LMW of the plurality of LMWs is separated from neighboring LMWs of the plurality of LMWs by a fixed step size.

14. The device of claim 11 , wherein the number of the best path changes is represented as follows:

c

⁡

(

w

)

=

∑

s

∈

S

⁢

δ

⁡

(

p

s

⁡

(

w

)

≠

p

s

⁡

(

w

+

ϵ

)

)

,

wherein c(w) denotes the number of the best path changes, s is an index of an utterance, S denotes utterances in a data set corresponding to the test audio, δ( ) denotes a Kronecker delta function, p s ( ) denotes a best path, w denotes the first LMW, ε denotes a fixed step size, and w+ε denotes the second LMW.

15. The device of claim 14 , wherein the first BCPR is represented as follows:

r

⁡

(

w

)

=

c

⁡

(

w

)

N

,

wherein r(w) denotes the first BCPR and N denotes a total number of the utterances in the data set.

16. The device of claim 15 , wherein a continuous function corresponding to the BCPR is represented as follows:

R (ω c )= I r (ω c ) +D r (ω c ),

wherein R(w c ) denotes the continuous function corresponding to the BCPR, wherein I r (w c ) denotes a normalized number of best paths which increase a word error rate after a LMW change, and D r (w c ) denotes a normalized number of best paths which decrease the word error rate after the LMW change.

17. The device of claim 16 , wherein I r (w c ) and D r (w c ) are represented as follows:

I r (ω c )| ω c ∈[ω opt −ε,ω opt +ε ] =−k*ω c +b I , and

D r (ω c )| ω c ∈[ω opt −ε,ω opt +ε ] =k*ω c +b D ,

wherein w opt denotes the optimal LMW, k denotes (k≥0) a slope, b I denotes an intercept for I r (w c ), and b D denotes an intercept for D r (w c ).

18. The device of claim 17 , wherein R(w c ) is represented as follows:

R (ω c )| ω c ∈[ω opt −ε,ω opt +ε ] =I r (ω c )+ D r (ω c )= b I +b D , and

wherein R′(w opt ) is represented as follows:

R

′

⁡

(

w

opt

)

=

dR

⁡

(

w

c

)

dw

c

⁢

|

w

opt

=

0.

19. The device of claim 18 , wherein the first LMW is determined to be the optimal LMW based on the first BCPR being closest to a minimum value of R′(w opt ) from among the plurality of BCPRs.

20. A non-transitory computer-readable medium storing instructions, the instructions comprising: one or more instructions that, when executed by one or more processors of a device for selecting an optimal language model weight (LMW) used to perform automatic speech recognition, cause the one or more processors to:

decode test audio into a lattice using a language model,

analyze the lattice using a first LMW of a plurality of LMWs to determine a first plurality of best paths,

analyze the lattice using a second LMW of the plurality of LMWs to determine a second plurality of best paths,

determine a first best path change rate (BCPR) corresponding to the first LMW based on a number of best path changes between the first plurality of best paths and the second plurality of best paths, and

determine the first LMW to be the optimal LMW based on the first BCPR being a lowest BCPR from among a plurality of BCPRs corresponding to the plurality of LMWs.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2020
From: WENG, CHAO
To: TENCENT AMERICA LLC
Reel/Frame 053931/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2020
From: YU, DONG
To: TENCENT AMERICA LLC
Reel/Frame 053932/0185 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2020
From: CUI, JIA
To: TENCENT AMERICA LLC
Reel/Frame 053932/0450 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2020
From: WANG, PEIDONG
To: TENCENT AMERICA LLC
Reel/Frame 053933/0186 →
Continuity (1)
Related Publication 20200265833A1 · Aug 20, 2020