IP Library › Granted Patent US 12,223,947
Granted Patent B2
US 12,223,947 · App. 17/761,217 · Granted Feb 11, 2025

Decoding network construction method, voice recognition method, device and apparatus, and storage medium

Inventors: Jianqing Gao (Anhui, CN); Zhiguo Wang (Anhui, CN); Guoping Hu (Anhui, CN)
Assignee: IFLYTEK CO., LTD.
G10L15/083G10L15/183G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,947
App. No.
17/761,217
Granted
Feb 11, 2025
Kind
B2
Abstract

A method for constructing a decoding network, a speech recognition method, a device, an apparatus, and a storage medium are provided. The method for constructing a decoding network includes: acquiring a general language model, a domain language model, and a general decoding network generated based on the general language model; generating a domain decoding network based on the domain language model and the general language model; and integrating the domain decoding network with the general decoding network to obtain a target decoding network. The speech recognition method includes: decoding to-be-recognized speech data by using a target decoding network to obtain a decoding path for the to-be-recognized speech data; and determining a speech recognition result for the to-be-recognized speech data based on the decoding path for the to-be-recognized speech data.

Claims (41)

1. A method for constructing a decoding network, the method comprising:

acquiring a universal language model, a domain language model, and a general decoding network generated based on the universal language model;

generating a domain decoding network based on the domain language model and the universal language model, the generating the domain decoding network based on the domain language model and the universal language model comprises:

performing interpolation on the universal language model and the domain language model, wherein a part on which the interpolation is performed comprises all parts in the domain language model and a part in the universal language model which also appears in the domain language model; and

generating the domain decoding network based on the part on which the interpolation is performed; and

integrating the domain decoding network with the general decoding network to obtain a target decoding network.

2. The method for constructing a decoding network according to claim 1 , wherein the integrating the domain decoding network with the general decoding network to obtain a target decoding network comprises:

cascading the domain decoding network and the general decoding network to obtain the target decoding network.

3. The method for constructing a decoding network according to claim 2 , wherein the cascading the domain decoding network and the general decoding network comprises:

adding virtual nodes for each of the general decoding network and the domain decoding network, wherein the virtual nodes comprise a start node and an end node; and

cascading the general decoding network and the domain decoding network by means of the start node and the end node.

4. The method for constructing a decoding network according to claim 3 , wherein the cascading the general decoding network and the domain decoding network by means of the start node and the end node comprises:

connecting the end node for the general decoding network and the start node for the domain decoding network in a direction from the end node for the general decoding network to the start node for the domain decoding network; and

connecting the end node for the domain decoding network and the start node for the general decoding network in a direction from the end node for the domain decoding network to the start node for the general decoding network.

5. A speech recognition method, comprising:

decoding to-be-recognized speech data by using a target decoding network to obtain a decoding path for the to-be-recognized speech data, wherein the target decoding network is constructed by using the method for constructing a decoding network according to claim 1 ; and

determining a speech recognition result for the to-be-recognized speech data based on the decoding path for the to-be-recognized speech data.

6. The speech recognition method according to claim 5 , wherein the determining a speech recognition result for the to-be-recognized speech data based on the decoding path for the to-be-recognized speech data comprises:

determining the speech recognition result for the to-be-recognized speech data based on a high-ordered language model obtained in advance and the decoding path for the to-be-recognized speech data, wherein the high-ordered language model is obtained by performing interpolation on the universal language model by using the domain language model.

7. The speech recognition method according to claim 5 , wherein a process of decoding to-be-recognized speech data by using a target decoding network to obtain a decoding path for the to-be-recognized speech data comprises:

inputting speech frames of the to-be-recognized speech data into the target decoding network sequentially for decoding, to obtain the decoding path for the to-be-recognized speech data, wherein

the speech frames of the to-be-recognized speech data enter, respectively via two start nodes in the target decoding network, the general decoding network and the domain decoding network in the target decoding network for decoding, and

in a case where a candidate decoding path in the general decoding network or the domain decoding network comprises an end node, the process jumps from the end node to at least one start node connected to the end node, the general decoding network and/or the domain decoding network is entered to continue decoding until an end of the speech frames.

8. A speech recognition apparatus, comprising:

a memory configured to store a program; and

a processor configured to execute the program to perform the speech recognition method according to claim 5 .

9. A non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when being executed by a processor, causes the processor to perform the speech recognition method according to claim 5 .

10. An apparatus for constructing a decoding network, the apparatus comprising:

a memory configured to store a program; and

a processor configured to execute the program to perform the method for constructing a decoding network according to claim 1 .

11. A non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when being executed by a processor, causes the processor to perform the method for constructing a decoding network according to claim 1 .

12. A device for constructing a decoding network, the device comprising a processor to:

acquire a universal language model, a domain language model, and a general decoding network generated based on the universal language model;

generate a domain decoding network based on the universal language model and the domain language model;

perform interpolation on the universal language model and the domain language model, wherein a part on which the interpolation is performed comprises all parts in the domain language model and a part in the universal language model which also appears in the domain language model;

generate the domain decoding network based on the part on which the interpolation is performed; and

integrate the domain decoding network with the general decoding network to obtain a target decoding network.

13. The device for constructing a decoding network according to claim 12 , wherein the processor is further configured to cascade the domain decoding network and the general decoding network to obtain the target decoding network.

14. A speech recognition device, the device comprising a processor to:

decode to-be-recognized speech data by using a target decoding network to obtain a decoding path for the to-be-recognized speech data, wherein the target decoding network is constructed by the device for constructing a decoding network according to claim 12 ; and

determine a speech recognition result for the to-be-recognized speech data based on the decoding path for the to-be-recognized speech data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2022
From: GAO, JIANQING; WANG, ZHIGUO; HU, GUOPING
To: IFLYTEK CO., LTD.
Reel/Frame 059289/0175 →
Priority Claims (1)
CN 201910983196.3 · Oct 16, 2019 · national
Continuity (1)
Related Publication 20220375459A1 · Nov 24, 2022
References Cited (43)
US 9460088B1 · Sak et al. · 2016 [cited by applicant]
US 10115393B1 · Kumar · 2018 [cited by examiner]
US 10402493B2 · Spencer · 2019 [cited by examiner]
US 20040138888A1 · Ramabadran · 2004 [cited by applicant]
US 20120053935A1 · Malegaonkar et al. · 2012 [cited by applicant]
US 20140236591A1 · Yue · 2014 [cited by examiner]
US 20150348542A1 · Pan et al. · 2015 [cited by applicant]
US 20170092266A1 · Wasserblat et al. · 2017 [cited by applicant]
US 20180204565A1 · Cohen et al. · 2018 [cited by applicant]
US 20180277103A1 · Wu et al. · 2018 [cited by applicant]
US 20190005947A1 · Kim · 2019 [cited by examiner]
US 20190122651A1 · Arik · 2019 [cited by examiner]
US 20190172466A1 · Lee · 2019 [cited by examiner]
US 20190189115A1 · Hori et al. · 2019 [cited by applicant]
US 20190279618A1 · Yadav · 2019 [cited by examiner]
US 20190279646A1 · Tian · 2019 [cited by examiner]
US 20190318725A1 · Le Roux · 2019 [cited by examiner]
US 20200020319A1 · Malhotra · 2020 [cited by examiner]
US 20200135174A1 · Cui · 2020 [cited by examiner]
CA 3082402A1 · 2019 [cited by examiner]
CN 103065630A · 2013 [cited by applicant]
CN 103077708A · 2013 [cited by applicant]
CN 103700369A · 2014 [cited by applicant]
CN 104064184A · 2014 [cited by applicant]
CN 104282301A · 2015 [cited by applicant]
CN 106294460A · 2017 [cited by applicant]
CN 108305634A · 2018 [cited by applicant]
CN 108538285A · 2018 [cited by applicant]
CN 108932944A · 2018 [cited by applicant]
CN 110120221A · 2019 [cited by applicant]
CN 110322884A · 2019 [cited by applicant]
CN 110428819B · 2020 [cited by examiner]
RU 2366007C2 · 2009 [cited by applicant]
WO 2014117577A1 · 2014 [cited by applicant]
WO 2014183373A1 · 2014 [cited by applicant]
WO WO2017114172A1 · 2017 [cited by examiner]
WO 2019103936A1 · 2019 [cited by applicant]
WO 2019116604A1 · 2019 [cited by applicant]
European Search Report issued in EP Application No. 19949233, mailed Oct. 13, 2023, 8 pages. [cited by applicant]
International Search Report and Written Opinion issued in PCT/CN2019/124790 mailed Jul. 1, 2020, 18 pages. [cited by applicant]
Guo et al., “Optimized Large Vocabulary WFST Speech Recognition System” IEEE, 2012, pp. 1243-1247. [cited by applicant]
Nanjing University of Science & Technology, Research on Chinese Speech:, 2012, pp. 1-68. [cited by applicant]
Russian First Office Action issued in 2022108762/28(018248) mailed Dec. 12, 2022, 20 pages. [cited by applicant]