IP Library Granted Patent US 12,183,326
Granted Patent B2
US 12,183,326 · App. 17/773,641 · Granted Dec 31, 2024

Speech recognition error correction method, related devices, and readable storage medium

Inventors: Li Xu (Anhui, CN); Jia Pan (Anhui, CN); Zhiguo Wang (Anhui, CN); Guoping Hu (Anhui, CN)
G10L15/06G10L15/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,183,326
App. No.
17/773,641
Filed
May 2, 2022
Granted
Dec 31, 2024
Kind
B2
Art Unit
2655
USPC
704/243
Abstract

A speech recognition error correction method and device, and a readable storage medium are provided. The method includes: acquiring to-be-recognized speech data and a first recognition result of the speech data, re-recognizing the speech data with reference to context information in the first recognition result to obtain a second recognition result, and determining a final recognition result based on the second recognition result. In the method, the speech data is re-recognized with reference to context information in the first recognition result, which fully considers context information in the recognition result and the application scenario of the speech data. If any error occurs in the first recognition result, the first recognition result is corrected based on the second recognition. Therefore, the accuracy of speech recognition can be improved.

Claims (41)

1. A speech recognition error correction method comprising:

acquiring to-be-recognized speech data and a first recognition result of the speech data;

extracting a correctly recognized field-specific word from the first recognition result as a keyword;

re-recognizing the speech data with reference to context information in the first recognition result and the keyword to obtain a second recognition result, comprising:

acquiring an acoustic feature of the speech data; and

inputting the acoustic feature of the speech data, the first recognition result and the keyword to a pre-trained speech recognition error correction model to obtain the second recognition result, wherein the speech recognition error correction model is obtained by training a preset model with an error-correction training data set, wherein the speech recognition error correction model is a neural network-based model, wherein the inputting the acoustic feature of the speech data, the first recognition result and the keyword to the pre-trained speech recognition error correction model to obtain the second recognition result comprises performing encoding and attention calculation on the acoustic feature of the speech data, the first recognition result and the keyword by using the speech recognition error correction model, to obtain the second recognition result based on a calculation result, wherein the performing encoding and attention calculation on the acoustic feature of the speech data, the first recognition result and the keyword by using the speech recognition error correction model to obtain the second recognition result based on the calculation result comprises:

performing encoding and attention calculation on each of the acoustic feature of the speech data, the first recognition result and the keyword by using an encoding layer and an attention layer of the speech recognition error correction model to obtain the calculation result; and

decoding the calculation result by using a decoding layer of the speech recognition error correction model to obtain the second recognition result; and

determining a final recognition result based on the second recognition result.

2. The method according to claim 1 , wherein

the error-correction training data set comprises at least one group of error-correction training data, and each group of error-correction training data comprises an acoustic feature of a piece of speech data, a text corresponding to the piece of speech data, a first recognition result corresponding to the piece of speech data, and a keyword in the first recognition result.

3. The method according to claim 1 , wherein the performing encoding and attention calculation on the acoustic feature of the speech data, the first recognition result and the keyword by using the speech recognition error correction model to obtain the second recognition result based on a calculation result comprises:

merging the acoustic feature of the speech data, the first recognition result and the keyword to obtain a merged vector;

performing, by an encoding layer and an attention layer of the speech recognition error correction model, encoding and attention calculation on the merged vector to obtain the calculation result; and

decoding, by a decoding layer of the speech recognition error correction model, the calculation result to obtain the second recognition result.

4. The method according to claim 3 , wherein the performing, by the encoding layer and the attention layer of the speech recognition error correction model, encoding and attention calculation on the merged vector to obtain the calculation result comprises:

encoding, by the encoding layer of the speech recognition error correction model, the merged vector to obtain an acoustic advanced feature of the merged vector;

performing, by the attention layer of the speech recognition error correction model, attention calculation on a previous semantic vector related to the merged vector and a previous output result of the speech recognition error correction model to obtain a hidden layer state related to the merged vector; and

performing, by the attention layer of the speech recognition error correction model, attention calculation on the acoustic advanced feature of the merged vector and the hidden layer state related to the merged vector to obtain a semantic vector related to the merged vector.

5. The method according to claim 1 , wherein the performing encoding and attention calculation on each of the acoustic feature of the speech data, the first recognition result and the keyword by using an encoding layer and an attention layer of the speech recognition error correction model to obtain the calculation result comprises:

for each target object:

encoding, by the encoding layer of the speech recognition error correction model, the target object to obtain an acoustic advanced feature of the target object;

performing, by the attention layer of the speech recognition error correction model, attention calculation on a previous semantic vector related to the target object and a previous output result of the speech recognition error correction model to obtain a hidden layer state related to the target object; and

performing, by the attention layer of the speech recognition error correction model, attention calculation on the acoustic advanced feature of the target object and the hidden layer state related to the target object to obtain a semantic vector related to the target object,

wherein the target object comprises the acoustic feature of the speech data, the first recognition result, and the keyword.

6. The method according to claim 1 , wherein the determining a final recognition result based on the second recognition result comprises:

acquiring a confidence of the first recognition result and a confidence of the second recognition result; and

determining one with a higher confidence between the first recognition result and the second recognition result as the final recognition result.

7. A speech recognition error correction system comprising:

a memory configured to store a program; and

a processor configured to execute the program to perform the speech recognition error correction method according to claim 1 .

8. A non-transitory machine readable storage medium storing a computer program, wherein the computer program, when being executed by a processor, implements the speech recognition error correction method according to claim 1 .

9. A speech recognition error correction device comprising:

an acquisition unit configured to acquire to-be-recognized speech data and a first recognition result of the speech data;

a keyword extraction unit configured to extract a correctly recognized field-specific word from the first recognition result as a keyword;

a second speech recognition unit configured to re-recognize the speech data with reference to context information in the first recognition result and the keyword to obtain a second recognition result, comprising:

acquiring an acoustic feature of the speech data; and

inputting the acoustic feature of the speech data, the first recognition result and the keyword to a pre-trained speech recognition error correction model to obtain the second recognition result, wherein the speech recognition error correction model is obtained by training a preset model with an error-correction training data set, wherein the speech recognition error correction model is a neural network-based model, wherein the inputting the acoustic feature of the speech data, the first recognition result and the keyword to the pre-trained speech recognition error correction model to obtain the second recognition result comprises performing encoding and attention calculation on the acoustic feature of the speech data, the first recognition result and the keyword by using the speech recognition error correction model, to obtain the second recognition result based on a calculation result, wherein the performing encoding and attention calculation on the acoustic feature of the speech data, the first recognition result and the keyword by using the speech recognition error correction model to obtain the second recognition result based on the calculation result comprises:

performing encoding and attention calculation on each of the acoustic feature of the speech data, the first recognition result and the keyword by using an encoding layer and an attention layer of the speech recognition error correction model to obtain the calculation result; and

decoding the calculation result by using a decoding layer of the speech recognition error correction model to obtain the second recognition result; and

a recognition result determination unit configured to determine a final recognition result based on the second recognition result.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2022
From: XU, LI; PAN, JIA; WANG, ZHIGUO; HU, GUOPING
To: IFLYTEK CO., LTD.
Reel/Frame 059769/0121 →
Priority Claims (1)
CN 201911167009.0 · Nov 25, 2019 · national
Continuity (1)
Related Publication 20220383853A1 · Dec 1, 2022
Cited By (2)
US 12,519,568 US 12,525,222