IP Library › Granted Patent US 12,033,615
Granted Patent B2
US 12,033,615 · App. 17/499,129 · Granted Jul 9, 2024

Method and apparatus for recognizing speech, electronic device and storage medium

Inventors: Yinlou Zhao (Beijing, CN); Liao Zhang (Beijing, CN); Zhengxiang Jiang (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
G10L15/005G10L15/142G10L15/16G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,033,615
App. No.
17/499,129
Granted
Jul 9, 2024
Kind
B2
Abstract

The disclosure provides a method and an apparatus for recognizing a speech, an electronic device and a storage medium. A speech to be recognized is obtained. An acoustic feature of the speech to be recognized and a language feature of the speech to be recognized are obtained. The speech to be recognized is input to a pronunciation difference statistics to generate a differential pronunciation pair corresponding to the speech to be recognized. The text information of the speech to be recognized is generated based on the differential pronunciation pair, the acoustic feature and the language feature.

Claims (69)

1. A method for recognizing a speech, performed by a speech recognition device including a search engine, an acoustic model, a language model and a decoder, the method comprising:

obtaining, through the search engine, a speech to be recognized;

obtaining an acoustic feature of the speech to be recognized by inputting the speech to be recognized into the acoustic model and obtaining a language feature of the speech to be recognized by inputting the speech to be recognized into the language model, wherein the acoustic model is a model trained by Gaussian Mixed Model (GMM)-Hidden Markov Model (HMI) or Deep Neural Network (DNN)-HMM, and the language model is a model trained by an N-Gram (which is a statistic-based language model) or a Neural Network Language Model (NNLM);

inputting the speech to be recognized to a pronunciation difference statistics to generate a differential pronunciation pair corresponding to the speech to be recognized; and

generating, through the decoder, text information of the speech to be recognized based on the differential pronunciation pair, the acoustic feature, and the language feature;

wherein the pronunciation difference statistics is trained by:

obtaining a target sample text under a target scene;

generating a sample recognition result by recognizing the target sample text;

obtaining a first audio corresponding to the target sample text and obtaining a second audio corresponding to the sample recognition result;

obtaining a differential pronunciation pair sample corresponding to a pronunciation difference between the first audio and the second audio being greater than a preset threshold; and

training the pronunciation difference statistics based on the differential pronunciation pair sample.

2. The method of claim 1 , wherein obtaining the target sample text under the target scene comprises:

obtaining a sample text;

determining whether the sample text belongs to a target scene by inputting the sample text to a target scene text classifier;

determining the sample text as the target sample text based on the sample text belonging to the target scene; and

discarding the sample text based on the sample text not belonging to the target scene.

3. The method of claim 2 , wherein, the target scene text classifier is trained by:

obtaining a target scene sample and a non-target scene sample;

obtaining a first word vector representation of the target scene sample and a second word vector representation of the non-target scene sample;

inputting the first word vector representation as a positive sample and the second word vector representation as a negative sample to an initial target scene text classifier to train the initial target scene text classifier.

4. The method of claim 1 , wherein generating the text information of the speech to be recognized based on the differential pronunciation pair sample, the acoustic feature, and the language feature comprises:

inputting the differential pronunciation pair, the acoustic feature, and the language feature to a decoder to generate the text information of the speech to be recognized.

5. An electronic device, comprising:

a search engine, an acoustic model, a language model, and a decoder;

at least one processor; and

a memory communicatively coupled to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, wherein the instructions are executed by the at least one processor, so that

the search engine obtains a speech to be recognized;

the processor obtains an acoustic feature of the speech to be recognized by inputting the speech to be recognized into the acoustic model and obtains a language feature of the speech to be recognized by inputting the speech to be recognized into the language model, wherein the acoustic model is a model trained by Gaussian Mixed Model (GMM)-Hidden Markov Model (HMM) or Deep Neural Network (DNN)-HMM, and the language model is a model trained by an N-Gram (which is a statistic-based language model) or a Neural Network Language Model (NNLM);

the processor inputs the speech to be recognized to a pronunciation difference statistics to generate a differential pronunciation pair corresponding to the speech to be recognized; and

the decoder generates text information of the speech to be recognized based on the differential pronunciation pair, the acoustic feature, and the language feature;

wherein the pronunciation difference statistics is trained by:

obtaining a target sample text under a target scene;

generating a sample recognition result by recognizing the target sample text;

obtaining a first audio corresponding to the target sample text and obtaining a second audio corresponding to the sample recognition result;

obtaining a differential pronunciation pair sample corresponding to a pronunciation difference between the first audio and the second audio being greater than a preset threshold; and

training the pronunciation difference statistics based on the differential pronunciation pair sample.

6. The electronic device of claim 5 , wherein the processor is further configured to:

obtain a sample text;

determine whether the sample text belongs to a target scene by inputting the sample text to a target scene text classifier;

determine the sample text as the target sample text based on the sample text belonging to the target scene; and

discard the sample text based on the sample text not belonging to the target scene.

7. The electronic device of claim 6 , wherein the target scene text classifier is trained by:

obtaining a target scene sample and a non-target scene sample;

obtaining a first word vector representation of the target scene sample and a second word vector representation of the non-target scene sample;

inputting the first word vector representation as a positive sample and the second word vector representation as a negative sample to an initial target scene text classifier to train the initial target scene text classifier.

8. The electronic device of claim 5 , wherein the processor is further configured to:

input the differential pronunciation pair, the acoustic feature, and the language feature to a decoder to generate the text information of the speech to be recognized.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to cause a computer to execute a method for recognizing a speech, wherein the computer includes a search engine, an acoustic mode, a language model and a decoder, the method comprising:

obtaining, through the search engine, a speech to be recognized;

obtaining an acoustic feature of the speech to be recognized by inputting the speech to be recognized into the acoustic model and obtaining a language feature of the speech to be recognized by inputting the speech to be recognized into the language model, wherein the acoustic model is a model trained by Gaussian Mixed Model (GMM)-Hidden Markov Model (HMI) or Deep Neural Network (DNN)-HMM, and the language model is a model trained by an N-Gram (which is a statistic-based language model) or a Neural Network Language Model (NNLM);

inputting the speech to be recognized to a pronunciation difference statistics to generate a differential pronunciation pair corresponding to the speech to be recognized; and

generating, through the decoder, text information of the speech to be recognized based on the differential pronunciation pair, the acoustic feature, and the language feature;

wherein the pronunciation difference statistics is trained by:

obtaining a target sample text under a target scene;

generating a sample recognition result by recognizing the target sample text;

obtaining a first audio corresponding to the target sample text and obtaining a second audio corresponding to the sample recognition result;

obtaining a differential pronunciation pair sample corresponding to a pronunciation difference between the first audio and the second audio being greater than a preset threshold; and

training the pronunciation difference statistics based on the differential pronunciation pair sample.

10. The non-transitory computer-readable storage medium of claim 9 , wherein obtaining the target sample text under the target scene comprises:

obtaining a sample text;

determining whether the sample text belongs to a target scene by inputting the sample text to a target scene text classifier;

determining the sample text as the target sample text based on the sample text belonging to the target scene; and

discarding the sample text based on the sample text not belonging to the target scene.

11. The non-transitory computer-readable storage medium of claim 10 , wherein the target scene text classifier is trained by:

obtaining a target scene sample and a non-target scene sample;

obtaining a first word vector representation of the target scene sample and a second word vector representation of the non-target scene sample;

inputting the first word vector representation as a positive sample and the second word vector representation as a negative sample to an initial target scene text classifier to train the initial target scene text classifier.

12. The non-transitory computer-readable storage medium of claim 9 , wherein generating the text information of the speech to be recognized based on the differential pronunciation pair sample, the acoustic feature, and the language feature comprises:

inputting the differential pronunciation pair, the acoustic feature, and the language feature to a decoder to generate the text information of the speech to be recognized.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2021
From: ZHAO, YINLOU; ZHANG, LIAO; JIANG, ZHENGXIANG
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 057806/0839 →
Priority Claims (1)
CN 202011219185.7 · Nov 4, 2020 · national
Continuity (1)
Related Publication 20220028370A1 · Jan 27, 2022
Cited By (1)
US 12,694,867