IP Library › Granted Patent US 11,688,391
Granted Patent B2
US 11,688,391 · App. 16/843,174 · Granted Jun 27, 2023

Mandarin and dialect mixed modeling and speech recognition

Inventor: Shenglong Yuan (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO.
G10L15/144G10L15/16G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,688,391
App. No.
16/843,174
Granted
Jun 27, 2023
Kind
B2
Abstract

The present disclosure provides a modeling method for speech recognition and a device. The method includes: determining N types of tags; training a neural network according to speech data of Mandarin to generate a recognition model whose outputs are the N types of tags; inputting speech data of each dialect into the recognition model to obtain an output tag of each frame of the speech data of each dialect; determining, according to the output tags and tagged true tags, error rates of the N types of tags for the each dialect, generating M types of target tags according to tags with error rates greater than a preset threshold; and training an acoustic model according to third speech data of Mandarin and third speech data of the P dialects, outputs of the acoustic model being the N types of tags and the M types of target tags corresponding to each dialect.

Claims (58)

1. A modeling method for speech recognition, comprising:

processing first speech data of Mandarin and first speech data of P dialects respectively based on a pre-trained alignment model to obtain a tag corresponding to each frame of the first speech data, and counting the obtained tags and performing deduplication on tags of each type to determine N types of tags, N being a positive integer and P being a positive integer;

training a neural network according to second speech data of Mandarin, and generating a recognition model when the neural network converges, wherein outputs of the recognition model are the N types of tags;

inputting second speech data of the P dialects into the recognition model for processing respectively to obtain an output tag of each frame of the second speech data of each dialect;

for each dialect i of the P dialects, determining, according to the output tags and tagged true tags of the second speech data of each dialect, an error rate of each type of the N types of tags, wherein the error rate of a specific type of tag is a ratio of a total number of misclassifications for the specific type of tag to a total number of the specific type of tag in the tagged true tags;

for each dialect i of the P dialects, generating, according to M i types of tags whose error rates are greater than a preset threshold, M i types of target tags specially used for the dialect i, where M i being an integer greater than or equal to zero, and 0=<i<=P−1; and

training an acoustic model according to third speech data of Mandarin and third speech data of the P dialects, wherein outputs of the acoustic model are the N types of tags and the M i types of target tags corresponding to each dialect i of the P dialects.

2. The method of claim 1 , wherein inputting the second speech data of the P dialects into the recognition model for processing respectively, to obtain the output tag of each frame of the second speech data of each dialect comprises:

extracting a filter bank coefficient characteristic of the second speech data of the P dialects, and determining N posterior probabilities of each frame of the second speech data of each dialect according to the filter bank coefficient characteristic; and

determining a tag corresponding to a maximum posterior probability in the N posterior probabilities as an output tag of a frame of the second speech data corresponding to the N posterior probabilities.

3. The method of claim 1 , wherein training the acoustic model according to the third speech data of Mandarin and the third speech data of the P dialects comprises:

generating training samples according to the third speech data of Mandarin, first tagged tags corresponding to the third speech data of Mandarin, the third speech data of the P dialects and second tagged tags corresponding to the third speech data of the P dialects;

for the third speech data of each dialect i of the P dialects, replacing the M i types of tags originally tagged whose error rates are greater than the preset threshold with corresponding M i types of target tags to obtain updated training samples; and

training a processing parameter of a preset model according to a preset objective function and the updated training samples to obtain the acoustic model.

4. The method of claim 1 , before processing the first speech data of Mandarin and the first speech data of the P dialects respectively based on the pre-trained alignment model, further comprising:

obtaining fourth speech data of Mandarin and corresponding text information; and

extracting a MFCC characteristic of each frame of the fourth speech data, and generating, according to the MFCC characteristic and the text information, the alignment model by training a parameter of a Gaussian mixture model based on maximum likelihood estimation.

5. The method of claim 1 , after generating the M i types of target tags specifically used for the dialect i according to the M i types of tags whose error rates are greater than the preset threshold, further comprising:

updating a decoding dictionary according to the M i types of target tags.

6. A computer device, comprising a processor and a memory,

wherein, the processor is configured to run a program corresponding to executable program codes by reading the executable program codes stored in the memory, to implement a modeling method for speech recognition, comprising:

processing first speech data of Mandarin and first speech data of P dialects respectively based on a pre-trained alignment model to obtain a tag corresponding to each frame of the first speech data, and counting the obtained tags and performing deduplication on tags of each type to determine N types of tags, N being a positive integer and P being a positive integer;

training a neural network according to second speech data of Mandarin, and generating a recognition model when the neural network converges, wherein outputs of the recognition model are the N types of tags;

inputting second speech data of the P dialects into the recognition model for processing respectively to obtain an output tag of each frame of the second speech data of each dialect;

for each dialect i of the P dialects, determining, according to the output tags and tagged true tags of the second speech data of each dialect, an error rate of each type of the N types of tags, wherein the error rate of a specific type of tag is a ratio of a total number of misclassifications for the specific type of tag to a total number of the specific type of tag in the tagged true tags;

for each dialect i of the P dialects, generating, according to M i types of tags whose error rates are greater than a preset threshold, M i types of target tags specially used for the dialect i, where M i being an integer greater than or equal to zero, and 0=<i<=P− 1 ; and

training an acoustic model according to third speech data of Mandarin and third speech data of the P dialects, wherein outputs of the acoustic model are the N types of tags and the M i types of target tags corresponding to each dialect I of the P dialects.

7. The computer device of claim 6 , wherein inputting the second speech data of the P dialects into the recognition model for processing respectively, to obtain the output tag of each frame of the second speech data of each dialect comprises:

extracting a filter bank coefficient characteristic of the second speech data of the P dialects, and determining N posterior probabilities of each frame of the second speech data of each dialect according to the filter bank coefficient characteristic; and

determining a tag corresponding to a maximum posterior probability in the N posterior probabilities as an output tag of a frame of the second speech data corresponding to the N posterior probabilities.

8. The computer device of claim 6 , wherein training the acoustic model according to the third speech data of Mandarin and the third speech data of the P dialects comprises:

generating training samples according to the third speech data of Mandarin, first tagged tags corresponding to the third speech data of Mandarin, the third speech data of the P dialects and second tagged tags corresponding to the third speech data of the P dialects;

for the third speech data of each dialect I of the P dialects, replacing the M i types of tags originally tagged whose error rates are greater than the preset threshold with corresponding M i types of target tags to obtain updated training samples; and

training a processing parameter of a preset model according to a preset objective function and the updated training samples to obtain the acoustic model.

9. The computer device of claim 6 , wherein, before processing the first speech data of Mandarin and the first speech data of the P dialects respectively based on the pre-trained alignment model, the method further comprises:

obtaining fourth speech data of Mandarin and corresponding text information; and

extracting a MFCC characteristic of each frame of the fourth speech data, and generating, according to the MFCC characteristic and the text information, the alignment model by training a parameter of a Gaussian mixture model based on maximum likelihood estimation.

10. The computer device of claim 6 , wherein, after generating the M i types of target tags specifically used for the dialect i according to the M i types of tags whose error rates are greater than the preset threshold, the method further comprises:

updating a decoding dictionary according to the M i types of target tags.

11. A non-transitory computer readable storage medium having a computer program stored thereon that, when the program is executed by a processor, a modeling method for speech recognition, comprising:

processing first speech data of Mandarin and first speech data of P dialects respectively based on a pre-trained alignment model to obtain a tag corresponding to each frame of the first speech data, and counting the obtained tags and performing deduplication on tags of each type to determine N types of tags, N being a positive integer and P being a positive integer;

training a neural network according to second speech data of Mandarin, and generating a recognition model when the neural network converges, wherein outputs of the recognition model are the N types of tags;

inputting second speech data of the P dialects into the recognition model for processing respectively to obtain an output tag of each frame of the second speech data of each dialect;

for each dialect I of the P dialects, determining, according to the output tags and tagged true tags of the second speech data of each dialect, an error rate of each type of the N types of tags, wherein the error rate of a specific type of tag is a ratio of a total number of misclassifications for the specific type of tag to a total number of the specific type of tag in the tagged true tags;

for each dialect i of the P dialects, generating, according to M i types of tags whose error rates are greater than a preset threshold, M i types of target tags specially used for the dialect i, where M i being an integer greater than or equal to zero, and 0=<i<=P−1; and

training an acoustic model according to third speech data of Mandarin and third speech data of the P dialects, wherein outputs of the acoustic model are the N types of tags and the M i types of target tags corresponding to each dialect i of the P dialects.

12. The storage medium of claim 11 , wherein inputting the second speech data of the P dialects into the recognition model for processing respectively, to obtain the output tag of each frame of the second speech data of each dialect comprises:

extracting a filter bank coefficient characteristic of the second speech data of the P dialects, and determining N posterior probabilities of each frame of the second speech data of each dialect according to the filter bank coefficient characteristic; and

determining a tag corresponding to a maximum posterior probability in the N posterior probabilities as an output tag of a frame of the second speech data corresponding to the N posterior probabilities.

13. The storage medium of claim 11 , wherein training the acoustic model according to the third speech data of Mandarin and the third speech data of the P dialects comprises:

generating training samples according to the third speech data of Mandarin, first tagged tags corresponding to the third speech data of Mandarin, the third speech data of the P dialects and second tagged tags corresponding to the third speech data of the P dialects;

for the third speech data of each dialect i of the P dialects, replacing the M i types of tags originally tagged whose error rates are greater than the preset threshold with corresponding M i types of target tags to obtain updated training samples; and

training a processing parameter of a preset model according to a preset objective function and the updated training samples to obtain the acoustic model.

14. The storage medium of claim 11 , wherein, before processing the first speech data of Mandarin and the first speech data of the P dialects respectively based on the pre-trained alignment model, the method further comprises:

obtaining fourth speech data of Mandarin and corresponding text information; and

extracting a MFCC characteristic of each frame of the fourth speech data, and generating, according to the MFCC characteristic and the text information, the alignment model by training a parameter of a Gaussian mixture model based on maximum likelihood estimation.

15. The storage medium of claim 11 , wherein, after generating the M i types of target tags specifically used for the dialect i according to the M i types of tags whose error rates are greater than the preset threshold, the method further comprises:

updating a decoding dictionary according to the M i types of target tags.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2020
From: YUAN, SHENGLONG
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO.
Reel/Frame 052344/0932 →
Priority Claims (1)
CN 201910297805.X · Apr 15, 2019 · national
Continuity (1)
Related Publication 20200327883A1 · Oct 15, 2020