IP Library › Granted Patent US 10,614,802
Granted Patent B2
US 10,614,802 · App. 15/860,706 · Granted Apr 7, 2020

Method and device for recognizing speech based on Chinese-English mixed dictionary

Inventors: Xiangang Li (Beijing, CN); Xuewei Zhang (Beijing, CN)
Assignee: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
G10L15/187G06F40/242G10L15/02G10L15/063G10L15/16G10L15/22G10L2015/025G10L2015/0631G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,614,802
App. No.
15/860,706
Granted
Apr 7, 2020
Kind
B2
Abstract

Embodiments of the present disclosure provide a method and a device for recognizing a speech based on a Chinese-English mixed dictionary. The method includes acquiring a Chinese-English mixed dictionary marked by an international phonetic alphabet, in which, the Chinese-English mixed dictionary includes a Chinese dictionary and an English dictionary revised by Chinglish; by taking the Chinese-English mixed dictionary as a training dictionary, taking a one-layer Convolutional Neural Network and a five-layer Long Short-Term Memory as a model, taking syllables or words as a target and taking a connectionist temporal classifier as a training criterion, training the model to obtain a trained CTC acoustic model; and performing a speech recognition on a Chinese-English mixed language based on the trained CTC acoustic model.

Claims (88)

1. A method for recognizing a speech based on a Chinese-English mixed dictionary, comprising:

acquiring a Chinese-English mixed dictionary marked by an International Phonetic Alphabet IPA, wherein, the Chinese-English mixed dictionary comprises a Chinese dictionary and an English dictionary revised by Chinglish;

by taking the Chinese-English mixed dictionary as a training dictionary, taking a one-layer Convolutional Neural Network CNN and a five-layer Long Short-Term Memory LSTM as a model, taking syllables or words as a target and taking a Connectionist Temporal Classifier CTC as a training criterion, training the model to obtain a trained CTC acoustic model; and

performing a speech recognition on a Chinese-English mixed language based on the trained CTC acoustic model.

2. The method according to claim 1 , wherein acquiring the Chinese-English mixed dictionary marked by the IPA comprises:

acquiring a Chinese dictionary marked by the IPA and an English dictionary marked by the IPA;

acquiring audio training data, wherein, the audio training data comprises a plurality of Chinglish sentences;

acquiring English words in the plurality of Chinglish sentences and Chinese pronunciations corresponding to the English words; and

adding the English words and the Chinese pronunciations corresponding to the English words to the English dictionary, to obtain the English dictionary revised by Chinglish.

3. The method according to claim 1 , wherein acquiring the Chinese-English mixed dictionary marked by the IPA comprises:

acquiring a Chinese dictionary marked by the IPA and an English dictionary marked by the IPA;

acquiring audio training data, wherein, the audio training data comprises a plurality of Chinglish sentences;

performing a phoneme decoding and a matching file division on the plurality of Chinglish sentences based on the English dictionary marked by the IPA, to obtain English words in the plurality of Chinglish sentences and Chinese pronunciations corresponding to the English words; and

generating the English dictionary revised by Chinglish based on the English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the English words and the English dictionary marked by the IPA.

4. The method according to claim 3 , wherein performing the phoneme decoding and the matching file division on the plurality of Chinglish sentences based on the English dictionary marked by the IPA, to obtain the English words in the plurality of Chinglish sentences and the Chinese pronunciations corresponding to the English words comprises:

performing the phoneme decoding on the plurality of Chinglish sentences based on the English dictionary marked by the IPA, to find an optimal path, to obtain frame positions corresponding to phonemes in the plurality of Chinglish sentences;

acquiring a matching file corresponding to the plurality of Chinglish sentences, wherein, the matching file comprises positions of respective phonemes in the plurality of Chinglish sentences and phonemes corresponding to the English words; and

determining positions of respective the English words in the plurality of Chinglish sentences based on the matching file and the frame positions corresponding to the phonemes in the plurality of Chinglish sentences, to divide to obtain the English words in the plurality of Chinglish sentences and the Chinese pronunciations corresponding to the English words.

5. The method according to claim 3 , before generating the English dictionary revised by Chinglish based on the English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the English words and the English dictionary marked by the IPA, further comprising:

acquiring word frequencies of respective phonemes in the English word for each English word in the plurality of Chinglish sentences;

acquiring high frequency phonemes whose corresponding word frequencies are greater than a preset word frequency and high frequency English words comprising the high frequency phonemes;

generating the English dictionary revised by Chinglish based on the English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the English words and the English dictionary marked by the IPA comprises:

generating the English dictionary revised by Chinglish based on the high frequency English words in the plurality of Chinglish sentences, Chinese pronunciations corresponding to the high frequency English words and the English dictionary marked by the IPA.

6. The method according to claim 3 , after generating the English dictionary revised by Chinglish based on the English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the English words and the English dictionary marked by the IPA, further comprising:

performing the phoneme decoding and the matching file division on the plurality of Chinglish sentences based on the English dictionary revised by Chinglish, to obtain new English words in the plurality of Chinglish sentences and Chinese pronunciations corresponding to the new English words; and

updating the English dictionary revised by Chinglish based on the new English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the new English words and the English dictionary revised by Chinglish.

7. The method according to claim 1 , wherein, by taking the Chinese-English mixed dictionary as the training dictionary, taking the one-layer CNN and the five-layer LSTM as the model, taking syllables or words as the target and taking the CTC as the training criterion, training the model to obtain the trained CTC acoustic model comprises:

extracting feature points in a Chinglish sentence using a filter bank FBANK as inputs of the model, and training the model by taking the one-layer CNN and the five-layer LSTM as the model, taking a matching file corresponding to the Chinglish sentence as the target and taking a Cross Entropy CE as the training criterion, to obtain an initial model; and

training the initial model by taking the Chinese-English mixed dictionary as the training dictionary, taking the initial model as the model, taking the syllables or words as the target and taking the CTC as the training criterion, to obtain the trained CTC acoustic model.

8. A device for recognizing a speech based on a Chinese-English mixed dictionary, comprising:

a memory;

a processor; and

computer programs stored in the memory and executable by the processor,

wherein when the processor executes the computer programs, a method for recognizing a speech based on a Chinese-English mixed dictionary is performed, the method comprising:

acquiring a Chinese-English mixed dictionary marked by an International Phonetic Alphabet IPA, wherein, the Chinese-English mixed dictionary comprises a Chinese dictionary and an English dictionary revised by Chinglish;

by taking the Chinese-English mixed dictionary as a training dictionary, taking a one-layer Convolutional Neural Network CNN and a five-layer Long Short-Term Memory LSTM as a model, taking syllables or words as a target and taking a Connectionist Temporal Classifier CTC as a training criterion, training the model to obtain a trained CTC acoustic model; and

performing a speech recognition on a Chinese-English mixed language based on the trained CTC acoustic model.

9. The device according to claim 8 , wherein acquiring the Chinese-English mixed dictionary marked by the IPA comprises:

acquiring a Chinese dictionary marked by the IPA and an English dictionary marked by the IPA;

acquiring audio training data, wherein, the audio training data comprises a plurality of Chinglish sentences;

acquiring English words in the plurality of Chinglish sentences and Chinese pronunciations corresponding to the English words; and

adding the English words and the Chinese pronunciations corresponding to the English words to the English dictionary, to obtain the English dictionary revised by Chinglish.

10. The device according to claim 8 , wherein acquiring the Chinese-English mixed dictionary marked by the IPA comprises:

acquiring a Chinese dictionary marked by the IPA and an English dictionary marked by the IPA;

acquiring audio training data, wherein, the audio training data comprises a plurality of Chinglish sentences;

performing a phoneme decoding and a matching file division on the plurality of Chinglish sentences based on the English dictionary marked by the IPA, to obtain English words in the plurality of Chinglish sentences and Chinese pronunciations corresponding to the English words; and

generating the English dictionary revised by Chinglish based on the English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the English words and the English dictionary marked by the IPA.

11. The device according to claim 10 , wherein the method further comprises:

performing the phoneme decoding on the plurality of Chinglish sentences based on the English dictionary marked by the IPA, to find an optimal path, to obtain frame positions corresponding to phonemes in the plurality of Chinglish sentences;

acquiring a matching file corresponding to the plurality of Chinglish sentences, wherein, the matching file comprises positions of respective phonemes in the plurality of Chinglish sentences and phonemes corresponding to the English words; and

determining positions of respective the English words in the plurality of Chinglish sentences based on the matching file and the frame positions corresponding to the phonemes in the plurality of Chinglish sentences, to divide to obtain the English words in the plurality of Chinglish sentences and the Chinese pronunciations corresponding to the English words.

12. The device according to claim 10 , wherein the method further comprises:

acquiring word frequencies of respective phonemes in the English word for each English word in the plurality of Chinglish sentences;

acquiring high frequency phonemes whose corresponding word frequencies are greater than a preset word frequency and high frequency English words comprising the high frequency phonemes;

generating the English dictionary revised by Chinglish based on the English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the English words and the English dictionary marked by the IPA comprises:

generating the English dictionary revised by Chinglish based on the high frequency English words in the plurality of Chinglish sentences, Chinese pronunciations corresponding to the high frequency English words and the English dictionary marked by the IPA.

13. The device according to claim 10 , wherein the method further comprises:

performing the phoneme decoding and the matching file division on the plurality of Chinglish sentences based on the English dictionary revised by Chinglish, to obtain new English words in the plurality of Chinglish sentences and Chinese pronunciations corresponding to the new English words; and

updating the English dictionary revised by Chinglish based on the new English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the new English words and the English dictionary revised by Chinglish.

14. The device according to claim 8 , wherein, by taking the Chinese-English mixed dictionary as the training dictionary, taking the one-layer CNN and the five-layer LSTM as the model, taking syllables or words as the target and taking the CTC as the training criterion, training the model to obtain the trained CTC acoustic model comprises:

extracting feature points in a Chinglish sentence using a filter bank FBANK as inputs of the model, and training the model by taking the one-layer CNN and the five-layer LSTM as the model, taking a matching file corresponding to the Chinglish sentence as the target and taking a Cross Entropy CE as the training criterion, to obtain an initial model; and

training the initial model by taking the Chinese-English mixed dictionary as the training dictionary, taking the initial model as the model, taking the syllables or words as the target and taking the CTC as the training criterion, to obtain the trained CTC acoustic model.

15. A non-transitory computer readable storage medium, configured to store computer programs, wherein when the computer programs are executed by a processor, a method for recognizing a speech based on a Chinese-English mixed dictionary is performed, the method comprising:

acquiring a Chinese-English mixed dictionary marked by an International Phonetic Alphabet IPA, wherein, the Chinese-English mixed dictionary comprises a Chinese dictionary and an English dictionary revised by Chinglish;

by taking the Chinese-English mixed dictionary as a training dictionary, taking a one-layer Convolutional Neural Network CNN and a five-layer Long Short-Term Memory LSTM as a model, taking syllables or words as a target and taking a Connectionist Temporal Classifier CTC as a training criterion, training the model to obtain a trained CTC acoustic model; and

performing a speech recognition on a Chinese-English mixed language based on the trained CTC acoustic model.

16. The non-transitory computer readable storage medium according to claim 15 , wherein acquiring the Chinese-English mixed dictionary marked by the IPA comprises:

acquiring a Chinese dictionary marked by the IPA and an English dictionary marked by the IPA;

acquiring audio training data, wherein, the audio training data comprises a plurality of Chinglish sentences;

acquiring English words in the plurality of Chinglish sentences and Chinese pronunciations corresponding to the English words; and

adding the English words and the Chinese pronunciations corresponding to the English words to the English dictionary, to obtain the English dictionary revised by Chinglish.

17. The non-transitory computer readable storage medium according to claim 15 , wherein acquiring the Chinese-English mixed dictionary marked by the IPA comprises:

acquiring a Chinese dictionary marked by the IPA and an English dictionary marked by the IPA;

acquiring audio training data, wherein, the audio training data comprises a plurality of Chinglish sentences;

performing a phoneme decoding and a matching file division on the plurality of Chinglish sentences based on the English dictionary marked by the IPA, to obtain English words in the plurality of Chinglish sentences and Chinese pronunciations corresponding to the English words; and

generating the English dictionary revised by Chinglish based on the English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the English words and the English dictionary marked by the IPA.

18. The non-transitory computer readable storage medium according to claim 17 , wherein the method further comprises:

performing the phoneme decoding on the plurality of Chinglish sentences based on the English dictionary marked by the IPA, to find an optimal path, to obtain frame positions corresponding to phonemes in the plurality of Chinglish sentences;

acquiring a matching file corresponding to the plurality of Chinglish sentences, wherein, the matching file comprises positions of respective phonemes in the plurality of Chinglish sentences and phonemes corresponding to the English words; and

determining positions of respective the English words in the plurality of Chinglish sentences based on the matching file and the frame positions corresponding to the phonemes in the plurality of Chinglish sentences, to divide to obtain the English words in the plurality of Chinglish sentences and the Chinese pronunciations corresponding to the English words.

19. The non-transitory computer readable storage medium according to claim 17 , wherein the method further comprises:

acquiring word frequencies of respective phonemes in the English word for each English word in the plurality of Chinglish sentences;

acquiring high frequency phonemes whose corresponding word frequencies are greater than a preset word frequency and high frequency English words comprising the high frequency phonemes;

generating the English dictionary revised by Chinglish based on the English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the English words and the English dictionary marked by the IPA comprises:

generating the English dictionary revised by Chinglish based on the high frequency English words in the plurality of Chinglish sentences, Chinese pronunciations corresponding to the high frequency English words and the English dictionary marked by the IPA.

20. The non-transitory computer readable storage medium according to claim 17 , wherein the method further comprises:

performing the phoneme decoding and the matching file division on the plurality of Chinglish sentences based on the English dictionary revised by Chinglish, to obtain new English words in the plurality of Chinglish sentences and Chinese pronunciations corresponding to the new English words; and

updating the English dictionary revised by Chinglish based on the new English words in the plurality of Chinglish sentences, the Chinese pronunciations corresponding to the new English words and the English dictionary revised by Chinglish.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2020
From: BAIDU.COM TIMES TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 051592/0589 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2020
From: LI, XIANGANG
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 051518/0553 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2020
From: ZHANG, XUEWEI
To: BAIDU.COM TIMES TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 051519/0001 →
Priority Claims (1)
CN 2017 1 0309321 · May 4, 2017 · national
Continuity (1)
Related Publication 20180322867A1 · Nov 8, 2018