IP Library Granted Patent US 12,217,739
Granted Patent B2
US 12,217,739 · App. 17/760,609 · Granted Feb 4, 2025

Speech recognition method and apparatus, and computer-readable storage medium

Inventor: Li Fu (Beijing, CN)
Assignee: JINGDONG TECHNOLOGY HOLDING CO., LTD.
G10L15/063G10L19/022G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,739
App. No.
17/760,609
Granted
Feb 4, 2025
Kind
B2
Abstract

A speech recognition method, including acquiring first linear frequency spectrums corresponding to audios to be trained with different sampling rates; determining the maximum sampling rate and other sampling rates; determining the maximum frequency domain sequence number of the first linear frequency spectrums as a first frequency domain sequence number and a second frequency domain sequence number; in the first linear frequency spectrums corresponding to the other sampling rate, configuring amplitude values corresponding to each frequency domain sequence number that is greater than the first frequency domain sequence number and less than or equal to the second frequency domain sequence number to be zero to obtain second linear frequency spectrums; determining first speech features and second voice features; and using the first speech features and the second speech features to train a machine learning model.

Claims (92)

1. A speech recognition method, comprising:

acquiring first linear spectrums corresponding to to-be-trained audios with different sampling rates, wherein an abscissa of the first linear spectrums is a spectrum-sequence serial number, an ordinate of the first linear spectrums is a frequency-domain serial number, and a value of a coordinate point determined by the abscissa and the ordinate is an original amplitude value corresponding to the to-be-trained audios;

determining a maximum sampling rate and other sampling rate than the maximum sampling rate in the different sampling rates;

determining a maximum frequency-domain serial number of the first linear spectrums corresponding to the other sampling rate as a first frequency-domain serial number;

determining a maximum frequency-domain serial number of the first linear spectrums corresponding to the maximum sampling rate as a second frequency-domain serial number;

setting, to zero, amplitude values corresponding to each frequency-domain serial number that is greater than the first frequency-domain serial number and less than or equal to the second frequency-domain serial number, in the first linear spectrums corresponding to the other sampling rate, to obtain second linear spectrums corresponding to the other sampling rate;

determining first speech features of the to-be-trained audios with the maximum sampling rate according to first Mel-spectrum features of the first linear spectrums corresponding to the maximum sampling rate;

determining second speech features of the to-be-trained audios with the other sampling rate according to second Mel-spectrum features of the second linear spectrums corresponding to the other sampling rate; and

training a machine learning model by using the first speech features and the second speech features,

wherein the determining second speech features of the to-be-trained audios with the other sampling rate comprises performing local normalization processing on the second Mel-spectrum features to obtain the second speech features and the local normalization processing comprises:

according to a maximum linear-spectrum frequency corresponding to the to-be-trained audios with the other sampling rate, acquiring a Mel-spectrum frequency corresponding to the maximum linear-spectrum frequency;

calculating a maximum Mel-filter serial number corresponding to the Mel-spectrum frequency;

acquiring first amplitude values corresponding to each other Mel-filter serial number in the second Mel-spectrum features, the other Mel-filter serial number being a Mel-filter serial number less than or equal to the maximum Mel-filter serial number;

respectively calculating a mean and a standard deviation of all first amplitude values as a local mean and a local standard deviation;

calculating a first difference between each of the first amplitude values and the local mean thereof;

calculating a ratio of each first difference to the local standard deviation as a normalized first amplitude value corresponding to each first amplitude value; and

replacing each first amplitude value in the second Mel-spectrum features with the normalized first amplitude value corresponding to each first amplitude value.

2. A speech recognition method, comprising:

acquiring first linear spectrums corresponding to to-be-trained audios with different sampling rates, wherein an abscissa of the first linear spectrums is a spectrum-sequence serial number, an ordinate of the first linear spectrums is a frequency-domain serial number, and a value of a coordinate point determined by the abscissa and the ordinate is an original amplitude value corresponding to the to-be-trained audios;

determining a maximum sampling rate and other sampling rate than the maximum sampling rate in the different sampling rates;

determining a maximum frequency-domain serial number of the first linear spectrums corresponding to the other sampling rate as a first frequency-domain serial number;

determining a maximum frequency-domain serial number of the first linear spectrums corresponding to the maximum sampling rate as a second frequency-domain serial number;

setting, to zero, amplitude values corresponding to each frequency-domain serial number that is greater than the first frequency-domain serial number and less than or equal to the second frequency-domain serial number, in the first linear spectrums corresponding to the other sampling rate, to obtain second linear spectrums corresponding to the other sampling rate;

determining first speech features of the to-be-trained audios with the maximum sampling rate according to first Mel-spectrum features of the first linear spectrums corresponding to the maximum sampling rate;

determining second speech features of the to-be-trained audios with the other sampling rate according to second Mel-spectrum features of the second linear spectrums corresponding to the other sampling rate; and

training a machine learning model by using the first speech features and the second speech features,

wherein the determining first speech features of the to-be-trained audios with the maximum sampling rate comprises performing global normalization processing on the first Mel-spectrum features to obtain the first speech features and the global normalization processing comprises:

acquiring second amplitude values corresponding to each Mel-filter serial number in the first Mel-spectrum features;

calculating a mean and a standard deviation of all second amplitude values as a global mean and a global standard deviation;

calculating a second difference between each of the second amplitude values and the global mean thereof;

calculating a ratio of each second difference to the global standard deviation as a normalized second amplitude value corresponding to each second amplitude value; and

replacing each second amplitude value in the first Mel-spectrum features with the normalized second amplitude value corresponding to each second amplitude value.

3. The speech recognition method according to claim 1 , wherein the acquiring first linear spectrums corresponding to to-be-trained audios with different sampling rates comprises:

respectively acquiring the first linear spectrums corresponding to the to-be-trained audios with the different sampling rates by using short-time Fourier transform.

4. The speech recognition method according to claim 1 , wherein the acquiring first linear spectrums corresponding to to-be-trained audios with different sampling rates comprises:

acquiring speech signal oscillograms of the to-be-trained audios with the different sampling rates;

respectively performing pre-emphasis processing on the speech signal oscillograms of the to-be-trained audios with the different sampling rates; and

acquiring the first linear spectrums corresponding to the to-be-trained audios with the different sampling rates according to the speech signal oscillograms after the pre-emphasis processing.

5. The speech recognition method according to claim 1 , further comprising:

respectively performing Mel-filtering transform on the first linear spectrums corresponding to the maximum sampling rate and the second linear spectrums corresponding to the other sampling rate by using a plurality of unit triangle filters, to obtain the first Mel-spectrum features and the second Mel-spectrum features.

6. The speech recognition method according to claim 1 , wherein the machine learning model comprises a deep neural network (DNN) model.

7. The speech recognition method according to claim 1 , wherein the different sampling rates comprise 16 kHZ and 8 kHZ.

8. The speech recognition method according to claim 1 , further comprising:

acquiring a to-be-recognized audio;

determining a speech feature of the to-be-recognized audio; and

inputting the speech feature of the to-be-recognized audio into the machine learning model, to obtain a speech recognition result.

9. The speech recognition method according to claim 8 , wherein the determining a speech feature of the to-be-recognized audio comprises:

determining a maximum frequency-domain serial number of a first linear spectrum of the to-be-recognized audio as a third frequency-domain serial number;

setting, to zero, amplitude values corresponding to each frequency-domain serial number that is greater than the third frequency-domain serial number and less than or equal to the second frequency-domain serial number in the first linear spectrum of the to-be-recognized audio, to obtain a second linear spectrum of the to-be-recognized audio; and

determining the speech feature of the to-be-recognized audio according to a Mel-spectrum feature of the second linear spectrum of the to-be-recognized audio.

10. A speech recognition apparatus, comprising:

a memory; and

a processor coupled to the memory, which is configured to executed the speech recognition method according to claim 1 .

11. A speech recognition apparatus, comprising:

a memory; and

a processor coupled to the memory, which is configured to executed the speech recognition method according to claim 2 .

12. A non-transitory computer-storable medium having stored thereon computer program instructions which, when executed by a processor, implement a speech recognition method, which comprises:

acquiring first linear spectrums corresponding to to-be-trained audios with different sampling rates, wherein an abscissa of the first linear spectrums is a spectrum-sequence serial number, an ordinate of the first linear spectrums is a frequency-domain serial number, and a value of a coordinate point determined by the abscissa and the ordinate is an original amplitude value corresponding to the to-be-trained audios;

determining a maximum sampling rate and other sampling rate than the maximum sampling rate in the different sampling rates;

determining a maximum frequency-domain serial number of the first linear spectrums corresponding to the other sampling rate as a first frequency-domain serial number;

determining a maximum frequency-domain serial number of the first linear spectrums corresponding to the maximum sampling rate as a second frequency-domain serial number;

setting, to zero, amplitude values corresponding to each frequency-domain serial number that is greater than the first frequency-domain serial number and less than or equal to the second frequency-domain serial number, in the first linear spectrums corresponding to the other sampling rate, to obtain second linear spectrums corresponding to the other sampling rate;

determining first speech features of the to-be-trained audios with the maximum sampling rate according to first Mel-spectrum features of the first linear spectrums corresponding to the maximum sampling rate;

determining second speech features of the to-be-trained audios with the other sampling rate according to second Mel-spectrum features of the second linear spectrums corresponding to the other sampling rate; and

training a machine learning model by using the first speech features and the second speech features,

wherein the determining second speech features of the to-be-trained audios with the other sampling rate comprises performing local normalization processing on the second Mel-spectrum features to obtain the second speech features and the local normalization processing comprises:

according to a maximum linear-spectrum frequency corresponding to the to-be-trained audios with the other sampling rate, acquiring a Mel-spectrum frequency corresponding to the maximum linear-spectrum frequency;

calculating a maximum Mel-filter serial number corresponding to the Mel-spectrum frequency;

acquiring first amplitude values corresponding to each other Mel-filter serial number in the second Mel-spectrum features, the other Mel-filter serial number being a Mel-filter serial number less than or equal to the maximum Mel-filter serial number;

respectively calculating a mean and a standard deviation of all first amplitude values as a local mean and a local standard deviation;

calculating a first difference between each of the first amplitude values and the local mean thereof;

calculating a ratio of each first difference to the local standard deviation as a normalized first amplitude value corresponding to each first amplitude value; and

replacing each first amplitude value in the second Mel-spectrum features with the normalized first amplitude value corresponding to each first amplitude value.

13. A non-transitory computer-storable medium having stored thereon computer program instructions which, when executed by a processor, implement the speech recognition method according to claim 2 .

14. The speech recognition method according to claim 2 , wherein the acquiring first linear spectrums corresponding to to-be-trained audios with different sampling rates comprises:

respectively acquiring the first linear spectrums corresponding to the to-be-trained audios with the different sampling rates by using short-time Fourier transform.

15. The speech recognition method according to claim 2 , wherein the acquiring first linear spectrums corresponding to to-be-trained audios with different sampling rates comprises:

acquiring speech signal oscillograms of the to-be-trained audios with the different sampling rates;

respectively performing pre-emphasis processing on the speech signal oscillograms of the to-be-trained audios with the different sampling rates; and

acquiring the first linear spectrums corresponding to the to-be-trained audios with the different sampling rates according to the speech signal oscillograms after the pre-emphasis processing.

16. The speech recognition method according to claim 2 , further comprising:

respectively performing Mel-filtering transform on the first linear spectrums corresponding to the maximum sampling rate and the second linear spectrums corresponding to the other sampling rate by using a plurality of unit triangle filters, to obtain the first Mel-spectrum features and the second Mel-spectrum features.

17. The speech recognition method according to claim 2 , wherein the machine learning model comprises a deep neural network (DNN) model.

18. The speech recognition method according to claim 2 , wherein the different sampling rates comprise 16 kHZ and 8 kHZ.

19. The speech recognition method according to claim 2 , further comprising:

acquiring a to-be-recognized audio;

determining a speech feature of the to-be-recognized audio; and

inputting the speech feature of the to-be-recognized audio into the machine learning model, to obtain a speech recognition result.

20. The speech recognition method according to claim 19 , wherein the determining a speech feature of the to-be-recognized audio comprises:

determining a maximum frequency-domain serial number of a first linear spectrum of the to-be-recognized audio as a third frequency-domain serial number;

setting, to zero, amplitude values corresponding to each frequency-domain serial number that is greater than the third frequency-domain serial number and less than or equal to the second frequency-domain serial number in the first linear spectrum of the to-be-recognized audio, to obtain a second linear spectrum of the to-be-recognized audio; and

determining the speech feature of the to-be-recognized audio according to a Mel-spectrum feature of the second linear spectrum of the to-be-recognized audio.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2022
From: FU, LI
To: JINGDONG TECHNOLOGY HOLDING CO., LTD.
Reel/Frame 059270/0044 →
Priority Claims (1)
CN 201910904271.2 · Sep 24, 2019 · national
Continuity (1)
Related Publication 20220343898A1 · Oct 27, 2022
References Cited (35)
US 5475792A · Stanford et al. · 1995 [cited by applicant]
US 8438026B2 · Fischer et al. · 2013 [cited by applicant]
US 10360899B2 · Zou et al. · 2019 [cited by applicant]
US 20080201139A1 · Yu et al. · 2008 [cited by applicant]
US 20080215322A1 · Fischer · 2008 [cited by examiner]
US 20140257804A1 · Li et al. · 2014 [cited by applicant]
US 20200380954A1 · Fan · 2020 [cited by applicant]
CN 101014997A · 2007 [cited by applicant]
CN 105513590 · 2016 [cited by examiner]
CN 105513590A · 2016 [cited by examiner]
CN 106997767A · 2017 [cited by applicant]
CN 107710323 · 2018 [cited by examiner]
CN 107710323A · 2018 [cited by examiner]
CN 108510979A · 2018 [cited by applicant]
CN 109061591A · 2018 [cited by examiner]
CN 109346061 · 2019 [cited by examiner]
CN 109346061A · 2019 [cited by examiner]
CN 110459205A · 2019 [cited by applicant]
EP 1229519A1 · 2002 [cited by applicant]
“First Office Action and English language translation”, CN Application No. 201910904271.2, Jul. 1, 2021, 17 pp. [cited by applicant]
“International Search Report and Written Opinion of the International Searching Authority with English language translation”, International Application No. PCT/CN2020/088229, Aug. 3, 2020, 17 pp. [cited by applicant]
Amodei, Dario , et al., “Deep Speech 2: End-to-End Speech Recognition in English and Mandarin”, https://arxiv.org/abs/1512.02595v1, Dec. 8, 2015, 28 pp. [cited by applicant]
Chan, William , et al., Presentation—“Listen, Attend and Spell—A Neural Network for Large Vocabulary Conversational Speech Recognition”, Carnegie Mellon University / Google, Sep. 13, 2016, 38 pp. [cited by applicant]
Gao, Jianqing , et al., “Mixed-Bandwidth Cross-Channel Speech Recognition via Joint Optimization of DNN-Based Bandwidth Expansion and Acoustic Modeling”, IEEE/ACM Transactions on Audio, Speech, and Language Processing, … [cited by applicant]
Graves, Alex , et al., “Speech Recognition With Deep Recurrent Neural Networks”, https://arxiv.org/abs/1303.5778v1, Mar. 22, 2013, 5 pp. [cited by applicant]
Jiangmosson , et al., “Micro Doppler Effects and Applications Thereof”, Beijing Aerospace Aviation University Press, Aug. 2013, 13 pp., including concise explanation of the relevance. [cited by applicant]
Seltzer, Michael L., et al., “Training Wideband Acoustic Models Using Mixed-Bandwidth Training Data for Speech Recognition”, IEEE Transactions on Audio, Speech, and Language Processing, Jan. 2007, pp. 235-245. [cited by applicant]
Shenchen , et al., “Developments and Applications of Cloud Computing-Based Big Data Processing Technology”, Electron and Technology University Press, Jan. 2019, 22 pp., including concise explanation of the relevance. [cited by applicant]
“Communication with Supplementary European Search Report”, EP Application No. 20868687.3, Jul. 12, 2023, 7 pp. [cited by applicant]
Fu, Li , et al., “Research on Modeling Units of Transformer Transducer for Mandarin Speech Recognition”, https://arxiv.org/abs/2004.13522, arXiv.org, Cornell University Library, Ithaca, NY, Apr. 26, 2020, 5 pp. [cited by applicant]
Li, Jinyu , et al., “Improving Wideband Speech Recognition Using Mixed-Bandwidth Training Data in CD-DNN-HMM”, 2012 IEEE Spoken Language Technology Workshop (SLT), Miami, Florida, USA, Dec. 2, 2012, pp. 131-136. [cited by applicant]
Mantena, Gautam , et al., “Bandwidth Embeddings for Mixed-Bandwidth Speech Recognition”, https://arxiv.org/abs/1909.02667, arXiv.org, Cornell University Library, Ithaca, NC, Sep. 5, 2019, 6 pp. [cited by applicant]
Yu, Dong , et al., “Feature Learning in Deep Neural Networks—Studies on Speech Recognition Tasks”, https://arxiv.org/abs/1301.3605v1, arXiv.org, Cornell University Library, Ithaca, NY, Jan. 16, 2013, 9 pp. [cited by applicant]
Yu, Dong , et al., “Feature Learning in Deep Neural Networks—Studies on Speech Recognition Tasks”, https://arxiv.org/abs/1301.3605v3, arXiv.org, Cornell University Library, Ithaca, NY, Mar. 8, 2013, 9 pp. [cited by applicant]
“Notice of Reasons for Refusal” and English language translation, JP Application No. 2022-516702, Jun. 17, 2024, 10 pp. [cited by applicant]