IP Library › Granted Patent US 12,609,109
Granted Patent B2
US 12,609,109 · App. 17/583,512 · Granted Apr 21, 2026

Speech recognition method and apparatus, and computer-readable storage medium

Inventors: Jun Wang (Shenzhen, CN); Wing Yip Lam (Shenzhen, CN)
Assignee: Tencent Technology (Shenzhen) Company Limited
G10L15/063G10L15/02G10L15/16G10L15/22G10L21/0272G10L25/18G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,609,109
App. No.
17/583,512
Filed
Jan 25, 2022
Granted
Apr 21, 2026
Kind
B2
Art Unit
2657
USPC
704/232
Abstract

This application relates to a speech recognition method and apparatus, and a computer-readable storage medium, and the method includes: obtaining a first loss function of a speech separation and enhancement model and a second loss function of a speech recognition model; performing back propagation based on the second loss function to train an intermediate model bridged between the speech separation and enhancement model and the speech recognition model, to obtain a representation model; fusing the first loss function and the second loss function, to obtain a target loss function; and jointly training the speech separation and enhancement model, the representation model, and the speech recognition model based on the target loss function, and ending the training when a preset convergence condition is met.

Claims (85)

1 . A speech recognition method, performed by a computer device, the method comprising:

obtaining a first loss function of a speech separation and enhancement model and a second loss function of a speech recognition model;

performing back propagation based on the second loss function to train an intermediate model bridging between the speech separation and enhancement model and the speech recognition model, to obtain a representation model;

fusing the first loss function and the second loss function using a machine learning algorithm, to obtain a target loss function; and

jointly training the speech separation and enhancement model, the representation model, and the speech recognition model based on the target loss function, and ending the joint training when a preset convergence condition is met.

2 . The method according to claim 1 , further comprising:

extracting an estimated spectrum and an embedding feature matrix of a sample speech stream based on a first neural network model;

determining an attractor corresponding to the sample speech stream according to the embedding feature matrix and a preset ideal masking matrix;

obtaining a target masking matrix of the sample speech stream by calculating a similarity between each matrix element in the embedding feature matrix and the attractor;

determining an enhanced spectrum corresponding to the sample speech stream according to the target masking matrix; and

training the first neural network model based on a mean-square error (MSE) loss between the estimated spectrum and the enhanced spectrum corresponding to the sample speech stream, to obtain the speech separation and enhancement model.

3 . The method according to claim 2 , wherein extracting the estimated spectrum and the embedding feature matrix of the sample speech stream based on the first neural network model comprises:

performing Fourier transform on the sample speech stream, to obtain a speech spectrum and a speech feature of each audio frame of the sample speech stream;

performing speech separation (SS) and speech enhancement (SE) on the speech spectrum based on the first neural network model, to obtain the estimated spectrum; and

mapping the speech feature to an embedding space based on the first neural network model, to obtain the embedding feature matrix.

4 . The method according to claim 3 , wherein determining the attractor corresponding to the sample speech stream according to the embedding feature matrix and the preset ideal masking matrix comprises:

determining an ideal masking matrix according to the speech spectrum and the speech feature;

filtering out noise elements in the ideal masking matrix based on a preset binary threshold matrix to obtain the preset ideal masking matrix; and

determining the attractor corresponding to the sample speech stream according to the embedding feature matrix and the preset ideal masking matrix.

5 . The method according to claim 1 , further comprising:

obtaining a second neural network model;

performing non-negative constraint processing on the second neural network model, to obtain a non-negative neural network model;

obtaining a differential model configured for performing auditory matching on an acoustic feature outputted by the non-negative neural network model; and

cascading the differential model and the non-negative neural network model, to obtain the intermediate model.

6 . The method according to claim 5 , wherein obtaining the differential model configured for performing auditory matching on the acoustic feature outputted by the non-negative neural network model comprises:

obtaining a logarithmic model configured for performing a logarithmic operation on a feature vector corresponding to the acoustic feature;

obtaining a difference model configured for performing a difference operation on the feature vector corresponding to the acoustic feature; and

constructing the differential model according to the logarithmic model and the difference model.

7 . The method according to claim 1 , further comprising:

obtaining a sample speech stream and corresponding phoneme categories that are annotated;

extracting a depth feature of each audio frame of the sample speech stream by using a third neural network model;

determining a center vector of the sample speech stream according to depth features corresponding to audio frames of all the phoneme categories;

determining a fusion loss between an inter-class confusion measurement index and an intra-class distance penalty index of each audio frame based on the depth features and the center vector; and

training the third neural network model based on the fusion losses, to obtain the speech recognition model.

8 . The method according to claim 7 , wherein determining a fusion loss between the inter-class confusion measurement index and the intra-class distance penalty index of each audio frame based on the depth features and the center vector comprises:

inputting the depth features into a cross entropy function and calculating the inter-class confusion measurement index of each audio frame of the sample speech stream;

inputting the depth features and the center vector into a center loss function and calculating the intra-class distance penalty index of each audio frame; and

performing a fusion operation on the inter-class confusion measurement index and the intra-class distance penalty index, to obtain the fusion loss.

9 . The method according to claim 1 , wherein jointly training the speech separation and enhancement model, the representation model, and the speech recognition model based on the target loss function comprises:

determining a global descent gradient generated by the target loss function; and

iteratively updating model parameters corresponding to the speech separation and enhancement model, the representation model, and the speech recognition model according to the global descent gradient, until a minimum loss value of the target loss function is obtained.

10 . A speech recognition method, performed by a computer device, the method comprising:

obtaining a target speech stream;

extracting an enhanced spectrum of each audio frame of the target speech stream based on a speech separation and enhancement model;

performing auditory matching on the enhanced spectrum based on a representation model to obtain a representation feature, wherein the representation model is obtained by performing back propagation based on a first loss function of a speech recognition model to train an intermediate model bridging between the speech separation and enhancement model and the speech recognition model; and

recognizing the representation feature based on the speech recognition model, to obtain a phoneme corresponding to each audio frame,

wherein the speech separation and enhancement model, the representation model, and the speech recognition model comprises neural networks with network parameters obtained by joint training by iteratively minimizing a fused loss function of the first loss function and a second loss function of the speech separation and enhancement model.

11 . The method according to claim 10 , wherein the speech separation and enhancement model comprises a first neural network model; and extracting the enhanced spectrum of each audio frame of the target speech stream based on the speech separation and enhancement model comprises:

extracting an embedding feature matrix of each audio frame of the target speech stream based on the first neural network model;

determining an attractor corresponding to the target speech stream according to the embedding feature matrix and a preset ideal masking matrix;

obtaining a target masking matrix of the target speech stream by calculating a similarity between each matrix element in the embedding feature matrix and the attractor; and

determining the enhanced spectrum corresponding to each audio frame of the target speech stream according to the target masking matrix.

12 . The method according to claim 10 , wherein the representation model comprises a second neural network model and a differential model; and performing the auditory matching on the enhanced spectrum based on the representation model to obtain the representation feature comprises:

extracting an acoustic feature from the enhanced spectrum based on the second neural network model;

performing non-negative constraint processing on the acoustic feature, to obtain a non-negative acoustic feature; and

performing a differential operation on the non-negative acoustic feature based on the differential model, to obtain the representation feature matching auditory habits of human ears.

13 . A speech recognition device, comprising a memory for storing computer instructions and a processor for executing the computer instructions to:

obtain a first loss function of a speech separation and enhancement model and a second loss function of a speech recognition model;

perform back propagation based on the second loss function to train an intermediate model bridging between the speech separation and enhancement model and the speech recognition model, to obtain a representation model;

fuse the first loss function and the second loss function using a machine learning algorithm, to obtain a target loss function; and

jointly train the speech separation and enhancement model, the representation model, and the speech recognition model based on the target loss function, and ending the joint training when a preset convergence condition is met.

14 . The speech recognition device of claim 13 , wherein the processor is further configured to execute the computer instructions to:

extract an estimated spectrum and an embedding feature matrix of a sample speech stream based on a first neural network model;

determine an attractor corresponding to the sample speech stream according to the embedding feature matrix and a preset ideal masking matrix;

obtain a target masking matrix of the sample speech stream by calculating a similarity between each matrix element in the embedding feature matrix and the attractor;

determine an enhanced spectrum corresponding to the sample speech stream according to the target masking matrix; and

train the first neural network model based on a mean-square error (MSE) loss between the estimated spectrum and the enhanced spectrum corresponding to the sample speech stream, to obtain the speech separation and enhancement model.

15 . The speech recognition device of claim 14 , wherein to extract the estimated spectrum and the embedding feature matrix of the sample speech stream based on the first neural network model, the processor is configured to execute the computer instruction to:

perform Fourier transform on the sample speech stream, to obtain a speech spectrum and a speech feature of each audio frame of the sample speech stream;

perform speech separation (SS) and speech enhancement (SE) on the speech spectrum based on the first neural network model, to obtain the estimated spectrum; and

map the speech feature to an embedding space based on the first neural network model, to obtain the embedding feature matrix.

16 . The speech recognition device of claim 13 , wherein the processor is further configured to execute the computer instructions to:

obtain a second neural network model;

perform non-negative constraint processing on the second neural network model, to obtain a non-negative neural network model;

obtain a differential model configured for performing auditory matching on an acoustic feature outputted by the non-negative neural network model; and

cascade the differential model and the non-negative neural network model, to obtain the intermediate model.

17 . The speech recognition device of claim 13 , wherein the processor is further configured to execute the computer instructions to:

obtain a sample speech stream and corresponding phoneme categories that are annotated;

extract a depth feature of each audio frame of the sample speech stream by using a third neural network model;

determine a center vector of the sample speech stream according to depth features corresponding to audio frames of all the phoneme categories;

determine a fusion loss between an inter-class confusion measurement index and an intra-class distance penalty index of each audio frame based on the depth features and the center vector; and

train the third neural network model based on the fusion losses, to obtain the speech recognition model.

18 . A computer device, comprising a memory and a processor, the memory storing computer-readable instructions, the computer-readable instructions, when executed by the processor, causing the processor to perform the method according to claim 10 .

19 . A computer-readable storage medium, storing computer-readable instructions, the computer-readable instructions, when executed by a processor, causing the processor to perform the method according to claim 1 .

20 . A computer-readable storage medium, storing computer-readable instructions, the computer-readable instructions, when executed by a processor, causing the processor to perform the method according to claim 10 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2022
From: WANG, JUN; LAM, WING YIP
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 058759/0403 →
Priority Claims (1)
CN 202010048780.2 · Jan 16, 2020 · national
Continuity (2)
Continuation PCTCN2020128392 · Nov 12, 2020
Related Publication 20220148571A1 · May 12, 2022
References Cited (31)
US 10147442B1 · Panchapagesan et al. · 2018 [cited by applicant]
US 20180053087A1 · Fukuda · 2018 [cited by examiner]
US 20180350351A1 · Kopys · 2018 [cited by examiner]
US 20190043516A1 · Germain et al. · 2019 [cited by applicant]
US 20190228776A1 · Yamashita · 2019 [cited by applicant]
US 20200335091A1 · Chang · 2020 [cited by examiner]
US 20210158799A1 · Zhang · 2021 [cited by examiner]
CN 109378010A · 2019 [cited by applicant]
CN 109637526A · 2019 [cited by applicant]
CN 110060660A · 2019 [cited by applicant]
CN 110070855A · 2019 [cited by applicant]
CN 110120227A · 2019 [cited by applicant]
CN 110517666A · 2019 [cited by applicant]
CN 110570845A · 2019 [cited by applicant]
CN 110600017A · 2019 [cited by examiner]
CN 110648659A · 2020 [cited by applicant]
CN 111261146A · 2020 [cited by applicant]
JP 2019078857A · 2019 [cited by applicant]
WO WO2019198265A · 2019 [cited by applicant]
Lam, M. W., Wang, J., Liu, X., Meng, H., Su, D., & Yu, D. (2019). Extract, Adapt and Recognize: An End-to-End Neural Network for Corrupted Monaural Speech Recognition. In Interspeech (pp. 2778-2782). (Year: 2019). [cited by examiner]
Chen, Z., Droppo, J., Li, J. and Xiong, W., 2017. Progressive joint modeling in unsupervised single-channel overlapped speech recognition. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 26(1), pp. 184-… [cited by examiner]
Zhang, X. and Wang, H., Jul. 2016. A joint model of intent determination and slot filling for spoken language understanding. In IJCAI (vol. 16, No. 2016, pp. 2993-2999) (Year: 2016). [cited by examiner]
Z. Chen, Y. Luo and N. Mesgarani, “Deep attractor network for single-microphone speaker separation,” 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), New Orleans, LA, USA, 2017, pp… [cited by examiner]
Schmidt, M.N. and Olsson, R.K., Sep. 2006. Single-channel speech separation using sparse non-negative matrix factorization. In Interspeech (vol. 2, pp. 2-5). (Year: 2006). [cited by examiner]
International Search Report and Written Opinion mailed Feb. 9, 2021 for International Application No. PCT/CN2020/128392. [cited by applicant]
Search Report and Office Action issued on Chinese application No. 202010048780.2 on May 23, 2022, 6 pages, in Chinese language. [cited by applicant]
Extended European Search Report issued on European application No. 20913796.7 on Oct. 11, 2022, 10 pages. [cited by applicant]
Notice of Reasons for Refusal issued on Japanese application 2022-520112 on Dec. 23, 2022, 3p, in Japanese language. [cited by applicant]
Lam, Max W.Y. et al., “Extract, Adapt and Recognize: an End-to-end Neural Network for Corrupted Monaural Speech Recognition”, Interspeech 2019, ISCA, Sep. 19, 2019, pp. 2778-2782, Graz, AT. [cited by applicant]
Wang, Zhong-Qiu et al., “A Joint Training Framework for Robust Automatic Speech Recognition”, IEEE/Transactions on Audio, Speech, Language Processing, vol. 24, No. 4, Apr. 2016, pp. 796-806. [cited by applicant]
Li, Chenda et al., “ESPNET-SE: End-to-End Speech Enhancement and Separation Toolkit Designed for ASR Integration”, arxiv.org, Cornell University Library, Cornell University, Nov. 7, 2020, 8p, Ithaca, NY. [cited by applicant]