IP Library › Granted Patent US 11,798,531
Granted Patent B2
US 11,798,531 · App. 17/077,141 · Granted Oct 24, 2023

Speech recognition method and apparatus, and method and apparatus for training speech recognition model

Inventors: Jun Wang (Shenzhen, CN); Dan Su (Shenzhen, CN); Dong Yu (Bothell, WA)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
G10L15/02G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,798,531
App. No.
17/077,141
Granted
Oct 24, 2023
Kind
B2
Abstract

A speech recognition method, a speech recognition apparatus, and a method and an apparatus for training a speech recognition model are provided. The speech recognition method includes: recognizing a target word speech from a hybrid speech, and obtaining, as an anchor extraction feature of a target speech, an anchor extraction feature of the target word speech based on the target word speech; obtaining a mask of the target speech according to the anchor extraction feature of the target speech; and recognizing the target speech according to the mask of the target speech.

Claims (71)

1. A speech recognition method, performed by at least one processor of an electronic device, the method comprising:

recognizing a target word speech from a hybrid speech, and obtaining, as an anchor extraction feature of a target speech, an anchor extraction feature of the target word speech based on the target word speech;

obtaining a mask of the target speech according to the anchor extraction feature of the target speech; and

recognizing the target speech according to the mask of the target speech,

wherein the obtaining the anchor extraction feature of the target word speech comprises:

determining, by the at least one processor, an embedding vector corresponding to each time-frequency window of the hybrid speech; and

obtaining, by the at least one processor, the anchor extraction feature of the target word speech according to determined embedding vectors and a preset anchor extraction feature.

2. The method according to claim 1 , wherein the obtaining the anchor extraction feature of the target word speech further comprises:

determining, according to determined embedding vectors and the preset anchor extraction feature, target word annotation information corresponding to the embedding vectors; and

obtaining the anchor extraction feature of the target word speech according to the embedding vectors, the preset anchor extraction feature, and the corresponding target word annotation information.

3. The method according to claim 1 , wherein the obtaining the mask comprises:

obtaining normalized embedding vectors corresponding to the embedding vectors according to the embedding vectors and the anchor extraction feature of the target speech; and

obtaining the mask of the target speech according to the normalized embedding vectors and a preset normalized anchor extraction feature.

4. The method according to claim 1 , wherein the determining the embedding vector comprises:

performing a short-time Fourier transform (STFT) on the hybrid speech, to obtain a frequency spectrum of the hybrid speech; and

mapping the frequency spectrum of the hybrid speech into an original embedding space of a fixed dimension, to obtain the embedding vector corresponding to each time-frequency window of the hybrid speech.

5. The method according to claim 2 , wherein the determining the target word annotation information comprises:

separately combining the embedding vectors with the preset anchor extraction feature; and

obtaining the target word annotation information corresponding to the embedding vectors by inputting combined embedding vectors into a pre-trained first forward network, wherein a value of target word annotation information corresponding to an embedding vector not comprising the target word speech is 0, and a value of target word annotation information corresponding to an embedding vector comprising the target word speech is 1.

6. The method according to claim 3 , wherein the obtaining the normalized embedding vectors comprises:

separately combining the embedding vectors with the anchor extraction feature of the target speech, to obtain combined 2K-dimensional vectors, wherein the embedding vectors and the anchor extraction feature of the target speech are K-dimensional vectors, respectively;

inputting the combined 2K-dimensional vectors into a pre-trained second forward network; and

mapping the combined 2K-dimensional vectors into a normalized embedding space of a fixed dimension based on the second forward network, to obtain, as the normalized embedding vectors of the corresponding embedding vectors, corresponding K-dimensional vectors outputted by the second forward network.

7. The method according to claim 3 , wherein the obtaining the mask of the target speech comprises:

obtaining distances between the normalized embedding vectors and the preset normalized anchor extraction feature, and obtaining the mask of the target speech according to the distances.

8. The method according to claim 1 , further comprising:

inputting the recognized target speech into a pre-trained target word determining module, which is implemented in computer code executable by the at least one processor, to determine whether the target speech comprises the target word speech;

adjusting the anchor extraction feature of the target speech to reduce a weight of a preset anchor extraction feature in response to determining that the target speech comprises the target word speech, or adjusting the anchor extraction feature of the target speech to increase the weight of the preset anchor extraction feature in response to determining that the target speech does not comprise the target word speech, wherein the anchor extraction feature of the target speech is obtained by using the preset anchor extraction feature; and

recognizing the target speech according to the adjusted anchor extraction feature of the target speech.

9. A method for training a speech recognition model, performed by at least one processor of an electronic device, the speech recognition model comprising a target speech extraction module and a target word determining module, each being implemented in computer code executable by the at least one processor, the method comprising:

obtaining a speech sample set, the speech sample set being any one or any combination of: a clean target word speech sample set, a positive and negative sample set of a noisy target word speech, and a noisy command speech sample set;

training the target speech extraction module by using the speech sample set as an input of the target speech extraction module and by using a recognized target speech as an output of the target speech extraction module, a target function of the target speech extraction module being to minimize a loss function between the recognized target speech and a clean target speech; and

training the target word determining module by using, as an input of the target word determining module, a target speech outputted by the target speech extraction module, and by using, as an output of the target word determining module, a target word determining probability, a target function of the a target word determining module being to minimize a cross entropy (CE) loss function of a target word determining result,

wherein the obtaining the speech sample comprises obtaining an embedding vector corresponding to each time-frequency window of any one or any combination of: the clean target word speech sample set, the positive and negative sample set of the noisy target word speech, and the noisy command n speech sample set, and obtaining an anchor extraction feature according to obtained embedding vectors and a preset anchor extraction feature, and

wherein, in the target speech extraction module, the target speech is recognized based on the anchor extraction feature.

10. A speech recognition apparatus, comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code comprising:

first obtaining code configured to cause at least one of the at least one processor to recognize a target word speech from a hybrid speech, and obtain, as an anchor extraction feature of a target speech, an anchor extraction feature of the target word speech based on the target word speech;

second obtaining code configured to cause at least one of the at least one processor to obtain a mask of the target speech according to the anchor extraction feature of the target speech; and

recognition code configured to cause at least one of the at least one processor to recognize the target speech according to the mask of the target speech,

wherein the first obtaining code is configured to cause at least one of the at least one processor to determine an embedding vector corresponding to each time-frequency window of the hybrid speech; and obtain the anchor extraction feature of the target word speech according to determined embedding vectors and a preset anchor extraction feature.

11. The apparatus according to claim 10 , wherein the first obtaining code is configured to cause at least one of the at least one processor to:

determine, according to determined embedding vectors and the preset anchor extraction feature, target word annotation information corresponding to the embedding vectors; and

obtain the anchor extraction feature of the target word speech according to the embedding vectors, the preset anchor extraction feature, and the corresponding target word annotation information.

12. The apparatus according to claim 10 , wherein the second obtaining code is configured to cause at least one of the at least one processor to:

obtain normalized embedding vectors corresponding to the embedding vectors according to the embedding vectors and the anchor extraction feature of the target speech; and

obtain the mask of the target speech according to the normalized embedding vectors and a preset normalized anchor extraction feature.

13. The apparatus according to claim 12 , wherein the second obtaining code is configured to cause at least one of the at least one processor to:

separately combine the embedding vectors with the anchor extraction feature of the target speech, to obtain combined 2K-dimensional vectors, wherein the embedding vectors and the anchor extraction feature of the target speech are K-dimensional vectors, respectively;

input the combined 2K-dimensional vectors into a pre-trained second forward network; and

map the combined 2K-dimensional vectors into a normalized embedding space of a fixed dimension based on the second forward network, to obtain, as the normalized embedding vectors of the corresponding embedding vectors, corresponding K-dimensional vectors outputted by the second forward network.

14. The apparatus according to claim 12 , wherein the second obtaining code is configured to cause at least one of the at least one processor to:

obtain distances between the normalized embedding vectors and the preset normalized anchor extraction feature, and obtain the mask of the target speech according to the distances.

15. The apparatus according to claim 10 , wherein the program code further comprises:

adjustment code configured to cause at least one of the at least one processor to input the recognized target speech into a pre-trained target word determining module, which is implemented in computer code executable by at least one of the at least one processor, to determine whether the target speech comprises the target word speech; and adjust the anchor extraction feature of the target speech to reduce a weight of a preset anchor extraction feature in response to determining that the target speech comprises the target word speech, or adjust the anchor extraction feature of the target speech to increase the weight of the preset anchor extraction feature in response to determining that the target speech does not comprise the target word speech, wherein the anchor extraction feature of the target speech is obtained by using the preset anchor extraction feature, and

wherein the target speech is recognized according to the adjusted anchor extraction feature of the target speech.

16. An apparatus for training a speech recognition model, the apparatus comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code to perform the method for training the speech recognition model according to claim 9 , the speech recognition model comprising a target speech extraction module and a target word determining module, each being implemented in computer code executable by the at least one processor,

the program code comprising:

obtaining code configured to cause at least one of the at least one processor to obtain a speech sample set, the speech sample set being any one or any combination: a clean target word speech sample set, a positive and negative sample set of a noisy target word speech, and a noisy command speech sample set; and

training code configured to cause at least one of the at least one processor to train the target speech extraction module by using the speech sample set as an input of the target speech extraction module and by using a recognized target speech as an output of the target speech extraction module, a target function of the target speech extraction module being to minimize a loss function between the recognized target speech and a clean target speech; and train the target word determining module by using, as an input of the target word determining module, a target speech outputted by the target speech extraction module, and by using, as an output of the target word determining module, a target word determining probability, a target function of the target word determining module being to minimize a cross entropy (CE) loss function of a target word determining result.

17. An electronic device, comprising:

at least one memory, configured to store computer-readable program instructions; and

at least one processor, configured to call the computer-readable program instructions stored in the at least one memory to perform the speech recognition method according to claim 1 .

18. A non-transitory computer-readable storage medium, storing computer-readable program instructions, the computer-readable program instructions being loaded by a processor to perform the method according to claim 1 .

19. An electronic device, comprising:

at least one memory, configured to store computer-readable program instructions; and

at least one processor, configured to call the computer-readable program instructions stored in the at least one memory to perform the method according to claim 9 .

20. A non-transitory computer-readable storage medium, storing computer-readable program instructions, the computer-readable program instructions being loaded by a processor to perform the method according to claim 9 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2020
From: WANG, JUN; SU, DAN; YU, DONG
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 054138/0681 →
Priority Claims (1)
CN 201811251081.7 · Oct 25, 2018 · national
Continuity (2)
Continuation PCTCN2019111905 · Oct 18, 2019
Related Publication 20210043190A1 · Feb 11, 2021