Method and system for constructing speech recognition model and speech processing
A method for constructing a speech recognition model includes obtaining a target keyword; determining a synonym group semantically associated with the target keyword; training a language model based on the target keyword and the synonym group; to obtain a target language model; generating a first decoding graph based on the target language model, where the first decoding graph indicates a plurality of decoding paths that satisfy a syntax constraint rule that is based on the target keyword and the synonym group; and determining the speech recognition model based on the first decoding graph.
1 . A method for constructing a speech recognition model, wherein the method comprises:
obtaining a target keyword, wherein the target keyword comprises a first keyword and a second keyword;
obtaining a synonym group semantically associated with the target keyword, wherein the synonym group comprises a first synonym group semantically associated with the first keyword and a second synonym group semantically associated with the second keyword;
training, based on the target keyword and the synonym group, a language model to obtain a target language model;
generating, based on the target language model, a first decoding graph indicating a plurality of decoding paths that satisfy a syntax constraint rule that is based on the target keyword and the synonym group; and
determining, based on the first decoding graph, the speech recognition model, wherein determining the speech recognition model comprises:
obtaining, from the first decoding graph, a first group of decoding paths and a second group of decoding paths, wherein the first group comprises first decoding paths corresponding to the first keyword and the first synonym group, and wherein the second group comprises second decoding paths corresponding to the second keyword and the second synonym group;
generating, based on the first group, a first subgraph;
generating, based on the second group, a second subgraph; and
determining, based on the first subgraph and the second subgraph, the speech recognition model.
2 . The method of claim 1 , wherein obtaining the synonym group comprises:
determining first semantics of the target keyword; and
determining the synonym group based on the first semantics, wherein a first difference between second semantics of each synonym in the synonym group and the first semantics is less than a difference threshold.
3 . The method of claim 2 , wherein determining the synonym group comprises determining the synonym group further based on a first length of the target keyword, and wherein a second difference between a second length of each synonym in the synonym group and the first length is less than a length threshold.
4 . The method of claim 2 , wherein determining the synonym group further comprises:
obtaining, based on the first semantics, a plurality of candidate synonyms;
providing, for a user, the candidate synonyms;
receiving, from the user, a user input indicating that at least one candidate synonym in the candidate synonyms is excluded or confirmed; and
determining, from the candidate synonyms and based on the user input, the synonym group.
5 . The method of claim 1 , wherein the first subgraph indicates a first decoding path corresponding to the first keyword and a second decoding path corresponding to a synonym in the first synonym group, and wherein the first decoding path and the second decoding path have a same weight in the first subgraph.
6 . The method of claim 1 , wherein obtaining the target keyword comprises:
obtaining, based on a pre-stored historical keyword and a received keyword, a first keyword group;
determining that a quantity of keywords in the first keyword group exceeds a predetermined threshold; and
obtaining, in response to determining that the quantity of keywords exceeds the predetermined threshold, from the first keyword group, and based on the predetermined threshold, the target keyword.
7 . The method of claim 6 , wherein obtaining the target keyword comprises obtaining, further based on an attribute of a keyword in the target keyword, the target keyword, and wherein a quantity of target keywords is the predetermined threshold.
8 . The method of claim 1 , further comprising indicating to provide the speech recognition model to a target computing device for deployment of the speech recognition model on the target computing device.
9 . The method of claim 1 , wherein the second subgraph indicates a third decoding path corresponding to the second keyword and a fourth decoding path corresponding to a synonym in the second synonym group, and wherein the third decoding path and the fourth decoding path have a same weight in the second subgraph.
10 . A method for speech processing comprising:
receiving a speech input;
obtaining a target keyword, wherein the target keyword comprises a first keyword and a second keyword;
obtaining a synonym group semantically associated with the target keyword, wherein the synonym group comprises a first synonym group semantically associated with the first keyword and a second synonym group semantically associated with the second keyword;
training, based on the target keyword and the synonym group, a language model to obtain a target language model;
generating, based on the target language model, a first decoding graph indicating a plurality of decoding paths that satisfy a syntax constraint rule that is based on the target keyword and the synonym group;
determining, based on the first decoding graph, a speech recognition model, wherein determining the speech recognition model comprises:
obtaining, from the first decoding graph, a first group of decoding paths and a second group of decoding paths, wherein the first group comprises first decoding paths corresponding to the first keyword and the first synonym group, and wherein the second group comprises second decoding paths corresponding to the second keyword and the second synonym group;
generating, based on the first group, a first subgraph;
generating, based on the second group, a second subgraph; and
determining, based on the first subgraph and the second subgraph, the speech recognition model; and
determining, using the speech recognition model, text representation associated with the speech input.
11 . The method of claim 10 , wherein obtaining the synonym group comprises:
determining first semantics of the target keyword; and
determining the synonym group based on the first semantics, wherein a first difference between second semantics of each synonym in the synonym group and the first semantics is less than a difference threshold.
12 . The method of claim 11 , wherein determining the synonym group comprises determining the synonym group further based on a first length of the target keyword, and wherein a second difference between a second length of each synonym in the synonym group and the first length is less than a length threshold.
13 . The method of claim 11 , wherein determining the synonym group further comprises:
obtaining, based on the first semantics, a plurality of candidate synonyms;
providing, for a user, the candidate synonyms;
receiving, from the user, a user input indicating that at least one candidate synonym in the candidate synonyms is excluded or confirmed; and
determining, from the candidate synonyms and based on the user input, the synonym group.
14 . The method of claim 10 , wherein the first subgraph indicates a first decoding path corresponding to the first keyword and a second decoding path corresponding to a synonym in the first synonym group, and wherein the first decoding path and the second decoding path have a same weight in the first subgraph.
15 . The method of claim 10 , wherein obtaining the target keyword comprises:
obtaining, based on a pre-stored historical keyword and a received keyword, a first keyword group;
determining that a quantity of keywords in the first keyword group exceeds a predetermined threshold; and
obtaining, in response to determining that the quantity of keywords exceeds the predetermined threshold, from the first keyword group, and based on the predetermined threshold, the target keyword.
16 . The method of claim 15 , wherein obtaining the target keyword further comprises obtaining, further based on an attribute of a keyword in the target keyword, the target keyword, and wherein a quantity of target keywords is the predetermined threshold.
17 . The method of claim 10 , further comprising performing an action corresponding to the text representation.
18 . The method of claim 10 , wherein the text representation corresponds to the target keyword or a synonym in the synonym group.
19 . The method of claim 10 , wherein the second subgraph indicates a third decoding path corresponding to the second keyword and a fourth decoding path corresponding to a synonym in the second synonym group, and wherein the third decoding path and the fourth decoding path have a same weight in the second subgraph.
20 . An electronic device, comprising:
a memory configured to store instructions; and
a processor coupled to the memory and configured to execute the instructions to cause the electronic device to:
obtain a target keyword, wherein the target keyword comprises a first keyword and a second keyword;
obtain a synonym group semantically associated with the target keyword, wherein the synonym group comprises a first synonym group semantically associated with the first keyword and a second synonym group semantically associated with the second keyword, and wherein obtaining the synonym group comprises:
determining first semantics of the target keyword; and
determining the synonym group based on the first semantics, wherein a first difference between second semantics of each synonym in the synonym group and the first semantics is less than a difference threshold, and wherein determining the synonym group comprises determining the synonym group further based on a first length of the target keyword, and wherein a second difference between a second length of each synonym in the synonym group and the first length is less than a length threshold;
train, based on the target keyword and the synonym group, a language model to obtain a target language model;
generate, based on the target language model, a first decoding graph indicating a plurality of decoding paths that satisfy a syntax constraint rule that is based on the target keyword and the synonym group; and
determine, based on the first decoding graph, a speech recognition model, wherein determining the speech recognition model comprises:
obtaining, from the first decoding graph, a first group of decoding paths and a second group of decoding paths, wherein the first group comprises first decoding paths corresponding to the first keyword and the first synonym group, and wherein the second group comprises second decoding paths corresponding to the second keyword and the second synonym group;
generating, based on the first group, a first subgraph;
generating, based on the second group, a second subgraph; and
determining, based on the first subgraph and the second subgraph, the speech recognition model.