Text generation including de-duplication of decoded word information to splice target word information into an information sequence
Text generation including de-duplication of decoded word information to splice target word information into an information sequence. In a specific scheme, a decoded-word information group set is determined on the basis of a text to be processed; the decoded-word information group set is de-duplicated to generate a candidate-word information group set; candidate-word information that meets a target condition as target-word information is selected from each candidate-word information group in the candidate-word information group set to obtain a target-word information set; if the target-word information meets a convergence condition is determined, the target-word information with a historical target-word information sequence corresponding to the target-word information is spliced on the basis of a preset-word list to generate a target text.
1 . A text generation method, comprising:
determining a decoded-word information group set on the basis of a text to be processed, wherein the text to be processed is used for describing a specified object;
de-duplicating the decoded-word information group set to generate a candidate-word information group set;
selecting, from each candidate-word information group in the candidate-word information group set, candidate-word information that meets a target condition as target-word information, to obtain a target-word information set;
for each piece of target-word information in the target-word information set, in response to determining that the target-word information meets a convergence condition, splicing, on the basis of a preset-word list, the target-word information with a historical target-word information sequence corresponding to the target-word information, to generate a target text,
wherein the decoded-word information in the decoded-word information group set includes: decoded words and decoded-word probability values corresponding to the decoded words; and
the de-duplicating the decoded-word information group set to generate a candidate-word information group set includes:
for each decoded-word information group in the decoded-word information group set, dividing, based on the decoded words, the decoded-word information group into a duplicated decoded-word information group and a single decoded-word information group;
selecting, from the duplicated decoded-word information group, duplicated decoded-word information that meets a preset condition as target duplicated decoded-word information;
splicing the target duplicated decoded-word information with the single decoded-word information group to generate the candidate-word information group.
2 . The method of claim 1 , wherein, the determining a decoded-word information group set on the basis of a text to be processed includes:
inputting the text to be processed into a text encoder to generate an encoding hidden layer vector;
inputting the encoding hidden layer vector into a decoder to generate a decoded-word information group set.
3 . The method of claim 2 , wherein, the candidate-word information in the candidate-word information group set includes: candidate words and candidate-word probability values corresponding to the candidate words; and
the selecting, from each candidate-word information group in the candidate-word information group set, candidate-word information that meets a target condition as target-word information includes:
for each candidate-word information group in the candidate-word information group set, selecting at least one piece of candidate-word information from the candidate-word information group in descending order of the candidate-word probability values;
placing candidate-word information with a largest probability value in the at least one piece of candidate-word information, in the historical target-word information set, determining at least one piece of initial target-word information of the candidate-word information with the largest probability value, and generating target-word information corresponding to the candidate-word information based on the at least one piece of initial target-word information; the initial target-word information in the at least one piece of initial target-word information includes: initial target words and initial target-word probability values corresponding to the initial target words.
4 . The method of claim 2 , wherein, the for each piece of target-word information in the target-word information set, in response to determining that the target-word information meets a convergence condition, splicing, on the basis of a preset-word list, the target-word information with a historical target-word information sequence corresponding to the target-word information, to generate a target text includes:
for each piece of target-word information in the target-word information set, in response to determining that the target-word information meets a preset convergence condition, splicing the target-word information with a historical target-word information sequence corresponding to the target-word information, to generate a spliced text;
selecting, from the preset-word list, the words matching the spliced text as conjunctions;
combining the conjunctions with the spliced text to generate the target text.
5 . The method of claim 1 , wherein, the for each decoded-word information group in the decoded-word information group set, dividing, based on the decoded words, the decoded-word information group into a duplicated decoded-word information group and a single decoded-word information group includes:
for each decoded-word information group in the decoded-word information group set, in response to the decoded-word information group having other decoded-word information containing decoded words of the said decoded-word information, placing the decoded-word information and the other decoded-word information in the duplicated decoded-word information group; otherwise, placing the decoded-word information in the single decoded-word information group.
6 . The method of claim 5 , wherein, the candidate-word information in the candidate-word information group set includes: candidate words and candidate-word probability values corresponding to the candidate words; and
the selecting, from each candidate-word information group in the candidate-word information group set, candidate-word information that meets a target condition as target-word information includes:
for each candidate-word information group in the candidate-word information group set, selecting at least one piece of candidate-word information from the candidate-word information group in descending order of the candidate-word probability values;
placing candidate-word information with a largest probability value in the at least one piece of candidate-word information, in the historical target-word information set, determining at least one piece of initial target-word information of the candidate-word information with the largest probability value, and generating target-word information corresponding to the candidate-word information based on the at least one piece of initial target-word information; the initial target-word information in the at least one piece of initial target-word information includes: initial target words and initial target-word probability values corresponding to the initial target words.
7 . The method of claim 1 , wherein, the candidate-word information in the candidate-word information group set includes: candidate words and candidate-word probability values corresponding to the candidate words; and
the selecting, from each candidate-word information group in the candidate-word information group set, candidate-word information that meets a target condition as target-word information includes:
for each candidate-word information group in the candidate-word information group set, selecting at least one piece of candidate-word information from the candidate-word information group in descending order of the candidate-word probability values;
placing candidate-word information with a largest probability value in the at least one piece of candidate-word information, in the historical target-word information set, determining at least one piece of initial target-word information of the candidate-word information with the largest probability value, and generating target-word information corresponding to the candidate-word information based on the at least one piece of initial target-word information; the initial target-word information in the at least one piece of initial target-word information includes: initial target words and initial target-word probability values corresponding to the initial target words.
8 . The method of claim 7 , wherein, the generating target-word information corresponding to the candidate-word information based on the at least one piece of initial target-word information includes:
in response to the initial target word corresponding to the initial target-word information with the largest initial target-word probability value not belonging to the historical target-word information set, setting the initial target-word information with the largest initial target-word probability value as the target-word information, and placing the target-word information in the historical target-word information set.
9 . The method of claim 7 , wherein, the generating target-word information corresponding to the candidate-word information based on the at least one piece of initial target-word information includes:
in response to the initial target word corresponding to the initial target-word information with the largest initial target-word probability value belonging to the historical target-word information set, determining a duplicated-word probability difference between the initial target-word information with the largest initial target-word probability value and the historical target-word information corresponding to the historical target-word information set;
in response to the duplicated-word probability difference being greater than the initial target-word probability value of other initial target-word information corresponding to the candidate-word information, setting the initial target-word information with the largest initial target-word probability value as the target-word information, and placing the target-word information in the historical target-word information set.
10 . The method of claim 9 , wherein, the generating target-word information corresponding to the candidate-word information based on the at least one piece of initial target-word information includes:
in response to the initial target word corresponding to the initial target-word information with the largest initial target-word probability value not belonging to the historical target-word information set, setting the initial target-word information with the largest initial target-word probability value as the target-word information, and placing the target-word information in the historical target-word information set.
11 . The method of claim 7 , wherein, the generating target-word information corresponding to the candidate-word information based on the at least one piece of initial target-word information includes:
in response to the initial target word corresponding to the initial target-word information with the largest initial target-word probability value belonging to the historical target-word information set, determining a duplicated-word probability difference between the initial target-word information with the largest initial target-word probability value and the historical target-word information corresponding to the historical target-word information set;
in response to the duplicated-word probability difference being less than or equal to the initial target-word probability value of other initial target-word information corresponding to the candidate-word information, setting the initial target-word information with the largest initial target-word probability value in other initial target-word information as the target-word information, and placing the target-word information in the historical target-word information set.
12 . The method of claim 11 , wherein, the generating target-word information corresponding to the candidate-word information based on the at least one piece of initial target-word information includes:
in response to the initial target word corresponding to the initial target-word information with the largest initial target-word probability value not belonging to the historical target-word information set, setting the initial target-word information with the largest initial target-word probability value as the target-word information, and placing the target-word information in the historical target-word information set.
13 . The method of claim 1 , wherein, the for each piece of target-word information in the target-word information set, in response to determining that the target-word information meets a convergence condition, splicing, on the basis of a preset-word list, the target-word information with a historical target-word information sequence corresponding to the target-word information, to generate a target text includes:
for each piece of target-word information in the target-word information set, in response to determining that the target-word information meets a preset convergence condition, splicing the target-word information with a historical target-word information sequence corresponding to the target-word information, to generate a spliced text;
selecting, from the preset-word list, the words matching the spliced text as conjunctions;
combining the conjunctions with the spliced text to generate the target text.
14 . The method of claim 1 , wherein, the candidate-word information in the candidate-word information group set includes: candidate words and candidate-word probability values corresponding to the candidate words; and
the selecting, from each candidate-word information group in the candidate-word information group set, candidate-word information that meets a target condition as target-word information includes:
for each candidate-word information group in the candidate-word information group set, selecting at least one piece of candidate-word information from the candidate-word information group in descending order of the candidate-word probability values;
placing candidate-word information with a largest probability value in the at least one piece of candidate-word information, in the historical target-word information set, determining at least one piece of initial target-word information of the candidate-word information with the largest probability value, and generating target-word information corresponding to the candidate-word information based on the at least one piece of initial target-word information; the initial target-word information in the at least one piece of initial target-word information includes: initial target words and initial target-word probability values corresponding to the initial target words.
15 . The method of claim 1 , wherein, the for each piece of target-word information in the target-word information set, in response to determining that the target-word information meets a convergence condition, splicing, on the basis of a preset-word list, the target-word information with a historical target-word information sequence corresponding to the target-word information, to generate a target text includes:
for each piece of target-word information in the target-word information set, in response to determining that the target-word information meets a preset convergence condition, splicing the target-word information with a historical target-word information sequence corresponding to the target-word information, to generate a spliced text;
selecting, from the preset-word list, the words matching the spliced text as conjunctions;
combining the conjunctions with the spliced text to generate the target text.
16 . An electronic device, comprising:
at least one processor;
a storage apparatus on which at least one program are stored, and
when at least one program are executed by at least one processor, the at least one processor implement a text generation method, comprising:
determining a decoded-word information group set on the basis of a text to be processed, wherein the text to be processed is used for describing a specified object;
de-duplicating the decoded-word information group set to generate a candidate-word information group set;
selecting, from each candidate-word information group in the candidate-word information group set, candidate-word information that meets a target condition as target-word information, to obtain a target-word information set;
for each piece of target-word information in the target-word information set, in response to determining that the target-word information meets a convergence condition, splicing, on the basis of a preset-word list, the target-word information with a historical target-word information sequence corresponding to the target-word information, to generate a target text,
wherein the decoded-word information in the decoded-word information group set includes: decoded words and decoded-word probability values corresponding to the decoded words; and
the de-duplicating the decoded-word information group set to generate a candidate-word information group set includes:
for each decoded-word information group in the decoded-word information group set, dividing, based on the decoded words, the decoded-word information group into a duplicated decoded-word information group and a single decoded-word information group;
selecting, from the duplicated decoded-word information group, duplicated decoded-word information that meets a preset condition as target duplicated decoded-word information;
splicing the target duplicated decoded-word information with the single decoded-word information group to generate the candidate-word information group.
17 . A non-transitory computer-readable medium, on which a computer program is stored, wherein a text generation method is implemented when the program is executed by a processor, the text generation method, comprising:
determining a decoded-word information group set on the basis of a text to be processed, wherein the text to be processed is used for describing a specified object;
de-duplicating the decoded-word information group set to generate a candidate-word information group set;
selecting, from each candidate-word information group in the candidate-word information group set, candidate-word information that meets a target condition as target-word information, to obtain a target-word information set;
for each piece of target-word information in the target-word information set, in response to determining that the target-word information meets a convergence condition, splicing, on the basis of a preset-word list, the target-word information with a historical target-word information sequence corresponding to the target-word information, to generate a target text,
wherein the decoded-word information in the decoded-word information group set includes: decoded words and decoded-word probability values corresponding to the decoded words; and
the de-duplicating the decoded-word information group set to generate a candidate-word information group set includes:
for each decoded-word information group in the decoded-word information group set, dividing, based on the decoded words, the decoded-word information group into a duplicated decoded-word information group and a single decoded-word information group;
selecting, from the duplicated decoded-word information group, duplicated decoded-word information that meets a preset condition as target duplicated decoded-word information;
splicing the target duplicated decoded-word information with the single decoded-word information group to generate the candidate-word information group.