Method, apparatus, electronic device and storage medium for text content matching
View Patent ↗Embodiments of the present disclosure provide a method, apparatus, electronic device, and storage medium for text content matching. The method of text content matching includes: in accordance with a collection of to-be-processed speech information, determining a to-be-processed acoustic feature corresponding to the to-be-processed speech information (S 110 ); processing, based on an audio following method, the to-be-processed acoustic feature to obtain a to-be-matched utterance corresponding to the to-be-processed acoustic feature (S 120 ); and determining a target utterance associated with the to-be-matched utterance in target text and differentiating a display of the target utterance in the target text (S 130 ).
1 . A method of text content matching, comprising:
in accordance with a collection of speech information, determining an acoustic feature corresponding to the speech information;
processing, based on an audio following method, the to be processed acoustic feature to obtain an utterance corresponding to the acoustic feature, wherein the processing the acoustic feature to obtain the utterance corresponding to the acoustic feature comprises:
processing the acoustic feature based on an acoustic model in the audio following method to obtain an acoustic posterior probability corresponding to the acoustic feature,
determining, based on the acoustic posterior probability and a decoder corresponding to text in the audio following method, a first utterance corresponding to the acoustic feature and a first confidence corresponding to the first utterance, wherein the decoder is determined based on an interpolation language model corresponding to the text, and the interpolation language model is determined based on a target language model and an ordinary language model corresponding to the text, and
in accordance with a determination that the first confidence satisfies a predetermined confidence threshold, determining the first utterance as the utterance; and
determining a target utterance associated with the utterance in the text and differentiating a display of the target utterance in the text.
2 . The method of claim 1 , further comprising:
uploading the text to determine, in accordance with the collection of the speech information, the target utterance associated with the utterance corresponding to the speech information in the text.
3 . The method of claim 1 , wherein the in accordance with a collection of speech information, determining an acoustic feature corresponding to the speech information comprises:
in accordance with a determination that a user interacts based on a real-time interactive interface, collecting speech information of a target user; and
performing, based on an audio feature extraction algorithm, feature extraction on the speech information, and obtaining the acoustic feature corresponding to the speech information.
4 . The method of claim 1 , wherein the processing, based on an audio following method, the acoustic feature to obtain an utterance corresponding to the acoustic feature comprises:
determining, based on a keyword detection system and the acoustic feature in the audio following method, a second utterance corresponding to the acoustic feature and a second confidence corresponding to the second utterance; wherein the keyword detection system matches the text; and
in accordance with a determination that the second confidence satisfies a predetermined confidence threshold, determining the second utterance as the utterance.
5 . The method of claim 1 , wherein the audio following method comprises a keyword detection system and a decoder and the processing, based on an audio following method, the acoustic feature to obtain an utterance corresponding to the acoustic feature comprises:
processing the acoustic feature based on the decoder and the keyword detection system, respectively, and in accordance with a determination that the first utterance and a second utterance are obtained, determining the utterance based on a first confidence of the first utterance and a second confidence of the second utterance.
6 . The method of claim 1 , wherein the differentiating a display of the target utterance in the text comprises:
highlighting the target utterance; or,
displaying the target utterance in bold; or,
displaying utterances in the target text other than the target utterance in a semi-transparent form; wherein transparency of a predetermined number of unmatched utterances adjacent to the target utterance is lower than transparency of other unmatched utterances.
7 . The method of claim 1 , wherein during the determining of the target utterance, the method further comprises:
determining an actual speech duration corresponding to the target utterance;
adjusting a predicted speech duration corresponding to an unmatched utterance based on the actual speech duration and the unmatched utterance in the text; and
displaying the predicted speech duration on a target client as a prompt to the target user.
8 . The method of claim 1 , further comprising:
in accordance with a reception of the text, performing utterance-segmentation marking on the text, and displaying an utterance-segmentation marking identification on a client, to cause a user to read the text based on the utterance-segmentation marking identification.
9 . The method of claim 1 , further comprising:
in accordance with a reception of the text, performing emotion marking on respective utterance in the text, and displaying an emotion marking identification on a client, to cause a user to read the text based on the emotional marking identification.
10 . An electronic device comprising:
at least one processor; and a storage device configured to store at least one program; wherein the at least one program, when executed by the at least one processor, causes the at least one processor to implement operations comprising:
in accordance with a collection of speech information, determining an acoustic feature corresponding to the speech information;
processing, based on an audio following method, the acoustic feature to obtain an utterance corresponding to the acoustic feature, wherein the processing the acoustic feature to obtain the utterance corresponding to the acoustic feature comprises:
processing the acoustic feature based on an acoustic model in the audio following method to obtain an acoustic posterior probability corresponding to the acoustic feature,
determining, based on the acoustic posterior probability and a decoder corresponding to the text in the audio following method, a first utterance corresponding to the acoustic feature and a first confidence corresponding to the first utterance; wherein the decoder is determined based on an interpolation language model corresponding to the text, and the interpolation language model is determined based on a target language model and an ordinary language model corresponding to the text, and
in accordance with a determination that the first confidence satisfies a predetermined confidence threshold, determining the first utterance as the utterance; and
determining a target utterance associated with the utterance in text and differentiating a display of the target utterance in the text.
11 . The electronic device of claim 10 , wherein the operations further comprise:
uploading the text to determine, in accordance with the collection of the speech information, the target utterance associated with the utterance corresponding to the speech information in the text.
12 . The electronic device of claim 10 , wherein the in accordance with a collection of speech information, determining an acoustic feature corresponding to the speech information comprises:
in accordance with a determination that a user interacts based on a real-time interactive interface, collecting speech information of a target user; and
performing, based on an audio feature extraction algorithm, feature extraction on the speech information, and obtaining the acoustic feature corresponding to the speech information.
13 . The electronic device of claim 10 , wherein the processing, based on an audio following method, the acoustic feature to obtain an utterance corresponding to the acoustic feature comprises:
determining, based on a keyword detection system and the acoustic feature in the audio following method, a second utterance corresponding to the acoustic feature and a second confidence corresponding to the second utterance; wherein the keyword detection system matches the text; and
in accordance with a determination that the second confidence satisfies a predetermined confidence threshold, determining the second utterance as the utterance.
14 . The electronic device of claim 10 , wherein the audio following method comprises a keyword detection system and a decoder and the processing, based on an audio following method, the acoustic feature to obtain an utterance corresponding to the acoustic feature comprises:
processing the acoustic feature based on the decoder and the keyword detection system, respectively, and in accordance with a determination that the first utterance and a second utterance are obtained, determining the utterance based on a first confidence of the first utterance and a second confidence of the second utterance.
15 . The electronic device of claim 10 , wherein the differentiating a display of the target utterance in the text comprises:
highlighting the target utterance; or,
displaying the target utterance in bold; or,
displaying utterances in the text other than the target utterance in a semi-transparent form; wherein transparency of a predetermined number of unmatched utterances adjacent to the target utterance is lower than transparency of other unmatched utterances.
16 . The electronic device of claim 10 , wherein during the determining of the target utterance, the operations further comprise:
determining an actual speech duration corresponding to the target utterance;
adjusting a predicted speech duration corresponding to an unmatched utterance based on the actual speech duration and the unmatched utterance in the text; and
displaying the predicted speech duration on a target client as a prompt to the target user.
17 . The electronic device of claim 10 , wherein the operations further comprise:
in accordance with a reception of the text, performing utterance-segmentation marking on the text, and displaying an utterance-segmentation marking identification on a client, to cause a user to read the text based on the utterance-segmentation marking identification.
18 . A non-transitory storage medium comprising computer-executable instructions, the computer-executable instructions, when executed by a processor, cause the processor to perform operations comprising:
in accordance with a collection of speech information, determining an acoustic feature corresponding to the speech information;
processing, based on an audio following method, the acoustic feature to obtain an utterance corresponding to the acoustic feature, wherein the processing the acoustic feature to obtain the utterance corresponding to the acoustic feature comprises:
processing the acoustic feature based on an acoustic model in the audio following method to obtain an acoustic posterior probability corresponding to the acoustic feature,
determining, based on the acoustic posterior probability and a decoder corresponding to the text in the audio following method, a first utterance corresponding to the acoustic feature and a first confidence corresponding to the first utterance; wherein the decoder is determined based on an interpolation language model corresponding to the text, and the interpolation language model is determined based on a target language model and an ordinary language model corresponding to the text, and
in accordance with a determination that the first confidence satisfies a predetermined confidence threshold, determining the first utterance as the utterance; and
determining a target utterance associated with the utterance in text and differentiating a display of the target utterance in the text.
19 . The non-transitory storage medium of claim 18 , the operations further comprising:
determining an actual speech duration corresponding to the target utterance;
adjusting a predicted speech duration corresponding to an unmatched utterance based on the actual speech duration and the unmatched utterance in the text; and
displaying the predicted speech duration as a prompt to a user.
20 . The non-transitory storage medium of claim 18 , the operations further comprising:
performing emotion marking on respective utterance in the text; and
displaying an emotion marking identification to cause a user to read the respective utterance in text based on the emotional marking identification.