Method and apparatus for improving performance of artificial intelligence model using speech recognition results as text input with delay times
View Patent ↗The present disclosure relates to a method and device for improving the performance of an AI model that uses voice recognition results as text input. A method of training an AI model according to an embodiment of the present disclosure may include: generating first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription; generating second time information by adding a pre-configured delay time to the first time information; generating a modified transcription based on an end time of a last word among the plurality of words and the second time information; and performing training of the AI model based on a second training sample including the voice and the modified transcription.
1 . A method of training an artificial intelligence (AI) model, the method comprising:
generating first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription;
generating second time information by adding a pre-configured delay time to the first time information;
generating a modified transcription based on an end time of a last word among the plurality of words and the second time information; and
performing training of the AI model based on a second training sample including the voice and the modified transcription,
wherein the pre-configured delay time is progressively increased in correspondence with an advancement of a stage of the training, and
wherein an amount of word removal for generating the modified transcription is increased in response to an increase of the pre-configured delay time.
2 . The method of claim 1 , wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.
3 . The method of claim 1 , wherein the first time information includes information on an end time for each word for the plurality of words.
4 . The method of claim 3 , wherein the second time information is generated by adding the pre-configured delay time to the end time for each word for the plurality of words.
5 . The method of claim 1 , wherein the modified transcription is generated by removing one or more words from among the plurality of words whose end time for each word is greater than or equal to the end time based on the second time information.
6 . The method of claim 1 , wherein the pre-configured delay time is set to a value of 0 in an initial section of the training.
7 . The method of claim 1 , wherein the pre-configured delay time is set between a 0 value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.
8 . The method of claim 1 , wherein the pre-configured delay time is set between a minimum delay time value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.
9 . The method of claim 1 , wherein the pre-configured delay time is set to increase up to a delay time of a voice recognizer as a stage of the training advanced.
10 . An apparatus for training an artificial intelligence (AI) model, the apparatus comprising:
a processor and a memory,
wherein the processor is configured to:
generate first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription;
generate second time information by adding a pre-configured delay time to the first time information;
generate a modified transcription based on an end time of a last word among the plurality of words and the second time information; and
perform training of the AI model based on a second training sample including the voice and the modified transcription,
wherein the pre-configured delay time is progressively increased in correspondence with an advancement of a stage of the training, and
wherein an amount of word removal for generating the modified transcription is increased in response to an increase of the pre-configured delay time.
11 . The apparatus of claim 10 , wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.
12 . The apparatus of claim 10 , wherein the first time information includes information on an end time for each word for the plurality of words.
13 . The apparatus of claim 12 , wherein the second time information is generated by adding the pre-configured delay time to the end time for each word for the plurality of words.
14 . The apparatus of claim 10 , wherein the modified transcription is generated by removing one or more words from among the plurality of words whose end time for each word is greater than or equal to the end time based on the second time information.
15 . The apparatus of claim 10 , wherein the pre-configured delay time is set to a value of 0 in an initial section of the training.
16 . The apparatus of claim 10 , wherein the pre-configured delay time is set between a 0 value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.
17 . The apparatus of claim 10 , wherein the pre-configured delay time is set between a minimum delay time value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.
18 . The apparatus of claim 10 , wherein the pre-configured delay time is set to increase up to a delay time of a voice recognizer as a stage of the training advanced.
19 . One or more non-transitory computer readable storage medium storing one or more instructions,
wherein the one or more instructions are executed by one or more processors and control an apparatus for training an artificial intelligence (AI) model to:
generate first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription;
generate second time information by adding a pre-configured delay time to the first time information;
generate a modified transcription based on an end time of a last word among the plurality of words and the second time information; and
perform training of the AI model based on a second training sample including the voice and the modified transcription,
wherein the pre-configured delay time is progressively increased in correspondence with an advancement of a stage of the training, and
wherein an amount of word removal for generating the modified transcription is increased in response to an increase of the pre-configured delay time.
20 . The computer readable storage medium of claim 19 , wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.