Information processing method, non-transitory recording medium, information processing apparatus, and information processing system
An information processing method includes obtaining speech data based on a distance between a sound collection device and a speaker, obtaining text data input in a service for exchanging messages, and outputting first learning data that is based on the speech data and second learning data that includes the text data.
1 . An information processing method, comprising:
generating first learning data from first audio data acquired based on a distance between a microphone and a speaker;
generating second learning data using text data that is independent from the first audio data;
storing third learning data corresponding to a speech recognition usage environment by pairing second audio data that is different from the first audio data with annotated text data corresponding to the second audio data;
performing pre-training of a speech recognition model using the first learning data and the second learning data to generate a first speech recognition model, the pre-training being performed independently of the speech recognition usage environment;
performing training by modifying parameters of the first speech recognition model based on the third learning data corresponding to the speech recognition usage environment to generate a second speech recognition model; and
performing speech recognition on sampled audio data from the speech recognition usage environment using the second speech recognition model and generating recognized text data.
2 . The information processing method according to claim 1 , further comprising:
obtaining the first audio data using a first microphone which is less than a predetermined distance from a mouth of the speaker;
obtaining the text data from a received text message; and
obtaining the second audio data using a second microphone which is greater than a predetermined distance from the mouth of the speaker.
3 . The information processing method of claim 2 , wherein:
the obtaining the first audio data utilizes a system that implements a remote conference via the Internet.
4 . The information processing method of claim 3 , wherein:
the obtaining the second audio data utilizes a system that implements a face-to-face conference.
5 . A non-transitory recording medium storing a plurality of instructions which, when executed by one or more processors, causes the processors to perform a method, the method comprising:
generating first learning data from first audio data acquired based on a distance between a microphone and a speaker;
generating second learning data using text data that is independent of the first audio data;
storing third learning data corresponding to a speech recognition usage environment by pairing second audio data that is different from the first audio data with annotated text data corresponding to the second audio data;
performing pre-training of a speech recognition model using the first learning data and the second learning data to generate a first speech recognition model, the pre-training being performed independently of the speech recognition usage environment;
performing by modifying parameters of the first speech recognition model based on the third learning data corresponding to the speech recognition usage environment to generate a second speech recognition model; and
performing speech recognition on sampled audio data from the speech recognition usage environment using the second speech recognition model and generating recognized text data.
6 . The non-transitory recording medium according to claim 5 , wherein the method further comprises:
obtaining the first audio data using a first microphone which is less than a predetermined distance from a mouth of the speaker;
obtaining the text data from a received text message; and
obtaining the second audio data using a second microphone which is greater than a predetermined distance from the mouth of the speaker.
7 . The non-transitory recording medium according to claim 6 , wherein:
the obtaining the first audio data utilizes a system that implements a remote conference via the Internet.
8 . The non-transitory recording medium according to claim 7 , wherein:
the obtaining the second audio data utilizes a system that implements a face-to-face conference.
9 . An information processing system, comprising:
circuitry configured to:
generate first learning data from first audio data acquired based on a distance between a microphone and a speaker;
generate second learning data using text data that is independent of the first audio data;
store third learning data corresponding to a speech recognition usage environment by pairing second audio data that is different from the first audio data with annotated text data corresponding to the second audio data;
perform pre-training of a speech recognition model using the first learning data and the second learning data to generate a first speech recognition model, the pre-training being performed independently of the speech recognition usage environment;
perform training by modifying parameters of the first speech recognition model based on the third learning data corresponding to the speech recognition usage environment to generate a second speech recognition model; and
perform speech recognition on sampled audio data from the speech recognition usage environment using the second speech recognition model and generate recognized text data.
10 . The information processing system according to claim 9 , wherein the circuitry is further configured to:
obtain the first audio data using a first microphone which is less than a predetermined distance from a mouth of the speaker;
obtain the text data from a received text message; and
obtain the second audio data using a second microphone which is greater than a predetermined distance from the mouth of the speaker.
11 . The information processing system according to claim 10 , wherein:
the obtaining the first audio data utilizes a remote conference system using the Internet.
12 . The information processing system according to claim 11 , wherein:
the obtaining the second audio data utilizes a system that implements a face-to-face conference.