Mobile terminal capable of processing voice and operation method therefor
A mobile terminal is disclosed. The mobile terminal comprises: a microphone configured to generate a voice signal in response to voices of speakers; a processor configured to generate a separated voice signal associated with each of the voices by separating the voice signal from a sound source on the basis of a sound source location of each of the voices, and output the result of translation for each of the voices, on the basis of the separated voice signal; and a memory configured to store source language information indicating source languages that are uttered languages of the voices of the speakers. The processor outputs the results of translations in which the languages of the voices of the speakers have been translated from the source languages into a target language, on the basis of the source language information and the separated voice signal.
1 . A mobile terminal comprising:
a microphone comprising a plurality of microphones disposed to form an array, wherein the plurality of microphones are configured to generate a plurality of voice signals in response to voices of speakers;
a processor configured to judge respective voice source positions of the voices based on a time delay among the plurality of voice signals generated from the plurality of microphones, wherein the respective voice source positions represent relative positions of the voices with respect to the mobile terminal, generate separated voice signals related to the respective voices by performing voice source separation of the voice signals based on the judged respective voice source positions of the voices, and output translation results for the respective voices based on the separated voice signals; and
a memory configured to store voice source position information representing positions of respective speakers and source language information representing information on source languages that are pronounced languages of the respective speakers, the source language information corresponding to the relative positions of the respective speakers with respect to the mobile terminal,
wherein the processor is configured to output the translation results in which the languages of the voices of the speakers have been translated from the source languages into target languages to be translated based on the source language information and the separated voice signals,
wherein the processor is configured to:
determine the source languages corresponding to the relative positions of the voices based on the voice source position information and the source language information by comparing the judged respective voice source positions of the voices with position information included in the voice source position information to identify the respective speakers and determining the source languages corresponding to the respective speakers included in the source language information,
read the determined source language information corresponding to the relative positions of the voices, and output the translation results for the respective voices in accordance with the determined source languages.
2 . The mobile terminal of claim 1 , further comprising a display configured to visually output the translation results.
3 . The mobile terminal of claim 1 , wherein the processor is configured to: generate voice source position information representing the voice source positions of the respective voices based on the time delay among the plurality of voice signals generated from the plurality of microphones, and match and store, in the memory, the voice source position information for the voices with the separated voice signals for the voices.
4 . The mobile terminal of claim 1 , further comprising a communication device configured to communicate with an external device,
wherein the communication device is configured to transmit the translation results output by the processor to the external device.
5 . An operation method of a mobile terminal capable of processing voices, the operation method comprising:
generating, by a plurality of microphones disposed to form an array, a plurality of voice signals in response to voices of speakers;
judging respective voice source positions of the voices based on a time delay among the plurality of voice signals generated from the plurality of microphones, wherein the respective voice source positions represent relative positions of the voices with respect to the mobile terminal;
performing voice source separation of the voice signals based on the judged respective voice source positions of the voices;
generating separated voice signals related to the respective voices in accordance with the result of the voice source separation; and
outputting translation results for the respective voices based on the separated voice signals,
wherein the outputting of the translation results includes:
storing voice source position information representing positions of respective speakers and source language information representing information on source languages that are pronounced languages of the respective speakers, the source language information corresponding to the relative positions of each of the respective speakers with respect to the mobile terminal; and
outputting the translation results in which the languages of the voices of the speakers have been translated from the source languages into target languages that are languages to be translated based on the source language information and the separated voice signals,
wherein the outputting of the translation results comprises:
determining the source languages corresponding to the relative positions of the voices based on the voice source position information and the source language information by comparing the judged respective voice source positions of the voices with position information included in the voice source position information to identify the respective speakers and determining the source languages corresponding to the respective speakers included in the source language information,
reading the determined source language information corresponding to the relative positions of the voices, and
outputting the translation results for the respective voices in accordance with the determined source languages.
6 . The operation method of claim 5 , further comprising:
generating voice source position information representing the voice source positions of the respective voices based on the time delay among the plurality of voice signals generated from the plurality of microphones; and
matching and storing the voice source position information for the voices with the separated voice signals for the voices.