Translation with audio spatialization
A method or an audio system for translation with audio spatialization. The audio system transcribes a first voice signal into first text in a first language. The first text is translated into second text in a second language. The audio system generates a second voice signal that corresponds to the second text in the second language. The first voice signal and the second voice signal are spatialized. The audio system presents the spatialized first voice signal and second voice signal to a user at a same time.
1 . A computer-implemented method, comprising:
transcribing, via a sensor of a headset frame, a first voice signal into first text in a first language;
translating the first text in the first language into second text in a second language;
generating, via an audio controller of the headset frame, a second voice signal that corresponds to the second text in the second language;
spatializing the first voice signal and the second voice signal, relative to a user wearing the headset frame; and
presenting, via speakers of the headset frame, the spatialized first voice signal and the second voice signal.
2 . The method of claim 1 , wherein the spatialized first voice signal and the second voice signal are presented at a same time.
3 . The method of claim 1 , wherein spatializing the first voice signal and the second voice signal comprises:
spatializing the first voice signal to sound as if a source of the first voice signal were at a location further away from a user than a source of the second voice signal.
4 . The method of claim 1 wherein spatializing the first signal and the second signal includes:
causing the second voice signal to sound louder than the first voice signal.
5 . The method of claim 1 , wherein generating a second voice signal comprises:
determining a frequency band associated with the first voice signal; and
generating the second voice signal based in part on the frequency band associated with the first voice signal.
6 . The method of claim 5 , wherein generating a second voice signal comprises:
selecting a voice from a plurality of voices that has a frequency band that is most close to frequency band of the first voice signal; and
generating the second voice signal based in part on the selected voice.
7 . The method of claim 1 , wherein generating a second voice signal comprises:
determining a speech speed associated with the first voice signal; and
generating the second voice signal based in part on the speech speed associated with the first voice signal, wherein the first voice signal and second voice signal are time synchronized.
8 . The method of claim 1 , wherein the method further comprises:
displaying the first text in the first language and the second text in the second language.
9 . The method of claim 1 , wherein the method further comprises:
receiving a third voice signal;
transcribing the third voice signal into third text in the first language;
translating the third text in the first language into fourth text in a second language;
generating a fourth voice signal that corresponds to the fourth text in the second language; and
presenting at least the second voice signal and the fourth voice signal.
10 . The method of claim 9 , wherein the method further comprises:
spatializing the second voice signal and the fourth voice signal based on locations of sources of the first voice signal and the third voice signal relative to a user.
11 . The method of claim 10 , wherein spatializing the second voice signal and the fourth voice signal comprises:
determining that the source of the first voice signal is closer to or further from the user compared to the source of the third voice signal; and
spatializing the second voice signal and the fourth voice signal based in part on the determination.
12 . The method of claim 10 , wherein the method further comprises:
tracking movement of eyes of a user;
determining whether the user is looking at the source of the first voice signal or the source of the second voice signal;
responsive to determining that the user is looking at the source of the first voice signal, causing the second voice signal to sound louder than the fourth voice signal; and
responsive to determining that the user is looking at the source of the third voice signal, causing the fourth voice signal to sound louder than the second voice signal.
13 . A non-transitory computer-readable medium having instructions encoded thereon that, when executed by a processor of a headset, cause the headset to:
transcribe, via a sensor of a headset frame, a first voice signal into first text in a first language;
translate the first text in the first language into second text in a second language;
generate, via an audio controller of the headset frame, a second voice signal that corresponds to the second text in the second language;
spatialize the first voice signal and the second voice signal, relative to a user wearing the headset frame; and
present, via speakers of the headset frame, the spatialized first voice signal and the second voice signal.
14 . The non-transitory computer-readable medium of claim 13 , wherein the first voice signal and the second voice signal are presented at a same time.
15 . The non-transitory computer-readable medium of claim 13 having additional instructions encoded thereon that, when executed by the processor, cause the headset to:
spatialize the first voice signal to sound as if a source of the first voice signal were at a location further away from a user than a source of the second voice signal.
16 . The non-transitory computer-readable medium of claim 13 having additional instructions encoded thereon that, when executed by the processor, cause the headset to:
cause the second voice signal to sound louder than the first voice signal.
17 . The non-transitory computer-readable medium of claim 13 having additional instructions encoded thereon that, when executed by the processor, cause the headset to:
determine a frequency band associated with the first voice signal; and
generate the second voice signal based in part on the frequency band associated with the first voice signal.
18 . The non-transitory computer-readable medium of claim 16 having additional instructions encoded thereon that, when executed by the processor, cause the processor to:
select a voice from a plurality of voices that has a frequency band that is most close to frequency band of the first voice signal; and
generate the second voice signal based in part on the selected voice.
19 . The non-transitory computer-readable medium of claim 13 having additional instructions encoded thereon that, when executed by the processor, cause the headset to:
receive a third voice signal;
transcribe the third voice signal into third text in the first language;
translate the third text in the first language into fourth text in a second language;
generate a fourth voice signal that corresponds to the fourth text in the second language; and
present at least the second voice signal and the fourth voice signal at a same time.
20 . An audio system comprising:
a transducer array configured to present sound to a user; and
an audio controller configured to:
translate first text in a first language into second text in a second language;
generate a first voice signal that corresponds to the first text in the first language;
generate, via the audio controller of a headset frame, a second voice signal that corresponds to the second text in the second language;
spatialize the first voice signal and the second voice signal, relative to a user wearing the headset frame; and
provide, via speakers of the headset frame, the spatialized first voice signal and the second voice signal to transducer array, causing the transducer array to present the spatialized first voice signal and the second voice signal to the user.