IP Library Granted Patent US 12,464,307
Granted Patent B2
US 12,464,307 · App. 18/132,851 · Granted Nov 4, 2025

Translation with audio spatialization

Inventors: Joy Westland (Seattle, WA); Amanda Barry (Kirkland, WA); Anjali Induchoodan Menon (Redmond, WA); Madeline Huberth (San Carlos, CA); Sangeetha Lalitha Ramachandran (Sammamish, WA)
Assignee: Meta Platforms Technologies, LLC
H04S7/303G06F3/013G06F40/58G10L15/26H04R5/02H04S2400/11H04S2400/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,464,307
App. No.
18/132,851
Granted
Nov 4, 2025
Kind
B2
Abstract

A method or an audio system for translation with audio spatialization. The audio system transcribes a first voice signal into first text in a first language. The first text is translated into second text in a second language. The audio system generates a second voice signal that corresponds to the second text in the second language. The first voice signal and the second voice signal are spatialized. The audio system presents the spatialized first voice signal and second voice signal to a user at a same time.

Claims (69)

1 . A computer-implemented method, comprising:

transcribing, via a sensor of a headset frame, a first voice signal into first text in a first language;

translating the first text in the first language into second text in a second language;

generating, via an audio controller of the headset frame, a second voice signal that corresponds to the second text in the second language;

spatializing the first voice signal and the second voice signal, relative to a user wearing the headset frame; and

presenting, via speakers of the headset frame, the spatialized first voice signal and the second voice signal.

2 . The method of claim 1 , wherein the spatialized first voice signal and the second voice signal are presented at a same time.

3 . The method of claim 1 , wherein spatializing the first voice signal and the second voice signal comprises:

spatializing the first voice signal to sound as if a source of the first voice signal were at a location further away from a user than a source of the second voice signal.

4 . The method of claim 1 wherein spatializing the first signal and the second signal includes:

causing the second voice signal to sound louder than the first voice signal.

5 . The method of claim 1 , wherein generating a second voice signal comprises:

determining a frequency band associated with the first voice signal; and

generating the second voice signal based in part on the frequency band associated with the first voice signal.

6 . The method of claim 5 , wherein generating a second voice signal comprises:

selecting a voice from a plurality of voices that has a frequency band that is most close to frequency band of the first voice signal; and

generating the second voice signal based in part on the selected voice.

7 . The method of claim 1 , wherein generating a second voice signal comprises:

determining a speech speed associated with the first voice signal; and

generating the second voice signal based in part on the speech speed associated with the first voice signal, wherein the first voice signal and second voice signal are time synchronized.

8 . The method of claim 1 , wherein the method further comprises:

displaying the first text in the first language and the second text in the second language.

9 . The method of claim 1 , wherein the method further comprises:

receiving a third voice signal;

transcribing the third voice signal into third text in the first language;

translating the third text in the first language into fourth text in a second language;

generating a fourth voice signal that corresponds to the fourth text in the second language; and

presenting at least the second voice signal and the fourth voice signal.

10 . The method of claim 9 , wherein the method further comprises:

spatializing the second voice signal and the fourth voice signal based on locations of sources of the first voice signal and the third voice signal relative to a user.

11 . The method of claim 10 , wherein spatializing the second voice signal and the fourth voice signal comprises:

determining that the source of the first voice signal is closer to or further from the user compared to the source of the third voice signal; and

spatializing the second voice signal and the fourth voice signal based in part on the determination.

12 . The method of claim 10 , wherein the method further comprises:

tracking movement of eyes of a user;

determining whether the user is looking at the source of the first voice signal or the source of the second voice signal;

responsive to determining that the user is looking at the source of the first voice signal, causing the second voice signal to sound louder than the fourth voice signal; and

responsive to determining that the user is looking at the source of the third voice signal, causing the fourth voice signal to sound louder than the second voice signal.

13 . A non-transitory computer-readable medium having instructions encoded thereon that, when executed by a processor of a headset, cause the headset to:

transcribe, via a sensor of a headset frame, a first voice signal into first text in a first language;

translate the first text in the first language into second text in a second language;

generate, via an audio controller of the headset frame, a second voice signal that corresponds to the second text in the second language;

spatialize the first voice signal and the second voice signal, relative to a user wearing the headset frame; and

present, via speakers of the headset frame, the spatialized first voice signal and the second voice signal.

14 . The non-transitory computer-readable medium of claim 13 , wherein the first voice signal and the second voice signal are presented at a same time.

15 . The non-transitory computer-readable medium of claim 13 having additional instructions encoded thereon that, when executed by the processor, cause the headset to:

spatialize the first voice signal to sound as if a source of the first voice signal were at a location further away from a user than a source of the second voice signal.

16 . The non-transitory computer-readable medium of claim 13 having additional instructions encoded thereon that, when executed by the processor, cause the headset to:

cause the second voice signal to sound louder than the first voice signal.

17 . The non-transitory computer-readable medium of claim 13 having additional instructions encoded thereon that, when executed by the processor, cause the headset to:

determine a frequency band associated with the first voice signal; and

generate the second voice signal based in part on the frequency band associated with the first voice signal.

18 . The non-transitory computer-readable medium of claim 16 having additional instructions encoded thereon that, when executed by the processor, cause the processor to:

select a voice from a plurality of voices that has a frequency band that is most close to frequency band of the first voice signal; and

generate the second voice signal based in part on the selected voice.

19 . The non-transitory computer-readable medium of claim 13 having additional instructions encoded thereon that, when executed by the processor, cause the headset to:

receive a third voice signal;

transcribe the third voice signal into third text in the first language;

translate the third text in the first language into fourth text in a second language;

generate a fourth voice signal that corresponds to the fourth text in the second language; and

present at least the second voice signal and the fourth voice signal at a same time.

20 . An audio system comprising:

a transducer array configured to present sound to a user; and

an audio controller configured to:

translate first text in a first language into second text in a second language;

generate a first voice signal that corresponds to the first text in the first language;

generate, via the audio controller of a headset frame, a second voice signal that corresponds to the second text in the second language;

spatialize the first voice signal and the second voice signal, relative to a user wearing the headset frame; and

provide, via speakers of the headset frame, the spatialized first voice signal and the second voice signal to transducer array, causing the transducer array to present the spatialized first voice signal and the second voice signal to the user.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE FIFTH INVENTOR'S NAME PREVIOUSLY RECORDED AT REEL: 063484 FRAME: 0750. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 8, 2023
From: WESTLAND, JOY; BARRY, AMANDA; MENON, ANJALI INDUCHOODAN; HUBERTH, MADELINE; RAMACHANDRAN, SANGEETHA LALITHA
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 063564/0058 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2023
From: WESTLAND, JOY; BARRY, AMANDA; MENON, ANJALI INDUCHOODAN; HUBERTH, MADELINE; RAMANCHANDRAN, SANGEETHA LALITHA
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 063484/0750 →
Continuity (1)
Related Publication 20240340604A1 · Oct 10, 2024
References Cited (48)
US 5899971A · De Vos · 1999 [cited by examiner]
US 6275789B1 · Moser · 2001 [cited by examiner]
US 6385586B1 · Dietz · 2002 [cited by examiner]
US 9301057B2 · Sprague · 2016 [cited by examiner]
US 9640181B2 · Parkinson · 2017 [cited by examiner]
US 9747282B1 · Baker · 2017 [cited by examiner]
US 10872605B2 · Adachi · 2020 [cited by examiner]
US 11328131B2 · Orlick · 2022 [cited by examiner]
US 11783137B2 · Tang · 2023 [cited by examiner]
US 12087291B2 · Park · 2024 [cited by examiner]
US 20020169592A1 · Aityan · 2002 [cited by examiner]
US 20030125927A1 · Seme · 2003 [cited by examiner]
US 20060285654A1 · Nesvadba et al. · 2006 [cited by applicant]
US 20080187143A1 · Mak-Fan · 2008 [cited by examiner]
US 20090080632A1 · Zhang · 2009 [cited by examiner]
US 20100198579A1 · Cunnington · 2010 [cited by examiner]
US 20100235161A1 · Kim · 2010 [cited by examiner]
US 20140142934A1 · Kim · 2014 [cited by examiner]
US 20160012465A1 · Sharp · 2016 [cited by examiner]
US 20160275076A1 · Ishikawa · 2016 [cited by examiner]
US 20170286407A1 · Chochowski · 2017 [cited by examiner]
US 20180046431A1 · Thagadur Shivappa · 2018 [cited by examiner]
US 20180054400A1 · Akopian · 2018 [cited by examiner]
US 20190347331A1 · Yu · 2019 [cited by examiner]
US 20200134026A1 · Lovitt · 2020 [cited by examiner]
US 20200137006A1 · Shannon · 2020 [cited by examiner]
US 20200142667A1 · Querze · 2020 [cited by examiner]
US 20200293622A1 · Orlick · 2020 [cited by examiner]
US 20210043066A1 · Wright · 2021 [cited by examiner]
US 20210160645A1 · Olivieri · 2021 [cited by examiner]
US 20210234611A1 · Bull · 2021 [cited by examiner]
US 20220022000A1 · Bruhn · 2022 [cited by examiner]
US 20220103963A1 · Satongar · 2022 [cited by examiner]
US 20220210531A1 · Carlson · 2022 [cited by examiner]
US 20220272477A1 · Stein · 2022 [cited by examiner]
US 20230093585A1 · Faundez Hoffmann · 2023 [cited by examiner]
US 20230114834A1 · Chadwick · 2023 [cited by examiner]
US 20230177948A1 · Wright · 2023 [cited by examiner]
US 20230186899A1 · Waibel · 2023 [cited by examiner]
US 20230209300A1 · Udesen · 2023 [cited by examiner]
US 20230300532A1 · Spittle · 2023 [cited by examiner]
US 20230368773A1 · Mishra · 2023 [cited by examiner]
US 20240056758A1 · Kronlachner · 2024 [cited by examiner]
US 20240311076A1 · Balsam · 2024 [cited by examiner]
US 20240340604A1 · Westland · 2024 [cited by examiner]
US 20250094211A1 · Spittle · 2025 [cited by examiner]
CN 113286217A · 2021 [cited by applicant]
European Search Report for European Patent Application No. 24162072.3, dated Jun. 14, 2024, 7 pages. [cited by applicant]