IP Library › Granted Patent US 12,737,954
Granted Patent B2
US 12,737,954 · App. 18/658,463 · Granted Sep 15, 2026

Spatial audio and avatar control at headset using audio signals

Inventors: Nadav Grossinger (Hillsborough, CA); Robert Hasbun (Placerville, CA)
Assignee: Meta Platforms Technologies, LLC
G06T13/205G02B27/0172G06N20/00G06T13/40G06T19/006G10L21/10H04R5/033G10L2021/105
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,954
App. No.
18/658,463
Granted
Sep 15, 2026
Kind
B2
Abstract

An audio system in a local area providing an audio signal to a headset of a remote user is presented herein. The audio system identifies sounds from a human sound source in the local area, based in part on sounds detected within the local area. The audio system generates an audio signal for presentation to a remote user within a virtual representation of the local area based in part on a location of the remote user within the virtual representation of the local area relative to a virtual representation of the human sound source within the virtual representation of the local area. The audio system provides the audio signal to a headset of the remote user, wherein the headset presents the audio signal as part of the virtual representation of the local area to the remote user.

Claims (42)

1 . A method comprising:

receiving audio data captured, by a first computing system, from a human sound source;

predicting a facial expression based at least in part on the audio data;

causing a second computing system, remote from the first computing system, to play audio, based on the audio data, in relation to a virtual representation of the human sound source; and

causing the second computing system to provide, in a virtual environment and in synchronization with the audio that is played, an additional virtual representation of the predicted expression in relation to the virtual representation of the human sound source, the additional virtual representation being provided responsive to an occluded portion of a face of the human sound source being within a threshold angle of a field of view of the second computing system.

2 . The method of claim 1 , wherein the predicting the facial expression comprises predicting a lip pose or movement for the representation of the human sound source.

3 . The method of claim 1 , wherein the predicting the facial expression comprises predicting the facial expression by applying a machine learning algorithm to the audio data.

4 . The method of claim 1 , further comprising:

selectively adjusting the audio data in response to one or more user inputs;

wherein the causing the second computing system to play the audio comprises causing the second computing system to play the audio based on the adjusted audio data.

5 . The method of claim 1 , wherein the audio that the second computing system is caused to play is modified based on a comparison between a location determined for the second computing system and a location determined for the representation of the human sound source.

6 . The method of claim 1 , wherein the captured audio data is received in response to:

generation of multiple captured audio data instances, from sound sources collocated with the human sound source; and

identifying one of the multiple captured audio data instances, as being from the human sound source, based on matching between the multiple captured audio data instances and data for the human sound source.

7 . The method of claim 1 :

wherein the captured audio data is associated with a location of the human sound source determined by performing beam-steering processing on the captured audio data; and

wherein the audio that the second computing system is caused to play is modified based on the location of the human sound source.

8 . The method of claim 1 , wherein the causing the second computing system to provide the predicted facial expression of the human sound source in synchronization with the played audio includes providing, to the second computing system via a network, visual information indicating the predicted facial expression with synchronization information for synchronizing the predicted facial expression with playing the audio.

9 . The method of claim 1 , wherein the method is performed by the first computing system.

10 . The method of claim 1 , wherein the method is performed by the second computing system.

11 . The method of claim 1 , wherein the method is performed by an intermediary system facilitating communication between the first computing system and the second computing system.

12 . A computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform a process comprising:

receiving audio data captured, by a first computing system, from a human sound source;

predicting a facial expression;

causing a second computing system, remote from the first computing system, to play audio, based on the audio data, in relation to a virtual representation of the human sound source; and

causing the second computing system to provide, in a virtual environment and in synchronization with the audio that is played, an additional virtual representation of the predicted expression in relation to the virtual representation of the human sound source, the additional virtual representation being provided responsive to an occluded portion of a face of the human sound source being within a threshold angle of a field of view of the second computing system.

13 . The computer-readable storage medium of claim 12 , wherein the predicting the facial expression comprises predicting the facial expression by applying a machine learning algorithm to the audio data.

14 . The computer-readable storage medium of claim 12 , wherein the predicting the facial expression comprises predicting a lip pose or movement for the representation of the human sound source.

15 . The computer-readable storage medium of claim 12 , wherein the audio that the second computing system is caused to play is modified based on a comparison between a location determined for the second computing system and a location determined for the representation of the human sound source.

16 . The computer-readable storage medium of claim 12 , wherein the process is performed by the second computing system.

17 . A computing system comprising:

one or more processors; and

one or more memories storing instructions that, when executed by the one or more processors, cause the computing system to perform a process comprising:

receiving audio data captured, by a first computing system;

predicting a facial expression;

causing a second computing system, remote from the first computing system, to play audio, based on the audio data, in relation to a virtual representation of a human sound source; and

causing the second computing system to provide, in a virtual environment and in synchronization with the audio that is played, an additional virtual representation of the predicted expression in relation to the virtual representation of the human sound source, the additional virtual representation being provided responsive to an occluded portion of a face of the human sound source being within a threshold angle of a field of view of the second computing system.

18 . The computing system of claim 17 :

wherein the captured audio data is associated with a location of the human sound source determined by performing beam-steering processing on the captured audio data; and

wherein the audio that the second computing system is caused to play is modified based on the location of the human sound source.

19 . The computing system of claim 17 , wherein the process is performed by the first computing system.

20 . The computing system of claim 17 , wherein the causing the second computing system to provide the predicted facial expression of the human sound source in conjunction with the played audio includes providing, to the second computing system, visual information indicating the predicted facial expression with synchronization information for synchronizing the predicted facial expression with playing the audio.

Continuity (5)
Continuation 18120808 · Mar 13, 2023
Continuation 17591181 · Feb 2, 2022
Continuation 16869925 · May 8, 2020
Provisional Application 62893052 · Aug 28, 2019
Related Publication 20240290020A1 · Aug 29, 2024
References Cited (70)
US 9942687B1 · Chemistruck et al. · 2018 [cited by applicant]
US 10206055B1 · Mindlin et al. · 2019 [cited by applicant]
US 10217286B1 · Angel et al. · 2019 [cited by applicant]
US 10225656B1 · Kratz et al. · 2019 [cited by applicant]
US 10602298B2 · Raghuvanshi et al. · 2020 [cited by applicant]
US 10674307B1 · Robinson et al. · 2020 [cited by applicant]
US 10755463B1 · Albuz et al. · 2020 [cited by applicant]
US 11113859B1 · Xiao · 2021 [cited by examiner]
US 11122385B2 · Robinson et al. · 2021 [cited by applicant]
US 11276215B1 · Grossinger et al. · 2022 [cited by applicant]
US 11523247B2 · Robinson et al. · 2022 [cited by applicant]
US 11605191B1 · Grossinger et al. · 2023 [cited by applicant]
US 11871198B1 · Robinson et al. · 2024 [cited by applicant]
US 12008700B1 · Grossinger et al. · 2024 [cited by applicant]
US 20080243278A1 · Dalton et al. · 2008 [cited by applicant]
US 20110069841A1 · Angeloff et al. · 2011 [cited by applicant]
US 20120093320A1 · Flaks et al. · 2012 [cited by applicant]
US 20120206452A1 · Geisner et al. · 2012 [cited by applicant]
US 20130041648A1 · Osman · 2013 [cited by applicant]
US 20130236040A1 · Crawford et al. · 2013 [cited by applicant]
US 20150187112A1 · Rozen · 2015 [cited by examiner]
US 20150373477A1 · Norris et al. · 2015 [cited by applicant]
US 20160125876A1 · Schroeter et al. · 2016 [cited by applicant]
US 20170039750A1 · Tong et al. · 2017 [cited by applicant]
US 20170223478A1 · Jot et al. · 2017 [cited by applicant]
US 20170280235A1 · Varerkar · 2017 [cited by examiner]
US 20170316115A1 · Lewis et al. · 2017 [cited by applicant]
US 20170352183A1 · Katz · 2017 [cited by examiner]
US 20170366896A1 · Adsumilli · 2017 [cited by examiner]
US 20180262849A1 · Farmani · 2018 [cited by examiner]
US 20180277133A1 · Deetz · 2018 [cited by examiner]
US 20180302738A1 · Di Censo · 2018 [cited by examiner]
US 20190116448A1 · Schmidt et al. · 2019 [cited by applicant]
US 20190130628A1 · Cao · 2019 [cited by examiner]
US 20190138096A1 · Lee · 2019 [cited by examiner]
US 20200037091A1 · Jeon et al. · 2020 [cited by applicant]
US 20200174734A1 · Gomes et al. · 2020 [cited by applicant]
US 20200265860A1 · Mouncer et al. · 2020 [cited by applicant]
US 20200292817A1 · Jones · 2020 [cited by examiner]
US 20200296521A1 · Wexler et al. · 2020 [cited by applicant]
US 20200314583A1 · Robinson et al. · 2020 [cited by applicant]
US 20210377690A1 · Robinson et al. · 2021 [cited by applicant]
CN 109416585A · 2019 [cited by applicant]
EP 3949447A1 · 2022 [cited by applicant]
JP 2005341092A · 2005 [cited by applicant]
JP 2020521454A · 2020 [cited by applicant]
JP 2020537849A · 2020 [cited by applicant]
WO 2016014254A1 · 2016 [cited by applicant]
WO 2018182274A1 · 2018 [cited by applicant]
WO 2019079523A1 · 2019 [cited by applicant]
WO 2020079485A2 · 2020 [cited by applicant]
WO 2020197839A1 · 2020 [cited by applicant]
Olszewski et al., “High-Fidelity Facial and Speech Animation for VR HMDs”, (Year: 2016). [cited by examiner]
Lee et al., “High-Fidelity Facial and Speech Animation for VR HMDs”, (Year: 2016). [cited by examiner]
Brandenburg K., et al., “Plausible Augmentation of Auditory Scenes Using Dynamic Binaural Synthesis for Personalized Auditory Realities,” Audio Engineering Society, Conference Paper, Aug. 20-22, 2018, 10 pages. [cited by applicant]
Final Office Action mailed May 3, 2022 for U.S. Appl. No. 16/508,648, filed Jul. 11, 2019, 23 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2020/023071, mailed Jul. 3, 2020, 13 Pages. [cited by applicant]
Non Final Office Action mailed Feb. 15, 2022 for U.S. Appl. No. 16/508,648, filed Jul. 11, 2019, 27 pages. [cited by applicant]
Non-Final Office Action mailed Nov. 8, 2022 for U.S. Appl. No. 17/591,181, filed Feb. 2, 2022, 8 pages. [cited by applicant]
Non-Final Office Action mailed Nov. 8, 2023 for U.S. Appl. No. 18/120,808, filed Mar. 13, 2023, 10 pages. [cited by applicant]
Non-Final Office Action mailed Jul. 18, 2022 for U.S. Appl. No. 17/402,012, filed Aug. 13, 2021, 10 pages. [cited by applicant]
Non-Final Office Action mailed Jul. 2, 2021 for U.S. Appl. No. 16/508,648, filed Jul. 11, 2019, 27 pages. [cited by applicant]
Non-Final Office Action mailed Jun. 23, 2020 for U.S. Appl. No. 16/259,990, filed Jan. 28, 2019, 27 pages. [cited by applicant]
Non-Final Office Action mailed Aug. 25, 2022 for U.S. Appl. No. 16/508,648, filed Jul. 11, 2019, 23 pages. [cited by applicant]
Non-Final Office Action mailed Sep. 29, 2021 for U.S. Appl. No. 16/869,925, filed May 8, 2020, 14 pages. [cited by applicant]
Office Action mailed Feb. 6, 2024 for Japanese Patent Application No. 2021-533833, filed on Mar. 17, 2020, 14 pages. [cited by applicant]
Office Action mailed Nov. 28, 2022 for Chinese Patent Application No. 202080022828.0, filed Sep. 18, 2021, 12 pages. [cited by applicant]
Plinge A., et al., “Six-Degrees-of-Freedom Binaural Audio Reproduction of First-Order Ambisonics with Distance Information,” Audio Engineering Society, Conference Paper, Aug. 20-22, 2018, 10 pages. [cited by applicant]
Traer J., et al., “Statistics of Natural Reverberation Enable Perceptual Separation of Sound and Space,” Proceedings of the National Academy of Sciences, Nov. 10, 2016, vol. 113 (48), pp. E7856-E7865. [cited by applicant]
Office Action mailed Jun. 20, 2024 for Korean Application No. 10-2021-7034826, filed Mar. 17, 2020, 5 pages. [cited by applicant]