IP Library Granted Patent US 12696045
Granted Patent B2
US 12696045 · App. 18/701,410 · Granted Jul 28, 2026

Voice analysis driven audio parameter modifications

Inventor: Brian Lloyd Schmidt (Bellevue, WA)
Assignee: MAGIC LEAP, INC.
H04S7/304G06F3/012G06F3/017H04R2201/107H04R2499/15H04S2420/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12696045
App. No.
18/701,410
Granted
Jul 28, 2026
Kind
B2
Abstract

Embodiments of the present disclosure can provide systems and methods for presenting audio signals based on an analysis of a voice of a speaker in an augmented reality or mixed reality environment. Methods according to embodiments of this disclosure can include receiving audio data from a microphone of a first wearable head device, the first wearable head device in communication with a virtual environment, the audio data comprising speech data. In some examples, the methods can include identifying a voice parameter based on the audio data. In some examples, the methods can include determining an acoustic parameter based on the voice parameter. In some examples, the methods can include applying the acoustic parameter to the audio data to generate a spatialized audio signal. In some examples, the methods can include presenting the spatialized audio signal to a second wearable head device in communication with the virtual environment.

Claims (57)

1 . A method comprising:

receiving audio data from a microphone of a first wearable head device, the first wearable head device in communication with a virtual environment, the audio data comprising speech data;

identifying a voice parameter based on the audio data;

determining an acoustic parameter based on the voice parameter, wherein the determining the acoustic parameter comprises determining one or more of a distance attenuation, a frequency dependent distance attenuation, a rolloff curve type, an environment model, an environment send level, a radiation based parameter, a frequency dependent radiation gain, and a maximum distance, and wherein the voice parameter indicates at least one of a normal tone, a lowered tone, a raised tone, a whisper, and a shout;

receiving from one or more image sensors of the first wearable head device, image data indicative of a gesture performed by a user of the first wearable head device;

determining gesture data based on the image data, wherein the gesture comprises at least one selected from a hand gesture, a body language gesture, and an interaction with an object in the environment, and wherein determining the acoustic parameter is further based on the gesture data;

determining an intended audience based on the acoustic parameter, wherein a second wearable head device corresponds to a member of the intended audience;

applying the acoustic parameter to the audio data to generate a spatialized audio signal, and

presenting the spatialized audio signal to the second wearable head device in communication with the virtual environment.

2 . The method of claim 1 , wherein the applying the acoustic parameter includes routing the audio data to one or more participants of the virtual environment.

3 . The method of claim 1 , further comprising receiving, from one or more tracking components of the first wearable head device, orientation data corresponding to an orientation of the first head wearable device.

4 . The method of claim 3 , further comprising determining one or more orientation parameters based on the orientation data.

5 . The method of claim 3 , wherein the determining the acoustic parameter is further based on the orientation data.

6 . The method of claim 1 , further comprising presenting a visual indicator to a user of the first wearable head device, wherein the visual indicator shows one or more participants who can hear the spatialized audio signal.

7 . The method of claim 1 , wherein the determining the acoustic parameter comprises determining one or more of a frequency dependent distance attenuation, a rolloff curve type, an environment model, an environment send level, a radiation based parameter, and a frequency dependent radiation gain.

8 . The method of claim 1 , wherein the voice parameter indicates a whisper.

9 . The method of claim 1 , wherein the voice parameter indicates a raised tone.

10 . The method of claim 1 , wherein the gesture data includes a hand gesture.

11 . The method of claim 1 , wherein the gesture data includes an interaction with the object in the environment.

12 . A method comprising:

receiving, at a second wearable head device, spatialized audio data generated via a first wearable head device, the first wearable head device in communication with a virtual environment and the second wearable head device in communication with the virtual environment;

wherein the generating the spatialized audio data via the first wearable head device comprises;

receiving, at the first wearable head device, audio data from a microphone of the first wearable head device, the audio data comprising speech data;

identifying a voice parameter based on the audio data;

determining an acoustic parameter based on the voice parameter, wherein the determining the acoustic parameter comprises determining one or more of a distance attenuation, a frequency dependent distance attenuation, a rolloff curve type, an environment model, an environment send level, a radiation based parameter, a frequency dependent radiation gain, and a maximum distance, and wherein the voice parameter indicates at least one of a normal tone, a lowered tone, a raised tone, a whisper, and a shout;

receiving from one or more image sensors of the first wearable head device, image data indicative of a gesture performed by a user of the first wearable head device;

determining gesture data based on the image data, wherein the gesture comprises at least one selected from a hand gesture, a body language gesture, and an interaction with an object in the environment, and wherein determining the acoustic parameter is further based on the gesture data;

determining an intended audience based on the acoustic parameter, wherein a second wearable head device corresponds to a member of the intended audience; and

applying the acoustic parameter to the audio data to generate a spatialized audio signal.

13 . The method of claim 12 , wherein the applying the acoustic parameter includes routing the audio data to one or more participants of the virtual environment.

14 . The method of claim 12 , wherein the generating the spatialized audio data further comprises receiving, from one or more tracking components of the first wearable head device, orientation data corresponding to an orientation of the first head wearable device.

15 . The method of claim 14 , wherein the generating the spatialized audio data further comprises determining one or more orientation parameters based on the orientation data.

16 . The method of claim 14 , wherein the determining the acoustic parameter is further based on the orientation data.

17 . The method of any of claim 12 , wherein the generating the spatialized audio data further comprises determining an intended audience based on the voice parameter, and wherein the second wearable head device corresponds to a member of the intended audience.

18 . A system comprising:

a memory; and

one or more processors configured to perform a method comprising:

receiving audio data from a microphone of a first wearable head device, the first wearable head device in communication with a virtual environment, the audio data comprising speech data;

identifying a voice parameter based on the audio data;

determining an acoustic parameter based on the voice parameter, wherein the determining the acoustic parameter comprises determining one or more of a distance attenuation, a frequency dependent distance attenuation, a rolloff curve type, an environment model, an environment send level, a radiation based parameter, a frequency dependent radiation gain, and a maximum distance, and wherein the voice parameter indicates at least one of a normal tone, a lowered tone, a raised tone, a whisper, and a shout;

receiving from one or more image sensors of the first wearable head device, image data indicative of a gesture performed by a user of the first wearable head device;

determining gesture data based on the image data, wherein the gesture comprises at least one selected from a hand gesture, a body language gesture, and an interaction with an object in the environment, and wherein determining the acoustic parameter is further based on the gesture data;

determining an intended audience based on the acoustic parameter, wherein a second wearable head device corresponds to a member of the intended audience;

applying the acoustic parameter to the audio data to generate a spatialized audio signal, and

presenting the spatialized audio signal to the second wearable head device in communication with the virtual environment.

19 . A system comprising:

a memory; and

one or more processors configured to perform a method comprising:

receiving, at a second wearable head device, spatialized audio data generated via a first wearable head device, the first wearable head device in communication with a virtual environment and the second wearable head device in communication with the virtual environment;

wherein the generating the spatialized audio data via the first wearable head device comprises:

receiving, at the first wearable head device, audio data from a microphone of the first wearable head device, the audio data comprising speech data;

identifying a voice parameter based on the audio data;

determining an acoustic parameter based on the voice parameter, wherein the determining the acoustic parameter comprises determining one or more of a distance attenuation, a frequency dependent distance attenuation, a rolloff curve type, an environment model, an environment send level, a radiation based parameter, a frequency dependent radiation gain, and a maximum distance, and wherein the voice parameter indicates at least one of a normal tone, a lowered tone, a raised tone, a whisper, and a shout;

receiving from one or more image sensors of the first wearable head device, image data indicative of a gesture performed by a user of the first wearable head device;

determining gesture data based on the image data, wherein the gesture comprises at least one selected from a hand gesture, a body language gesture, and an interaction with an object in the environment, and wherein determining the acoustic parameter is further based on the gesture data;

determining an intended audience based on the acoustic parameter, wherein a second wearable head device corresponds to a member of the intended audience; and

applying the acoustic parameter to the audio data to generate a spatialized audio signal.