IP Library Granted Patent US 12711993
Granted Patent B1
US 12711993 · App. 17/814,740 · Granted Aug 18, 2026

Generation of audio-visual filters for virtual spaces

Inventors: Lauren E. Buchanan (Quincy, MA); Matthew Eliot Neutra (Sherborn, MA); Jocelyn Scheirer (Newton, MA); James Ryan Mattingly Strong (Brattleboro, VT)
Assignee: ImmerSphere, Inc.
G11B27/02G06F3/165G06V10/40G06V20/50G10K15/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711993
App. No.
17/814,740
Granted
Aug 18, 2026
Kind
B1
Abstract

A method for creating an audio-video performance as though performed at a given location is described. The method includes receiving location data. The location data is analyzed to identify acoustic features and visual features. An audio adjustment profile and/or a video adjustment profile is generated based on the visual features. The method also includes receiving a sound file and a video file. The audio adjustment profile is applied to the sound file to create a sound performance and the video adjustment profile is applied to the video file to create a video performance. The sound performance and the video performance are combined to create an audio-video performance so that it appears as though the performance were performed at the location. The method includes playing the audio-video performance.

Claims (42)

1 . A method, comprising:

receiving visual data depicting a physical environment of a location, without receiving audio captured from the location;

analyzing, using a trained computer-based system, the visual data to identify environmental structural features of the location that correspond to acoustic reflection and absorption behavior of the physical environment;

generating, based on the identified environmental structural features, a simulated acoustic response profile of the location that includes at least one of reverberation decay time, reflection delay, or impulse response characteristics of the location; and

applying the simulated acoustic response profile to an audio file that is not recorded at the location, such that playback of the audio file is modified to audibly simulate performance within the physical environment of the location.

2 . The method of claim 1 , wherein generating the simulated acoustic response profile of the location includes reverberation decay time, reflection delay, and impulse response characteristics of the location.

3 . The method of claim 1 , further comprising:

analyzing the visual data to identify visual features; and

generating a video adjustment profile based on the visual features.

4 . The method of claim 3 , further comprising:

receiving a video file;

applying the video adjustment profile to the video file to create a video performance, wherein video from the video performance is adjusted to appear as though performed at the location; and

playing the video performance.

5 . The method of claim 1 , wherein analyzing the visual data comprises determining values for at least one of the following location description tags:

inside/outsideness, vastness, reflectiveness, and absorbness.

6 . The method of claim 1 , wherein analyzing the visual data comprises determining values for at least one of the following sound adjustment filters: warbliness, tonal distortion, delay, echoic overlay, and reverb.

7 . The method of claim 1 , wherein analyzing the visual data comprises determining at least one characteristic for the location.

8 . The method of claim 7 , wherein the at least one characteristic comprises one of:

indoors, outdoors, underwater and outer space.

9 . The method of claim 1 , wherein the visual data is a 360-degree visual file.

10 . A method for using a trained network to generate an audio performance in a volume of space based on visual data, the method comprising the steps of:

training, by a computer, a network based on input data and a selected training algorithm to generate a trained network, wherein the input data comprises audio data and visual data of events occurring in a plurality of spaces;

acquiring visual data depicting a physical environment for of the volume of space;

analyzing, using the trained network, the visual data to identify environmental structural features of the volume of space that correspond to acoustic reflection and absorption behavior of the physical environment;

generating, based on the identified environmental structural features, a simulated acoustic response profile of the volume of space that includes at least one of reverberation decay time, reflection delay, or impulse response characteristics of the volume of space; and

applying the simulated acoustic response profile to an audio file that is not recorded at the volume of space, such that playback of the audio file is modified to audibly simulate performance within the physical environment of the volume of space.

11 . The method of claim 10 , wherein the step of analyzing the visual data comprises determining values for at least one of the following location description tags: inside/outsideness, vastness, reflectiveness, and absorbness.

12 . The method of claim 10 , wherein the step of analyzing the visual data comprises determining values for at least one of the following sound adjustment filters: warbliness, tonal distortion, delay, echoic overlay, and reverb.

13 . The method of claim 10 , wherein the visual data is one of a 360-degree visual file and a two-dimensional visual file.

14 . The method of claim 10 , wherein analyzing the visual data includes determining at least one characteristic of the volume of space, the at least one characteristic comprising comprises one of:

indoors, outdoors, underwater and outer space.

15 . A system for providing a virtual experience of a performance in a volume of space, the system comprising:

a processing unit comprising a network trained on historical visual data of a plurality of volumes of space, each of the plurality of volumes of space having features that affect acoustic properties of each of the plurality of volumes of space, wherein the trained network is configured to:

receive visual data depicting a physical environment for the volume of space, without receiving audio captured from the volume of space;

extrapolate from the visual data, using the trained network, environmental structural features of the volume of space that correspond to acoustic reflection and absorption behavior of the physical environment;

generate, based on the extrapolated environmental structural features, a simulated acoustic response profile of the volume of space that includes at least one of reverberation decay time, reflection delay, or impulse response characteristics of the volume of space; and

apply the simulated acoustic response profile to a file comprising audio data such that playback of audio from the file is modified to audibly simulate performance within the physical environment of the volume of space.

16 . The system of claim 15 , wherein the trained neural network is configured to analyze the visual data and determine values for at least one of the following location description tags: inside/outsideness, vastness, reflectiveness, and absorbness.

17 . The method of claim 15 , wherein the trained network is configured to analyze the visual data and determine values for at least one of the following sound adjustment filters: warbliness, tonal distortion, delay, echoic overlay, and reverb.

18 . The method of claim 15 , wherein the visual data is one of a 360-degree visual file and a two-dimensional visual file.

19 . The method of claim 15 , wherein the at least one of the extrapolated features of the volume of space comprises one of:

indoors, outdoors, underwater and outer space.