IP Library Granted Patent US 12,470,889
Granted Patent B2
US 12,470,889 · App. 18/391,527 · Granted Nov 11, 2025

Future visual frame prediction for autonomous vehicles using long-range acoustic beamforming and synthetic aperture expansion

Inventors: Felix Heide (Blackburg, VA); Jim Aldon D'Souza (Blacksburg, VA)
Assignee: Torc Robotics, Inc.
H04S7/40B60W60/00B60W60/001G01S5/18H04R1/406H04S7/302B60W2420/40B60W2420/403B60W2420/54B60W2556/35B60W2556/40H04R3/005H04R2201/401H04R2499/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,470,889
App. No.
18/391,527
Granted
Nov 11, 2025
Kind
B2
Abstract

An autonomous vehicle including a microphone array of a plurality of microphones, a visual sensor network configured to receive visual signals, at least one processor, and at least one memory storing instructions is disclosed. The instructions, when executed by the at least one processor, cause the at least one processor to: (i) generate spatial beamforming maps locating a sound source using a beamforming model corresponding to acoustic signals received at the plurality of microphones of the microphone array; (ii) apply a synthetic aperture expansion to the acoustic signals to increase resolution of the spatial beamforming maps; and (iii) generate a future visual frame based at least partially upon temporal information extracted from the spatial beamforming maps and visualization maps generated based on the visual signals received by the visual sensor network.

Claims (33)

1 . An autonomous vehicle, comprising:

a microphone array of a plurality of microphones;

a visual sensor network configured to receive visual signals;

at least one processor; and

at least one memory storing instructions, which, when executed by the at least one processor, cause the at least one processor to:

generate spatial beamforming maps locating a sound source using a beamforming model corresponding to acoustic signals received at the plurality of microphones of the microphone array;

apply a synthetic aperture expansion to the acoustic signals to increase resolution of the spatial beamforming maps; and

generate a future visual frame based at least partially upon temporal information extracted from the spatial beamforming maps and visualization maps generated based on the visual signals received by the visual sensor network.

2 . The autonomous vehicle of claim 1 , wherein the temporal information is extracted from the spatial beamforming maps and visualization maps using a neural network trained to extract the temporal information according to future =( perc + adv ) (O RGB t+k , I RGB T+kT ), wherein, O RGB t+kT =ƒ future (O BF t−n+1, . . . , t,t+kT , I RGB t−n+1, . . . , t,t+kT ) n is integer time steps, and perc and adv are perceptual and adversarial losses, respectively.

3 . The autonomous vehicle of claim 1 , wherein to generate the future visual frame, the instructions further cause the at least one processor to extrapolate a previous RGB frame I RGB t at time t using signals O BF t+kT in which k is current beamforming sample modulo sampling rate, T is sampling time of a microphone of the microphone array, and t+kT is the current time.

4 . The autonomous vehicle of claim 3 , wherein the instructions further cause the at least one processor to generate a plurality of future visual frames by cascading predictions of previously predicted visual frames and the corresponding measured audio inputs.

5 . The autonomous vehicle of claim 3 , wherein the plurality of microphones of the microphone array is arranged in a grid pattern of 32×32.

6 . The autonomous vehicle of claim 1 , wherein a sampling rate of a microphone of the microphone array is different from a sampling rate of a visual sensor of the visual sensor network.

7 . The autonomous vehicle of claim 1 , wherein the future visual frame is a red, green, and blue (RGB) frame.

8 . A computer-implemented method, comprising:

generating spatial beamforming maps locating a sound source using a beamforming model corresponding to acoustic signals received at a plurality of microphones of a microphone array;

applying a synthetic aperture expansion to the acoustic signals to increase resolution of the spatial beamforming maps; and

generating a future visual frame based at least partially upon temporal information extracted from the spatial beamforming maps and visualization maps generated based on visual signals received by a visual sensor network.

9 . The computer-implemented method of claim 8 , further comprising extracting the temporal information from the spatial beamforming maps and visualization maps using a neural network trained to extract the temporal information according to future =( perc + adv ) (O RGB t+kT , I RGB t+kT ), wherein, O RGB t+kT =ƒ future (O BF t−n+1, . . . , t,t+kT , I RGB t−n+1, . . . , t,t+kT , n is integer time steps, and perc and adv are perceptual and adversarial losses, respectively.

10 . The computer-implemented method of claim 8 , wherein the generating the future visual frame comprises extrapolating a previous RGB frame I RGB t at time t using signals O BF t+kT in which k is current beamforming sample modulo sampling rate, T is sampling time of a microphone of the microphone array, and t+kT is the current time.

11 . The computer-implemented method of claim 10 , further comprising generating a plurality of future visual frames by cascading predictions of previously predicted visual frames and the corresponding measured audio inputs.

12 . The computer-implemented method of claim 10 , wherein the plurality of microphones of the microphone array is arranged in a grid pattern of 32×32.

13 . The computer-implemented method of claim 8 , wherein a sampling rate of a microphone of the microphone array is different from a sampling rate of a visual sensor of the visual sensor network.

14 . The computer-implemented method of claim 8 , wherein the future visual frame is a red, green, and blue (RGB) frame.

15 . A non-transitory computer-readable medium (CRM) embodying programmed instructions which, when executed by at least one processor of an autonomous vehicle, cause the at least one processor to perform operations comprising:

generating spatial beamforming maps locating a sound source using a beamforming model corresponding to acoustic signals received at a plurality of microphones of a microphone array;

applying a synthetic aperture expansion to the acoustic signals to increase resolution of the spatial beamforming maps; and

generating a future visual frame based at least partially upon temporal information extracted from the spatial beamforming maps and visualization maps generated based on visual signals received by a visual sensor network.

16 . The non-transitory CRM of claim 15 , wherein the operations further comprising extracting the temporal information from the spatial beamforming maps and visualization maps using a neural network trained to extract the temporal information according to future =( perc + adv ) (O RGB t+kT , I RGB t+kT ) wherein, O RGB t+kT =ƒ future (O BF t−n+1, . . . , t,t+kT , I RGB t−n+, . . . , t,t+kT ) n is integer time steps, and perc and adv are perceptual and adversarial losses, respectively.

17 . The non-transitory CRM of claim 15 , wherein the generating the future visual frame comprises extrapolating a previous RGB frame I RGB t at time t using signals O BF t+kT in which k is current beamforming sample modulo sampling rate, T is sampling time of a microphone of the microphone array, and t+kT is the current time.

18 . The non-transitory CRM of claim 17 , wherein the operations further comprising generating a plurality of future visual frames by cascading predictions of previously predicted visual frames and the corresponding measured audio inputs.

19 . The non-transitory CRM of claim 15 , wherein the plurality of microphones of the microphone array is arranged in a grid pattern of 32×32.

20 . The non-transitory CRM of claim 15 , wherein a sampling rate of a microphone of the microphone array is different from a sampling rate of a visual sensor of the visual sensor network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: HEIDE, FELIX; D'SOUZA, JIM ALDON
To: TORC ROBOTICS, INC.
Reel/Frame 065926/0511 →
Continuity (2)
Provisional Application 63508784 · Jun 16, 2023
Related Publication 20240416953A1 · Dec 19, 2024
References Cited (18)
US 6005610A · Pingali · 1999 [cited by applicant]
US 6914854B1 · Heberley et al. · 2005 [cited by applicant]
US 7030905B2 · Carlbom et al. · 2006 [cited by applicant]
US 8098842B2 · Florencio et al. · 2012 [cited by applicant]
US 9444558B1 · Carbone et al. · 2016 [cited by applicant]
US 9598076B1 · Jain et al. · 2017 [cited by applicant]
US 10045120B2 · Adsumilli et al. · 2018 [cited by applicant]
US 11368652B1 · Johnson et al. · 2022 [cited by applicant]
US 11417041B2 · Li et al. · 2022 [cited by applicant]
US 11553159B1 · Rothschild et al. · 2023 [cited by applicant]
US 11620903B2 · Xu et al. · 2023 [cited by applicant]
US 20070025183A1 · Zimmerman et al. · 2007 [cited by applicant]
US 20180284246A1 · LaChapelle · 2018 [cited by applicant]
US 20190302232A1 · Harrison · 2019 [cited by applicant]
US 20200234579A1 · Silver · 2020 [cited by examiner]
US 20220208205A1 · Macoskey · 2022 [cited by examiner]
US 20220219736A1 · Xu et al. · 2022 [cited by applicant]
International Search Report and Written Opinion dated Sep. 3, 2024 for International Patent Application No. PCT/US24/32569. [cited by applicant]