IP Library › Granted Patent US 12,713,200
Granted Patent B2
US 12,713,200 · App. 18/391,537 · Granted Aug 18, 2026

Object tracking for autonomous vehicles using long-range acoustic beamforming combined with RGB visual data

Inventors: Felix Heide (Blacksburg, VA); Jim Aldon D'Souza (Blacksburg, VA)
Assignee: TORC Robotics, Inc.
H04S7/40B60W60/00B60W60/001G01S5/18H04R1/406H04S7/302B60W2420/40B60W2420/403B60W2420/54B60W2556/35B60W2556/40H04R3/005H04R2201/401H04R2499/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,713,200
App. No.
18/391,537
Filed
Dec 20, 2023
Granted
Aug 18, 2026
Kind
B2
Art Unit
2695
USPC
381/302
Abstract

An autonomous vehicle including a network of sensors including a plurality of acoustic sensors and a plurality of visual sensors, at least one processor, and at least one memory storing instructions is disclosed. The instructions, when executed by the at least one processor, cause the at least one processor to: (i) generate spatial beamforming maps locating a sound source based upon acoustic signals received at the plurality of acoustic sensors; (ii) identify a type of an object generating the acoustic signals received at the plurality of acoustic sensors based upon comparison of the acoustic signals with a plurality of acoustic signals and respective objects stored in a dataset; and (iii) generate feature maps for an application in an autonomous vehicle driving by enhancing visualization maps generated based upon visual signals received by the plurality of visual sensors.

Claims (30)

1 . An autonomous vehicle, comprising:

a network of sensors including a plurality of acoustic sensors and a plurality of visual sensors;

at least one processor; and

at least one memory storing instructions, which, when executed by the at least one processor, cause the at least one processor to:

generate spatial beamforming maps locating a sound source based upon acoustic signals received at the plurality of acoustic sensors;

identify a type of an object generating the acoustic signals received at the plurality of acoustic sensors based upon comparison of the acoustic signals with a plurality of acoustic signals and respective objects stored in a dataset, wherein the dataset is a multimodal long-range beamforming dataset including audio visual data corresponding to a plurality of objects; and

generate feature maps for an application in an autonomous vehicle driving by enhancing visualization maps generated based upon visual signals received by the plurality of visual sensors.

2 . The autonomous vehicle of claim 1 , wherein the application in the autonomous vehicle driving is an object detection application or a future RGB frame detection application for planning and behavior control of the autonomous vehicle.

3 . The autonomous vehicle of claim 1 , wherein the plurality of visual sensors includes one or more RGB cameras, one or more serial cameras, or one or more lidar sensors.

4 . The autonomous vehicle of claim 1 , wherein the plurality of acoustic sensors is arranged in a grid pattern.

5 . The autonomous vehicle of claim 4 , wherein a grid spacing between the plurality of acoustic sensors in the grid pattern is selected based on upper frequency bounds of the acoustic signals.

6 . The autonomous vehicle of claim 1 , wherein the audio visual data is synchronized with a global navigation satellite system as a time reference.

7 . A computer-implemented method, comprising:

generating spatial beamforming maps locating a sound source based upon acoustic signals received at a plurality of acoustic sensors of a network of sensors;

identifying a type of an object generating the acoustic signals received at the plurality of acoustic sensors based upon comparison of the acoustic signals with a plurality of acoustic signals and respective objects stored in a dataset, wherein the dataset is a multimodal long-range beamforming dataset including audio visual data corresponding to a plurality of objects; and

generating feature maps for an application in an autonomous vehicle driving by enhancing visualization maps generated based upon visual signals received by a plurality of visual sensors of the network of sensors.

8 . The computer-implemented method of claim 7 , wherein the application in the autonomous vehicle driving is an object detection application or a future RGB frame detection application for planning and behavior control of the autonomous vehicle.

9 . The computer-implemented method of claim 7 , wherein the plurality of visual sensors includes one or more RGB cameras, one or more serial cameras, or one or more lidar sensors.

10 . The computer-implemented method of claim 7 , wherein the plurality of acoustic sensors is arranged in a grid pattern.

11 . The computer-implemented method of claim 10 , wherein a grid spacing between the plurality of acoustic sensors in the grid pattern is selected based on upper frequency bounds of the acoustic signals.

12 . The computer-implemented method of claim 7 , wherein the audio visual data is synchronized with a global navigation satellite system as a time reference.

13 . A non-transitory computer-readable medium (CRM) embodying programmed instructions which, when executed by at least one processor of an autonomous vehicle, cause the at least one processor to perform operations comprising:

generating spatial beamforming maps locating a sound source based upon acoustic signals received at a plurality of acoustic sensors of a network of sensors;

identifying a type of an object generating the acoustic signals received at the plurality of acoustic sensors based upon comparison of the acoustic signals with a plurality of acoustic signals and respective objects stored in a dataset, wherein the dataset is a multimodal long-range beamforming dataset including audio visual data corresponding to a plurality of objects; and

generating feature maps for an application in an autonomous vehicle driving by enhancing visualization maps generated based upon visual signals received by a plurality of visual sensors of the network of sensors.

14 . The non-transitory CRM of claim 13 , wherein the application in the autonomous vehicle driving is an object detection application or a future RGB frame detection application for planning and behavior control of the autonomous vehicle.

15 . The non-transitory CRM of claim 13 , wherein the plurality of visual sensors includes one or more RGB cameras, one or more serial cameras, or one or more lidar sensors.

16 . The non-transitory CRM of claim 13 , wherein the plurality of acoustic sensors is arranged in a grid pattern.

17 . The non-transitory CRM of claim 16 , wherein a grid spacing between the plurality of acoustic sensors in the grid pattern is selected based on upper frequency bounds of the acoustic signals.

18 . The non-transitory CRM of claim 13 , wherein the audio visual data is synchronized with a global navigation satellite system as a time reference.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: HEIDE, FELIX; D'SOUZA, JIM ALDON
To: TORC ROBOTICS, INC.
Reel/Frame 065926/0617 →
Continuity (2)
Provisional Application 63508784 · Jun 16, 2023
Related Publication 20240422502A1 · Dec 19, 2024
References Cited (30)
US 6005610A · Pingali · 1999 [cited by applicant]
US 6614386B1 · Moore et al. · 2003 [cited by applicant]
US 6914854B1 · Heberley et al. · 2005 [cited by applicant]
US 7030905B2 · Carlbom et al. · 2006 [cited by applicant]
US 8098842B2 · Florencio et al. · 2012 [cited by applicant]
US 9444558B1 · Carbone et al. · 2016 [cited by applicant]
US 9598076B1 · Jain et al. · 2017 [cited by applicant]
US 10045120B2 · Adsumilli et al. · 2018 [cited by applicant]
US 11368652B1 · Johnson et al. · 2022 [cited by applicant]
US 11417041B2 · Li et al. · 2022 [cited by applicant]
US 11553159B1 · Rothschild et al. · 2023 [cited by applicant]
US 11601749B1 · Graham et al. · 2023 [cited by applicant]
US 11620903B2 · Xu et al. · 2023 [cited by applicant]
US 20070025183A1 · Zimmerman et al. · 2007 [cited by applicant]
US 20110194700A1 · Hetherington · 2011 [cited by examiner]
US 20160059418A1 · Nakamura et al. · 2016 [cited by applicant]
US 20160065323A1 · Zemp · 2016 [cited by applicant]
US 20160277863A1 · Cahill · 2016 [cited by examiner]
US 20180284246A1 · LaChapelle · 2018 [cited by applicant]
US 20190302232A1 · Harrison · 2019 [cited by examiner]
US 20200234579A1 · Silver et al. · 2020 [cited by applicant]
US 20220157165A1 · Dantrey et al. · 2022 [cited by applicant]
US 20220208205A1 · Macoskey et al. · 2022 [cited by applicant]
US 20220214423A1 · Markish et al. · 2022 [cited by applicant]
US 20220219736A1 · Xu et al. · 2022 [cited by applicant]
CN 102207548A · 2011 [cited by applicant]
CN 106950569A · 2017 [cited by applicant]
International Search Report and Written Opinion dated Sep. 3, 2024 for International Patent Application No. PCT/US24/32569. [cited by applicant]
Watanabe et al., “An Ultrasonic Robot Eye Using Neural Networks”, 1991, Acoustical Imaging, vol. 18, pp. 83-95 (Year: 1991). [cited by applicant]
Chakrabarty et al., “Multi-Speaker DOA Estimation Using Deep Convolutional Networks Trained with Noise Signals”, 2019, IEEE Journal of Selected Topics in Signal Processing, vol. 13 (Year: 2019). [cited by applicant]