IP Library Granted Patent US 12,293,773
Granted Patent B2
US 12,293,773 · App. 18/052,507 · Granted May 6, 2025

Automatically selecting a sound recognition model for an environment based on audio data and image data associated with the environment

Inventors: Luca Bondi (Pittsburgh, PA); Irtsam Ghazi (Pittsburgh, PA)
Assignee: Robert Bosch GmbH
G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,293,773
App. No.
18/052,507
Granted
May 6, 2025
Kind
B2
Abstract

A system for automatically selecting a sound recognition model for an environment based on audio data and image data associated with the environment. The system includes a camera, a microphone, a memory including a plurality of sound recognition models, and an electronic processor. The electronic processor is configured to receive the audio data associated with the environment from the microphone, receive the image data associated with the environment from the camera, and determine one or more characteristics of the environment based on the audio data and the image data. The electronic processor is also configured to select the sound recognition model from the plurality of sound recognition models based on the one or more characteristics of the environment, receive additional audio data associated with the environment from the microphone, and analyze the additional audio data using the sound recognition model to perform a sound recognition task.

Claims (48)

1. A system for automatically selecting a sound recognition model for an environment based on audio data and image data associated with the environment, the system comprising:

a camera;

a microphone;

a memory including a plurality of sound recognition models; and

an electronic processor configured to

receive, via an input device, a selection of one or more sound recognition tasks;

receive the audio data associated with the environment from the microphone;

receive the image data associated with the environment from the camera;

determine one or more characteristics of the environment based on the audio data and the image data;

for each of the one or more sound recognition tasks selected, select the sound recognition model from the plurality of sound recognition models based on the one or more characteristics of the environment and the selected sound recognition task;

receive additional audio data associated with the environment from the microphone; and

analyze the additional audio data using the sound recognition model to perform a sound recognition task, wherein the sound recognition task includes generating a prediction regarding the additional audio data.

2. The system according to claim 1 , wherein the one or more characteristics include one selected from a group consisting of a depth map of the environment, an acoustic impulse response associated with the environment, a reverberation time associated with the environment, an acoustic property of a surface included in the environment, an acoustic absorption coefficient of a surface included in the environment, a signal-to-noise ratio associated with the environment, a direct-to-reverberant ratio associated with the environment, a clarity index associated with the environment, a dimensional measurement of the environments, and an acoustic scene.

3. The system according to claim 1 , wherein the electronic processor is configured to determine one or more characteristics of the environment based on the audio data and the image data using one or more deep learning models.

4. The system according to claim 1 , wherein each of the plurality of sound recognition models associated with an environment and is trained to perform a sound recognition task.

5. The system according to claim 1 , wherein the electronic processor is configured to select a sound recognition model from the plurality of sound recognition models based on the one or more characteristics of the environment using a heuristic model.

6. The system according to claim 1 , wherein the system includes a server and the electronic processor is further configured to

send the one or more characteristics of the environment to the server to be used by the server to retrain one or more sound recognition models the plurality of sound recognition models.

7. The system according to claim 6 , wherein the electronic processor is further configured to

receive the one or more retrained sound recognition models from the server.

8. The system according to claim 1 , wherein the system includes an electronic device installed in the environment and the camera, the microphone, the memory including the plurality of trained sound recognition models, and the electronic processor are included in the electronic device.

9. A method for automatically selecting a sound recognition model for an environment based on audio data and image data associated with the environment, the method comprising:

receiving, via an input device, a selection of one or more sound recognition tasks;

receiving the audio data associated with the environment from a microphone;

receiving the image data associated with the environment from a camera;

determining one or more characteristics of the environment based on the audio data and the image data;

for each of the one or more sound recognition tasks selected, selecting the sound recognition model from a plurality of sound recognition models based on the one or more characteristics of the environment and the selected sound recognition task;

receiving additional audio data associated with the environment from the microphone; and

analyzing the additional audio data using the sound recognition model to perform a sound recognition task, wherein the sound recognition task includes generating a prediction regarding the additional audio data.

10. The method according to claim 9 , wherein the one or more characteristics include one selected from a group consisting of a depth map of the environment, an acoustic impulse response associated with the environment, a reverberation time associated with the environment, an acoustic property of a surface included in the environment, an acoustic absorption coefficient of a surface included in the environment, a signal-to-noise ratio associated with the environment, a direct-to-reverberant ratio associated with the environment, a clarity index associated with the environment, a dimensional measurement of the environment, and an acoustic scene.

11. The method according to claim 9 , wherein determining one or more characteristics of the environment based on the audio data and the image data includes using one or more deep learning models to determine the one or more characteristics of the environment based on the audio data and the image data.

12. The method according to claim 9 , wherein each of the plurality of sound recognition models associated with an environment and is trained to perform a sound recognition task.

13. The method according to claim 9 , wherein selecting a sound recognition model from the plurality of sound recognition models based on the one or more characteristics of the environment includes using a heuristic model to selecting the trained sound recognition model from the plurality of sound recognition models based on the one or more characteristics of the environment.

14. The method according to claim 9 , the method further comprising

sending the one or more characteristics of the environment to a server to be used by the server to retrain one or more sound recognition models of the plurality of sound recognition models.

15. The method according to claim 14 , the method further comprising

receiving the one or more retrained sound recognition models from the server.

16. A system for automatically selecting a sound recognition model for an environment based on image data associated with the environment, the system comprising:

a camera;

a microphone;

a memory including a plurality of sound recognition models; and

an electronic processor configured to

receive, via an input device, a selection of one or more sound recognition tasks;

receive the image data associated with the environment from the camera;

determine one or more characteristics of the environment based on the image data;

for each of the one or more sound recognition tasks selected, select the sound recognition model from the plurality of sound recognition models based on the one or more characteristics of the environment and the selected sound recognition task;

receive audio data associated with the environment from the microphone; and

analyze the audio data using the sound recognition model to perform a sound recognition task, wherein a sound recognition task includes generating a prediction regarding the audio data.

Assignments (3)
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL 73363, FRAME 0729 Recorded May 4, 2026
From: GLAS TRUST CORPORATION LIMITED
To: ELECTRO-VOICE DYNACORD LLC (F/K/A BOSCH SECURITY SYSTEMS, LLC, F/K/A BOSCH SECURITY SYSTEMS, INC.)
Reel/Frame 075499/0896 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Oct 28, 2025
From: BOSCH SECURITY SYSTEMS, LLC (FKA BOSCH SECURITY SYSTEMS, INC.)
To: GLAS TRUST CORPORATION LIMITED
Reel/Frame 073363/0729 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: BONDI, LUCA; GHAZI, IRTSAM
To: BOSCH SECURITY SYSTEMS, INC.; ROBERT BOSCH GMBH
Reel/Frame 061651/0229 →
Continuity (1)
Related Publication 20240153524A1 · May 9, 2024
References Cited (26)
US 10248744B2 · Schissler et al. · 2019 [cited by applicant]
US 10628714B2 · Pradeep et al. · 2020 [cited by applicant]
US 20150248884A1 · Ljolje · 2015 [cited by examiner]
US 20160109284A1 · Hammershøi et al. · 2016 [cited by applicant]
US 20180285767A1 · Chew · 2018 [cited by examiner]
US 20200013427A1 · Boulanger · 2020 [cited by examiner]
US 20200302951A1 · Deng et al. · 2020 [cited by applicant]
US 20210240431A1 · Gorzel et al. · 2021 [cited by applicant]
US 20210312943A1 · Kulkarni · 2021 [cited by examiner]
US 20220093089A1 · Chen · 2022 [cited by examiner]
US 20220101623A1 · Walsh et al. · 2022 [cited by applicant]
US 20220164662A1 · Saki et al. · 2022 [cited by applicant]
US 20220343917A1 · Tang et al. · 2022 [cited by applicant]
US 20230336694A1 · Wexler · 2023 [cited by examiner]
US 20240112661A1 · Blewett · 2024 [cited by examiner]
JP 2007264328A · 2007 [cited by examiner]
Niklaus et al., “3D Ken Burns Effect from a Single Image,” ACM Transactions on Graphics, 2019, 38(6), 15 pages. [cited by applicant]
Luo et al., “Consistent Video Depth Estimation,” ACM Transactions on Graphics, 2020, 39(4), 13 pages. [cited by applicant]
Singh et al., “Image2Reverb: Cross-Modal Reverb Impulse Response Synthesis,” Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, 22 pages. [cited by applicant]
Gamper et al., “Blind Reverberation Time Estimation Using a Convolutional Neural Network,” 16th International Workshop on Acoustic Signal Enhancement, 2018, 5 pages. [cited by applicant]
Richard et al., “Deep Impulse Responses: Estimating and Parameterizing Filters with Deep Networks,” IEEE International Conference on Acoustics, Speech and Signal Processing, 2022, 5 pages. [cited by applicant]
Zhang et al., “SAP-Net: Deep Learning to Predict Sound Absorption Performance of Metaporous Materials,” Materials & Design, 2021, 212, 7 pages. [cited by applicant]
Hu et al., “A Two-Stage Approach to Device-Robust Acoustic Scene Classification,” IEEE International Conference on Acoustics, Speech and Signal Processing, 2021, 5 pages. [cited by applicant]
Zhang et al., “Acoustic Scene Classification Using Deep CNN with Fine-Resolution Feature,” Expert Systems with Applications, 2020, 143, 9 pages. [cited by applicant]
Papa et al., “A Data-Driven Approach for Acoustic Parameter Similarity Estimation of Speech Recording,” IEEE International Conference on Acoustics, Speech and Signal Processing, 2022, 5 pages. [cited by applicant]
International Search Report for Application No. PCT/EP2023/079746 dated Jan. 29, 2024 (5 pages). [cited by applicant]