IP Library › Granted Patent US 12,432,492
Granted Patent B2
US 12,432,492 · App. 18/513,796 · Granted Sep 30, 2025

Device, method and system for detecting objects of interest using soundscaped signatures

Inventors: Chew How Lim (Sungai Petani, MY); Chun Seng Song (Georgetown, MY); Mohd Hizami Abdul Hamid (Seberang Jaya, MY); Cheah Min Wong (Bayan Lepas, MY)
Assignee: MOTOROLA SOLUTIONS, INC.
H04R3/005H04R1/406H04R3/04G06V20/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,432,492
App. No.
18/513,796
Granted
Sep 30, 2025
Kind
B2
Abstract

A device, method and system for detecting objects of interest using soundscaped signatures is provided. A device identifies, using first video or first audio from a first camera, an object of interest (OOI) and a sound type of a sound made by the OOI, and extracts a sound signature of the OOI from the first audio, using the sound type and a first soundscaping prediction model for a first environment of the first camera; generating a soundscaped signature of the OOI which predicts a modification of the sound signature in a second environment of a second camera, the soundscaped signature generated by inputting the sound signature and the sound type into a second soundscaping prediction model for the second environment. The device detects the soundscaped signature of the OOI in second audio from the second camera, and generates a notification that the OOI was detected at the second camera location.

Claims (48)

1. A method comprising:

identifying, via at least one computing device, using first video or first audio from a first camera, an object of interest (OOI) and a sound type of a sound made by the OOI;

extracting, via the at least one computing device, a sound signature of the OOI from the first audio, using the sound type and a first soundscaping prediction model for a first environment of the first camera;

generating, via the at least one computing device, a soundscaped signature of the OOI which predicts a modification of the sound signature in a second environment of a second camera, the soundscaped signature generated by inputting the sound signature and the sound type into a second soundscaping prediction model for the second environment;

detecting, via the at least one computing device, the soundscaped signature of the OOI in second audio from the second camera; and

generating, via the at least one computing device, a notification that the OOI was detected at a location of the second camera.

2. The method of claim 1 , further comprising:

adjusting a configuration of the second camera based on the soundscaped signature of the OOI.

3. The method of claim 1 , further comprising:

adjusting a configuration of the second camera based on the soundscaped signature of the OOI by adjusting sensitivity to frequencies identified in the soundscaped signature.

4. The method of claim 1 , further comprising:

extracting the sound signature of the OOI from the first audio by inputting the first audio and the sound type into the first soundscaping prediction model, the first soundscaping prediction model comprising a first machine learning model trained to extract sound signatures of given types from audio based on first environmental acoustic modifier features present at the first environment of the first camera.

5. The method of claim 1 , wherein the second soundscaping prediction model comprises a second machine learning model trained to output soundscaped signatures of given types from audio based on second environmental acoustic modifier features present at the second environment of the second camera.

6. The method of claim 1 , further comprising:

generating the first soundscaping prediction model and the second soundscaping prediction model based on environmental acoustic modifier features present at respective locations of the first camera and the second camera.

7. The method of claim 1 , further comprising:

generating the first soundscaping prediction model and the second soundscaping prediction model by detecting environmental acoustic modifier features present at respective locations of the first camera and the second camera, the detecting occurring using one or more of: respective video from the first camera and the second camera; microphones and multidirectional speakers at the respective locations; and respective sensors at the respective locations.

8. The method of claim 1 , further comprising:

detecting in the second audio from the second camera, the soundscaped signature of the OOI, by comparing the soundscaped signature with the second audio.

9. The method of claim 1 , further comprising:

generating a score associated with detecting the soundscaped signature of the OOI in the second audio from the second camera; and

generating the notification only when the score is greater than a threshold score.

10. The method of claim 1 , wherein the OOI is absent in second video of the second camera and the soundscaped signature is present in the second audio of the second camera.

11. A device comprising:

a communication interface; and

a controller in communication with a first camera and second camera, the controller configured to:

identify, using first video or first audio from the first camera, an object of interest (OOI) and a sound type of a sound made by the OOI;

extract a sound signature of the OOI from the first audio, using the sound type and a first soundscaping prediction model for a first environment of the first camera;

generating a soundscaped signature of the OOI which predicts a modification of the sound signature in a second environment of the second camera, the soundscaped signature generated by inputting the sound signature and the sound type into a second soundscaping prediction model for the second environment;

detect the soundscaped signature of the OOI in second audio from the second camera; and

generate a notification that the OOI was detected at a location of the second camera.

12. The device of claim 11 , wherein the controller is further configured to:

adjust a configuration of the second camera based on the soundscaped signature of the OOI.

13. The device of claim 11 , wherein the controller is further configured to:

adjust a configuration of the second camera based on the soundscaped signature of the OOI by adjusting sensitivity to frequencies identified in the soundscaped signature.

14. The device of claim 11 , wherein the controller is further configured to:

extract the sound signature of the OOI from the first audio by inputting the first audio and the sound type into the first soundscaping prediction model, the first soundscaping prediction model comprising a first machine learning model trained to extract sound signatures of given types from audio based on first environmental acoustic modifier features present at the first environment of the first camera.

15. The device of claim 11 , wherein the second soundscaping prediction model comprises a second machine learning model trained to output soundscaped signatures of given types from audio based on second environmental acoustic modifier features present at the second environment of the second camera.

16. The device of claim 11 , wherein the controller is further configured to:

generate the first soundscaping prediction model and the second soundscaping prediction model based on environmental acoustic modifier features present at respective locations of the first camera and the second camera.

17. The device of claim 11 , wherein the controller is further configured to:

generate the first soundscaping prediction model and the second soundscaping prediction model by detecting environmental acoustic modifier features present at respective locations of the first camera and the second camera, the detecting occurring using one or more of: respective video from the first camera and the second camera; microphones and multidirectional speakers at the respective locations; and respective sensors at the respective locations.

18. The device of claim 11 , wherein the controller is further configured to:

detect, in the second audio from the second camera, the soundscaped signature of the OOI, by comparing the soundscaped signature with the second audio.

19. The device of claim 11 , wherein the controller is further configured to:

generate a score associated with detecting the soundscaped signature of the OOI in the second audio from the second camera; and

generate the notification only when the score is greater than a threshold score.

20. The device of claim 11 , wherein the OOI is absent in second video of the second camera and the soundscaped signature is present in the second audio of the second camera.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2023
From: LIM, CHEW HOW; SONG, CHUN SENG; HAMID, MOHD HIZAMI ABDUL; WONG, CHEAH MIN
To: MOTOROLA SOLUTIONS, INC.
Reel/Frame 065616/0345 →
Continuity (1)
Related Publication 20250168561A1 · May 22, 2025
References Cited (13)
US 9197976B2 · Park · 2015 [cited by applicant]
US 10685075B2 · Blanco et al. · 2020 [cited by applicant]
US 10866950B2 · Sabripour et al. · 2020 [cited by applicant]
US 11722763B2 · Kee et al. · 2023 [cited by applicant]
US 20230089648A1 · Peleg · 2023 [cited by examiner]
KR 101794733B1 · 2017 [cited by applicant]
WO WO2020089917A1 · 2020 [cited by examiner]
Yue, Ran, et al.. “A visualized soundscape prediction model for design processes in urban parks”, Building Simulation, vol. 16, No. 3, pp. 337-356, published Nov. 17, 2022—https://link.springer.com/article/10.1007/s1227… [cited by applicant]
Wirecutter—https://www.nytimes.com/wirecutter/blog/automatic-room-correction/. [cited by applicant]
https://developer.nvidia.com/vrworks/vrworks-audio—VRWorks—Audio. [cited by applicant]
Savioja, Lauri, “Overview of Geometrical Room Acoustic Modeling Techniques”, The Journal of the Acoustical Society of America/ AIP Publishing, Aug. 10, 2015. [cited by applicant]
Mehra, Ravish, et al. “Wave: Interactive Wave-based Sound Propagation for Virtual Environments”, IEEE Transactions on Visualization and Computer Graphics, vol. 21, No. 4, pp. 434-442, Apr. 2015. [cited by applicant]
Luo, Andrew, et al., “Learning Neural Acoustic Fields”—https://www.andrew.cmu.edu/user/afluo/Neural_Acoustic_Fields/—downloaded Sep. 27, 2023. [cited by applicant]