IP Library Granted Patent US 12,437,539
Granted Patent B2
US 12,437,539 · App. 18/122,956 · Granted Oct 7, 2025

System and method for identifying activity in an area using a video camera and an audio sensor

Inventor: Hisao Chang (Medina, MN)
Assignee: HONEYWELL INTERNATIONAL INC.
G06V20/46G06F16/634G06F16/635G06V10/12G06V20/10G06V20/41G06V20/53G10L25/51G10L25/72G06V20/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,539
App. No.
18/122,956
Granted
Oct 7, 2025
Kind
B2
Abstract

Identifying activity in an area even during periods of poor visibility using a video camera and an audio sensor are disclosed. The video camera is used to identify visible events of interest and the audio sensor is used to capture audio occurring temporally with the identified visible events of interest. A sound profile is determined for each of the identified visible events of interest based on sounds captured by the audio sensor during the corresponding identified visible event of interest. Then, during a time of poor visibility, a subsequent sound event is identified in a subsequent audio stream captured by the audio sensor. One or more sound characteristics of the subsequent sound event are compared with the sound profiles associated with each of the identified visible events of interest, and if there is a match, one or more matching sound profiles are filtered out from the subsequent audio stream.

Claims (52)

1. A method for identifying activity in an area, the method comprising:

identifying a sound event in an audio stream captured by an audio sensor;

comparing one or more sound characteristics of the sound event with one or more predetermined sound profiles stored in an audio library, wherein each of the one or more predetermined sound profiles is associated with one or more events of interest, and when there is a match, identifying a matching event of interest;

removing from the audio stream one or more sound characteristics associated with the matching event of interest, including filtering out one or more spectral components from the audio stream that are associated with the matching event of interest while passing the remaining spectral components, resulting in a modified audio stream to determine whether one or more abnormal sounds are present in the audio stream now that the one or more sound characteristics associated with the matching event of interest are removed;

analyzing the modified audio stream for an abnormal sound remaining in the modified audio stream; and

issuing an alert when the abnormal sound is detected in the modified audio stream.

2. The method of claim 1 , comprising:

determining whether a legible video was captured by a video camera of the identified sound event;

when no legible video was captured by the video camera of the identified sound event, only then performing the identifying, comparing, removing, analyzing and issuing steps of claim 1 .

3. The method of claim 1 , wherein the abnormal sound is at least partially masked by one or more sound characteristics associated with the matching event of interest, and wherein removing at least some of the one or more sound characteristics associated with the matching event of interest at least partially unmasks the abnormal sound.

4. The method of claim 1 , wherein each of the one or more predetermined sound profiles of the audio library comprises an audio feature vector representative of the sound that occurred during the corresponding one or more events of interests.

5. The method of claim 4 , comprising:

computing an audio feature vector that is representative of the sound that occurred during the identified sound event; and

comparing one or more sound characteristics of the identified sound event with one or more predetermined sound profiles of the audio library includes:

comparing the audio feature vector that is representative of the sound that occurred during the identified sound event with the audio feature vector representative of the sound that occurred during each of the one or more events of interest, and when there is a match, identifying the matching event of interest.

6. The method of claim 5 , wherein:

the audio feature vector representative of the sound that occurred during the identified sound event identifies one or more spectral components of the sound that occurred during the identified sound event; and

the audio feature vector representative of the sound that occurred during the corresponding one or more events of interests identifies one or more spectral components of the sound that occurred during the corresponding one or more events of interests.

7. The method of claim 1 , wherein the abnormal sound comprises one or more of talking, shouting, chanting, screaming, laughing, sneezing, coughing, walking footsteps, and running footsteps.

8. The method of claim 1 comprising:

generating each of one or more of the predetermined sound profiles for the audio library by:

capturing a legible video using a video camera;

processing the legible video to identify one or more events of interest;

determining a corresponding sound profile for each of the identified events of interest based on sounds captured by the audio sensor during the corresponding identified event of interest; and

associating each of the identified events of interest with the corresponding sound profile.

9. The method of claim 8 , wherein at least one of the one or more of the predetermined sound profiles are generated locally in the area to account for at least some of the local acoustics associated with the area.

10. The method of claim 1 , wherein identifying the sound event in an audio stream comprises identifying a sound in the audio stream that exceeds a threshold sound level.

11. A method for revealing one or more sounds of interest captured in an area by an audio stream, wherein the one or more sounds of interest are at least partially masked by one or more obstructing sounds associated with one or more obstructing events occurring in the area, the method comprising:

comparing sounds in the audio stream to one or more predetermined sound profiles, wherein each of the one or more predetermined sound profiles is associated with one or more obstructing events, and when there is a match, identifying a matching obstructing event;

removing from the audio stream one or more sound characteristics associated with the matching obstructing event, including filtering out one or more spectral components from the audio stream that are associated with the matching obstructing event while passing the remaining spectral components, resulting in a modified audio stream to determine whether the one or more sounds of interest are present in the audio stream now that the one or more sound characteristics associated with the matching obstructing event are removed;

analyzing the modified audio stream for the one or more sounds of interest; and

issuing an alert when one or more sounds of interest are detected.

12. The method of claim 11 , comprising:

identifying a sound event in the audio stream;

determining whether a legible video was captured by a video camera of the identified sound event, only then performing the comparing, removing, analyzing and issuing steps of claim 11 .

13. The method of claim 11 , wherein the one or more sounds of interest comprises one or more of talking, shouting, chanting, screaming, laughing, sneezing, coughing, walking footsteps, and running footsteps.

14. The method of claim 11 , wherein the one or more obstructing events occurring in the area correspond to a moving object in the area.

15. The method of claim 14 , wherein the moving object comprises one or more of a plane, an automobile, a bus, a train and a truck.

16. The method of claim 11 , wherein the one or more obstructing events occurring in the area correspond to one or more of wind, rain, sleet, snow, and thunder.

17. The method of claim 11 , wherein the one or more obstructing events occurring in the area correspond to construction work.

18. The method of claim 11 , wherein the one or more obstructing events occurring in the area correspond to ambient background noise.

19. A system for identifying activity in an area even during periods of poor visibility, the system comprising:

a video camera;

an audio sensor;

a processor operatively coupled to the video camera and the audio sensor, the processor configured to:

identify a sound event in an audio stream captured by the audio sensor;

determine whether a legible video was captured by the video camera of the identified sound event;

when no legible video was captured by the video camera of the identified sound event:

compare one or more sound characteristics of the sound event with one or more predetermined sound profiles in an audio library, wherein each of the one or more predetermined sound profiles is associated with one or more events of interest, and when there is a match, identifying a matching event of interest;

remove from the audio stream one or more sound characteristics associated with the matching event of interest, including filtering out one or more spectral components from the audio stream that are associated with the matching event of interest while passing the remaining spectral components, resulting in a modified audio stream to determine whether one or more abnormal sounds are present in the audio stream now that the one or more sound characteristics associated with the matching event of interest are removed;

analyze the modified audio stream for an abnormal sound remaining in the modified audio stream;

issue an alert when the abnormal sound is detected in the modified audio stream.

Continuity (2)
Continuation 17208542 · Mar 22, 2021
Related Publication 20230222799A1 · Jul 13, 2023
References Cited (48)
US 6775642B2 · Rembowski et al. · 2004 [cited by applicant]
US 7683929B2 · Elazar et al. · 2010 [cited by applicant]
US 8643539B2 · Pauly et al. · 2014 [cited by applicant]
US 8938404B2 · Capman et al. · 2015 [cited by applicant]
US 9244042B2 · Rank · 2016 [cited by applicant]
US 9658100B2 · Park · 2017 [cited by applicant]
US 9740940B2 · Chattopadhyay et al. · 2017 [cited by applicant]
US 10354655B1 · White et al. · 2019 [cited by applicant]
US 10475468B1 · Yelchuru et al. · 2019 [cited by applicant]
US 10615995B2 · Yu · 2020 [cited by applicant]
US 10755730B1 · Maurer et al. · 2020 [cited by applicant]
US 11076274B1 · Buentello · 2021 [cited by examiner]
US 20050004797A1 · Azencott · 2005 [cited by applicant]
US 20060227237A1 · Kienzle et al. · 2006 [cited by applicant]
US 20080309761A1 · Kienzle et al. · 2008 [cited by applicant]
US 20120008821A1 · Sharon et al. · 2012 [cited by applicant]
US 20120245927A1 · Bondy · 2012 [cited by applicant]
US 20130057761A1 · Bloom et al. · 2013 [cited by applicant]
US 20160091398A1 · Pluemer · 2016 [cited by applicant]
US 20160191268A1 · Diebel · 2016 [cited by applicant]
US 20160316293A1 · Klimanis · 2016 [cited by examiner]
US 20160327522A1 · Tanaka et al. · 2016 [cited by applicant]
US 20160330062A1 · Alloin et al. · 2016 [cited by applicant]
US 20180040222A1 · Findlay et al. · 2018 [cited by applicant]
US 20180358052A1 · Miller et al. · 2018 [cited by applicant]
US 20190228229A1 · Cotoros · 2019 [cited by applicant]
US 20190246075A1 · Khadloya · 2019 [cited by examiner]
US 20190259378A1 · Khadloya · 2019 [cited by examiner]
US 20200020328A1 · Gordon · 2020 [cited by examiner]
US 20200066257A1 · Smith et al. · 2020 [cited by applicant]
US 20200301378A1 · McQueen et al. · 2020 [cited by applicant]
US 20200302951A1 · Deng · 2020 [cited by examiner]
US 20210084389A1 · Young et al. · 2021 [cited by applicant]
US 20220125021A1 · Herborn · 2022 [cited by examiner]
US 20220189267A1 · Sun · 2022 [cited by examiner]
CN 103366738B · 2016 [cited by applicant]
CN 205600145U · 2016 [cited by applicant]
EP 3193317A1 · 2017 [cited by applicant]
Saimurugan, et al; “Intelligent Fault Diagnosis for Rotating Machinery Based on Fusion of Sound Signal”, International Journal of Prognostics and Health Management, 10 pages, 2016. [cited by applicant]
Pan, et al; “Cognitive Acoustic Analytics Service for Internet of Things”, 2017 IEEE International Conference on Cognitive Computing (ICCC), 8 pages, Jun. 25-30, 2017. [cited by applicant]
Scardapane et al; “Microphone Array Based Classification for Security Monitoring in Unstructured Environments”, AEU _ International Journal of Electronics and Communications, vol. 69, Issue 11, 9 pages, Nov. 2015. [cited by applicant]
Ntalampiras, et al; “On Acoustic Surveillance of Hazardous Situations”, 2009 IEEE International Conference on Acoustics, Speech and Signal Processing, 5 pages, Apr. 19-24, 2009. [cited by applicant]
Maijala, et al; “Environmental Noise Monitoring Using Source Classification in Sensors”, Applied Acoustics, vol. 129, 10 pages, Jan. 2018. [cited by applicant]
Dey, et al; “Smart City Surveillance: Leveraging Benefits of Cloud Data Stores”, First IEEE International Workshop on GLObal Trends in Smart Cities, go SMART 2012, pp. 868-876, Clearwater, 2012. [cited by applicant]
Foggia et al; “Audio Surveillance of Roads: A System for Detecting Anamalous Sounds,” IEEE Transactions on Intelligent Transportation Systems, vol. 17, No. 1, pp. 278-288, Jan. 2016. [cited by applicant]
Sound effects-Royalty Free FX Library/Pond 5 https://www.pond5.com/sound-effects/ Accessed Mar. 22, 2021. [cited by applicant]
Street Corner Videos/Royalty-Free Stock Footage, Pond 5 Inc. 2021. [cited by applicant]
Street Comer Stock Video Footage, Royalty Free Street Corner Videos, Pond5, retrieved from https://www.pond5.com/stock-footage/tag/street-corner/ on Nov. 17, 2022 (26 pages). [cited by applicant]