IP Library Granted Patent US 11,902,656
Granted Patent B2
US 11,902,656 · App. 17/852,718 · Granted Feb 13, 2024

Audio sensors for controlling surveillance video data capture

Inventors: Ramanathan Muthiah (Bangalore, IN); Akhilesh Yadav (Bangalore, IN)
Assignee: Western Digital Technologies, Inc.
H04N23/667H04N7/188H04N23/69H04N23/695H04R1/326
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,902,656
App. No.
17/852,718
Filed
Jun 29, 2022
Granted
Feb 13, 2024
Kind
B2
Art Unit
2638
USPC
348/222.1
Abstract

Systems, video cameras, and methods for using audio sensors to control surveillance video capture are described. A video camera and audio sensor are deployed so that the audio sensor has an audio field that is at least partially outside the field of view of the video camera. The audio sensor collects audio data from the audio field and a controller for the video camera uses audio events from the audio data for modifying the video capture operations of the video camera. Video data is then captured based on the modified video capture operations, such as initiating video capture, changing the video capture rate, or changing the camera position.

Claims (89)

1. A system, comprising:

a video camera configured for a plurality of video capture rates;

an audio sensor, wherein:

the audio sensor is configured to collect audio data from an audio field; and

the audio field is at least partially outside a field of view of the video camera; and

a controller configured to:

receive audio data from the audio sensor;

determine, from the audio data, an audio event;

select, responsive to the audio event, a first video capture rate from the plurality of video capture rates;

modify, responsive to the audio event, a video capture operation of the video camera using the first video capture rate during a first operating period; and

capture, using the video camera, video data based on the modified video capture operation.

2. The system of claim 1 , wherein the controller is further configured to select a second video capture rate from the plurality of video capture rates during a second operating period.

3. The system of claim 1 , wherein the controller is further configured to:

suspend video capture during a second operating period; and

initiate, responsive to the audio event, video capture at the first video capture rate to modify the video capture operation during the first operating period.

4. The system of claim 1 , wherein:

the audio event is associated with a video object of interest; and

the audio event precedes the video object being detectable in the field of view of the video camera.

5. The system of claim 1 , wherein:

the audio sensor comprises at least one directional microphone configured with a direction and an audio range to detect sound sources outside the field of view of the video camera; and

the controller is further configured to determine, based on the audio data, a direction of movement of a sound source that intercepts the field of view of the video camera.

6. The system of claim 5 , wherein the at least one directional microphone is configured as an audio tripwire for the sound source approaching the field of view of the video camera.

7. The system of claim 1 , further comprising:

an analytics engine configured to:

receive the audio data from the audio sensor;

determine, in the audio data, the audio event, wherein determining the audio event is based on:

an audio recognition value meeting an audio recognition threshold; and

an audio duration value meeting an audio duration threshold; and

return the audio event for use by the controller.

8. The system of claim 7 , wherein:

the analytics engine is further configured to use an audio recognition model to determine the audio recognition value;

the audio recognition model is configured to classify the audio data using at least one audio source type identifier; and

the controller is further configured to use the at least one audio source type identifier to determine a modification of the video capture operation of the video camera.

9. The system of claim 7 , wherein:

the analytics engine is further configured to use an audio recognition model to determine the audio recognition value;

the audio recognition model is configured to determine a location and a direction of movement of a sound source; and

the controller is further configured to send, responsive to the location and the direction of movement of the sound source, a pan-tilt-zoom position control signal to the video camera to adjust the field of view of the video camera.

10. The system of claim 7 , wherein:

the analytics engine is further configured to use an audio recognition model to determine the audio recognition value;

the audio recognition model is a machine learning model trained with audio reference data corresponding to known sound sources;

the controller is further configured to:

detect, using the video data, at least one data object in the field of view of the video camera; and

determine, based on correlations of the audio event and detecting at least one data object, additional audio reference data; and

the analytics engine is further configured to retrain the machine learning model using the additional audio reference data.

11. A computer-implemented method, comprising:

collecting, by an audio sensor, audio data from an audio field, wherein the audio field is at least partially outside a field of view of a video camera;

receiving the audio data from the audio sensor;

determining, based on the audio data, an audio recognition value;

determining, from the audio data, an audio event based on the audio recognition value meeting an audio recognition threshold;

modifying, responsive to the audio event, a video capture operation of the video camera; and

capturing, using the video camera, video data based on the modified video capture operation.

12. The computer-implemented method of claim 11 , further comprising:

selecting a first video capture rate from a plurality of video capture rates for the video camera during a first operating period; and

selecting, responsive to the audio event, a second video capture rate to modify the video capture operation during a second operating period.

13. The computer-implemented method of claim 11 , further comprising:

suspending video capture during a first operating period; and

initiating, responsive to the audio event, video capture at a selected video capture rate to modify the video capture operation during a second operating period.

14. The computer-implemented method of claim 11 , wherein:

the audio event is associated with a video object of interest; and

the audio event precedes the video object being detectable in the field of view of the video camera.

15. The computer-implemented method of claim 11 , further comprising:

determining, based on the audio data, a direction of movement of a sound source that intercepts the field of view of the video camera, wherein the audio sensor comprises at least one directional microphone configured with a direction and an audio range to detect sound sources outside the field of view of the video camera.

16. The computer-implemented method of claim 15 , further comprising configuring the at least one directional microphone as an audio tripwire for the sound source approaching the field of view of the video camera.

17. The computer-implemented method of claim 11 , further comprising:

determining the audio recognition value using an audio recognition model;

classifying, using the audio recognition model, the audio data using at least one audio source type identifier; and

determining, using the at least one audio source type identifier, a modification of the video capture operation of the video camera.

18. The computer-implemented method of claim 11 , further comprising:

determining the audio recognition value using an audio recognition model;

determining, using the audio recognition model, a location and a direction of movement of a sound source; and

adjusting, responsive to the location and the direction of movement of the sound source, the field of view of the video camera using a pan-tilt-zoom position control signal.

19. The computer-implemented method of claim 11 , further comprising:

determining the audio recognition value using an audio recognition model;

training, using a machine learning model and audio reference data corresponding to known sound sources, the audio recognition model;

detecting, using the video data, at least one data object in the field of view of the video camera;

determining, based on correlations of the audio event and detecting at least one data object, additional audio reference data; and

retraining, using the machine learning model and the additional audio reference data, the audio recognition model.

20. A storage system, comprising:

a video camera;

an audio sensor, wherein:

the audio sensor is configured to collect audio data from an audio field;

the audio field is at least partially outside a field of view of the video camera; and

the audio sensor comprises at least one directional microphone configured with a direction and an audio range to detect sound sources outside the field of view of the video camera;

a processor;

a memory;

means for collecting, by the audio sensor, audio data from the audio field;

means for determining, from the audio data, an audio event based on determining a direction of movement of a sound source that intercepts the field of view of the video camera;

means for modifying, responsive to the audio event, a video capture operation of the video camera; and

means for capturing, using the video camera, video data based on the modified video capture operation.

Assignments (8)
PARTIAL RELEASE OF SECURITY INTERESTS Recorded Apr 25, 2025
From: JPMORGAN CHASE BANK, N.A., AS AGENT
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 071382/0001 →
SECURITY AGREEMENT Recorded Apr 25, 2025
From: SANDISK TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 071050/0001 →
PATENT COLLATERAL AGREEMENT Recorded Aug 23, 2024
From: SANDISK TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A., AS THE AGENT
Reel/Frame 068762/0494 →
CHANGE OF NAME Recorded Jun 27, 2024
From: SANDISK TECHNOLOGIES, INC.
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 067982/0032 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2024
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: SANDISK TECHNOLOGIES, INC.
Reel/Frame 067567/0682 →
PATENT COLLATERAL AGREEMENT - A&R LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 064715/0001 →
PATENT COLLATERAL AGREEMENT - DDTL LOAN AGREEMENT Recorded Aug 21, 2023
From: WESTERN DIGITAL TECHNOLOGIES, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 067045/0156 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2022
From: MUTHIAH, RAMANATHAN; YADAV, AKHILESH
To: WESTERN DIGITAL TECHNOLOGIES, INC.
Reel/Frame 060354/0216 →
Continuity (1)
Related Publication 20240007744A1 · Jan 4, 2024
Cited By (6)
US 12,289,528 US 12,395,794 US 12,563,158 US 12,621,572 US 12,677,068 US 12,689,818