IP Library › Granted Patent US 12,254,659
Granted Patent B2
US 12,254,659 · App. 17/662,050 · Granted Mar 18, 2025

Hybrid video analytics for small and specialized object detection

Inventors: Peter L. Venetianer (McLean, VA); Burak Kakillioglu (Syracuse, NY); Aleksey Lipchin (Newton, MA); Xiao Xiao (Winchester, MA)
Assignee: MOTOROLA SOLUTIONS, INC.
G06V10/255G06T7/194G06V10/82G06V20/52
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,659
App. No.
17/662,050
Granted
Mar 18, 2025
Kind
B2
Abstract

One aspect provides a method for object detection including detecting, using an electronic processor, a plurality of candidate objects in a video using a convolutional neural network detection process and a background subtraction detection process and identifying, using the electronic processor, a candidate object from the plurality of candidate objects. The candidate object detected by the background subtraction detection process in a location of the video with no candidate objects detected by the convolutional neural network detection process. The method also includes determining, using the electronic processor, a background subtraction confidence level of the candidate object and categorizing, using the electronic processor, the candidate object as a detected object in the video in response to the background subtraction confidence level satisfying a background subtraction confidence threshold.

Claims (87)

1. A video surveillance system comprising:

a video camera configured to capture a video; and

an object detector in communication with the video camera and including an electronic processor configured to

receive the video from the video camera,

detect a plurality of candidate objects in the video using a convolutional neural network detection process and a background subtraction detection process;

identify a candidate object from the plurality of candidate objects, the candidate object detected by the background subtraction detection process in a location of the video with no candidate objects detected by the convolutional neural network detection process;

determine a background subtraction confidence level of the candidate object;

categorize the candidate object as a detected object in the video in response to the background subtraction confidence level satisfying a background subtraction confidence threshold; and

in response to the background subtraction confidence level not satisfying the background subtraction confidence threshold

perform an additional convolutional neural network detection process on a location of the candidate object;

determine an additional convolutional neural network confidence level of the candidate object; and

categorize the candidate object as the detected object in the video in response to the additional convolutional neural network confidence level satisfying an additional convolutional neural network confidence threshold.

2. The video surveillance system of claim 1 , wherein the electronic processor is further configured to discard the candidate object in response to the additional convolutional neural network confidence level not satisfying the additional convolutional neural network confidence threshold.

3. The video surveillance system of claim 1 , wherein the electronic processor is further configured to

determine whether a convolutional neural network confidence level of a second candidate object from the plurality of candidate objects satisfies a convolutional neural network confidence threshold, the second candidate object detected by the convolutional neural network detection process;

categorize the second candidate object as a second detected object in the video in response to the convolutional neural network confidence level satisfying the convolutional neural network confidence threshold;

in response to the convolutional neural network confidence level not satisfying the convolutional neural network confidence threshold

determine whether the convolutional neural network confidence level satisfies a second convolutional neural network confidence threshold lower than the convolutional neural network confidence threshold,

in response to the convolutional neural network confidence level satisfying the second convolutional neural network confidence threshold

determine whether a size parameter of the second candidate object satisfies a size threshold for detection,

in response to the size parameter not satisfying the size threshold for detection

determine whether a motion of the second candidate object is isolated by the background subtraction detection process,

in response to the motion of the second candidate object being isolated by the background subtraction detection process

 determine a boosted convolutional neural network confidence level of the second candidate object based on the convolutional neural network confidence level and an overlap ratio of a second location of the second candidate object detected by the convolutional neural network detection process and a third location of the second candidate object detected by the background subtraction detection process; and

discard the second candidate object in response to the boosted convolutional neural network confidence level not satisfying the convolutional neural network confidence threshold.

4. The video surveillance system of claim 3 , wherein the electronic processor further configured to

in response to the boosted convolutional neural network confidence level satisfying the convolutional neural network confidence threshold

categorize the second candidate object as the second detected object.

5. The video surveillance system of claim 1 , wherein the background subtraction detection process includes foreground detection for detecting candidate objects.

6. The video surveillance system of claim 1 , wherein the additional convolutional neural network detection process is different from the convolutional neural network detection process.

7. The video surveillance system of claim 1 , wherein performing the additional convolutional neural network detection process on the location of the candidate object includes performing the additional convolutional neural network detection process on a cropped portion of the location of the video including the candidate object.

8. An object detector comprising:

an electronic processor configured to

detect a plurality of candidate objects in a video using a convolutional neural network detection process and a background subtraction detection process;

identify a candidate object from the plurality of candidate objects, the candidate object detected by the background subtraction detection process in a location of the video with no candidate objects detected by the convolutional neural network detection process;

determine a background subtraction confidence level of the candidate object; and

categorize the candidate object as a detected object in the video in response to the background subtraction confidence level satisfying a background subtraction confidence threshold; and

in response to the background subtraction detection level not satisfying the background subtraction confidence threshold

perform an additional convolutional neural network detection process on a location of the candidate object;

determine an additional convolutional neural network confidence level of the candidate object; and

categorize the candidate object as the detected object in the video in response to the additional convolutional neural network confidence level satisfying an additional convolutional neural network confidence threshold.

9. The object detector of claim 8 , wherein the electronic processor is further configured to discard the candidate object in response to the additional convolutional neural network confidence level not satisfying the additional convolutional neural network confidence threshold.

10. The object detector of claim 8 , wherein the electronic processor is further configured to

determine whether a convolutional neural network confidence level of a second candidate object from the plurality of candidate objects satisfies a convolutional neural network confidence threshold, the second candidate object detected by the convolutional neural network detection process;

categorize the second candidate object as a second detected object in the video in response to the convolutional neural network confidence level satisfying the convolutional neural network confidence threshold;

in response to the convolutional neural network confidence level not satisfying the convolutional neural network confidence threshold

determine whether the convolutional neural network confidence level satisfies a second convolutional neural network confidence threshold lower than the convolutional neural network confidence threshold;

in response to the convolutional neural network confidence level satisfying the second convolutional neural network confidence threshold

determine whether a size parameter of the second candidate object satisfies a size threshold for detection,

in response to the size parameter not satisfying the size threshold for detection

determine whether a motion of the second candidate object is isolated by the background subtraction detection process,

in response to the motion of the second candidate object being isolated by the background subtraction detection process

 determine a boosted convolutional neural network confidence level of the second candidate object based on the convolutional neural network confidence level and an overlap ratio of a second location of the second candidate object detected by the convolutional neural network detection process and a third location of the second candidate object detected by the background subtraction detection process; and

discard the second candidate object in response to the boosted convolutional neural network confidence level not satisfying the convolutional neural network confidence threshold.

11. The object detector of claim 10 , wherein the electronic processor further configured to

in response to the boosted convolutional neural network confidence level satisfying the convolutional neural network confidence threshold

categorize the second candidate object as the second detected object in response to detecting movement of the second candidate object between multiple frames of the video.

12. The object detector of claim 8 , wherein the background subtraction detection process includes foreground detection for detecting candidate objects.

13. The object detector of claim 8 , wherein the additional convolutional neural network detection process is be different from the convolutional neural network detection process.

14. The object detector of claim 8 , wherein performing the additional convolutional neural network detection process on a location of the candidate object includes performing the additional convolutional neural network detection process on a cropped portion of the location of the video including the candidate object.

15. A method for object detection comprising:

detecting, using an electronic processor, a plurality of candidate objects in a video using a convolutional neural network detection process and a background subtraction detection process;

identifying, using the electronic processor, a candidate object from the plurality of candidate objects, the candidate object detected by the background subtraction detection process in a location of the video with no candidate objects detected by the convolutional neural network detection process;

determining, using the electronic processor, a background subtraction confidence level of the candidate object;

categorizing, using the electronic processor, the candidate object as a detected object in the video in response to the background subtraction confidence level satisfying a background subtraction confidence threshold; and

in response to the background subtraction detection level not satisfying the background subtraction confidence threshold

performing an additional convolutional neural network detection process on a location of the candidate object;

determining an additional convolutional neural network confidence level of the candidate object; and

categorizing the candidate object as the detected object in the video in response to the additional convolutional neural network confidence level satisfying an additional convolutional neural network confidence threshold.

16. The method of claim 15 , further comprising discarding the candidate object in response to the additional convolutional neural network confidence level not satisfying the additional convolutional neural network confidence threshold.

17. The method of claim 15 , further comprising:

determining whether a convolutional neural network confidence level of a second candidate object from the plurality of candidate objects satisfies a convolutional neural network confidence threshold, the second candidate object detected by the convolutional neural network detection process;

categorizing the second candidate object as a second detected object in the video in response to the convolutional neural network confidence level satisfying the convolutional neural network confidence threshold;

in response to the convolutional neural network confidence level not satisfying the convolutional neural network confidence threshold

determining whether the convolutional neural network confidence level satisfies a second convolutional neural network confidence threshold lower than the convolutional neural network confidence threshold;

in response to the convolutional neural network confidence level satisfying the second convolutional neural network confidence threshold

determining whether a size parameter of the second candidate object satisfies a size threshold for detection;

in response to the size parameter not satisfying the size threshold for detection

determining whether a motion of the second candidate object is isolated by the background subtraction detection process;

in response to the motion of the second candidate object being isolated by the background subtraction detection process

 determining a boosted convolutional neural network confidence level of the second candidate object based on the convolutional neural network confidence level and an overlap ratio of a second location of the second candidate object detected by the convolutional neural network detection process and a third location of the second candidate object detected by the background subtraction detection process; and

discarding the second candidate object in response to the boosted convolutional neural network confidence level not satisfying the convolutional neural network confidence threshold.

18. The method of claim 17 , further comprising:

in response to the boosted convolutional neural network confidence level satisfying the convolutional neural network confidence threshold

categorizing the second candidate object as the second detected object in response to detecting movement of the second candidate object between multiple frames of the video.

19. The method of claim 15 , wherein the background subtraction detection process includes foreground detection for detecting candidate objects.

20. The method of claim 15 , wherein the additional convolutional neural network detection process different from the convolutional neural network detection process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2022
From: VENETIANER, PETER L; KAKILLIOGLU, BURAK; LIPCHIN, ALEKSEY; XIAO, XIAO
To: MOTOROLA SOLUTIONS, INC.
Reel/Frame 059958/0370 →
Continuity (1)
Related Publication 20230360355A1 · Nov 9, 2023
References Cited (19)
US 10817739B2 · Gupta · 2020 [cited by examiner]
US 10878578B2 · Cheng · 2020 [cited by examiner]
US 11461919B2 · Wolf · 2022 [cited by examiner]
US 20190130580A1 · Chen · 2019 [cited by examiner]
US 20190130586A1 · Zhou · 2019 [cited by examiner]
US 20200364466A1 · Latapie et al. · 2020 [cited by applicant]
US 20210374941A1 · Liu et al. · 2021 [cited by applicant]
US 20230360355A1 · Venetianer · 2023 [cited by examiner]
WO 2019083738A1 · 2019 [cited by applicant]
Zhou et al., “Moving Object Detection by Detecting Contiguous Outliers in the Low-Rank Representation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2013, 35(3): 597-610. [cited by applicant]
Zivkovic, “Improved Adaptive Gaussian Mixture Model for Background Subtraction,” Proceedings of the 17th International Conference on Pattern Recognition, 2004, 4 pages. [cited by applicant]
Xin et al., “Background Subtraction via Generalized Fused Lasso Foreground Modeling,” IEEE Conference on Computer Vision and Pattern Recognition, 2015, 4676-4684. [cited by applicant]
Stauffer et al., “Adaptive Background Mixture Models for Real-time Tracking,” 1999 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 1999, 246-252. [cited by applicant]
Zivkovic et al., “Efficient Adaptive Density Estimation per Image Pixel for the Task of Background Subtraction,” Pattern Recognition Letters, 2006, 27(7): 773-780. [cited by applicant]
Lu et al., “Hybrid Deep Learning Based Moving Object Detection via Motion Prediction,” 2018 Chinese Automation Congress, 2018, 1862-1867. [cited by applicant]
Camplani et al., “Multi-sensor Background Subtraction by Fusing Multiple Region-based Probabilistic Classifiers,” Pattern Recognition Letters, 2014, 11 pages. [cited by applicant]
Grosgeorge et al., “Concurrent Segmentation and Object Detection CNNS for Aircraft Detection and Identification in Satellite Images,” IGARSS 2020-2020 IEEE International Geoscience and Remote Sensing Symposium, 2020, 4 … [cited by applicant]
Ferariu et al., “Fusing Faster R-CNN and Background Subtraction Based on the Mixture of Gaussians Model,” 24th International Conference on System Theory, Control and Computing, 2020, 367-372. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2023/017759 dated Jun. 16, 2023 (12 pages). [cited by applicant]