IP Library › Granted Patent US 12,340,523
Granted Patent B2
US 12,340,523 · App. 17/845,245 · Granted Jun 24, 2025

Relevant motion detection in video

Inventors: Ruichi Yu (Greenbelt, MD); Hongcheng Wang (Arlington, VA)
Assignee: Comcast Cable Communications, LLC
G06T7/246G06N5/046G06T7/254G06T2207/20081G06T2207/20084G06T2207/30232
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,523
App. No.
17/845,245
Granted
Jun 24, 2025
Kind
B2
Abstract

Methods, systems, and/or apparatuses are described for detecting relevant motion of objects of interest (e.g., persons and vehicles) in surveillance videos. As described herein input data based on a plurality of captured images and/or video is received. The input data may then be pre-processed and used as an input into a convolution network that may, in some instances, have elements that perform both spatial-wise max pooling and temporal-wise max pooling. The convolution network may be used to generate a plurality of prediction results of relevant motion of the objects of interest.

Claims (41)

1. A method comprising:

receiving, by a computing device, a sequence of images; and

outputting data associated with motion of one or more objects in the sequence of images, wherein the motion of the one or more objects is determined based on application of a sequential filter to a spatially filtered version of the sequence of images, and wherein the sequential filter reduces, in a sequential dimension, the spatially filtered version of the sequence of images.

2. The method of claim 1 , wherein the application of the sequential filter comprises sampling a quantity of images in the spatially filtered version of the sequence of images.

3. The method of claim 1 , wherein the spatially filtered version of the sequence of images has reduced spatial dimensions relative to spatial dimensions of the received sequence of images.

4. The method of claim 1 , wherein the spatially filtered version of the sequence of images is based on subtraction of one or more reference images from the received sequence of images and based on application of a spatial filter to the reference image-subtracted received sequence of images.

5. The method of claim 1 , wherein the spatially filtered version of the sequence of images is based on application of a spatial filter to the received sequence of images and based on application of a filter, associated with one or more probabilities of motion of one or more types of objects, to the spatially filtered received sequence of images.

6. The method of claim 1 , wherein the data associated with the motion of the one or more objects comprises one or more of:

an indication of the motion of the one or more objects;

an indication of a type of the one or more objects; or

a location, in the sequence of images, associated with the motion of the one or more objects.

7. The method of claim 1 , wherein the sequence of images is based on video content received from a security system.

8. The method of claim 1 , wherein the outputting the data comprises outputting the data to a mobile device.

9. An apparatus comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the apparatus to:

receive a sequence of images; and

output data associated with motion of one or more objects in the sequence of images, wherein the motion of the one or more objects is determined based on application of a sequential filter to a spatially filtered version of the sequence of images, and wherein the sequential filter reduces, in a sequential dimension, the spatially filtered version of the sequence of images.

10. The apparatus of claim 9 , wherein the application of the sequential filter comprises sampling a quantity of images in the spatially filtered version of the sequence of images.

11. The apparatus of claim 9 , wherein the spatially filtered version of the sequence of images has reduced spatial dimensions relative to spatial dimensions of the received sequence of images.

12. The apparatus of claim 9 , wherein the spatially filtered version of the sequence of images is based on subtraction of one or more reference images from the received sequence of images and based on application of a spatial filter to the reference image-subtracted received sequence of images.

13. The apparatus of claim 9 , wherein the spatially filtered version of the sequence of images is based on application of a spatial filter to the received sequence of images and based on application of a filter, associated with one or more probabilities of motion of one or more types of objects, to the spatially filtered received sequence of images.

14. The apparatus of claim 9 , wherein the data associated with the motion of the one or more objects comprises one or more of:

an indication of the motion of the one or more objects;

an indication of a type of the one or more objects; or

a location, in the sequence of images, associated with the motion of the one or more objects.

15. The apparatus of claim 9 , wherein the sequence of images is based on video content received from a security system.

16. The apparatus of claim 9 , wherein the instructions, when executed by the one or more processors, cause the apparatus to output the data by outputting the data to a mobile device.

17. A non-transitory computer-readable medium having stored thereon computer-executable instructions that, when executed, cause:

receiving a sequence of images; and

outputting data associated with motion of one or more objects in the sequence of images, wherein the motion of the one or more objects is determined based on application of a sequential filter to a spatially filtered version of the sequence of images, and wherein the sequential filter reduces, in a sequential dimension, the spatially filtered version of the sequence of images.

18. The non-transitory computer-readable medium of claim 17 , wherein the application of the sequential filter comprises sampling a quantity of images in the spatially filtered version of the sequence of images.

19. The non-transitory computer-readable medium of claim 17 , wherein the spatially filtered version of the sequence of images has reduced spatial dimensions relative to spatial dimensions of the received sequence of images.

20. The non-transitory computer-readable medium of claim 17 , wherein the spatially filtered version of the sequence of images is based on subtraction of one or more reference images from the received sequence of images and based on application of a spatial filter to the reference image-subtracted received sequence of images.

21. The non-transitory computer-readable medium of claim 17 , wherein the spatially filtered version of the sequence of images is based on application of a spatial filter to the received sequence of images and based on application of a filter, associated with one or more probabilities of motion of one or more types of objects, to the spatially filtered received sequence of images.

22. The non-transitory computer-readable medium of claim 17 , wherein the data associated with the motion of the one or more objects comprises one or more of:

an indication of the motion of the one or more objects;

an indication of a type of the one or more objects; or

a location, in the sequence of images, associated with the motion of the one or more objects.

23. The non-transitory computer-readable medium of claim 17 , wherein the sequence of images is based on video content received from a security system.

24. The non-transitory computer-readable medium of claim 17 , wherein the outputting the data comprises outputting the data to a mobile device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2022
From: YU, RUICHI; WANG, HONGCHENG
To: COMCAST CABLE COMMUNICATIONS, LLC
Reel/Frame 060396/0455 →
Continuity (4)
Continuation 17088203 · Nov 3, 2020
Continuation 16125203 · Sep 7, 2018
Provisional Application 62555501 · Sep 7, 2017
Related Publication 20220319017A1 · Oct 6, 2022
References Cited (62)
US 9430923B2 · Kniffen et al. · 2016 [cited by applicant]
US 9984466B1 · Duran · 2018 [cited by applicant]
US 10026007B1 · Morton · 2018 [cited by applicant]
US 10270642B2 · Zhang · 2019 [cited by examiner]
US 20070263905A1 · Chang et al. · 2007 [cited by applicant]
US 20100034423A1 · Zhao et al. · 2010 [cited by applicant]
US 20100220939A1 · Tourapis · 2010 [cited by examiner]
US 20100290710A1 · Gagvani · 2010 [cited by examiner]
US 20160171311A1 · Case · 2016 [cited by examiner]
US 20180374233A1 · Zhou · 2018 [cited by examiner]
US 20190057588A1 · Savvides · 2019 [cited by examiner]
M. S. A. Akbarian, F. Saleh, B. Fernando, M. Salzmann, L. Petersson, and L. Andersson, “Deep Action- and Context-Aware Sequence Learning for Activity Recognition and Anticipation”, Australian National University, Smart … [cited by applicant]
M. Babaee, D. Dinh, and G. Rigoll, “A Deep Convolutional Neural Network for Background Subtraction”, Institute for Human-Machine Communication, Technical Univ. of Munich, Germany, Feb. 2017. Retrieved from: https://arxi… [cited by applicant]
L. Bertinetto, J. Valmadre, J. F. Henriques, A. Vedaldi, and P. H. S. Torr, “Fully-Convolutional Siamese Networks for Object Tracking”, Deparatment of Engineering Science, University of Oxford, submitted Jun. 2016, revi… [cited by applicant]
A. Bewley, Z. Ge, L. Ott, F. Ramos, and B. Upcroft, “Simple Online and Realtime Tracking”, Queensland University of Technology, University of Sydney, submitted Feb. 2016, revised Jul. 2017. Retrieved from: https://arxiv… [cited by applicant]
D. Comaniciu, V. Ramesh, and P. Meer, “Kernel-Based Object Tracking”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 25(5): 564-577, May 2003. Retrieved from: https://pdfs.semanticscholar.org/7e16/50ef9… [cited by applicant]
R. Cucchiara, C. Grana, A. Prati, and R. Vezzani, “Probabilistic Posture Classification for Human-Behavior Analysis”, IEEE Transactions on Systems, Man, and Cybernetics-Part A, 35(1):42-54, Jan. 2005. Retrieved from: ht… [cited by applicant]
J. Donahue, L. A. Hendricks, S. Guadarrama, M. Rohrbach, S. Venugopalan, K. Saenko, and T. Darrell, “Long-term Recurrent Convolutional Networks for Visual Recognition and Description”, submitted Nov. 2014, revised May 2… [cited by applicant]
R. B. Girshick, “Fast R-CNN”, Microsoft Research, submitted Apr. 2015, revised Sep. 2015. Retrieved from: https://arxiv.org/abs/1504.08083. [cited by applicant]
R. B. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation”, UC Berkeley, submitted Nov. 2013, revised Oct. 2014. Retrieved from: https://arxi… [cited by applicant]
M. Hasan and A. K. Roy-Chowdhury, “Context Aware Active Learning of Activity Recognition Models”, 2015 Proceedings of the IEEE International Conference on Computer Vision, pp. 4543-4551, 2015. Retrieved from: https://ww… [cited by applicant]
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition”, Microsoft Research, Dec. 2015. Retrieved from: https://arxiv.org/abs/1512.03385. [cited by applicant]
J. Heikkila and O. Silvén, “A Real-Time System for Monitoring of Cyclists and Pedestrians”, Proceedings Second IEEE Workshop on Visual Surveillance (VS'99), Jun. 1999. Retrieved from: http://www.ee.oulu.fi/mvg/giva/moti… [cited by applicant]
D. Held et al., “Learning to track at 100 FPS with deep regression networks”, submitted Apr. 2015, revised Aug. 2016. Retrieved from: https://arxiv.org/abs/1604.01802. [cited by applicant]
A. G. Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications”, Google Inc., Apr. 2017. Retrieved from: https://arxiv.org/abs/1704.04861. [cited by applicant]
J. Huang et al., “Speed/accuracy trade-offs for modern convolutional object detectors”, Google Research, submitted Nov. 2016, revised Apr. 2017. Retrieved from: https://arxiv.org/abs/1611.10012. [cited by applicant]
F. N. Iandola et al., “SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size”, DeepScale & UC Berkeley, Stanford University, submitted Feb. 2016, revised Nov. 2016. Retrieved from: https://a… [cited by applicant]
D. Kang et al., “NoScope: Optimizing Neural Network Queries over Video at Scale”, Stanford InfoLab, submitted Mar. 2017, revised Aug. 2017. Retrieved from: https://arxiv.org/abs/1703.02529. [cited by applicant]
W. Liu et al., “SSD: Single Shot MultiBox Detector”, submitted Dec. 2015, revised Dec. 2016. Retrieved from: https://arxiv.org/pdf/1512.02325.pdf. [cited by applicant]
L. Nan et al., “An Improved Motion Detection Method for Real-Time Surveillance”, IAENG International Journal of Computer Science, 35(1), Feb. 2008. Retrieved from: https://pdfs.semanticscholar.org/b47c/e6112b2ca7d5b59d8… [cited by applicant]
P. Ondruska and I. Posner, “Deep Tracking: Seeing Beyond Seeing Using Recurrent Neural Networks”, presented at AAAI-16 Conference, Feb. 2016. Retrieved from: https://arxiv.org/abs/1602.00991. [cited by applicant]
J. Redmon et al., “You Only Look Once: Unified, Real-Time Object Detection”, submitted Jun. 2015, revised May 2016. Retrieved from: https://arxiv.org/abs/1506.02640. [cited by applicant]
J. Redmon and A. Farhadi, “YOLO9000: Better, Faster, Strong”, Dec. 2016. Retrieved from: https://arxiv.org/abs/1612.08242. [cited by applicant]
S. Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks” In Advances in Neural Information Processing Systems 28, Curran Associates, Inc., 2015. Retrieved from: https://papers.nips… [cited by applicant]
W. Shuigen et al., “Motion Detection Based on Temporal Difference Method and Optical Flow field”, 2009 Second International Symposium on Electronic Commerce and Security, pp. 85-88. Retrieved from IEEE. [cited by applicant]
K. Simonyan and A. Zisserman, “Two-Stream Convolutional Networks for Action Recognition in Videos”, Visual Geometry Group, University of Oxfrod, submitted Jun. 2014, revised Nov. 2014. Retrieved from: https://arxiv.org/… [cited by applicant]
A. Sobral and T. Bouwmans, “BGS Library: A Library Framework for Algorithm's Evaluation in Foreground/Background Segmentation” in Background Modeling and Foreground Detection for Video Surveillance: Traditional and Rece… [cited by applicant]
C. Stauffer and W. E. L. Grimson, “Adaptive background mixture models for real-time tracking”, in CVPR, pp. 2246-2252, IEEE Computer Society, 1999. Retreived from: http://www.ai.mit.edu/projects/vsam/Publications/stauff… [cited by applicant]
C. Szegedy et al., “Rethinking the Inception Architecture for Computer Vision”, Dec. 2015. Retrieved from: https://arxiv.org/abs/1512.00567. [cited by applicant]
D. Tran et al., “C3D: Generic Features for Video Analysis”, Dec. 2014. Retrieved from: https://arxiv.org/abs/1412.0767v1. [cited by applicant]
M. Wang, B. Ni, and X. Yang, “Recurrent Modeling of Interaction Context for Collective Activity Recognition”, in The IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017. Retrieved from: http://o… [cited by applicant]
Y. Wang et al., “CDnet 2014: An Expanded Change Detection Benchmark Dataset”, in Proc. IEEE Workshop on Change Detection (CDW-2014) at CVPR-2014, pp. 387-394, 2014. Retrieved from: http://openaccess.thecvf.com/content_c… [cited by applicant]
C. Wren et al., “Pfinder: Real-time tracking of the human body”, IEEE Transactions on Pattern Analysis and Machine Intelligence, 19:780-785, 1997. [cited by applicant]
D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014. [cited by applicant]
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. E. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich. Going deeper with convolutions. CoRR, abs/1409.4842, 2014. [cited by applicant]
B. Wu, F. Iandola, P. H. Jin, and K. Keutzer. Squeezedet: Unified, small, low power fully convolutional neural networks for real-time object detection for autonomous driving. In The IEEE Conference on Computer Vision an… [cited by applicant]
P. Xu, M. Ye, X. Li, Q. Liu, Y. Yang, and J. Ding. Dynamic background learning through deep auto-encoder networks. In Proceedings of the 22Nd ACM International Conference on Multimedia, MM '14, pp. 107-116, New York, NY… [cited by applicant]
R. Y. D. Xu, J. G. Allen, and J. S. Jin. Robust real-time tracking of non-rigid objects. In Proceedings of the Pan-Sydney Area Workshop on Visual Information Processing, VIP '05, pp. 95-98, Darlinghurst, Australia, Aust… [cited by applicant]
R. Yu, X. Chen, V. I. Morariu, and L. S. Davis. The role of context selection in object detection. In British Machine Vision Conference (BMVC), 2016. [cited by applicant]
Z. Zivkovic. Improved adaptive gaussian mixture model for background subtraction. In Proceedings of the Pattern Recognition, 17th International Conference on (ICPR'04) vol. 2—vol. 02, ICPR '04, pp. 28-31, Washington, DC… [cited by applicant]
Lan, Y., Ji, Z., Gao, J. and Wang, Y., 2014, May. Robot fish detection based on a combination method of three-frame-difference and background subtraction. In The 26th Chinese Control and Decision Conference (2014 CCDC) … [cited by applicant]
Gupta, P., Singh, Y. and Gupt, M., 2014. Moving object detection using frame difference, background subtraction and sobs for video surveillance application. In The Proceedings of the 3rd International Conference System … [cited by applicant]
Hussain, Z., Naaz, A. and Uddin, N., 2016. Moving object detection based on background subtraction and frame differencing technique. International Journal of Advanced Research in Computer and Communication Engineering, … [cited by applicant]
Zhang, Y., Wang, X. and Qu, B., 2012. Three-frame difference algorithm research based on mathematical morphology. Procedia Engineering, 29, pp. 2705-2709. [cited by applicant]
Open CV, Wikipedia, retrieved from https://en.wikipedia.org/wiki/OpenCV on Jul. 9, 2020. [cited by applicant]
How to Use Background Subtraction Methods, docs.OpenCV.com, retrieved from https://docs.opencv.org/master/d1/dc5/tutorial_background_subtraction.html#:˜:text=OpenCV%3A%20How%20to%20Use%20Background%20Subtraction%20Metho… [cited by applicant]
Background Subtraction, Open CV-Python Tutorials, retrieved from https://opencv-python-tutroals.readthedocs.io/en/latest/py_tutorials/py_video/py_bg_subtraction/py_bg_subtraction.html on Jul. 9, 2020. [cited by applicant]
Python | Background subtraction using Open CV, GeeksforGeeks, retrieved from https://www.geeksforgeeks.org/python-background-subtraction-using-opencv/ on Jul. 9, 2020. [cited by applicant]
A Beginner's Guide to Convolutional Neural Networks (CNNs), Skymind, retrieved from https://skymind.ai/wiki/convolutional-network on Aug. 20, 2019. [cited by applicant]
Xu, J. et al., “MSR-VTT: A Large Video Description Dataset for Bridging Video and Language Supplementary Material,” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 5288-5296, https://do… [cited by applicant]
Singh, B., et al., ‘A Multi-stream Bi-Directional Recurrent Neural Network for Fine-Grained Action Detection,’ 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 1961-1970, https://doi.org… [cited by applicant]
Dec. 6, 2024—Canadian Office Action—CA App. No. 3,016,953. [cited by applicant]