IP Library › Granted Patent US 12,361,700
Granted Patent B2
US 12,361,700 · App. 18/236,747 · Granted Jul 15, 2025

Image processing and object detecting system, image processing and object detecting method, and program storage medium

Inventor: Hiroyoshi Miyano (Tokyo, JP)
Assignee: NEC CORPORATION
G06V10/82G06T7/215G06T7/251G06T7/70G06V10/28G06V20/52H04N7/18G06T2207/10016G06T2207/10024G06T2207/20076G06T2207/20081G06T2207/30196G06T2207/30232
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,700
App. No.
18/236,747
Granted
Jul 15, 2025
Kind
B2
Abstract

Provided is an image processing system, an image processing method, and a program for preferably detecting a mobile object. The image processing system includes: an image input unit for receiving an input for some image frames having different times in a plurality of image frames constituting a picture, which is of a pixel on which the mobile object appears or a pixel on which the mobile object does not appear, for selected arbitrary one or more pixels in an image frame at the time of processing; and a mobile object detection model constructing unit for learning a parameter for detecting the mobile object based on the input.

Claims (56)

1. An image processing system comprising:

at least one memory storing instructions; and

at least one processor configured to process the instructions to control the image processing system to:

receive an input of a plurality of image frames having different capturing times; and

perform one or more convolution calculations by using values of a background model of a neighboring region of one or more target pixels in the plurality of image frames to learn a parameter of a detection model for detecting a moving object, wherein

the plurality of background models include three background models, and these three background models are different from each other in the number of image frames to be processed;

the first background model has a largest number of image frames to be processed, the second background model has a second largest number of image frames to be processed, and the third background model has a smallest number of image frames to be processed;

the at least one processor configured to process the instructions to control the image processing system to:

calculate a first distance between the first background model and the second background model, a second distance between the second background model and the third background model, and a third distance the first background model and the third background model for each pixel or unit region of the plurality of image frames;

determine, based on the first distance, the second distance, and the third distance, whether or not the moving object is shown in a region corresponding to each pixel of the plurality of image frames; and

based on user input specifying whether some pixels in the plurality of image frames are pixels that correspond to the moving object or a pixel that does not show the moving object, learn parameters for detecting the moving object in the plurality of image frames.

2. The image processing system according to claim 1 , wherein

the at least one processor is configured to process the instructions to control the image processing system to:

learn, as the parameter, the one or more convolution calculations, and a threshold compared with a value obtained as a result of the one or more convolution calculations.

3. The image processing system according to claim 1 , wherein

the at least one processor is configured to process the instructions to control the image processing system to:

identify whether pixels that are selected at random from the plurality of image frames are a background region or an object region based on additional user inputs.

4. The image processing system according to claim 1 , wherein

the at least one processor is configured to process the instructions to control the image processing system to:

display an image frame including the moving object;

place a first icon on a background region of the image frame based on a first user input; and

place a second icon on an object region of the image frame based on a second user input.

5. An image processing method performed by at least one computer, the method comprising:

receiving an input of a plurality of image frames having different capturing times; and

performing one or more convolution calculations by using values of a background model of a neighboring region of one or more target pixels in the plurality of image frames to learn a parameter of a detection model for detecting a moving object, wherein

the plurality of background models include three background models, and these three background models are different from each other in the number of image frames to be processed;

the first background model has a largest number of image frames to be processed, the second background model has a second largest number of image frames to be processed, and the third background model has a smallest number of image frames to be processed;

the method comprises:

calculating a first distance between the first background model and the second background model, a second distance between the second background model and the third background model, and a third distance the first background model and the third background model for each pixel or unit region of the plurality of image frames;

determining, based on the first distance, the second distance, and the third distance, whether or not the moving object is shown in a region corresponding to each pixel of the plurality of image frames; and

based on user input specifying whether some pixels in the plurality of image frames are pixels that correspond to the moving object or a pixel that does not show the moving object, learning parameters for detecting the moving object in the plurality of image frames.

6. The image processing method according to claim 5 , wherein the method comprises:

learning, as the parameter, the one or more convolution calculations, and a threshold compared with a value obtained as a result of the one or more convolution calculations.

7. The image processing method according to claim 5 , wherein the method comprises:

identifying whether pixels that are selected at random from the plurality of image frames are a background region or an object region based on additional user inputs.

8. The image processing method according to claim 5 , wherein the method comprises:

displaying an image frame including the moving object;

placing a first icon on a background region of the image frame based on a first user input; and

placing a second icon on an object region of the image frame based on a second user input.

9. A non-transitory computer readable recording medium storing program instructions for causing a computer to perform:

receiving an input of a plurality of image frames having different capturing times; and

performing one or more convolution calculations by using values of a background model of a neighboring region of one or more target pixels in the plurality of image frames to learn a parameter of a detection model for detecting a moving object, wherein

the plurality of background models include three background models, and these three background models are different from each other in the number of image frames to be processed;

the first background model has a largest number of image frames to be processed, the second background model has a second largest number of image frames to be processed, and the third background model has a smallest number of image frames to be processed;

the program instructions causes the computer to perform:

calculating a first distance between the first background model and the second background model, a second distance between the second background model and the third background model, and a third distance the first background model and the third background model for each pixel or unit region of the plurality of image frames;

determining, based on the first distance, the second distance, and the third distance, whether or not the moving object is shown in a region corresponding to each pixel of the plurality of image frames; and

based on user input specifying whether some pixels in the plurality of image frames are pixels that reflect the moving object or a pixel that does not show the moving object, learning parameters for detecting the moving object in the plurality of image frames.

10. The non-transitory computer readable recording medium according to claim 9 , wherein the program instructions cause the computer to perform:

learning, as the parameter, the one or more convolution calculations, and a threshold compared with a value obtained as a result of the one or more convolution calculations.

11. The non-transitory computer readable recording medium according to claim 9 , wherein the program instructions cause the computer to perform:

identifying whether pixels that are selected at random from the plurality of image frames are a background region or an object region based on additional user inputs.

12. The non-transitory computer readable recording medium according to claim 9 , wherein the program instructions cause the computer to perform:

displaying an image frame including the moving object;

placing a first icon on a background region of the image frame based on a first user input; and

placing a second icon on an object region of the image frame based on a second user input.

Priority Claims (1)
JP 2014-115205 · Jun 3, 2014 · national
Continuity (4)
Continuation 18139111 · Apr 25, 2023
Continuation 16289745 · Mar 1, 2019
Continuation 15314572
Related Publication 20230394808A1 · Dec 7, 2023
References Cited (91)
US 6128396A · Hasegawa et al. · 2000 [cited by applicant]
US 6324532B1 · Spence · 2001 [cited by examiner]
US 7295700B2 · Schiller et al. · 2007 [cited by applicant]
US 7813528B2 · Porikli et al. · 2010 [cited by applicant]
US 8891864B2 · Pettigrew et al. · 2014 [cited by applicant]
US 9292929B2 · Hayata · 2016 [cited by applicant]
US 10192129B2 · Price et al. · 2019 [cited by applicant]
US 10540771B2 · Chefd'hotel et al. · 2020 [cited by applicant]
US 10586102B2 · Ren et al. · 2020 [cited by applicant]
US 20040125115A1 · Takeshima et al. · 2004 [cited by applicant]
US 20040202368A1 · Lee · 2004 [cited by examiner]
US 20050089216A1 · Schiller · 2005 [cited by examiner]
US 20080247599A1 · Porikli · 2008 [cited by examiner]
US 20090097709A1 · Tsukamoto · 2009 [cited by applicant]
US 20100119147A1 · Blake · 2010 [cited by examiner]
US 20100272363A1 · Steinberg et al. · 2010 [cited by applicant]
US 20110170769A1 · Sakimura et al. · 2011 [cited by applicant]
US 20110216976A1 · Rother · 2011 [cited by examiner]
US 20110293247A1 · Bhagavathy et al. · 2011 [cited by applicant]
US 20120210274A1 · Pettigrew et al. · 2012 [cited by applicant]
US 20120263346A1 · Datta et al. · 2012 [cited by applicant]
US 20120288153A1 · Tojo et al. · 2012 [cited by applicant]
US 20130050502A1 · Saito et al. · 2013 [cited by applicant]
US 20130071032A1 · Nishino · 2013 [cited by examiner]
US 20130084006A1 · Zhang et al. · 2013 [cited by applicant]
US 20130155229A1 · Thornton et al. · 2013 [cited by applicant]
US 20130342705A1 · Huang · 2013 [cited by examiner]
US 20140078170A1 · Ohki · 2014 [cited by applicant]
US 20140205141A1 · Gao et al. · 2014 [cited by applicant]
US 20140226052A1 · Kang et al. · 2014 [cited by applicant]
US 20140355882A1 · Hayata · 2014 [cited by applicant]
US 20150063689A1 · Datta et al. · 2015 [cited by applicant]
US 20160209927A1 · Yamagishi et al. · 2016 [cited by applicant]
US 20160328856A1 · Mannino et al. · 2016 [cited by applicant]
US 20170140236A1 · Price et al. · 2017 [cited by applicant]
US 20170344860A1 · Sachs et al. · 2017 [cited by applicant]
US 20180357472A1 · Dreessen · 2018 [cited by applicant]
US 20190236394A1 · Price et al. · 2019 [cited by applicant]
US 20200202533A1 · Cohen et al. · 2020 [cited by applicant]
US 20210027083A1 · Cohen et al. · 2021 [cited by applicant]
US 20230177824A1 · Price et al. · 2023 [cited by applicant]
JP H10285581A · 1998 [cited by applicant]
JP 2004180259A · 2004 [cited by applicant]
JP 2008092471A · 2008 [cited by applicant]
JP 2008257693A · 2008 [cited by applicant]
JP 2011059898A · 2011 [cited by applicant]
JP 2011145791A · 2011 [cited by applicant]
JP 2012059224A · 2012 [cited by applicant]
JP 2012104872A · 2012 [cited by applicant]
JP 2012517647A · 2012 [cited by applicant]
JP 5058010A · 2012 [cited by applicant]
JP 2012257173A · 2012 [cited by applicant]
JP 201358063A · 2013 [cited by applicant]
JP 2013065151A · 2013 [cited by applicant]
JP 2013125529A · 2013 [cited by applicant]
JP 2014059691A · 2014 [cited by applicant]
Ramesh, Visvanathan. “Background modeling and subtraction of dynamic scenes.” Proceedings ninth IEEE international conference on computer vision. IEEE, 2003. (Year: 2003). [cited by examiner]
Bai, Xue, and Guillermo Sapiro. “A geodesic framework for fast interactive image and video segmentation and matting.” 2007 IEEE 11th International Conference on Computer Vision. IEEE, 2007. (Year: 2007). [cited by examiner]
Akihiro Matsuda et al., “Automatic Extraction of Moving Object from Video Frames by Graph Cut”, The Institute of Electronics, Information and Communication Engineers, IEICE Technical Report, Jan. 14, 2010, Japan, vol. 1… [cited by applicant]
Communication dated May 28, 2019, from the Japanese Patent Office in counterpart Application No. 2016-525694. [cited by applicant]
Bai, Xue, et al., “A Geodesic Framework for Fast Interactive Image and Video Segmentation and Matting”, Dec. 26, 2007, 8 pages, IEEE Xplore, USA (Year:2007). [cited by applicant]
Japanese Office Action for JP Application No. 2020-209692 mailed on Dec. 21, 2021 with English translation. [cited by applicant]
Dan Ring et al., “Feature-Cut: Video object segmentation through local feature correspondences”. 2009 IEEE 12th International Conference on Computer Vision Workshops. ICCV Workshops, USA, IEEE. Sep. 27, 2009, pp. 617-62… [cited by applicant]
Xue Bai et al., “Distancecut: Interactive Segmentation and Matting of images and Videos”, 2007 IEEE international Conference on Image Processing, USA, IEEE, Sep. 16, 2007, pp. 249-252. [cited by applicant]
Satomi Kudo et al., “Accuracy Enhancement in Automatic Video Segmentation based on Graph-Cut using the SURF features”. IE1CE Technical Report, Japan, The Institute of Electronics, information and Communication Engineers… [cited by applicant]
Packt Video, OpenCV Tutorial: Segmenting an Image Using the Grabcut Algorithm, Screen shot 0.41s https://www.youtube.com/ watch?v=aOqOwM-Qb&t=31s, 2013 (Year: 2013). [cited by applicant]
US Notice of Allowance for U.S. Appl. No. 15/314,572 mailed on Jan. 22, 2021. [cited by applicant]
Yang et al. “Semi-supervised Learning of Feature Hierarchies for Object Detection in a Video”, pp. 1650-1657. [cited by applicant]
Nikolaos Doulamis et al. “Semi-supervised deep learning for object tracking and classification.” 2014 IEEE International Conference on Image Processing (ICIP). IEEE, 2014. [cited by applicant]
US Office Action for U.S. Appl. No. 15/314,572 mailed on Aug. 7, 2020. [cited by applicant]
Japanese Office Action for JP Application No. 2019-105217 mailed on Nov. 17, 2020 with English Translation. [cited by applicant]
Jiani Wang et al., “Study on the Creation of a Scenery Category Map using Geo-tagged Photographic Images”, Technical Report of IEICE, The Institute of Electronics, Information and Communication Engineers, Oct. 14, 2010,… [cited by applicant]
Japanese Office Action for JP Application No. 2019-105217 mailed on Sep. 15, 2020 with English Translation. [cited by applicant]
Atsushi Kawabata, Shinya Tanifuji, Yasuo Morooka, An Image Extraction Method for Moving Object, Information Processing Society of Japan, vol. 28, No. 4, pp. 395-402, 1987. [cited by applicant]
C. Stauffer and W.E.L. Crimson, “Adaptive background mixture models for real-time tracking”. Proceedings of CVPR, The Artificial Intelligence Laboratory, Massachusetts Institute of Technology, Cambridge, MA, USA vol. 2,… [cited by applicant]
International Search Report for PCT Application No. PCT/JP2015/002768, mailed on Aug. 11, 2015. [cited by applicant]
English translation of Written opinion for PCT Application No. PCT/JP2015/002768. [cited by applicant]
US Office Action for U.S. Appl. No. 15/314,572 mailed on Sep. 19, 2019. [cited by applicant]
US Office Action for U.S. Appl. No. 15/314,572 on Jun. 19, 2019. [cited by applicant]
US Office Action for U.S. Appl. No. 15/314,572 mailed on Mar. 8, 2019. [cited by applicant]
Japanese Office Action for JP Application No. 2019-105217 mailed on Jul. 7, 2020 with English Translation. [cited by applicant]
Communication dated Apr. 25, 2023 issued by the Japanese Intellectual Property Office in counterpart Japanese Application No. 2022-091955. [cited by applicant]
Minglun Gong et al., “Foreground Segmentation of Live Videos using Locally Competing 1SVMs”, IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2011, IEEE, Jun. 20, 2011, 8 pages total. [cited by applicant]
Weiwei Du et al., “Image and Video Matting by Membership Propagation”, IEICE Technical Report, Institute of Electronics, vol. 106, No. 428, Dec. 2006, 9 pages total. [cited by applicant]
JP Office Action for JP Application No. 2022-091955, mailed on Jul. 4, 2023 with English Translation. [cited by applicant]
Christoph Sommer et al., “Machine learning in cell biology-teaching computers to recognize phenotypes”, Journal of Cell Science, United Kingdom, The Company of Biologists Ltd, Dec. 15, 2013, vol. 126, No. 24, pp. 5529-5… [cited by applicant]
US Office Action for U.S. Appl. No. 18/236,772, mailed on Sep. 27, 2024. [cited by applicant]
Pei Xu et al., “Dynamic background learning through deep auto-encoder networks”, Proceedings of the 22nd ACM international conference on Multimedia (MM '14), pp. 107-116, 2014. (Year: 2014). [cited by applicant]
US Office Action for U.S. Appl. No. 18/139,111, mailed on Sep. 25, 2024. [cited by applicant]
Xu, Ning, et al. “Deep interactive object selection”. Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. (Year: 2016). [cited by applicant]
Liu, Ziwei, et al. “Semantic image segmentation via deep parsing network”. Proceedings of the IEEE international conference on computer vision. 2015. (Year: 2015). [cited by applicant]