IP Library Granted Patent US 12,307,756
Granted Patent B2
US 12,307,756 · App. 17/663,035 · Granted May 20, 2025

Fast object detection in video via scale separation

Inventors: Ishay Goldin (Tel-Aviv, IL); Netanel Stein (Tel-Aviv, IL); Alexandra Dana (Tel-Aviv, IL); Alon Intrater (Tel-Aviv, IL); David Tsidkiahu (Tel-Aviv, IL); Nathan Levy (Tel-Aviv, IL); Omer Shabtai (Tel-Aviv, IL); Ran Vitek (Tel-Aviv, IL); Tal Heller (Tel-Aviv, IL); Yaron Ukrainitz (Tel-Aviv, IL); Yotam Platner (Tel-Aviv, IL); Zuf Pilosof (Tel-Aviv, IL)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06V10/96G06T3/4046G06V10/774G06V10/82G06V20/41G06V20/49G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,756
App. No.
17/663,035
Granted
May 20, 2025
Kind
B2
Abstract

Techniques and apparatuses enabling high accuracy video object detection using reduced system resource requirements (e.g., reduced computational load, shallower neural network designs, etc.) are described. For example, a search domain of an object detection scheme (e.g., a target object class, a target object size, a target object rotation angle, etc.) may be separated into subdomains (e.g., such as subdomains of object classes, subdomains of object sizes, subdomains object rotation angles, etc.). Specialized, subdomain-level object detection/segmentation tasks may then be separated across sequential video frames. As such, different subdomain-level processing techniques (e.g., via specialized neural networks) may be implemented across different frames of a video sequence. Moreover, redundancy information of consecutive video frames may be leveraged, such that specialized object detection tasks combined with visual object tracking across consecutive frames may enable more efficient (e.g., more accurate, less computationally intensive, etc.) full domain object detection and object segmentation schemes.

Claims (64)

1. A method for image processing, comprising:

dividing a video into a plurality of frames including a first frame and a second frame;

selecting the first frame for processing with a first neural network and the second frame for processing with a second neural network based on an ordering of the plurality of frames;

generating first object information for a first object in the first frame by performing a first computer vision task using the first neural network;

generating second object information for a second object in the second frame by performing a second computer vision task using the second neural network; and

generating label data for the video including a first label based on the first object information and a second label based on the second object information.

2. The method of claim 1 , further comprising:

alternately applying the first neural network to a first subset of the plurality of frames to obtain first label data corresponding to the first computer vision task and applying the second neural network to a second subset of the plurality of frames to obtain second label data corresponding to the second computer vision task, wherein the first subset and the second subset alternate according to a temporal ordering of the plurality of frames.

3. The method of claim 1 , further comprising:

reducing a resolution of the second frame to obtain a low-resolution version of the second frame, wherein the second neural network takes the low-resolution version as input and the first neural network takes a high-resolution version of the first frame as input.

4. The method of claim 1 , further comprising:

detecting a third object in a third frame of the video using a third neural network trained to detect objects by performing a third computer vision task different from the first computer vision task and second computer vision task.

5. The method of claim 1 , wherein:

the first computer vision task generates object information for objects that are small in size relative to objects which have object information generated by the second computer vision task.

6. The method of claim 1 , wherein:

the first computer vision task and the second computer vision task overlap,

the first frame is a high-resolution frame,

the first computer vision task is performed based on the first frame to detect objects that are small in size,

the second frame is a low-resolution frame, and

the second computer vision task is performed based on the second frame to detect objects that are large in size.

7. The method of claim 1 , wherein:

the second computer vision task is different from the first computer vision task.

8. The method of claim 1 , further comprising:

identifying available system resources for generating the label data; and

selecting a frequency for applying the first neural network and the second neural network based on the available system resources.

9. A method for training a machine learning model, the method comprising:

identifying first training data corresponding to a first computer vision task;

training a first neural network to perform the first computer vision task using the first training data;

identifying second training data corresponding to a second computer vision task;

training a second neural network to perform the second computer vision task using the second training data; and

alternately applying the first neural network and the second neural network to different frames of a video to generate label data corresponding to the first computer vision task and the second computer vision task, respectively.

10. The method of claim 9 , wherein:

the second training data has a lower resolution than the first training data.

11. The method of claim 9 , wherein:

the first neural network and the second neural network are trained independently.

12. The method of claim 9 , further comprising:

identifying third training data corresponding to a third computer vision task;

training a third neural network to perform the third computer vision task using the third training data; and

alternately applying the third neural network to the frames of the video to generate the label data, wherein the label data includes data corresponding to the third computer vision task.

13. The method of claim 9 , wherein:

the first training data includes first label data for objects that are small in size relative to objects for which label data is generated by the second computer vision task.

14. The method of claim 9 , wherein:

the first training data and the second training data overlap,

the first computer vision task is performed based on a first frame which is a high-resolution frame to detect objects that are small in size, and

the second computer vision task is performed based on a second frame which is a low-resolution frame to detect objects that are large in size.

15. A system for image processing, comprising:

a frame selection component to divide a video into a plurality of frames including a first frame and a second frame based on a temporal ordering of the plurality of frames;

a first neural network configured to generate first object information for a first object in the first frame by performing a first computer vision task;

a second neural network configured to generate second object information for a second object in the second frame by performing a second computer vision task; and

a data synthesis component configured to generate label data for the video including a first label based on the first object information and a second label based on the second object information.

16. The system of claim 15 , further comprising:

a third neural network configured to generate third object information for a third object in a third frame, wherein the third neural network is trained to perform a third computer vision task different from the first computer vision task and second computer vision task.

17. The system of claim 15 , further comprising:

a resolution reduction component configured to reduce a resolution of the second frame to obtain a low-resolution version of the second frame, wherein the second neural network takes the low-resolution version as input and the first neural network takes a high-resolution version of the first frame as input.

18. The system of claim 15 , further comprising:

a resource optimization component configured to identify available system resources for generating the label data and to select a frequency for applying the first neural network and the second neural network based on the available system resources.

19. The system of claim 15 , wherein:

the second neural network has fewer parameters than the first neural network,

the first frame is a high-resolution frame,

the first computer vision task is performed based on the first frame to detect objects that are small in size,

the second frame is a low-resolution frame, and

the second computer vision task is performed based on the second frame to detect objects that are large in size.

20. The system of claim 15 , further comprising:

a training component configured to train the first neural network and the second neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2022
From: GOLDIN, ISHAY; STEIN, NETANEL; DANA, ALEXANDRA; INTRATER, ALON; TSIDKIAHU, DAVID; LEVY, NATHAN; SHABTAI, OMER; VITEK, RAN; HELLER, TAL; UKRAINITZ, YARON; PLATNER, YOTAM; PILOSOF, ZUF
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 059986/0698 →
Continuity (1)
Related Publication 20230368520A1 · Nov 16, 2023
References Cited (46)
US 8977000B2 · Yano · 2015 [cited by applicant]
US 10628961B2 · Sundaresan · 2020 [cited by examiner]
US 10701394B1 · Caballero et al. · 2020 [cited by applicant]
US 10733714B2 · El-Khamy · 2020 [cited by examiner]
US 10891525B1 · Olgiati et al. · 2021 [cited by applicant]
US 10984246B2 · Ramaswamy et al. · 2021 [cited by applicant]
US 20020102024A1 · Jones et al. · 2002 [cited by applicant]
US 20180225820A1 · Liang · 2018 [cited by examiner]
US 20190251702A1 · Chandler · 2019 [cited by examiner]
US 20190311202A1 · Lee · 2019 [cited by examiner]
US 20200293783A1 · Ramaswamy · 2020 [cited by examiner]
US 20200293828A1 · Wang · 2020 [cited by examiner]
US 20200349686A1 · Shapovalova et al. · 2020 [cited by applicant]
US 20200364554A1 · Wang · 2020 [cited by examiner]
US 20210012503A1 · Cho · 2021 [cited by examiner]
US 20210046861A1 · Li · 2021 [cited by examiner]
US 20210073525A1 · Weinzaepfel · 2021 [cited by examiner]
US 20210100526A1 · Schein · 2021 [cited by examiner]
US 20210149441A1 · Bartscherer · 2021 [cited by examiner]
US 20210152739A1 · Lu et al. · 2021 [cited by applicant]
US 20210213973A1 · Carillo Peña · 2021 [cited by examiner]
US 20210374947A1 · Shin · 2021 [cited by examiner]
US 20220035684A1 · Gupte · 2022 [cited by examiner]
US 20220055689A1 · Mandlekar · 2022 [cited by examiner]
US 20220215232A1 · Pardeshi · 2022 [cited by examiner]
US 20220238012A1 · Ghadiok · 2022 [cited by examiner]
US 20220269888A1 · Stoeva · 2022 [cited by examiner]
US 20220414928A1 · Venkataraman · 2022 [cited by examiner]
US 20230021926A1 · Zhao · 2023 [cited by examiner]
US 20230042004A1 · Kumar · 2023 [cited by examiner]
JP 2021072592 · 2021 [cited by applicant]
WO 2014101036 · 2017 [cited by applicant]
Xiang Yan et al., “Self-Supervised Learning to Detect Key Frames in Videos,” Dec. 4, 2020, Sensors 2020, 20, 6941; doi: 10.3390/s20236941, pp. 1-14. [cited by examiner]
Chen Zhang et al.,“Video Object Detection With Two-Path Convolutional LSTM Pyramid,” Aug. 28, 2020, IIEEE Access, vol. 8, 2020,pp. 151681-151689. [cited by examiner]
Chun-Han Yao et al.,“Video Object Detection via Object-Level Temporal Aggregation,” Nov. 13, 2020, Computer Vision—ECCV 2020 ,pp. 160-166. [cited by examiner]
Mason Liu et al.,“Looking Fast and Slow: Memory-Guided Mobile Video Object Detection,” Mar. 25, 2019, pararXiv: 1903.10172v1 [cs.CV],pp. 1-8. [cited by examiner]
Hao Luo et al.,“Detect or Track: Towards Cost-Effective Video Object Detection/Tracking,” Jul. 17, 2029, The Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), vol. 33 No. 01: AAAI-19, IAAI-19, EAAI-20 ,… [cited by examiner]
Ozgur Erkent et al.,“Semantic Grid Estimation with a Hybrid Bayesian and Deep Neural Network Approach,” Jan. 6, 2019,2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) Madrid, Spain, Oct. 1-… [cited by examiner]
Hao Long et al.,“Object Detection in Aerial Images Using Feature Fusion Deep Networks,” Mar. 25, 2019,IEEE Access , vol. 7,2019, pp. 30980-30988. [cited by examiner]
Kai Chen et al.,“Optimizing Video Object Detection via a Scale-Time Lattice,” Jun. 2018, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018,pp. 7814-7820. [cited by examiner]
Chin, et al., “Adascale: Towards Real-Time Video Object Detection Using Adaptive Scaling”, Proceedings of the 2nd SysML Conference, arXiv preprint arXiv:1902.02910v1 [cs.CV] Feb. 8, 2019 (18 pages). [cited by applicant]
Luo, et al., “Detect or Track: Towards Cost-Effective Video Object Detection/Tracking”, arXiv preprint arXiv:1811.05340v1 [cs.CV] Nov. 13, 2018, 9 pages. [cited by applicant]
Yang, et al., “Face Detection through Scale-Friendly Deep Convolutional Networks” arXiv preprint arXiv:1706.02863v1 [cs.CV] Jun. 9, 2017, 12 pages. [cited by applicant]
Chen, et al., “Optimizing Video Object Detection via a Scale-Time Lattice”, arXiv preprint arXiv:1804.05472v1 [cs.CV] Apr. 16, 2018 (10 pages). [cited by applicant]
Viola, et al., “Rapid Object Detection using a Boosted Cascade of Simple Features”, Accepted Conference on Computer Vision and Pattern Recognition 2001, pp. 1-9. [cited by applicant]
Yau, et al., “Video Object Detection via Object-level Temporal Aggregation” (17 pages). [cited by applicant]