IP Library Patent Application 16575297
Patent Application
App. No. 16/575,297

REGION PROPOSAL WITH TRACKER FEEDBACK

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/575,297
Abstract

Methods, systems, and techniques for object detection and tracking are provided. A system may include a module configured to generate a plurality of region proposals, each region proposal comprising a part of a video frame, a CNN pre-trained for object detection, the plurality of region proposals being input to the CNN; a tracker for tracking one or more targets based on outputs from the CNN across the series of video frames and generating tracking information on the one or more targets; and a module further configured to refine the plurality of region proposals to be input to the CNN, based on the tracking information.

Claims (41)

1 . A method comprising:

generating a plurality of region proposals, each region proposal comprising a part of a video frame, the plurality of region proposals being input to a convolutional neural network (CNN) pre-trained for object detection;

detecting, using the CNN, one or more objects in a series of video frames;

tracking one or more targets based on outputs from the CNN across the series of video frames and generating tracking information on the one or more targets; and

refining the plurality of region proposals to be input to the CNN, based on the tracking information.

2 . The method of claim 1 , wherein the outputs from the CNN comprise a bounding box and a classification score for each detected object, wherein each bounding box is defined by a location and vertical and horizontal dimensions.

3 . The method of claim 1 , further comprising:

categorizing each of the one or more targets into a status category, based on the outputs from the CNN, the status category of a target indicating the time since the target was likely detected by the CNN.

4 . The method of claim 3 , further comprising:

for each of the one or more targets, identifying a region of the region proposals likely containing the target or a new region likely containing the target.

5 . The method of claim 4 , further comprising:

calculating a region priority score for each of the plurality of region proposals based on a priority score of the target that is likely within the identified region proposal, the priority score of the target being determined based on the corresponding status category.

6 . The method of claim 5 , wherein refining the region proposals comprises:

sorting the plurality of region proposals in a descending order by the region priorities scores, and selecting Nmax regions as final region proposals to be input to the CNN, wherein Nmax represents an upper threshold number.

7 . The method of claim 1 , wherein the region proposals include a non-zero motion vector and are selected from a plurality of predefined regions covering the frame.

8 . The method of claim 7 , wherein the predefined regions are generated from an object size map, the object size map's value for a given location in the object size map representing an estimated object size in pixels.

9 . The method of claim 7 , wherein the total number of the region proposals satisfies an upper threshold number criterion.

10 . The method of claim 7 , wherein generating region proposals comprises:

adding an additional region to the region proposals until the number of the region proposals satisfies an upper threshold number criterion.

11 . The method of claim 10 , wherein the additional region is determined using a default region or a last checking time map, the last checking time map describing the time since the local region was provided to the CNN.

12 . The method of claim 9 , wherein generating region proposals comprises:

merging at least two of the selected region proposals based on a motion vector density, the motion vector density defined as a percentage of pixels inside of a region proposal that have non-zero motion vectors.

13 . The method of claim 2 , further comprising:

for each of the one or more targets, identifying a region of the region proposals in which the target is likely contained based on the corresponding bounding box.

14 . The method of claim 13 , further comprising:

creating a new region for a target that is not likely within any of the region proposals and is likely within the new region.

15 . The method of claim 14 , further comprising:

categorizing each of the one or more targets into a status category, based on the outputs from the CNN, the status category of a target indicating the time since the target was likely detected by the CNN; and

calculating a region priority score for each region that likely contains a target based on the priority score of the target, the priority score of the target based on the corresponding status category.

16 . The method of claim 15 , wherein refining the region proposals comprises:

sorting regions including any region that likely contains a target and the region proposals, in a descending order of the region priority scores, and selecting Nmax regions as final proposal regions, wherein Nmax represents an upper threshold number.

17 . A computer readable medium storing instructions, which when executed by a computer cause the computer to perform a method comprising:

generating a plurality of region proposals, each region proposal comprising a part of a video frame, the plurality of region proposals being input to a CNN pre-trained for object detection;

detecting, using the CNN, one or more objects in a series of video frames;

tracking one or more targets based on outputs from the CNN across the series of video frames and generating tracking information on the one or more targets; and

refining the plurality of region proposals to be input to the CNN, based on the tracking information.

18 . A system comprising:

a module for generating a plurality of region proposals, each region proposal comprising a part of a video frame;

a CNN pre-trained for object detection, the plurality of region proposals being input to the CNN;

a tracker for tracking one or more targets based on outputs from the CNN across the series of video frames and generating tracking information on the one or more targets; and

a module further configured to refine the plurality of region proposals to be input to the CNN, based on the tracking information.

Assignments (2)
NUNC PRO TUNC ASSIGNMENT Recorded Aug 31, 2022
From: AVIGILON CORPORATION
To: MOTOROLA SOLUTIONS, INC.
Reel/Frame 061361/0905 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2020
From: LIPCHIN, ALEKSEY; LEE, CHIA YING; WANG, YIN; XIAO, XIAO; ZHANG, HAO
To: AVIGILON CORPORATION
Reel/Frame 053394/0473 →