Feature selection for object tracking using motion mask, motion prediction, or both
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for feature selection for object tracking. One of the methods includes: obtaining first feature points of an object in a first image of a scene captured by a camera; obtaining a second image of the scene captured by the camera after the first image was captured; determining whether a motion prediction of the object is available that indicates an area of the second image where the object is likely located; in response to determining that the motion prediction of the object is available, identifying, in the area of the second image where the object is likely located, second feature points that satisfy a similarity threshold for the first feature points in the first image; and detecting the object in the second image using the identified second feature points.
1 . A computer-implemented method comprising:
obtaining, by a camera, first feature points of an object in a first image of a scene captured by the camera;
obtaining, by the camera, a second image of the scene captured by the camera after the first image was captured;
determining, by the camera, whether a motion prediction of the object based on one or more images other than the second image is available and can be used to predict an area of the second image where the object is likely located;
in response to determining that the motion prediction of the object is available and can be used to predict the area of the second image where the object is likely located, predicting, by the camera and using (i) the motion prediction of the object and (ii) an area of the object in the first image, the area of the second image where the object is likely located;
identifying, by the camera and in the area of the second image where the object is likely located, second feature points that satisfy a similarity threshold for the first feature points in the first image;
detecting, by the camera, the object in the second image using the identified second feature points, the detecting comprising tracking, by the camera, the object across multiple images that include the first image and the second image based on the first feature points in the first image and the second feature points in the second image; and
performing, by the camera, one or more actions using a result of the detection of the object in the second image.
2 . The method of claim 1 , wherein obtaining the first feature points of the object in the first image comprises:
obtaining the first image of the scene captured by the camera;
identifying a bounding object around the object in the first image;
identifying areas of motion in the first image; and
selecting the first feature points that are within both the bounding object and the areas of motion in the first image.
3 . The method of claim 1 , comprising:
generating the motion prediction of the object using a Kalman filter algorithm.
4 . The method of claim 1 , comprising:
generating the motion prediction of the object using the first image and one or more images captured by the camera before the first image was captured.
5 . The method of claim 1 , wherein:
the motion prediction comprises a predicted trajectory of the object in the second image, and
identifying the second feature points that satisfy the similarity threshold for the first feature points in the first image comprises:
determining, using the predicted trajectory of the object in the second image, a candidate region of a bounding object around the object in the second image; and
identifying, in the candidate region of the bounding object, the second feature points that satisfy the similarity threshold for the first feature points in the first image.
6 . The method of claim 5 , wherein identifying, in the candidate region of the bounding object, the second feature points comprises:
determining, a search area: (i) centered at a center of the candidate region of the bounding object in the second image, and (ii) being larger than the candidate region by a predetermined percentage; and
identifying, in the search area, the second feature points.
7 . The method of claim 1 , wherein obtaining the first feature points of the object in the first image comprises:
identifying a bounding object around the object in the first image;
identifying areas of motion in the first image;
calculating an amount of motion in the bounding object using the area of motion in the first image;
determining whether the amount of motion in the bounding object satisfies a threshold; and
in response to determining that the amount of motion in the bounding object satisfies the threshold, selecting the first feature points from the areas of motion inside the bounding object in the first image.
8 . A system comprising one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:
obtaining first feature points of an object in a first image of a scene captured by a camera;
obtaining a second image of the scene captured by the camera after the first image was captured;
determining, by the camera, whether a motion prediction of the object based on one or more images other than the second image is available and can be used to predict an area of the second image where the object is likely located;
in response to determining that the motion prediction of the object is available and can be used to predict the area of the second image where the object is likely located, predicting, by the camera and using (i) the motion prediction of the object and (ii) an area of the object in the first image, the area of the second image where the object is likely located;
identifying, by the camera and in the area of the second image where the object is likely located, second feature points that satisfy a similarity threshold for the first feature points in the first image;
detecting the object in the second image using the identified second feature points;
obtaining a third image of the scene captured by the camera after the first image was captured;
determining whether the motion prediction of the object can be used to predict an area of the third image where the object is likely located;
in response to determining that the motion prediction of the object cannot be used to predict the area of the third image where the object is likely located, identifying, in the third image, third feature points that satisfy the similarity threshold for the first feature points in the first image;
detecting the object in the third image using the identified third feature points; and
performing, by the camera, one or more actions using a result of the detection of the object in the third image.
9 . The system of claim 8 , wherein obtaining the first feature points of the object in the first image comprises:
obtaining the first image of the scene captured by the camera;
identifying a bounding object around the object in the first image;
identifying areas of motion in the first image; and
selecting the first feature points that are within both the bounding object and the areas of motion in the first image.
10 . The system of claim 8 , the operations comprise:
generating the motion prediction of the object using a Kalman filter algorithm.
11 . The system of claim 8 , the operations comprise:
generating the motion prediction of the object using the first image and one or more images captured by the camera before the first image was captured.
12 . The system of claim 8 , wherein the motion prediction comprises a predicted trajectory of the object in the second image, wherein identifying the second feature points that satisfy the similarity threshold for the first feature points in the first image comprises:
determining, using the predicted trajectory of the object in the second image, a candidate region of a bounding object around the object in the second image; and
identifying, in the candidate region of the bounding object, the second feature points that satisfy the similarity threshold for the first feature points in the first image.
13 . The system of claim 8 , wherein obtaining the first feature points of the object in the first image comprises:
identifying a bounding object around the object in the first image;
identifying areas of motion in the first image;
calculating an amount of motion in the bounding object using the area of motion in the first image;
determining whether the amount of motion in the bounding object satisfies a threshold; and
in response to determining that the amount of motion in the bounding object satisfies the threshold, selecting the first feature points from the areas of motion inside the bounding object in the first image.
14 . A non-transitory computer storage medium encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations comprising:
obtaining, by a camera, first feature points of an object in a first image of a scene captured by the camera;
obtaining, by the camera, a second image of the scene captured by the camera after the first image was captured;
determining, by the camera, whether a motion prediction of the object based on one or more images other than the second image is available and can be used to predict an area of the second image where the object is likely located;
in response to determining that the motion prediction of the object is available and can be used to predict the area of the second image where the object is likely located, predicting, by the camera and using (i) the motion prediction of the object and (ii) an area of the object in the first image, the area of the second image where the object is likely located;
identifying, by the camera and in the area of the second image where the object is likely located, second feature points that satisfy a similarity threshold for the first feature points in the first image;
detecting, by the camera, the object in the second image using the identified second feature points, the detecting comprising tracking, by the camera, the object across multiple images that include the first image and the second image based on the first feature points in the first image and the second feature points in the second image; and
performing, by the camera, one or more actions using a result of the detection of the object in the second image.
15 . The non-transitory computer storage medium of claim 14 , wherein obtaining the first feature points of the object in the first image comprises:
obtaining the first image of the scene captured by the camera;
identifying a bounding object around the object in the first image;
identifying areas of motion in the first image; and
selecting the first feature points that are within both the bounding object and the areas of motion in the first image.
16 . The non-transitory computer storage medium of claim 14 , the operations comprise:
generating the motion prediction of the object using a Kalman filter algorithm.
17 . The non-transitory computer storage medium of claim 14 , the operations comprise:
generating the motion prediction of the object using the first image and one or more images captured by the camera before the first image was captured.
18 . The non-transitory computer storage medium of claim 14 , the operations comprise:
obtaining a third image of the scene captured by the camera after the first image was captured;
determining whether the motion prediction of the object is available that indicates an area of the third image where the object is likely located;
in response to determining that the motion prediction of the object is not available, identifying, in the third image, third feature points that satisfy the similarity threshold for the first feature points in the first image; and
detecting the object in the third image using the identified third feature points.
19 . The non-transitory computer storage medium of claim 14 , wherein the motion prediction comprises a predicted trajectory of the object in the second image, wherein identifying the second feature points that satisfy the similarity threshold for the first feature points in the first image comprises:
determining, using the predicted trajectory of the object in the second image, a candidate region of a bounding object around the object in the second image; and
identifying, in the candidate region of the bounding object, the second feature points that satisfy the similarity threshold for the first feature points in the first image.
20 . The non-transitory computer storage medium of claim 14 , wherein obtaining the first feature points of the object in the first image comprises:
identifying a bounding object around the object in the first image;
identifying areas of motion in the first image;
calculating an amount of motion in the bounding object using the area of motion in the first image;
determining whether the amount of motion in the bounding object satisfies a threshold; and
in response to determining that the amount of motion in the bounding object satisfies the threshold, selecting the first feature points from the areas of motion inside the bounding object in the first image.