Systems and methods for performing self-improving visual odometry
In an example method of training a neural network for performing visual odometry, the neural network receives a plurality of images of an environment, determines, for each image, a respective set of interest points and a respective descriptor, and determines a correspondence between the plurality of images. Determining the correspondence includes determining one or point correspondences between the sets of interest points, and determining a set of candidate interest points based on the one or more point correspondences, each candidate interest point indicating a respective feature in the environment in three-dimensional space). The neural network determines, for each candidate interest point, a respective stability metric and a respective stability metric. The neural network is modified based on the one or more candidate interest points.
1 . A method of training a neural network for performing visual odometry, the method comprising:
accessing, by the neural network implemented using one or more computer systems, a plurality of images of an environment;
determining, by the neural network based on the plurality of images, a plurality of three-dimensional points in the environment;
tracking, by the neural network, the plurality of three-dimensional points across the plurality of images;
selecting a subset of the three-dimensional points, wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:
determining a stability metric of the three-dimensional point based on (i) a number of images of the plurality of images that depict the three-dimensional point, and (ii) a re-projection error associated with the three-dimensional point, and
determining whether to select the three-dimensional point based on the stability metric; and
modifying the neural network based on the subset of the three-dimensional points,
wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:
classifying, based on the stability metric, the three-dimensional point as one of:
a first classification representing stable points,
a second classification representing unstable points, or
a third classification other than the first classification and the second classification,
wherein classifying the three-dimensional point as the first classification comprises:
determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to a threshold number, and
determining that the re-projection error associated with the three-dimensional point is less than or equal to a first threshold error level, and
wherein classifying the three-dimensional point as the second classification comprises:
determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to the threshold number, and
determining that the re-projection error associated with the three-dimensional point is greater than or equal to a second threshold error level, wherein the second threshold error level is different from the first threshold error level.
2 . The method of claim 1 , wherein selecting the subset of the three-dimensional points comprises:
selecting the three-dimensional points having the first classification or the second classification.
3 . The method of claim 2 , wherein selecting the subset of the three-dimensional points comprises:
refraining from selecting the three-dimensional points of the plurality of three-dimensional points having the third classification.
4 . The method of claim 1 , wherein classifying the three-dimensional point as the third classification comprises at least one of:
determining that the three-dimensional point is depicted in a number of images of the plurality of images that is less than the threshold number, or
determining that the re-projection error associated with three-dimensional point is between the first threshold error level and the second threshold error level.
5 . The method of claim 1 , wherein the plurality of images comprise two-dimensional images extracted from a video sequence.
6 . The method of claim 5 , wherein the plurality of images correspond to non-contiguous frames of the video sequence.
7 . The method of claim 1 , further comprising:
subsequent to modifying the neural network, receiving, by the neural network, a second plurality of images of a second environment from a head-mounted display device; and
determining, by the neural network based on the second plurality of images, a second plurality of three-dimensional points in the second environment.
8 . The method of claim 7 , wherein performing visual odometry with respect to the second environment comprises determining a position and orientation of the head-mounted display device using the second plurality of three-dimensional points as landmarks.
9 . A system comprising:
one or more processors;
one or more non-transitory computer-readable media including one or more sequences of instructions which, when executed by the one or more processors, causes the one or more processors to perform operations comprising:
accessing, by a neural network, a plurality of images of an environment;
determining, by the neural network based on the plurality of images, a plurality of three-dimensional points in the environment;
tracking, by the neural network, the plurality of three-dimensional points across the plurality of images;
selecting a subset of the three-dimensional points, wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:
determining a stability metric of the three-dimensional point based on (i) a number of images of the plurality of images that depict the three-dimensional point, and (ii) a re-projection error associated with the three-dimensional point, and
determining whether to select the three-dimensional point based on the stability metric; and
modifying the neural network based on the subset of the three-dimensional points,
wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:
classifying, based on the stability metric, the three-dimensional point as one of:
a first classification representing stable points,
a second classification representing unstable points, or
a third classification other than the first classification and the second classification,
wherein classifying the three-dimensional point as the first classification comprises:
determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to a threshold number, and
determining that the re-projection error associated with the three-dimensional point is less than or equal to a first threshold error level,
wherein classifying the three-dimensional point as the second classification comprises:
determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to the threshold number, and
determining that the re-projection error associated with the three-dimensional point is greater than or equal to a second threshold error level, wherein the second threshold error level is different from the first threshold error level.
10 . The system of claim 9 , wherein selecting the subset of the three-dimensional points comprises:
selecting the three-dimensional points having the first classification or the second classification.
11 . The system of claim 10 , wherein selecting the subset of the three-dimensional points comprises:
refraining from selecting the three-dimensional points of the plurality of three-dimensional points having the third classification.
12 . The system of claim 9 , wherein classifying the three-dimensional point as the third classification comprises at least one of:
determining that the three-dimensional point is depicted in a number of images of the plurality of images that is less than the threshold number, or
determining that the re-projection error associated with three-dimensional point is between the first threshold error level and the second threshold error level.
13 . The system of claim 9 , further comprising:
subsequent to modifying the neural network, receiving, by the neural network, a second plurality of images of a second environment from a head-mounted display device; and
determining, by the neural network based on the second plurality of images, a second plurality of three-dimensional points in the second environment.
14 . One or more non-transitory computer-readable media including one or more sequences of instructions which, when executed by one or more processors, causes the one or more processors to perform operations comprising:
accessing, by a neural network, a plurality of images of an environment;
determining, by the neural network based on the plurality of images, a plurality of three-dimensional points in the environment;
tracking, by the neural network, the plurality of three-dimensional points across the plurality of images;
selecting a subset of the three-dimensional points, wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:
determining a stability metric of the three-dimensional point based on (i) a number of images of the plurality of images that depict the three-dimensional point, and (ii) a re-projection error associated with the three-dimensional point, and
determining whether to select the three-dimensional point based on the stability metric; and
modifying the neural network based on the subset of the three-dimensional points,
wherein selecting the subset of the three-dimensional points comprises, for each of the three-dimensional points:
classifying, based on the stability metric, the three-dimensional point as one of:
a first classification representing stable points,
a second classification representing unstable points, or
a third classification other than the first classification and the second classification,
wherein classifying the three-dimensional point as the first classification comprises:
determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to a threshold number, and
determining that the re-projection error associated with the three-dimensional point is less than or equal to a first threshold error level, and
wherein classifying the three-dimensional point as the second classification comprises:
determining that the three-dimensional point is depicted in a number of images of the plurality of images that is greater than or equal to the threshold number, and
determining that the re-projection error associated with the three-dimensional point is greater than or equal to a second threshold error level, wherein the second threshold error level is different from the first threshold error level.