Systems and methods for combining multiple depth maps
Certain embodiments of the present disclosure are relating to devices, systems and/or methods that may be used for determining scene information using depth information obtained from more than one point of view and combining that depth information to improve the accuracy of the scene information.
1 . A method of combining depth information from at least two depth maps to produce refined depth information comprising:
a. providing at least two source depth maps of a scene from at least two points of view and forwarding the at least two source depth maps to a computer system configured to:
b. receive the at least two source depth maps;
c. select a target viewpoint in order to generate a target depth map;
d. select at least a first selected location on the target depth map, and locate corresponding locations and depth values in at least one of the at least two source depth maps and by using zero or at least one of the corresponding depth values determine at least one refined depth value for the first selected location;
e. repeat the process in step (d) for at least a substantial portion of locations in the target depth map to determine refined depth values for at least the substantial portion of locations in the target depth map;
f. repeat steps (c), (d), and (e) to produce at least two refined target depth maps;
g. repeat steps (b), (c), (d), (e), and (f) using the at least two refined target depth maps as the at least two source depth maps to further improve the at least two refined target depth maps;
h. repeat step (g) until a termination condition is reached; and
i. output the refined depth information representative of the scene.
2 . The method of claim 1 , wherein at least two depth sensors are used to collect the at least two source depth maps.
3 . The method of claim 1 , wherein at least three depth sensors are used.
4 . The method of claim 1 , wherein at least four depth sensors are used.
5 . The method of claim 1 , wherein at least three source depth maps are generated of the scene from at least three viewpoints.
6 . The method of claim 1 , wherein at least four source depth maps are generated of the scene from at least four viewpoints.
7 . The method of claim 1 , wherein the at least two source depth maps is one or more of the following: the at least two source depth maps; the at least two refined target depth maps from a previous iteration; the at least two source depth maps used in at least one previous iteration; the at least two refined target depth maps produced by at least one previous iteration; and one or more additional depth maps available during a current iteration.
8 . The method of claim 1 , wherein a number of target depth maps produced at each iteration is at least two, irrespective of a number of at least two source depth maps.
9 . The method of claim 1 , wherein a number of target depth maps produced in steps (c) and (d) of claim 1 is at least two, irrespective of a number of source depth maps used in an iteration of said steps (c) and (d).
10 . The method of claim 2 , wherein the at least two source depth maps are subject to a pre-processing step at least one time so that depth information within the at least two source depth maps is transformed to generate transformed depth information that is representative of the at least two depth sensors being in a different location to their actual location in 3D space.
11 . The method of claim 1 , wherein the determined at least one refined depth value for a location is no value.
12 . The method of claim 1 , wherein portions of at least one of the at least two refined target depth maps have an improved accuracy.
13 . The method of claim 1 , wherein substantial portions of at least one of the at least two refined target depth maps have an improved accuracy.
14 . The method of claim 1 , wherein the target depth map is selected to have a centre that is substantially coincident with a centre of at least one depth sensor.
15 . The method of claim 14 , wherein the target depth map is selected to have a centre that is not substantially coincident with any centres of the at least one depth sensor.
16 . The method of claim 1 , wherein at least one of the source depth maps is transformed so that depth information is representative of depth information from a depth sensor whose centre in 3D space is not substantially coincident with any centres of at least one sensor centre.
17 . The method of claim 1 , wherein at least one of the source depth maps is transformed so that depth information is representative of depth information from a depth sensor whose centre in 3D space is substantially coincident with at least one sensor centre.
18 . The method of claim 1 , wherein the determination of the at least one refined depth value is based on one or more of the following: no value, a value from a corresponding location on one or more source depth maps, a value from earlier refined target depth maps, a value from the refined target depth map being constructed, a value from neighbouring cells on earlier refined depth maps, a value from neighbouring cells on the refined target depth map being constructed, and a value from neighbouring cells on one or more source depth maps.
19 . The method of claim 18 , wherein the at least one refined depth value is determined using one or more of the following processes on corresponding depth data points: a numerical average (mean), a trimmed mean, a median, a trimmed median, a weighted mean, a weighted median, a bootstrapped median, and a bootstrapped mean.
20 . The method of claim 1 , wherein the output of the refined depth information is one or more of the following: at least one refined target depth map, at least one point cloud, information representative of the scene that is isomorphic to the refined target depth map, information representative of the scene that is isomorphic to a final refined target depth map, information representative of the scene that is isomorphic to a plurality of refined target depth maps, information representative of the scene that is isomorphic to a plurality of the final refined target depth maps, and a surface mesh that consists at least in part of connected triangles or other geometric models of surfaces in the scene.
21 . One or more computer-readable non-transitory storage media embodying software that is operable when executed to implement claim 1 .