Depth map completion in visual content using semantic and three-dimensional information
Certain aspects of the present disclosure provide techniques for generating fine depth maps for images of a scene based on semantic segmentation and segment-based refinement neural networks. An example method generally includes generating, through a segmentation neural network, a segmentation map based on an image of a scene. The segmentation map generally comprises a map segmenting the scene into a plurality of regions, and each region of the plurality of regions is generally associated with one of a plurality of categories. A first depth map of the scene is generated through a first depth neural network based on a depth measurement of the scene. A second depth map of the scene is generated through a depth refinement neural network based on the segmentation map and the first depth map. One or more actions are taken based on the second depth map of the scene.
1. A method, comprising:
generating, through a segmentation neural network, a segmentation map based on an image of a scene, wherein the segmentation map comprises a plurality of segments, and wherein each segment of the plurality of segments is associated with one of a plurality of categories;
generating, through a first depth neural network, a first depth map of the scene based on a depth measurement of the scene;
generating a plurality of masks based on the segmentation map and the first depth map, each mask corresponding to one of the plurality of segments;
generating a plurality of enhanced depth masks based on the plurality of masks and three-dimensional coordinate information derived from the first depth map as inputs into a depth refinement neural network, wherein each enhanced depth mask of the plurality of enhanced depth masks corresponds to one of the plurality of segments;
combining the plurality of enhanced depth maps to form a second depth map of the scene;
taking one or more actions based on the second depth map of the scene.
2. The method of claim 1 , wherein the depth refinement neural network generates each output mask of the plurality of enhanced depth masks based on a convolutional kernel configured to apply weights based on a proximity of pixels in the scene in a three-dimensional space.
3. The method of claim 1 , wherein the depth measurement of the scene comprises depth measurements spanning a portion of a height of the image of the scene on a first axis and spanning a length of the image on a second axis.
4. The method of claim 1 , wherein the first depth map comprises a depth map of the scene having a lower resolution than a resolution of the image of the scene.
5. The method of claim 1 , wherein the taking one or more actions comprises controlling a motor vehicle based on the second depth map.
6. The method of claim 1 , wherein the taking one or more actions comprises:
generating an extended reality scene combining the scene and one or more virtual scenes based on the second depth map; and
displaying the extended reality scene on a display device.
7. The method of claim 1 , wherein the taking one or more actions comprises controlling a robot to interact with one or more physical objects in the scene based on the second depth map.
8. The method of claim 1 , wherein the first depth map comprises a coarse depth map and the second depth map comprises a finer depth map having finer detail than the coarse depth map.
9. An apparatus, comprising:
a memory having executable instructions stored thereon; and
a processor configured to execute the executable instructions to cause the apparatus to:
generate, through a segmentation neural network, a segmentation map based on an image of a scene, wherein the segmentation map comprises a plurality of segments, and wherein each segment of the plurality of segments being associated with one of a plurality of categories;
generate, through a first depth neural network, a first depth map of the scene based on a depth measurement of the scene;
generate a plurality of masks based on the segmentation map and the first depth map, each mask corresponding to one of the plurality of segments;
generate a plurality of enhanced depth masks based on the plurality of masks and three-dimensional coordinate information derived from the first depth map as inputs into a depth refinement neural network, wherein each enhanced depth mask of the plurality of enhanced depth masks corresponds to one of the plurality of segments;
combine the plurality of enhanced depth maps to form a second depth map of the scene;
take one or more actions based on the second depth map of the scene.
10. The apparatus of claim 9 , wherein the depth refinement neural network generates each output mask of the plurality of enhanced depth masks based on a convolutional kernel configured to apply weights based on a proximity of pixels in the scene in a three-dimensional space.
11. The apparatus of claim 9 , wherein the depth measurement of the scene comprises depth measurements spanning a portion of a height of the image of the scene on a first axis and spanning a length of the image on a second axis.
12. The apparatus of claim 9 , wherein the first depth map comprises a depth map of the scene having a lower resolution than a resolution of the image of the scene.
13. The apparatus of claim 9 , wherein in order to take the one or more actions, the processor is configured to control a motor vehicle based on the second depth map.
14. The apparatus of claim 9 , wherein in order to take the one or more actions, the processor is configured to:
generate an extended reality scene combining the scene and one or more virtual scenes based on the second depth map; and
display the extended reality scene on a display device.
15. The apparatus of claim 9 , wherein in order to take the one or more actions, the processor is configured to control a robot to interact with one or more physical objects in the scene based on the second depth map.
16. The apparatus of claim 9 , wherein the first depth map comprises a coarse depth map and the second depth map comprises a finer depth map having finer detail than the coarse depth map.
17. An apparatus, comprising:
means for generating, through a segmentation neural network, a segmentation map based on an image of a scene, wherein the segmentation map comprises a map segmenting the scene into a plurality of segments, each segment of the plurality of segments being associated with one of a plurality of categories;
means for generating, through a first depth neural network, a first depth map of the scene based on a depth measurement of the scene;
means for generating a plurality of masks based on the segmentation map and the first depth map, each mask corresponding to one of the plurality of segments;
means for generating a plurality of enhanced depth masks based on the plurality of masks and three-dimensional coordinate information derived from the first depth map as inputs into a depth refinement neural network, wherein each enhanced depth mask of the plurality of enhanced depth masks corresponds to one of the plurality of segments;
means for combining the plurality of enhanced depth maps to form a second depth map of the scene;
means for taking one or more actions based on the second depth map of the scene.
18. The apparatus of claim 17 , wherein the depth refinement neural network is configured to generate each output mask of the plurality of enhanced depth masks based on a convolutional kernel configured to apply weights based on a proximity of pixels in the scene in a three-dimensional space.
19. The apparatus of claim 17 , wherein the depth measurement of the scene comprises depth measurements spanning a portion of a height of the image of the scene on a first axis and spanning a length of the image on a second axis.
20. The apparatus of claim 17 , wherein the first depth map comprises a depth map of the scene having a lower resolution than a resolution of the image of the scene.
21. The apparatus of claim 17 , wherein the means for taking one or more actions comprises means for controlling a motor vehicle based on the second depth map.
22. The apparatus of claim 17 , wherein the means for taking one or more actions comprises:
means for generating an extended reality scene combining the scene and one or more virtual scenes based on the second depth map; and
means for displaying the extended reality scene on a display device.
23. The apparatus of claim 17 , wherein the means for taking one or more actions comprises means for controlling a robot to interact with one or more physical objects in the scene based on the second depth map.
24. The apparatus of claim 17 , wherein the first depth map comprises a coarse depth map and the second depth map comprises a finer depth map having finer detail than the coarse depth map.
25. A non-transitory computer-readable medium having executable instructions stored thereon which, when executed by a processor, causes the processor to:
generate, through a segmentation neural network, a segmentation map based on an image of a scene, wherein the segmentation map comprises a map segmenting the scene into a plurality of segments, each segment of the plurality of segments being associated with one of a plurality of categories;
generate, through a first depth neural network, a first depth map of the scene based on a depth measurement of the scene;
generate a plurality of masks based on the segmentation map and the first depth map, each mask corresponding to one of the plurality of segments;
generate a plurality of enhanced depth masks based on the plurality of masks and three-dimensional coordinate information derived from the first depth map as inputs into a depth refinement neural network, wherein each enhanced depth mask of the plurality of enhanced depth masks corresponds to one of the plurality of segments;
combine the plurality of enhanced depth maps to form a second depth map of the scene;
take one or more actions based on the second depth map of the scene.
26. The non-transitory computer-readable medium of claim 25 , wherein the depth refinement neural network generates each output mask of the plurality of enhanced depth masks based on a convolutional kernel configured to apply weights based on a proximity of pixels in the scene in a three-dimensional space.