Enhanced three dimensional visualization using artificial intelligence
Apparatus and methods for enhanced 3D visualization includes receiving a plurality of images of an image scene from a plurality of image sensors. Depth information at locations of the image scene is received from a plurality of depth sensors. The depth information is combined with the plurality of images of the image scene using a machine learning model. A 3D representation of the image scene is generated based on the combined depth and image information.
1. A method for enhanced three dimensional (3D) visualization comprising:
receiving a plurality of frames of an image scene from a plurality of image sensors, wherein the plurality of frames includes an object of interest that is moving within the image scene;
receiving depth information at locations of the image scene from a plurality of depth sensors;
combining the depth information with the plurality of frames of the image scene using a machine learning model, wherein the machine learning model is configured to perform image enhancement using an interpolation technique, wherein the image enhancement includes using the depth information to generate artificially rendered frames of the image scene along a path between or beyond the plurality of frames; and
generating a 3D representation of the object of interest moving in the image scene, based on the plurality of frames of the image scene and the artificially rendered frames of the image scene.
2. The method of claim 1 , further comprising:
generating, from the depth information, a depth map identifying depth discontinuities.
3. The method of claim 1 , wherein generating the 3D representation of the image scene further comprises:
generating a polygon model of the object of interest, each polygon in the polygon model having a plurality of vertices; and
triangulating each polygon of the polygon model using vertex triangulation.
4. The method of claim 1 , further comprising isolating the object of interest from the image scene based on the depth information, wherein generating the 3D representation of the image scene comprises generating the 3D representation of the object of interest.
5. The method of claim 4 , further comprising determining a position of the object of interest based on the combining of the depth information and the plurality of frames.
6. The method of claim 1 , further comprising detecting motion activities within a predefined depth range of the image scene using the machine learning model.
7. The method of claim 1 , further comprising analyzing the 3D representation of the image scene and generating a notification to indicate a detected event based on the analyzing of the 3D representation of the image scene.
8. The method of claim 7 , wherein the detected event comprises detection of an unrecognized person and wherein the notification includes a 3D representation of a face of the unrecognized person.
9. The method of claim 7 , wherein the detected event comprises detection of an unattended object and wherein the notification includes a 3D representation of the unattended object.
10. The method of claim 1 , wherein the path follows a non-linear trajectory.
11. The method of claim 1 , wherein to generate the artificially rendered frames of the image scene along the path between or beyond the plurality of frames comprises generating an artificially rendered frame of a location outside of a trajectory between two frames.
12. A system for enhanced three dimensional (3D) visualization comprising a hardware processor configured to:
receive a plurality of frames of an image scene from a plurality of image sensors, wherein the plurality of frames includes an object of interest that is moving within the image scene;
receive depth information at locations of the image scene from a plurality of depth sensors;
combine the depth information with the plurality of images frames of the image scene using a machine learning model, wherein the machine learning model is configured to perform image enhancement using an interpolation technique, wherein the image enhancement includes using the depth information to generate artificially rendered frames of the image scene along a path between or beyond the plurality of frames; and
generate a 3D representation of the object of interest moving in the image scene, based on the plurality of frames of the image scene and the artificially rendered frames of the image scene.
13. The system of claim 12 , wherein the hardware processor is further configured to:
generate, from the depth information, a depth map identifying depth discontinuities.
14. The system of claim 12 , wherein the hardware processor configured to generate the 3D representation of the image scene is further configured to:
generate a polygon model of the object of interest, each polygon in the polygon model having a plurality of vertices; and
triangulate each polygon of the polygon model using vertex triangulation.
15. The system of claim 12 , wherein the hardware processor is further configured to isolate the object of interest from the image scene based on the depth information, and wherein the hardware processor configured to generate the 3D representation of the image scene is further configured to generate the 3D representation of the object of interest.
16. The system of claim 15 , wherein the hardware processor is further configured to determine a position of the object of interest based on the hardware processor combining the depth information and the plurality of frames.
17. The system of claim 12 , wherein the hardware processor is further configured to detect motion activities within a predefined depth range of the image scene using the machine learning model.
18. The system of claim 12 , wherein the hardware processor is further configured to analyze the 3D representation of the image scene and to generate a notification to indicate a detected event based on an analysis of the 3D representation of the image scene.
19. The system of claim 18 , wherein the detected event comprises detection of an unrecognized person and wherein the notification includes a 3D representation of a face of the unrecognized person.
20. The system of claim 18 , wherein the detected event comprises detection of an unattended object and wherein the notification includes a 3D representation of the unattended object.