Systems and methods for 2D to 3D conversion
According to some embodiments, a method includes accessing a video generated by a first physical camera in a physical environment. The method further includes identifying an object of interest in the video frame that corresponds to a physical object in the physical environment. The method further includes displaying a virtual 3D environment that corresponds to the physical environment. The method further includes configuring a plurality of settings of a first virtual camera to match a plurality of settings of a first physical camera and configuring a plurality of settings of a second virtual camera to match a plurality of settings of a second physical camera. The method further includes projecting the identified object of interest into the virtual 3D environment using the configured first and second virtual cameras.
1. A system comprising:
one or more memory units; and
one or more computer processors communicatively coupled to the one or more memory units and configured to:
access a video generated by a first physical camera located within a physical environment;
identify, by analyzing a video frame of the video, an object of interest in the video frame, the object of interest corresponding to a physical object that is physically located within the physical environment;
display, in a graphical user interface, a virtual three-dimensional (3D) environment that corresponds to the physical environment, the virtual 3D environment comprising:
a first virtual camera that corresponds to the first physical camera; and
a second virtual camera that corresponds to a second physical camera located within the physical environment;
access, from a calibrations database, a plurality of settings of the first physical camera and a plurality of settings of the second physical camera;
configure a plurality of settings of the first virtual camera to match the plurality of settings of the first physical camera such that a field of view of the first virtual camera in the virtual 3D environment is identical to a field of view of the first physical camera in the physical environment;
configure a plurality of settings of the second virtual camera to match the plurality of settings of the second physical camera such that a field of view of the second virtual camera in the virtual 3D environment is identical to a field of view of the second physical camera in the physical environment; and
project the identified object of interest into the virtual 3D environment using the configured first and second virtual cameras.
2. The system of claim 1 , the one or more memory units are further configured to:
create a first synthetic depth map for the first virtual camera, wherein the first synthetic depth map provides a one-to-one mapping of each 2D pixel created by the first physical camera and a corresponding space of the first virtual camera in the virtual 3D environment; and
create a second synthetic depth map for the second virtual camera, wherein the second synthetic depth map provides a one-to-one mapping of each 2D pixel created by the second physical camera and a corresponding space of the second virtual camera in the virtual 3D environment;
create a camera matrix comprising the first and second synthetic depth maps.
3. The system of claim 2 , wherein projecting the identified object of interest into the virtual 3D environment comprises:
determining a 2D coordinate of the identified object of interest; and
converting the 2D coordinate to a location within the virtual 3D environment using the camera matrix.
4. The system of claim 2 , wherein the camera matrix enables the first and second physical cameras to be aware of detection capabilities of other physical cameras using a shared unified coordinate system.
5. The system of claim 1 , wherein identifying the object of interest comprises utilizing a convolution neural network architecture.
6. The system of claim 1 , wherein:
the first virtual camera is placed in the virtual 3D environment at a same latitude, longitude, and altitude as the first physical camera is positioned in the physical environment; and
the second virtual camera is placed in the virtual 3D environment at a same latitude, longitude, and altitude as the second physical camera is positioned in the physical environment.
7. The system of claim 1 , wherein the settings of the first and second virtual cameras and the first and second physical cameras each comprise:
a surge setting;
a sway setting;
a heave setting;
a roll setting;
a pitch setting; and
a yaw setting.
8. The system of claim 7 , wherein the settings of the first and second virtual cameras and the first and second physical cameras further comprise:
a radial distortion; and
a tangential distortion.
9. A method by a computing system, the method comprising:
accessing a video generated by a first physical camera located within a physical environment;
identifying, by analyzing a video frame of the video, an object of interest in the video frame, the object of interest corresponding to a physical object that is physically located within the physical environment;
displaying, in a graphical user interface, a virtual three-dimensional (3D) environment that corresponds to the physical environment, the virtual 3D environment comprising:
a first virtual camera that corresponds to the first physical camera; and
a second virtual camera that corresponds to a second physical camera located within the physical environment;
accessing, from a calibrations database, a plurality of settings of the first physical camera and a plurality of settings of the second physical camera;
configuring a plurality of settings of the first virtual camera to match the plurality of settings of the first physical camera such that a field of view of the first virtual camera in the virtual 3D environment corresponds to a field of view of the first physical camera in the physical environment;
configuring a plurality of settings of the second virtual camera to match the plurality of settings of the second physical camera such that a field of view of the second virtual camera in the virtual 3D environment corresponds to a field of view of the second physical camera in the physical environment; and
projecting the identified object of interest into the virtual 3D environment using the configured first and second virtual cameras.
10. The method of claim 9 , further comprising:
creating a first synthetic depth map for the first virtual camera, wherein the first synthetic depth map provides a one-to-one mapping of each 2D pixel created by the first physical camera and a corresponding space of the first virtual camera in the virtual 3D environment;
creating a second synthetic depth map for the second virtual camera, wherein the second synthetic depth map provides a one-to-one mapping of each 2D pixel created by the second physical camera and a corresponding space of the second virtual camera in the virtual 3D environment; and
creating a camera matrix comprising the first and second synthetic depth maps.
11. The method of claim 10 , wherein projecting the identified object of interest into the virtual 3D environment comprises:
determining a 2D coordinate of the identified object of interest; and
converting the 2D coordinate to a location within the virtual 3D environment using the camera matrix.
12. The method of claim 9 , wherein identifying the object of interest comprises utilizing a convolution neural network architecture.
13. The method of claim 9 , wherein:
the first virtual camera is placed in the virtual 3D environment at a same latitude, longitude, and altitude as the first physical camera is positioned in the physical environment; and
the second virtual camera is placed in the virtual 3D environment at a same latitude, longitude, and altitude as the second physical camera is positioned in the physical environment.
14. The method of claim 9 , wherein the settings of the first and second virtual cameras and the first and second physical cameras each comprise:
a surge setting;
a sway setting;
a heave setting;
a roll setting;
a pitch setting; and
a yaw setting.
15. The method of claim 14 , wherein the settings of the first and second virtual cameras and the first and second physical cameras further comprise:
a radial distortion; and
a tangential distortion.
16. One or more computer-readable non-transitory storage media embodying instructions that, when executed by a processor, cause the processor to perform operations comprising:
accessing a video generated by a first physical camera located within a physical environment;
identifying, by analyzing a video frame of the video, an object of interest in the video frame, the object of interest corresponding to a physical object that is physically located within the physical environment;
displaying, in a graphical user interface, a virtual three-dimensional (3D) environment that corresponds to the physical environment, the virtual 3D environment comprising:
a first virtual camera that corresponds to the first physical camera; and
a second virtual camera that corresponds to a second physical camera located within the physical environment;
accessing, from a calibrations database, a plurality of settings of the first physical camera and a plurality of settings of the second physical camera;
configuring a plurality of settings of the first virtual camera to match the plurality of settings of the first physical camera such that a field of view of the first virtual camera in the virtual 3D environment corresponds to a field of view of the first physical camera in the physical environment;
configuring a plurality of settings of the second virtual camera to match the plurality of settings of the second physical camera such that a field of view of the second virtual camera in the virtual 3D environment corresponds to a field of view of the second physical camera in the physical environment; and
projecting the identified object of interest into the virtual 3D environment using the configured first and second virtual cameras.
17. The one or more computer-readable non-transitory storage media of claim 15 , the operations further comprising:
creating a first synthetic depth map for the first virtual camera, wherein the first synthetic depth map provides a one-to-one mapping of each 2D pixel created by the first physical camera and a corresponding space of the first virtual camera in the virtual 3D environment;
creating a second synthetic depth map for the second virtual camera, wherein the second synthetic depth map provides a one-to-one mapping of each 2D pixel created by the second physical camera and a corresponding space of the second virtual camera in the virtual 3D environment; and
creating a camera matrix comprising the first and second synthetic depth maps.
18. The one or more computer-readable non-transitory storage media of claim 17 , wherein projecting the identified object of interest into the virtual 3D environment comprises:
determining a 2D coordinate of the identified object of interest; and
converting the 2D coordinate to a location within the virtual 3D environment using the camera matrix.
19. The one or more computer-readable non-transitory storage media of claim 15 , wherein:
the first virtual camera is placed in the virtual 3D environment at a same latitude, longitude, and altitude as the first physical camera is positioned in the physical environment; and
the second virtual camera is placed in the virtual 3D environment at a same latitude, longitude, and altitude as the second physical camera is positioned in the physical environment.
20. The one or more computer-readable non-transitory storage media of claim 15 , wherein the settings of the first and second virtual cameras and the first and second physical cameras each comprise:
a surge setting;
a sway setting;
a heave setting;
a roll setting;
a pitch setting;
a yaw setting;
a radial distortion; and
a tangential distortion.