METHODS AND APPARATUS FOR SCALE RECOVERY FROM MONOCULAR VIDEO
Methods, apparatus, systems and articles of manufacture are disclosed for scale recovery from monocular video. An example non-transitory computer readable medium comprises instructions that, when executed, cause a machine to at least segment an input image from a monocular video to detect an object in the camera field, estimate camera parameters from the segmented input image, iteratively refine the estimated camera parameters using known object heights, calculate a scale for the video, iteratively refine the scale based on a user input, and report the scaling results for visualization.
1 .- 25 . (canceled)
26 . A non-transitory computer readable medium comprising instructions that, when executed, cause a machine to at least:
segment an input image from a monocular video to detect an object in a camera field;
estimate camera parameters from the segmented input image;
iteratively refine the estimated camera parameters using object heights;
calculate a scale for the video;
iteratively refine the scale based on a user input; and
report scaling results for visualization.
27 . The non-transitory computer readable medium of claim 1 , wherein the input image segmentation is performed using a segmentation backbone network.
28 . The non-transitory computer readable medium of claim 1 , wherein the video scale is calculated using a first and second camera parameter.
29 . The non-transitory computer readable medium of claim 3 , wherein the first and second camera parameters are adjusted according to a projection model.
30 . The non-transitory computer readable medium of claim 1 , wherein the scaling results are reported via a graphical user interface.
31 . The non-transitory computer readable medium of claim 1 , wherein the object heights are used to train a branch of a neural network model.
32 . The non-transitory computer readable medium of claim 6 , wherein the branch of the neural network model is trained to adjust at least one of a first camera parameter or a second camera parameter.
33 . The non-transitory computer readable medium of claim 1 , wherein the user input for iterative scale refinement is provided via a graphical user interface.
34 . An apparatus to recover scale from monocular video comprising:
interface circuitry;
machine readable instructions; and
programmable circuitry to at least one of instantiate or execute the machine-readable instructions to:
segment an input image from the monocular video to detect an object in a camera field;
estimate camera parameters from the segmented input image; and
iteratively refine the estimated camera parameters using object heights;
calculate a scale for the video;
iteratively refine the scale based on a user input; and
report scaling results for visualization.
35 . The apparatus of claim 9 , wherein the input image segmentation is performed using a segmentation backbone network.
36 . The apparatus of claim 9 , wherein the video scale is calculated using a first and second camera parameter.
37 . The apparatus of claim 11 , wherein the first and second camera parameters are adjusted according to a projection model.
38 . The apparatus of claim 9 , wherein the scaling results are reported via a graphical user interface.
39 . The apparatus of claim 9 , wherein the object heights are obtained from a dataset.
40 . The apparatus of claim 14 , wherein the object heights are used to train a branch of a neural network model.
41 . The apparatus of claim 15 , wherein the branch of the neural network model is trained to adjust at least one of a first camera parameter or a second camera parameter.
42 . A method for scale recovery from monocular video, the method comprising:
segmenting an input image from the monocular video to detect an object in a camera field;
estimating camera parameters from the monocular input video;
iteratively refining the estimated camera parameters;
calculating a scale for relative depth;
iteratively refining the scale with provided user input; and
reporting scaling results for visualization.
43 . The method of claim 17 , wherein the input image segmentation is performed using a segmentation backbone network.
44 . The method of claim 17 , wherein the video scale is calculated using a first and second camera parameter.
45 . The method of claim 19 , wherein the first and second camera parameters are adjusted according to a projection model.