Vision-based 6DOF camera pose estimation in bronchoscopy
Methods and systems provide improved navigation through tubular networks such as lung airways by providing improved estimation of location and orientation information of a medical instrument (e.g., an endoscope) within the tubular network. Various input data such as image data and CT data, are used to model the tubular networks, and the model information is used to generate a camera pose representing a specific site location within the tubular network and/or to determine navigation information including position and orientation for the medical instrument.
1 . A method, comprising:
transforming video image data of an anatomical structure into a first depth map;
receiving a plurality of second depth maps based at least in part on transformed computed tomography (CT) image data, each of the second depth maps representing a virtual model of the anatomical structure; and
generating a camera pose representative of a spatial location within the anatomical structure, including:
finding a minimum of a scalar similarity function relative to a comparison of each of the plurality of second depth maps to the first depth map; and
generating the camera pose as a function of a second depth map associated with the minimum of the scalar similarity function.
2 . The method of claim 1 , further comprising referencing the video image data to a coordinate system of the CT image data prior to the transforming.
3 . The method of claim 1 , further comprising applying a convolutional neural network (CNN) to transform the video image data into the first depth map.
4 . The method of claim 1 , further comprising comparing each of the plurality of second depth maps to the first depth map for similarity of shape.
5 . The method of claim 1 , further comprising comparing each of the plurality of second depth maps to the first depth map using a normalized-cross correlation or mutual information.
6 . The method of claim 1 , further comprising generating the camera pose as a function of a second depth map associated to a given time point.
7 . The method of claim 1 , further comprising initializing the scalar similarity function with a previously generated camera pose.
8 . The method of claim 1 , further comprising generating a camera pose for each of the plurality of second depth maps.
9 . The method of claim 1 , wherein the first depth map comprises a three-dimensional virtual model of the anatomical structure.
10 . The method of claim 1 , wherein the first depth map and the plurality of second depth maps contain three-dimensional position and orientation information.
11 . The method of claim 1 , wherein the similarity function provides a minimum difference between the first depth map at a given time and one of the second depth maps at the given time.
12 . A method, comprising:
transforming video image data of a bronchial airway into a first depth map;
receiving a plurality of second depth maps based at least in part on transformed computed tomography (CT) image data, each of the second depth maps representing a virtual model of the bronchial airway; and
generating a camera pose representative of a spatial location within the bronchial airway, including:
finding a minimum of a scalar similarity function relative to a comparison of each of the plurality of second depth maps to the first depth map; and
generating the camera pose as a function of a second depth map associated with the minimum of the scalar similarity function.
13 . The method of claim 12 , further comprising generating the camera pose to include an orientation within the bronchial airway.
14 . The method of claim 12 , further comprising using an artificial neural network architecture to find the minimum of the scalar similarity function.
15 . The method of claim 12 , further comprising estimating a camera pose by passing each of the plurality of second depth maps and the first depth map into a spatial transformation network that regresses the relative transformation between each of the plurality of second depth maps and the first depth map.
16 . The method of claim 12 , wherein the generating the camera pose is an iterative or continuous process that includes using a camera pose of a previous iteration to initialize a current iteration.
17 . A medical system, comprising:
a means for transforming video image data of an anatomical structure into a first depth map;
a means for receiving a plurality of second depth maps based at least in part on transformed computed tomography (CT) image data, each of the second depth maps representing a virtual model of the anatomical structure;
a means for generating a camera pose representative of a spatial location within the anatomical structure, including:
a means for finding a minimum of a scalar similarity function relative to a comparison of each of the plurality of second depth maps to the first depth map; and
a means for generating the camera pose as a function of a second depth map associated with the minimum of the scalar similarity function.
18 . The medical system of claim 17 , further comprising a means for initializing the scalar similarity function.
19 . The medical system of claim 17 , further comprising a means for comparing each of the plurality of second depth maps to the first depth map for similarity of identified features.
20 . The medical system of claim 17 , further comprising a means for communicating the camera pose in real time, relative to the anatomical structure.