Tracking multiple surgical tools in a surgical video
Disclosed are various systems and techniques for tracking surgical tools in a surgical video. In one aspect, the system begins by receiving one or more established tracks for one or more previously-detected surgical tools in the surgical video. The system then processes a current frame of the surgical video to detect one or more objects using a first deep-learning model. Next, for each detected object in the one or more detected objects, the system further performs the flowing steps to assign the detected object to a right track: (1) computing a semantic similarity between the detected object and each of the one or more established tracks; (2) computing a spatial similarity between the detected object and the latest predicted location for each of the one or more established tracks; and (3) attempting to assign the detected object to one of the one or more established tracks based on the computed semantic similarity and the spatial similarity metric.
1 . A computer-implemented method for tracking surgical tools in a surgical video, the method comprising:
receiving a current video frame of the surgical video for processing;
receiving one or more established tracks for one or more previously-detected surgical tools in the surgical video, wherein each established track comprises a first array that has a plurality of previously-predicted locations of a previously-detected surgical tool and a second array of semantic features associated with the previously-detected surgical tool across a plurality of previously-processed video frames of the surgical video;
processing the current video frame of the surgical video to detect one or more objects using a first deep-learning model; and
for each detected object in the one or more detected objects,
computing a semantic similarity metric between the detected object and each of the one or more established tracks by:
using a second deep-learning model to extract a set of semantic features of the detected object;
determining whether an established track comprises a detection gap in which a corresponding previously-detected surgical tool is temporarily out of view from a first subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the current video frame over a predefined time interval; and
responsive to the established track comprising the detection gap, comparing the set of semantic features with one or more semantic features of the second array across a second subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the first subset of previously-processed video frames to determine whether the detected object and the corresponding previously-detected surgical tool are visually similar;
computing a spatial similarity metric between the detected object and a latest predicted location in the first array for each of the one or more established tracks; and
assigning the detected object to one of the one or more established tracks based on the computed semantic similarity metric and the computed spatial similarity metric.
2 . The computer-implemented method of claim 1 , wherein prior to processing the current video frame, the method further comprises:
converting a frame rate of the surgical video so that the converted frame rate is greater or equal to a predetermined frame rate; and
resizing the current video frame into a predetermined image size.
3 . The computer-implemented method of claim 1 , wherein the first deep-learning model is a Faster-RCNN model trained to detect and classify a set of diverse types of surgical tools within a given video frame of the surgical video.
4 . The computer-implemented method of claim 1 , wherein the one or more previously-detected surgical tools include a left-hand tool and a right-hand tool.
5 . The computer-implemented method of claim 1 , wherein the set of semantic features forms a feature vector of 128 dimensions.
6 . The computer-implemented method of claim 1 , wherein the second array of semantic features are associated with a number of previously-detected images of the surgical tool over a predetermined time period.
7 . The computer-implemented method of claim 1 , wherein prior to computing the spatial similarity metric, the method further comprises:
receiving a location on a generated bounding box for the detected object; and
generating the latest predicted location for the established track by applying a Kalman filter to the received location of the detected object and a last known location of the established track within the first array.
8 . The computer-implemented method of claim 7 , wherein the location of the generated bounding box is a center of the generated bounding box.
9 . The computer-implemented method of claim 1 , wherein assigning the detected object to one of the one or more established tracks includes using a data association technique on the computed semantic similarity metric and the computed spatial similarity metric between the detected object and each of the one or more established tracks.
10 . The computer-implemented method of claim 9 , wherein the data association technique employs a Hungarian method that is configured to identify a correct track assignment within the one or more established tracks for the detected object by minimizing a cost function of a track assignment between the detected object and the one or more established tracks.
11 . The computer-implemented method of claim 10 , wherein the Hungarian method employs a bipartite graph to solve the cost function associated with the detected object.
12 . The computer-implemented method of claim 10 , wherein the method further comprises assigning a weight greater than a threshold to the track assignment if a corresponding computed spatial similarity metric is greater than a predetermined distance threshold to prohibit the track assignment.
13 . The computer-implemented method of claim 10 , wherein the cost function additionally includes a track ID associated with the established track and a class ID assigned to the detected object.
14 . The computer-implemented method of claim 13 , wherein the method further comprises assigning a weight greater than a threshold to the track assignment if a corresponding track ID associated with the established track does not match the class ID of the detected object to prohibit the track assignment.
15 . The computer-implemented method of claim 1 , wherein the method further comprises:
responsive to a determination that a previously detected object is undetectable, identify one of the one or more established tracks as an inactive established track;
after the previously detected object has been undetectable for a plurality of video frames of the surgical video, receiving an unassigned object in the one or more detected objects that cannot be assigned to any track in the one or more established tracks;
determining if a location of the unassigned object is sufficiently close to a last known location of the inactive established track in the one or more established tracks and if a class ID of the unassigned object matches a track ID of the inactive established track; and
if so, re-assigning the unassigned object to the inactive established track to reactivate the inactive established track.
16 . The computer-implemented method of claim 1 , wherein, responsive to the established track not including the detection gap, comparing the set of semantic features with the one or more semantic features of the second array across a third subset of previously-processed video frames of the plurality of previously-processed video frames immediately before the current video frame.
17 . An apparatus, comprising:
one or more processors; and
a memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors, cause the apparatus to:
receive a current video frame of a surgical video for processing;
receive one or more established tracks for one or more previously-detected surgical tools in the surgical video, wherein each established track comprises a first array that has a plurality of previously-predicted locations of a previously-detected surgical tool and a second array of semantic features associated with the previously-detected surgical tool across a plurality of previously-processed video frames of the surgical video;
process the current video frame of the surgical video to detect one or more objects using a first deep-learning model; and
for each detected object in the one or more detected objects,
compute a semantic similarity metric between the detected object and each of the one or more established tracks by:
using a second deep-learning model to extract a set of semantic features of the detected object;
determining whether an established track comprises a detection gap in which a corresponding previously-detected surgical tool is temporarily out of view from a first subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the current video frame over a predefined time interval; and
responsive to the established track comprising the detection gap, comparing the set of semantic features with one or more semantic features of the second array across a second subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the first subset of previously-processed video frames to determine whether the detected object and the corresponding previously-detected surgical tool are visually similar;
compute a spatial similarity metric between the detected object and a latest predicted location in the first array for each of the one or more established tracks; and
assign the detected object to one of the one or more established tracks based on the computed semantic similarity metric and the computed spatial similarity metric.
18 . The apparatus of claim 17 , wherein the memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors, cause the apparatus to:
responsive to a determination that a previously detected object is undetectable, identifying one of the one or more established tracks as an inactive established track;
after the previously detected object has been undetectable for a plurality of video frames of the surgical video, receive an unassigned object in the one or more detected objects that cannot be assigned to any track in the one or more established tracks;
determine if a location of the unassigned object is sufficiently close to a last known location of an inactive track in the one or more established tracks and if a class ID of the unassigned object matches a track ID of the inactive track; and
if so, re-assign the unassigned object to the inactive established track to reactivate the inactive established track.
19 . A system, comprising:
one or more processors; and
a memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors, cause the system to:
receive a current video frame of a surgical video for processing;
receive one or more established tracks for one or more previously-detected surgical tools in the surgical video, wherein each established track comprises a first array that has a plurality of previously-predicted locations of a previously-detected surgical tool and a second array of semantic features associated with the previously-detected surgical tool across a plurality of previously-processed video frames of the surgical video;
process the current video frame of the surgical video to detect one or more objects using a first deep-learning model; and
for each detected object in the one or more detected objects,
compute a semantic similarity metric between the detected object and each of the one or more established tracks by:
using a second deep-learning model to extract a set of semantic features of the detected object;
determining whether an established track comprises a detection gap in which a corresponding previously-detected surgical tool is temporarily out of view from a first subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the current video frame over a predefined time interval; and
responsive to the established track comprising the detection gap, comparing the set of semantic features with one or more semantic features of the second array across a second subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the first subset of previously-processed video frames to determine whether the detected object and the corresponding previously-detected surgical tool are visually similar;
compute a spatial similarity metric between the detected object and a latest predicted location in the first array for each of the one or more established tracks; and
assign the detected object to one of the one or more established tracks based on the computed semantic similarity metric and the computed spatial similarity metric.
20 . The system of claim 19 , wherein the memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors, cause the system to:
responsive to a determination that a previously detected object is undetectable, identify one of the one or more established tracks as an inactive established track;
after the previously detected object has been undetectable for a plurality of video frames of the surgical video, receive an unassigned object in the one or more detected objects that cannot be assigned to any track in the one or more established tracks;
determine if a location of the unassigned object is sufficiently close to a last known location of the inactive established track in the one or more established tracks and if a class ID of the unassigned object matches a track ID of the inactive established track; and
if so, re-assign the unassigned object to the inactive established track to reactivate the inactive established track.