Video-based surgical skill assessment using tool tracking
A process for classifying a surgeon's technical skill in performing a surgery receives a tool-motion track comprising a sequence of detected tool motions of a surgeon performing a surgery with a surgical tool. A sequence of multi-channel feature matrices to mathematically represent the tool-motion track is then generated. Next, a one-dimensional (1D) convolution operation is performed on the sequence of multi-channel feature matrices to generate a sequence of context-aware multi-channel feature representations of the tool-motion track. The sequence of context-aware multi-channel feature representations is subsequently processed by a transformer model to generate a skill classification, wherein the transformer model is trained to focus on a subset of tool motions in the sequence of detected tool motions that are most relevant to the skill classification. Other aspects are also described and claimed.
1 . A computer-implemented method for classifying a surgeon's technical skill in performing a surgery, the method comprising:
receiving a tool-motion track comprising a sequence of detected tool motions of a surgeon performing a surgery with a surgical tool;
generating a sequence of multi-channel feature matrices to mathematically represent the tool-motion track; and
processing the sequence of multi-channel feature matrices using a deep-learning model to generate a skill classification for the surgeon performing the surgery, wherein the deep-learning model is a transformer model, and wherein processing the sequence of multi-channel feature matrices using the deep-learning model comprises:
performing a one-dimensional (1D) convolution operation on the sequence of multi-channel feature matrices by convolving each multi-channel feature matrix within the sequence of multi-channel feature matrices with a kernel having a predetermined time length to generate a context-aware multi-channel feature representation of the multi-channel feature matrix, as part of a sequence of context-aware multi-channel feature representations; and
processing the sequence of context-aware multi-channel feature representations of the tool-motion track, by the transformer model, to generate the skill classification.
2 . The computer-implemented method of claim 1 wherein the deep-learning model has been trained to identify and focus on a subset of tool motions in the sequence of detected tool motions that are most relevant to the skill classification and wherein the transformer model identifies and focuses on the subset of tool motions that are most relevant to the skill classification by using a self-attention technique.
3 . The computer-implemented method of claim 1 wherein convolving each multi-channel feature matrix with the kernel involves separately convolving each channel of the multi-channel feature matrix with the kernel.
4 . The computer-implemented method of claim 1 wherein the 1D convolution operation compares the multi-channel feature matrix at a given time-step with a number of adjacent time-steps both before and after the given time-step; and
wherein the context-aware multi-channel feature representation embeds an amount of learned relationships to the number of adjacent time-steps both before and after the given time-step.
5 . The computer-implemented method of claim 1 , wherein the tool-motion track is generated based on a sequence of locations of the tool detected within a sequence of video frames captured at a set of time-steps; and
wherein each multi-channel feature matrix within the sequence of multi-channel feature matrices is generated at a corresponding time-step in the set of time-steps.
6 . The computer-implemented method of claim 5 , wherein the multi-channel feature matrix is composed of at least the following signal channels:
a time-step;
a (X, Y) coordinates of the detected tool location within the corresponding video frame detected at the time-step; and
a size of a bounding box of the detected tool within the corresponding video frame detected at the time-step.
7 . The computer-implemented method of claim 6 , wherein the multi-channel feature matrix additionally includes a temporal mask channel indicating the tool present/absence in the corresponding video frame.
8 . A surgeon-skill classification system, comprising:
one or more processors;
a memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors:
receive a motion track comprising a sequence of detected tool motions of a surgeon performing a surgery with a surgical tool;
generate a sequence of multi-channel feature matrices to mathematically represent the motion track; and
process the sequence of multi-channel feature matrices using a deep-learning model to generate a skill classification for the surgeon performing the surgery, wherein the deep-learning model is a transformer model, and wherein to process the sequence of multi-channel feature matrices using the deep-learning model:
a one-dimensional (1D) convolution operation is performed on the sequence of multi-channel feature matrices by convolving each multi-channel feature matrix within the sequence of multi-channel feature matrices with a kernel of a predetermined time length to generate a context-aware multi-channel feature representation of the multi-channel feature matrix as part of a sequence of context-aware multi-channel feature representations; and
the sequence of context-aware multi-channel feature representations of the motion track is processed by the transformer model to generate the skill classification.
9 . The surgeon-skill classification system of claim 8 wherein the memory further stores instructions that, when executed by the one or more processors, cause the system to use the transformer model to identify and focus on a subset of tool motions in the sequence of detected tool motions by using a self-attention technique.
10 . The surgeon-skill classification system of claim 8 wherein the memory further stores instructions that, when executed by the one or more processors, cause the system to convolve each multi-channel feature matrix with the kernel by separately convolving each channel of the multi-channel feature matrix with the kernel.
11 . The surgeon-skill classification system of claim 8 wherein the 1D convolution operation compares the multi-channel feature matrix at a given time-step with a number of adjacent time-steps both before and after the given time-step; and
wherein the context-aware multi-channel feature representation embeds an amount of learned relationships to the number of adjacent time-steps both before and after the given time-step.
12 . The surgeon-skill classification system of claim 8 , wherein the motion track is generated based on a sequence of locations of the tool detected within a sequence of video frames captured at a set of time-steps; and
wherein each multi-channel feature matrix within the sequence of multi-channel feature matrices is generated at a corresponding time-step in the set of time-steps.
13 . The surgeon-skill classification system of claim 12 , wherein the multi-channel feature matrix is composed of some or all of the following signal channels:
a time-step;
a (X, Y) coordinates of the detected tool location within the corresponding video frame detected at the time-step;
a size of a bounding box of the detected tool within the corresponding video frame detected at the time-step; and
a temporal mask indicating the tool present/absence in the corresponding video frame.