IP Library › Granted Patent US 12,614,287
Granted Patent B2
US 12,614,287 · App. 18/077,912 · Granted Apr 28, 2026

Tracking multiple surgical tools in a surgical video

Inventors: Mona Fathollahi Ghezelghieh (Sunnyvale, CA); Jocelyn Barker (San Jose, CA)
Assignee: Auris Health, Inc.
G06T7/246A61B34/20G06V10/26G06V10/761G06V10/764G06V10/7715G06V10/774G06V10/82A61B2034/2065G06T2207/10016G06T2207/10068G06V2201/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,287
App. No.
18/077,912
Granted
Apr 28, 2026
Kind
B2
Abstract

Disclosed are various systems and techniques for tracking surgical tools in a surgical video. In one aspect, the system begins by receiving one or more established tracks for one or more previously-detected surgical tools in the surgical video. The system then processes a current frame of the surgical video to detect one or more objects using a first deep-learning model. Next, for each detected object in the one or more detected objects, the system further performs the flowing steps to assign the detected object to a right track: (1) computing a semantic similarity between the detected object and each of the one or more established tracks; (2) computing a spatial similarity between the detected object and the latest predicted location for each of the one or more established tracks; and (3) attempting to assign the detected object to one of the one or more established tracks based on the computed semantic similarity and the spatial similarity metric.

Claims (70)

1 . A computer-implemented method for tracking surgical tools in a surgical video, the method comprising:

receiving a current video frame of the surgical video for processing;

receiving one or more established tracks for one or more previously-detected surgical tools in the surgical video, wherein each established track comprises a first array that has a plurality of previously-predicted locations of a previously-detected surgical tool and a second array of semantic features associated with the previously-detected surgical tool across a plurality of previously-processed video frames of the surgical video;

processing the current video frame of the surgical video to detect one or more objects using a first deep-learning model; and

for each detected object in the one or more detected objects,

computing a semantic similarity metric between the detected object and each of the one or more established tracks by:

using a second deep-learning model to extract a set of semantic features of the detected object;

determining whether an established track comprises a detection gap in which a corresponding previously-detected surgical tool is temporarily out of view from a first subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the current video frame over a predefined time interval; and

responsive to the established track comprising the detection gap, comparing the set of semantic features with one or more semantic features of the second array across a second subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the first subset of previously-processed video frames to determine whether the detected object and the corresponding previously-detected surgical tool are visually similar;

computing a spatial similarity metric between the detected object and a latest predicted location in the first array for each of the one or more established tracks; and

assigning the detected object to one of the one or more established tracks based on the computed semantic similarity metric and the computed spatial similarity metric.

2 . The computer-implemented method of claim 1 , wherein prior to processing the current video frame, the method further comprises:

converting a frame rate of the surgical video so that the converted frame rate is greater or equal to a predetermined frame rate; and

resizing the current video frame into a predetermined image size.

3 . The computer-implemented method of claim 1 , wherein the first deep-learning model is a Faster-RCNN model trained to detect and classify a set of diverse types of surgical tools within a given video frame of the surgical video.

4 . The computer-implemented method of claim 1 , wherein the one or more previously-detected surgical tools include a left-hand tool and a right-hand tool.

5 . The computer-implemented method of claim 1 , wherein the set of semantic features forms a feature vector of 128 dimensions.

6 . The computer-implemented method of claim 1 , wherein the second array of semantic features are associated with a number of previously-detected images of the surgical tool over a predetermined time period.

7 . The computer-implemented method of claim 1 , wherein prior to computing the spatial similarity metric, the method further comprises:

receiving a location on a generated bounding box for the detected object; and

generating the latest predicted location for the established track by applying a Kalman filter to the received location of the detected object and a last known location of the established track within the first array.

8 . The computer-implemented method of claim 7 , wherein the location of the generated bounding box is a center of the generated bounding box.

9 . The computer-implemented method of claim 1 , wherein assigning the detected object to one of the one or more established tracks includes using a data association technique on the computed semantic similarity metric and the computed spatial similarity metric between the detected object and each of the one or more established tracks.

10 . The computer-implemented method of claim 9 , wherein the data association technique employs a Hungarian method that is configured to identify a correct track assignment within the one or more established tracks for the detected object by minimizing a cost function of a track assignment between the detected object and the one or more established tracks.

11 . The computer-implemented method of claim 10 , wherein the Hungarian method employs a bipartite graph to solve the cost function associated with the detected object.

12 . The computer-implemented method of claim 10 , wherein the method further comprises assigning a weight greater than a threshold to the track assignment if a corresponding computed spatial similarity metric is greater than a predetermined distance threshold to prohibit the track assignment.

13 . The computer-implemented method of claim 10 , wherein the cost function additionally includes a track ID associated with the established track and a class ID assigned to the detected object.

14 . The computer-implemented method of claim 13 , wherein the method further comprises assigning a weight greater than a threshold to the track assignment if a corresponding track ID associated with the established track does not match the class ID of the detected object to prohibit the track assignment.

15 . The computer-implemented method of claim 1 , wherein the method further comprises:

responsive to a determination that a previously detected object is undetectable, identify one of the one or more established tracks as an inactive established track;

after the previously detected object has been undetectable for a plurality of video frames of the surgical video, receiving an unassigned object in the one or more detected objects that cannot be assigned to any track in the one or more established tracks;

determining if a location of the unassigned object is sufficiently close to a last known location of the inactive established track in the one or more established tracks and if a class ID of the unassigned object matches a track ID of the inactive established track; and

if so, re-assigning the unassigned object to the inactive established track to reactivate the inactive established track.

16 . The computer-implemented method of claim 1 , wherein, responsive to the established track not including the detection gap, comparing the set of semantic features with the one or more semantic features of the second array across a third subset of previously-processed video frames of the plurality of previously-processed video frames immediately before the current video frame.

17 . An apparatus, comprising:

one or more processors; and

a memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors, cause the apparatus to:

receive a current video frame of a surgical video for processing;

receive one or more established tracks for one or more previously-detected surgical tools in the surgical video, wherein each established track comprises a first array that has a plurality of previously-predicted locations of a previously-detected surgical tool and a second array of semantic features associated with the previously-detected surgical tool across a plurality of previously-processed video frames of the surgical video;

process the current video frame of the surgical video to detect one or more objects using a first deep-learning model; and

for each detected object in the one or more detected objects,

compute a semantic similarity metric between the detected object and each of the one or more established tracks by:

using a second deep-learning model to extract a set of semantic features of the detected object;

determining whether an established track comprises a detection gap in which a corresponding previously-detected surgical tool is temporarily out of view from a first subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the current video frame over a predefined time interval; and

responsive to the established track comprising the detection gap, comparing the set of semantic features with one or more semantic features of the second array across a second subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the first subset of previously-processed video frames to determine whether the detected object and the corresponding previously-detected surgical tool are visually similar;

compute a spatial similarity metric between the detected object and a latest predicted location in the first array for each of the one or more established tracks; and

assign the detected object to one of the one or more established tracks based on the computed semantic similarity metric and the computed spatial similarity metric.

18 . The apparatus of claim 17 , wherein the memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors, cause the apparatus to:

responsive to a determination that a previously detected object is undetectable, identifying one of the one or more established tracks as an inactive established track;

after the previously detected object has been undetectable for a plurality of video frames of the surgical video, receive an unassigned object in the one or more detected objects that cannot be assigned to any track in the one or more established tracks;

determine if a location of the unassigned object is sufficiently close to a last known location of an inactive track in the one or more established tracks and if a class ID of the unassigned object matches a track ID of the inactive track; and

if so, re-assign the unassigned object to the inactive established track to reactivate the inactive established track.

19 . A system, comprising:

one or more processors; and

a memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors, cause the system to:

receive a current video frame of a surgical video for processing;

receive one or more established tracks for one or more previously-detected surgical tools in the surgical video, wherein each established track comprises a first array that has a plurality of previously-predicted locations of a previously-detected surgical tool and a second array of semantic features associated with the previously-detected surgical tool across a plurality of previously-processed video frames of the surgical video;

process the current video frame of the surgical video to detect one or more objects using a first deep-learning model; and

for each detected object in the one or more detected objects,

compute a semantic similarity metric between the detected object and each of the one or more established tracks by:

using a second deep-learning model to extract a set of semantic features of the detected object;

determining whether an established track comprises a detection gap in which a corresponding previously-detected surgical tool is temporarily out of view from a first subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the current video frame over a predefined time interval; and

responsive to the established track comprising the detection gap, comparing the set of semantic features with one or more semantic features of the second array across a second subset of previously-processed video frames of the plurality of previously-processed video frames immediately before to the first subset of previously-processed video frames to determine whether the detected object and the corresponding previously-detected surgical tool are visually similar;

compute a spatial similarity metric between the detected object and a latest predicted location in the first array for each of the one or more established tracks; and

assign the detected object to one of the one or more established tracks based on the computed semantic similarity metric and the computed spatial similarity metric.

20 . The system of claim 19 , wherein the memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors, cause the system to:

responsive to a determination that a previously detected object is undetectable, identify one of the one or more established tracks as an inactive established track;

after the previously detected object has been undetectable for a plurality of video frames of the surgical video, receive an unassigned object in the one or more detected objects that cannot be assigned to any track in the one or more established tracks;

determine if a location of the unassigned object is sufficiently close to a last known location of the inactive established track in the one or more established tracks and if a class ID of the unassigned object matches a track ID of the inactive established track; and

if so, re-assign the unassigned object to the inactive established track to reactivate the inactive established track.

Assignments (2)
MERGER Recorded Jan 27, 2026
From: VERB SURGICAL INC.
To: AURIS HEALTH, INC.
Reel/Frame 073602/0979 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2023
From: FATHOLLAHI GHEZELGHIEH, MONA; BARKER, JOCELYN
To: VERB SURGICAL INC.
Reel/Frame 065243/0475 →
Continuity (2)
Provisional Application 63287477 · Dec 8, 2021
Related Publication 20230177703A1 · Jun 8, 2023
References Cited (16)
US 20150297313A1 · Reiter et al. · 2015 [cited by applicant]
US 20210290317A1 · Sen et al. · 2021 [cited by applicant]
CN 112037263A · 2020 [cited by applicant]
Fujie, Hiroki, et al. “Detecting and Tracking Surgical Tools for Recognizing Phases of the Awake Brain Tumor Removal Surgery.” ICPRAM. 2019. (Year: 2019). [cited by examiner]
Robu, Maria, et al. “Towards real-time multiple surgical tool tracking.” Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization 9.3 (Dec. 3, 2020): 279-285. (Year: 2020). [cited by examiner]
Liu, Liqiang, and Jianzhong Cao. “End-to-end learning interpolation for object tracking in low frame-rate video.” IET Image Processing 14.6 (Apr. 2020): 1066-1072. (Year: 2020). [cited by examiner]
Bewley, Simple Online and Realtime Tracking, 2017, https://arxiv.org/abs/1602.00763v2 (Year: 2017). [cited by examiner]
Bergmann, P., “Tracking without bells and whistles”, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 941-951. [cited by applicant]
Harmans, A., et al., “In Defense of the Triplet Loss for Person Re-Identification”, arXiv:1703.07737v4 [cs.CV], Nov. 21, 2017, 17 pages. [cited by applicant]
Lee, D., et al., “Evaluation of Surgical Skills during Robotic Surgery by Deep Learning-Based Multiple Surgical Instrument Tracking in Training and Actual Operations”, Journal of Clinical Medicine, vol. 9, No. 6, 15 pag… [cited by applicant]
Robu, M., et al., “Towards real-time multiple surgical tool tracking”, Computer Methods in Biomechanics and Biomedical Engineering: Imaging & Visualization>, vol. 9, No. 3, 12 Pages. [cited by applicant]
PCT/IB2022/061944, “PCT Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, or the Declaration”, Mailed Mar. 15, 2023, 7 pages. [cited by applicant]
Supplementary European Search Report received for European Patent Application No. 22903708.0, mailed Oct. 22, 2025, 11 pages. [cited by applicant]
Anonymous: “Hungarian Algorithm;” WIKIPEDIA, The Free Encyclopedia; retrieved online: https://simple.wikipedia.org/wiki/Hungarian_algorithm; Nov. 15, 2021; 15 pages. [cited by applicant]
Fathollahi et al.; “Video-based Surgical Skills Assessment using Long term Tool Tracking;” retrieved online: https://doi.org/10.48550/arXiv.2207.02247; Jul. 5, 2022; 11 pages. [cited by applicant]
Wojke et al.; “Simple Online and Realtime Tracking with a Deep Association Metric;” retrieved online: https://doi.org/10.48550/arXiv.1703.07402v1; Mar. 21, 2017; 5 pages. [cited by applicant]