IP Library › Granted Patent US 12,731,389
Granted Patent B2
US 12,731,389 · App. 18/186,798 · Granted Sep 8, 2026

Video-based surgical skill assessment using tool tracking

Inventors: Mona Fathollahi Ghezelghieh (Sunnyvale, CA); Mohammad Hasan Sarhan (Santa Clara, CA); Jocelyn Barker (San Jose, CA); Lela Dimonte (Mountain View, CA)
Assignee: Auris Health, Inc.
G06V10/82G06V10/764G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,731,389
App. No.
18/186,798
Granted
Sep 8, 2026
Kind
B2
Abstract

A process for classifying a surgeon's technical skill in performing a surgery receives a tool-motion track comprising a sequence of detected tool motions of a surgeon performing a surgery with a surgical tool. A sequence of multi-channel feature matrices to mathematically represent the tool-motion track is then generated. Next, a one-dimensional (1D) convolution operation is performed on the sequence of multi-channel feature matrices to generate a sequence of context-aware multi-channel feature representations of the tool-motion track. The sequence of context-aware multi-channel feature representations is subsequently processed by a transformer model to generate a skill classification, wherein the transformer model is trained to focus on a subset of tool motions in the sequence of detected tool motions that are most relevant to the skill classification. Other aspects are also described and claimed.

Claims (36)

1 . A computer-implemented method for classifying a surgeon's technical skill in performing a surgery, the method comprising:

receiving a tool-motion track comprising a sequence of detected tool motions of a surgeon performing a surgery with a surgical tool;

generating a sequence of multi-channel feature matrices to mathematically represent the tool-motion track; and

processing the sequence of multi-channel feature matrices using a deep-learning model to generate a skill classification for the surgeon performing the surgery, wherein the deep-learning model is a transformer model, and wherein processing the sequence of multi-channel feature matrices using the deep-learning model comprises:

performing a one-dimensional (1D) convolution operation on the sequence of multi-channel feature matrices by convolving each multi-channel feature matrix within the sequence of multi-channel feature matrices with a kernel having a predetermined time length to generate a context-aware multi-channel feature representation of the multi-channel feature matrix, as part of a sequence of context-aware multi-channel feature representations; and

processing the sequence of context-aware multi-channel feature representations of the tool-motion track, by the transformer model, to generate the skill classification.

2 . The computer-implemented method of claim 1 wherein the deep-learning model has been trained to identify and focus on a subset of tool motions in the sequence of detected tool motions that are most relevant to the skill classification and wherein the transformer model identifies and focuses on the subset of tool motions that are most relevant to the skill classification by using a self-attention technique.

3 . The computer-implemented method of claim 1 wherein convolving each multi-channel feature matrix with the kernel involves separately convolving each channel of the multi-channel feature matrix with the kernel.

4 . The computer-implemented method of claim 1 wherein the 1D convolution operation compares the multi-channel feature matrix at a given time-step with a number of adjacent time-steps both before and after the given time-step; and

wherein the context-aware multi-channel feature representation embeds an amount of learned relationships to the number of adjacent time-steps both before and after the given time-step.

5 . The computer-implemented method of claim 1 , wherein the tool-motion track is generated based on a sequence of locations of the tool detected within a sequence of video frames captured at a set of time-steps; and

wherein each multi-channel feature matrix within the sequence of multi-channel feature matrices is generated at a corresponding time-step in the set of time-steps.

6 . The computer-implemented method of claim 5 , wherein the multi-channel feature matrix is composed of at least the following signal channels:

a time-step;

a (X, Y) coordinates of the detected tool location within the corresponding video frame detected at the time-step; and

a size of a bounding box of the detected tool within the corresponding video frame detected at the time-step.

7 . The computer-implemented method of claim 6 , wherein the multi-channel feature matrix additionally includes a temporal mask channel indicating the tool present/absence in the corresponding video frame.

8 . A surgeon-skill classification system, comprising:

one or more processors;

a memory coupled to the one or more processors, wherein the memory stores instructions that, when executed by the one or more processors:

receive a motion track comprising a sequence of detected tool motions of a surgeon performing a surgery with a surgical tool;

generate a sequence of multi-channel feature matrices to mathematically represent the motion track; and

process the sequence of multi-channel feature matrices using a deep-learning model to generate a skill classification for the surgeon performing the surgery, wherein the deep-learning model is a transformer model, and wherein to process the sequence of multi-channel feature matrices using the deep-learning model:

a one-dimensional (1D) convolution operation is performed on the sequence of multi-channel feature matrices by convolving each multi-channel feature matrix within the sequence of multi-channel feature matrices with a kernel of a predetermined time length to generate a context-aware multi-channel feature representation of the multi-channel feature matrix as part of a sequence of context-aware multi-channel feature representations; and

the sequence of context-aware multi-channel feature representations of the motion track is processed by the transformer model to generate the skill classification.

9 . The surgeon-skill classification system of claim 8 wherein the memory further stores instructions that, when executed by the one or more processors, cause the system to use the transformer model to identify and focus on a subset of tool motions in the sequence of detected tool motions by using a self-attention technique.

10 . The surgeon-skill classification system of claim 8 wherein the memory further stores instructions that, when executed by the one or more processors, cause the system to convolve each multi-channel feature matrix with the kernel by separately convolving each channel of the multi-channel feature matrix with the kernel.

11 . The surgeon-skill classification system of claim 8 wherein the 1D convolution operation compares the multi-channel feature matrix at a given time-step with a number of adjacent time-steps both before and after the given time-step; and

wherein the context-aware multi-channel feature representation embeds an amount of learned relationships to the number of adjacent time-steps both before and after the given time-step.

12 . The surgeon-skill classification system of claim 8 , wherein the motion track is generated based on a sequence of locations of the tool detected within a sequence of video frames captured at a set of time-steps; and

wherein each multi-channel feature matrix within the sequence of multi-channel feature matrices is generated at a corresponding time-step in the set of time-steps.

13 . The surgeon-skill classification system of claim 12 , wherein the multi-channel feature matrix is composed of some or all of the following signal channels:

a time-step;

a (X, Y) coordinates of the detected tool location within the corresponding video frame detected at the time-step;

a size of a bounding box of the detected tool within the corresponding video frame detected at the time-step; and

a temporal mask indicating the tool present/absence in the corresponding video frame.

Assignments (2)
MERGER Recorded Jan 27, 2026
From: VERB SURGICAL INC.
To: AURIS HEALTH, INC.
Reel/Frame 073602/0979 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 29, 2024
From: FATHOLLAHI GHEZELGHIEH, MONA; SARHAN, MOHAMMAD HASAN; BARKER, JOCELYN; DIMONTE, LELA
To: VERB SURGICAL INC.
Reel/Frame 066608/0566 →
Continuity (2)
Provisional Application 63322166 · Mar 21, 2022
Related Publication 20230298336A1 · Sep 21, 2023
References Cited (36)
US 20120253360A1 · White et al. · 2012 [cited by applicant]
US 20200367974A1 · Khalid · 2020 [cited by examiner]
US 20210313051A1 · Asselmann et al. · 2021 [cited by applicant]
US 20230177703A1 · Ghezelghieh et al. · 2023 [cited by applicant]
WO 2012060901A1 · 2012 [cited by applicant]
WO 2017083768A1 · 2017 [cited by applicant]
Lee et al, Evaluation of Surgical Skills during Robotic Surgery by Deep Learning-Based Multiple Surgical Instrument Tracking in Training and Actual Operations, J. Clin. Med. 2020, 9(6), 1964 (Year: 2020). [cited by examiner]
Bergmann et al, Tracking without bells and whistles, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 941-951. (Year: 2019). [cited by examiner]
Oropesa et al, Laparoscopic instrument tracking based on endoscopic video analysis for psychomotor skills assessment. Surg. Endosc, 27:1029-1039 (Year: 2013). [cited by examiner]
Pérez-Escamirosa et al, Objective classifcation of psychomotor laparoscopic skills of surgeons based on three dfferent approaches, International Journal of Computer Assisted Radiology and Surgery 15:27-40 (Year: 2020). [cited by examiner]
Bergmann, P., “Tracking without bells and whistles”, Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 941-951. [cited by applicant]
Harmans, A., et al., “In Defense of the Triplet Loss for Person Re-Identification”, arXiv:1703.07737v4 [cs.CV], Nov. 21, 2017, 17 pages. [cited by applicant]
Supplementary European Search Report received for EP Patent Application No. 23774110.3, mailed Jul. 25, 2025, 9 pages. [cited by applicant]
International Search Report and The Written Opinion of The Searching Authority received for PCT Patent Application No. PCT/IB2023/052782, mailed on Sep. 28, 2023, 8 pages. [cited by applicant]
International Preliminary Report of Patentability received for PCT Patent Application No. PCT/IB2023/052782, mailed on Oct. 3, 2024, 6 pages. [cited by applicant]
David P. Azari, et al.; “Modeling surgical technical skill using expert assessment for automated computer rating,” retrieved online: Ann Surg. Mar. 2019 ; 269(3): 574-581. doi:10.1097/SLA.0000000000002478; HHS Public Ac… [cited by applicant]
Ariel Kate Dubin, et al.; “A model for predicting the GEARS score from virtual reality surgical simulator metrics;” retrieved online: https://doi.org/10.1007/s00464-018-6082-7; Surgical Endoscopy (2018) 32:3576-3581; Fe… [cited by applicant]
Sean Estrada, et al.; “On the Development of Objective Metrics for Surgical Skills Evaluation Based on Tool Motion;” retrieved online: https://ieeexplore.ieee.org/abstract/document/6974411/DOI: 10.1109/SMC.2014.6974411;… [cited by applicant]
Mahtab J. Fard, et al.; “Automated robot-assisted surgical skill evaluation: Predictive analytics approach;” retrieved online: https://doi.org/10.1002/rcs.1850; wileyonlinelibrary.com/journal/rcs; Int J Med Robotics Com… [cited by applicant]
Mahtab J. Fard, et al.“Soft Boundary Approach for Unsupervised Gesture Segmentation in Robotic-Assisted Surgery;” retrieved online: IEEE Xplore.com; IEEE Robotics and Automation Letters, vol. 2, No. 1, Jan. 2016; 8 page… [cited by applicant]
Andrew J. Hung, et al.; “Automated Performance Metrics and Machine Learning Algorithms toMeasure Surgeon Performance and Anticipate Clinical Outcomes in Robotic Surgery;” retrieved online: jamasurgery.com; JAMA Surgery … [cited by applicant]
H.W. Kuhn; “The Hungarian Method for the Assignment Problem;” retrieved online:https://onlinelibrary.wiley.com/doi/abs/10.1002/nav.3800020109; https://doi.org/10.1002/nav.3800020109; Mar. 1955;15 pages. [cited by applicant]
Hei Law, et al. ; “Surgeon Technical Skill Assessment using Computer Vision based Analysis;” retrieved online: https://proceedings.mlr.press/v68/law17a/law17a; Proceedings of Machine Learning for Healthcare 2017; JMLR W… [cited by applicant]
Marc Levin et al.; “Automated Methods of Technical Skill Assessment in Surgery: A Systematic Review;” retrieved online: doi:10.1016/j.jsurg.2019.06.011.Journal; Journal of Surgical Education; vol. 76; No. 6; Nov./Dec. 2… [cited by applicant]
Anton Milan, et al.; “MOT16: A Benchmark for Multi-Object Tracking;” retrieved online: arXiv:1603.00831v2 [cs. CV] May 3, 2016; 12 pages. [cited by applicant]
Chinedu Innocent Nwoyea, et al.; “Rendezvous: Attention Mechanisms for the Recognition of Surgical Action Triplets in Endoscopic Videos;” retrieved online: arXiv:2109.03223v2 [cs.CV] 3; https://endovissub-workflowandski… [cited by applicant]
Irene Rivas-Blanco, et al.; “A surgical dataset from the da Vinci Research Kit for task automation and recognition;” retrieved online: arXiv:2102.03643v2 [cs.RO] Jun. 29, 2023; Proceedings of the International Conferenc… [cited by applicant]
Somayeh B Shafiei et al.; “Using Two-Third Power Law for Segmentation of Hand Movement in Robotic Assisted Surgery;” retrieved online:http://asmedigitalcollection.asme.org/IDETC-CIE/proceedings-pdf/IDETC-CIE2015/57144/V… [cited by applicant]
Andru P. Twinanda, et al.; “EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos;” retrieved online arXiv:1602.03012v2 [cs.CV] May 23, 2016; 11 pages. [cited by applicant]
Melina C. Vassiliou, M.D., et al.; “A Global Assessment Tool For Evaluation of Intraoperative Laparoscopic Skills;” retrieved online: doi:10.1016/j.amjsurg.2005.04.004; The American Journal of Surgery (Excerpta Medica I… [cited by applicant]
Ashish Vaswani, et al.; “Attention Is All You Need;” retrieved online: arXiv:1706.03762v7 [cs.CL] Aug. 2, 2023; 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA; 15 pages. [cited by applicant]
Qingsong Wen, et al.; “Transformers in Time Series: A Survey;” retrieved online: arXiv:2202.07125v5 [cs. LG] May 11, 2023; 9 pages. [cited by applicant]
Benjaminde Witte, et al.; “Cost-Efficient Laparoscopic HapticTrainer based on Affine VelocityAnalysis;” retrieved online: HAL Id:hal-01563262, https://hal.science/hal-01563262v1; Feb. 4, 2019; 2 pages. [cited by applicant]
Yi Wu, et al.; “Online Object Tracking: A Benchmark;” retrieved online: https://ieeexplore.ieee.org/document/6619156; DOI: 10.1109/CVPR.2013.312; Oct. 3, 2013; 8 pages. [cited by applicant]
Yifu Zhang et al.; “ByteTrack: Multi-Object Tracking by Associating Every Detection Box;” retrieved online: https://doi.org/10.48550/arXiv.2110.06864v3; Apr. 7, 2022; 19 pages. [cited by applicant]
Aneeq Zia, et al.; “Automated Surgical Skill Assessment in RMIS Training;” retrieved online: https://doi.org/10.48550/arXiv.1712.08604v1, Dec. 22, 2017; 12 pages. [cited by applicant]