IP Library › Granted Patent US 12,548,332
Granted Patent B2
US 12,548,332 · App. 18/130,445 · Granted Feb 10, 2026

Cascade stage boundary awareness networks for surgical workflow analysis

Inventors: Jinglu Zhang (Bournemouth, GB); Abdolrahim Kadkhodamohammadi (London, GB); Imanol Luengo Muntion (London, GB); Danail V. Stoyanov (London, GB); Santiago Barbarisi (London, GB)
Assignee: DIGITAL SURGERY LIMITED
G06V20/41G06V10/82G06V20/46G16H30/40G06V2201/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,332
App. No.
18/130,445
Granted
Feb 10, 2026
Kind
B2
Abstract

Techniques are described for improving computer-assisted surgical (CAS) systems, particularly, to recognize surgical phases in a video of a surgical procedure. A CAS system includes cameras that provide video stream of a surgical procedure. According to one or more aspects the surgical phases are automatically detected in the video stream using a machine learning model. Particularly, the machine learning model includes a boundary aware cascade stage network to perform surgical phase recognition.

Claims (32)

1 . A system comprising:

a memory device; and

one or more processors coupled with the memory device, the one or more processors configured to:

encode a frame of a video of a surgical procedure into a plurality of features;

provide the features to a boundary supervision branch of a model;

provide the features to a cascade of temporal stages of the model, wherein the cascade of temporal stages comprises a dilated convolution stage in series with one or more reweighted dilated convolution stages; and

perform fusion of an output of the cascade of temporal stages with an output of the boundary supervision branch to predict a surgical phase of the surgical procedure depicted in the frame.

2 . The system of claim 1 , wherein the boundary supervision branch is trained to detect a transition condition between surgical phases.

3 . The system of claim 2 , wherein the boundary supervision branch is trained in parallel with the cascade of temporal stages according to a loss function that changes over a plurality of epochs.

4 . The system of claim 1 , wherein the frame is encoded by an encoder trained using self-supervised learning.

5 . The system of claim 4 , wherein the self-supervised learning comprises a student network that attempts to predict from an augmented image and a target signal generated by a teacher network for a same image under different augmentation, wherein only the student network is updated through backpropagation, while an exponential moving average is used to update the teacher network.

6 . The system of claim 5 , wherein phase labels are used to update the student network and a contribution of phase supervision is reduced as training progresses according to a classification loss function.

7 . The system of claim 1 , wherein the video of the surgical procedure is captured by an endoscopic camera from inside of a patient's body.

8 . The system of claim 1 , wherein the video of the surgical procedure is captured by a camera from outside of a patient's body.

9 . A computer-implemented method comprising:

encoding a frame of a video of a surgical procedure into a plurality of features;

providing the features to a boundary supervision branch of a model;

providing the features to a cascade of temporal stages of the model, wherein the cascade of temporal stages comprises a dilated convolution stage in series with one or more reweighted dilated convolution stages, and the boundary supervision branch is trained to detect a transition condition between surgical phases; and

performing fusion of an output of the cascade of temporal stages with an output of the boundary supervision branch to predict the surgical phase of the surgical procedure depicted in the frame.

10 . The method of claim 9 , wherein the boundary supervision branch is trained in parallel with the cascade of temporal stages according to a loss function that changes over a plurality of epochs for both the boundary supervision branch and the cascade of temporal stages.

11 . The method of claim 10 , wherein the frame is encoded by an encoder trained using self-supervised learning.

12 . The method of claim 11 , wherein the self-supervised learning comprises a student network that attempts to predict from an augmented image and a target signal generated by a teacher network for a same image under different augmentation.

13 . The method of claim 12 , wherein only the student network is updated through backpropagation, while an exponential moving average is used to update the teacher network, and phase labels are used to update the student network.

14 . The method of claim 13 , wherein a contribution of phase supervision is reduced as training progresses according to a classification loss function.

15 . The method of claim 9 , wherein the video of the surgical procedure is captured by an endoscopic camera from inside of a patient's body.

16 . The method of claim 9 , wherein the video of the surgical procedure is captured by a camera from outside of a patient's body.

17 . A computer program product comprising a non-transitory memory device with computer-readable instructions stored thereon, wherein executing the computer-readable instructions by one or more processing units causes the one or more processing units to perform a plurality of operations comprising:

encoding a frame of a video of a surgical procedure into a plurality of features, wherein the features comprise one or more structures and/or events in the surgical procedure;

providing the features to a boundary supervision branch trained to detect a transition between surgical phases;

providing the features to a cascade of temporal stages of a model configured to adjust a weight for one or more frames based on a confidence score from a previous stage; wherein the cascade of temporal stages comprises a dilated convolution stage in series with one or more reweighted dilated convolution stages; and

performing fusion of an output of the cascade of temporal stages with an output of the boundary supervision branch to predict the surgical phase of the surgical procedure depicted in the frame.

18 . The computer program product of claim 17 , wherein the boundary supervision branch is trained in parallel with the cascade of temporal stages according to a loss function that changes over a plurality of epochs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 4, 2023
From: ZHANG, JINGLU; KADKHODAMOHAMMADI, ABDOLRAHIM; LUENGO MUNTION, IMANOL; STOYANOV, DANAIL V.; BARBARISI, SANTIAGO
To: DIGITAL SURGERY LIMITED
Reel/Frame 063212/0764 →
Continuity (2)
Provisional Application 63328712 · Apr 7, 2022
Related Publication 20230326207A1 · Oct 12, 2023
References Cited (18)
US 10791301B1 · Garcia Kilroy · 2020 [cited by examiner]
US 20200349711A1 · Duke · 2020 [cited by examiner]
US 20200405406A1 · Harris · 2020 [cited by examiner]
US 20220076115A1 · Nam · 2022 [cited by examiner]
US 20220160433A1 · Rafii-Tari · 2022 [cited by examiner]
US 20220296334A1 · Fathollahi Ghezelghieh · 2022 [cited by examiner]
Lei et al, (“Temporal Deformable Residual Networks for Action Segmentation in Videos”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4742-4751) (Year: 2018). [cited by examiner]
Beatrice Van Amsterdam et al., “Gesture Recognition in Robotic Surgery: A Review”, arxiv.org, Jan. 29, 2021, 15 pages. [cited by applicant]
Extended European Search Report for Application No. 23166760.1-1207; Mailing Date: Aug. 24, 2023; 13 pages. [cited by applicant]
Odysseas Zisimopoulos et al., “DeepPhase: Surgical Phase Recognition in Cataracts Videos”, arxiv.org, Jul. 17, 2018, 8 pages. [cited by applicant]
Peng Lei et al, “Temporal Deformable Residual Networks for Action segmentation in Videos”, IEEE CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, 10 pages. [cited by applicant]
Tobias Czempiel et al., “TeCNO: Surgical Phase Recognition with Multi-Stage Temporal Convolutional Networks”, arxiv.org, Mar. 24, 2020, 10 pages. [cited by applicant]
Tobias Ross et al., “Exploiting the potential of unlabeled endoscopic video data with self-supervised learning”, Inter. Journal of Computer Assisted Radiology and Surbery, Apr. 27, 2018, Springer, 9 pages. [cited by applicant]
Kinpeng Ding et al., “Exploring Segment-level Semantics for Online Phase Recognition from Surgical Videos” arxiv.org, Mar. 20, 2022, 11 pages. [cited by applicant]
Zhang, Jinglu et al., “Self-Knowledge Distillation for Surgical Phase Recognition”, arxiv.org, Jun. 15, 2023, pp. 1-13. [cited by applicant]
Farha et al., “MS-TCN: Multi-Stage Temporal Convolutional Network for ActionSegmentation”, Proceedings of the IEEE/CVF Conference on Computer Vision and PatternRecognition, 2019; 10 pages. [cited by applicant]
Hu et al., “Squeeze-and-Excitation Networks”, Proceedings of the IEEEconference on computer vision and pattern recognition, 2018; 13 pages. [cited by applicant]
Wang et al, “Boundary-Aware Cascade Networks for Temporal Action Segmentation”, European Conference on Computer Vision. Springer, Cham, 2020.; 17 pages. [cited by applicant]