IP Library › Granted Patent US 12,682,637
Granted Patent B2
US 12,682,637 · App. 18/281,968 · Granted Jul 14, 2026

Adaptive visualization of contextual targets in surgical video

Inventors: Ricardo Sanchez-Matilla (London, GB); Maria Ruxandra Robu (London, GB); Imanol Luengo Muntion (London, GB); Danail V. Stoyanov (London, GB); David Owen (London, GB); Maria Grammatikopoulou (London, GB)
Assignee: DIGITAL SURGERY LIMITED
G06V20/41A61B90/37G06T7/20G06V10/25G06V10/62G06V10/774G06V10/806G06V10/82G06V20/20G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/20104G06T2207/30004G06T2207/30201G06V2201/03
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,682,637
App. No.
18/281,968
Filed
Sep 14, 2023
Granted
Jul 14, 2026
Kind
B2
Art Unit
2615
USPC
345/581
Abstract

An aspect includes a computer-implemented method that predicts a proposed region of interest in an image from a video of a surgical procedure based on one or more contextual targets. An image adjustment is synthesized based on the proposed region of interest and the image. A modified visualization of the surgical procedure is generated by incorporating the image adjustment in a real-time output of the video of the surgical procedure. The video of the surgical procedure is displayed with the modified visualization.

Claims (39)

1 . A computer-implemented method comprising:

extracting, by a system comprising a machine-learning processing system and a procedural control system, as a prediction of a first machine-learning model of the machine-learning processing system or as a user input received by the procedural control system, a proposed region of interest in an image from a video of a surgical procedure based on one or more contextual targets, wherein the procedural control system is configured to receive the user input as a drawing input from a connected device to identify the proposed region of interest as an alternate input source, wherein the first machine-learning model comprises a surgical phase and structure network configured to determine a phase of the surgical procedure in the image, and the one or more contextual targets are determined based on one or more outputs of the surgical phase and structure network;

overlaying a contour around the proposed region of interest as an outline without modifying the image within the proposed region of interest;

modifying the proposed region of interest based on a change to an input that generates the contour;

synthesizing an image adjustment, by a second machine-learning model, based on the proposed region of interest and the image after modifying the proposed region of interest;

generating a modified visualization of the surgical procedure by incorporating the image adjustment in a real-time output of the video of the surgical procedure; and

displaying the video of the surgical procedure with the modified visualization.

2 . The computer-implemented method of claim 1 , wherein the first machine learning model uses weak labels, and the second machine learning model uses weak labels and joint detection and segmentation.

3 . The computer-implemented method of claim 1 , further comprising:

determining motion based on temporal data associated with the image; and

determining a current area of focus as at least a portion of the proposed region of interest based on the motion.

4 . The computer-implemented method of claim 1 , further comprising:

using a depth map to refine the proposed region of interest.

5 . The computer-implemented method of claim 1 , wherein the first machine-learning model is trained based on a training dataset of a plurality of temporally aligned annotated data streams comprising temporal annotations, spatial annotations, and sensor annotations.

6 . The computer-implemented method of claim 1 , further comprising:

performing feature fusion to combine one or more task-specific features of the surgical procedure with one or more temporally aligned features spanning two or more frames.

7 . The computer-implemented method of claim 1 , wherein the user input comprises a drawing input received from one or more devices.

8 . The computer-implemented method of claim 1 , further comprising:

performing eye tracking of a surgeon during the surgical procedure; and

predicting the proposed region of interest based at least in part on a detected area of focus from the eye tracking of the surgeon.

9 . A system comprising:

a data collection system configured to capture a video of a surgical procedure;

a model execution system configured to execute one or more machine-learning models to predict a proposed region of interest in an image from the video of the surgical procedure based on one or more contextual targets, wherein the one or more machine-learning models are configured to use a depth map to refine the proposed region of interest by providing the depth map as an input to an encoder to perform feature fusion with a feature space of an encoding of the image prior to a feature decoder that is used to predict the proposed region of interest; and

an output generator configured to generate a modified visualization of the surgical procedure in a real-time output of the video of the surgical procedure based on the proposed region of interest, wherein the modified visualization comprises an adjustment to one or more enhancement options of: a brightness, a contrast, sharpness, and a size ratio within the proposed region of interest relative to one or more background structures, and the one or more enhancement options are user selectable through a user input.

10 . The system of claim 9 , wherein the system is further configured to determine motion based on temporal data and determine a current area of focus as at least a portion of the proposed region of interest based on the motion.

11 . The system of claim 9 , wherein the one or more machine-learning models are configured to perform feature fusion to combine one or more task-specific features of the surgical procedure with one or more temporally aligned features spanning two or more frames of the video.

12 . The system of claim 9 , further comprising a display configured to output the modified visualization comprising an image adjustment within the proposed region of interest.

13 . A computer program product comprising a non-transitory memory device having computer executable instructions stored thereon, which when executed by one or more processors cause the one or more processors to perform a method comprising:

identifying a proposed region of interest in an image from a video of a surgical procedure;

synthesizing an image adjustment, by one or more machine-learning models, based on the proposed region of interest and the image, wherein the image adjustment is concentrated with a greater intensity near a centroid of the proposed region of interest to appear as a virtual light source, and wherein the one or more machine-learning models are configured to use a depth map to refine the proposed region of interest, the depth map comprises an estimate of three-dimensional depth based on one or more two-dimensional images, and the depth map is provided as an input to an encoder to perform feature fusion with a feature space of an encoding of the image prior to a feature decoder that is used to identify the proposed region of interest; and

generating a modified visualization of the surgical procedure by incorporating the image adjustment in an output of the video of the surgical procedure.

14 . A computer program product comprising a non-transitory memory device having computer executable instructions stored thereon, which when executed by one or more processors cause the one or more processors to perform a method comprising:

identifying a proposed region of interest in an image from a video of a surgical procedure based on one or more contextual targets;

determining motion based on temporal data associated with the image and a current area of focus as at least a portion of the proposed region of interest based on the motion, wherein the one or more contextual targets are determined based on one or more outputs of a surgical phase and structure network;

synthesizing an image adjustment, by one or more machine-learning models, based on the proposed region of interest and the image, wherein the image adjustment is concentrated with a greater intensity near a centroid of the proposed region of interest to appear as a virtual light source; and

generating a modified visualization of the surgical procedure by incorporating the image adjustment in an output of the video of the surgical procedure.

15 . The computer program product of claim 13 , wherein the one or more machine-learning models are trained to perform feature fusion to combine one or more task-specific features of the surgical procedure with one or more temporally aligned features spanning two or more frames, and the feature fusion is based on one or more transform-domain fusion algorithms to implement an image fusion neural network.

16 . The computer program product of claim 13 , wherein the proposed region of interest is identified based on a user input received as a drawing input from a connected device.

17 . The computer program product of claim 13 , wherein execution of the computer executable instructions causes the one or more processors to perform eye tracking of a surgeon during the surgical procedure to track a position of gaze relative to a region of an image being observed and predict the proposed region of interest based at least in part on a detected area of focus from the eye tracking of the surgeon.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2023
From: STOYANOV, DANAIL V.; LUENGO MUNTION, IMANOL; OWEN, DAVID; GRAMMATIKOPOULOU, MARIA; SANCHEZ-MATILLA, RICARDO; ROBU, MARIA RUXANDRA
To: DIGITAL SURGERY LIMITED
Reel/Frame 064898/0835 →
Continuity (4)
Provisional Application 63212157 · Jun 18, 2021
Provisional Application 63163425 · Mar 19, 2021
Provisional Application 63163417 · Mar 19, 2021
Related Publication 20240303984A1 · Sep 12, 2024
References Cited (120)
US 9123155B2 · Cunningham et al. · 2015 [cited by applicant]
US 9788907B1 · Alvi et al. · 2017 [cited by applicant]
US 10152789B2 · Carnes et al. · 2018 [cited by applicant]
US 10251714B2 · Carnes et al. · 2019 [cited by applicant]
US 10356385B2 · Petrichkovich · 2019 [cited by examiner]
US 10383694B1 · Venkataraman · 2019 [cited by applicant]
US 10579878B1 · Bauer · 2020 [cited by applicant]
US 10956790B1 · Victoroff · 2021 [cited by examiner]
US 11189379B2 · Giataganas et al. · 2021 [cited by applicant]
US 11246664B2 · Andrews · 2022 [cited by examiner]
US 11263772B2 · Siemionow et al. · 2022 [cited by applicant]
US 11311342B2 · Parihar et al. · 2022 [cited by applicant]
US 11322248B2 · Grantcharov · 2022 [cited by examiner]
US 11370113B2 · Habbecke et al. · 2022 [cited by applicant]
US 11464583B2 · Sano et al. · 2022 [cited by applicant]
US 11497557B2 · Haslam · 2022 [cited by examiner]
US 11564756B2 · Shelton, IV et al. · 2023 [cited by applicant]
US 11602398B2 · Regensburger · 2023 [cited by examiner]
US 11622818B2 · Siemionow et al. · 2023 [cited by applicant]
US 11910995B2 · Hendriks · 2024 [cited by examiner]
US 11915378B2 · Koza · 2024 [cited by examiner]
US 11977998B2 · Stiller et al. · 2024 [cited by applicant]
US 12062442B2 · Shelton, IV · 2024 [cited by applicant]
US 12136220B2 · Vasilev · 2024 [cited by examiner]
US 12150719B2 · Wright · 2024 [cited by examiner]
US 12161430B2 · Wright · 2024 [cited by examiner]
US 12198330B2 · Blau · 2025 [cited by examiner]
US 12217427B2 · Schreckenberg · 2025 [cited by examiner]
US 12220174B2 · Khan · 2025 [cited by examiner]
US 12226163B2 · Besier · 2025 [cited by examiner]
US 12290938B2 · Wright · 2025 [cited by examiner]
US 12327351B2 · Pai Raikar et al. · 2025 [cited by applicant]
US 12360351B2 · Segev et al. · 2025 [cited by applicant]
US 12380998B2 · Giataganas et al. · 2025 [cited by applicant]
US 20070136218A1 · Bauer · 2007 [cited by applicant]
US 20090036902A1 · Dimaio · 2009 [cited by applicant]
US 20100167248A1 · Ryan · 2010 [cited by applicant]
US 20120020547A1 · Zhao et al. · 2012 [cited by applicant]
US 20130288214A1 · Kesavadas · 2013 [cited by applicant]
US 20130295540A1 · Kesavadas · 2013 [cited by applicant]
US 20140107471A1 · Haider et al. · 2014 [cited by applicant]
US 20140286533A1 · Luo · 2014 [cited by applicant]
US 20140350391A1 · Prisco et al. · 2014 [cited by applicant]
US 20150297313A1 · Reiter · 2015 [cited by applicant]
US 20170020395A1 · Malchano et al. · 2017 [cited by applicant]
US 20170105713A1 · Frimer · 2017 [cited by applicant]
US 20180032130A1 · Meglan · 2018 [cited by applicant]
US 20180071032A1 · De Almeida Barreto · 2018 [cited by applicant]
US 20180243906A1 · Hourtash · 2018 [cited by applicant]
US 20180247128A1 · Alvi et al. · 2018 [cited by applicant]
US 20180271603A1 · Nir · 2018 [cited by applicant]
US 20180296075A1 · Meglan et al. · 2018 [cited by applicant]
US 20180310811A1 · Meglan et al. · 2018 [cited by applicant]
US 20180310875A1 · Meglan et al. · 2018 [cited by applicant]
US 20180322949A1 · Mohr et al. · 2018 [cited by applicant]
US 20180325604A1 · Atarot · 2018 [cited by applicant]
US 20190008587A1 · Allison et al. · 2019 [cited by applicant]
US 20190029766A1 · Griffiths et al. · 2019 [cited by applicant]
US 20190038362A1 · Nash et al. · 2019 [cited by applicant]
US 20190046276A1 · Inglese · 2019 [cited by examiner]
US 20190053872A1 · Meglan · 2019 [cited by applicant]
US 20190088162A1 · Meglan · 2019 [cited by applicant]
US 20190090954A1 · Kotian et al. · 2019 [cited by applicant]
US 20190099226A1 · Hallen · 2019 [cited by applicant]
US 20190251723A1 · Coppersmith, III · 2019 [cited by examiner]
US 20190325574A1 · Jin et al. · 2019 [cited by applicant]
US 20200036797A1 · Kudelski · 2020 [cited by applicant]
US 20200226751A1 · Jin et al. · 2020 [cited by applicant]
US 20210015554A1 · Chow · 2021 [cited by applicant]
US 20210183124A1 · Benditte-Klepetko · 2021 [cited by examiner]
US 20220000565A1 · Gururaj · 2022 [cited by examiner]
US 20220044440A1 · Blau · 2022 [cited by examiner]
US 20230024942A1 · Wright · 2023 [cited by examiner]
US 20230029184A1 · Wright · 2023 [cited by examiner]
US 20230047100A1 · Arcadu · 2023 [cited by examiner]
US 20230074481A1 · Pai Raikar · 2023 [cited by examiner]
US 20230098859A1 · Kitamura et al. · 2023 [cited by applicant]
US 20230105111A1 · Nickel · 2023 [cited by examiner]
US 20230268051A1 · Wachs et al. · 2023 [cited by applicant]
US 20240156547A1 · Luengo Muntion et al. · 2024 [cited by applicant]
US 20240161497A1 · Luengo Muntion et al. · 2024 [cited by applicant]
US 20240169579A1 · Luengo Muntion et al. · 2024 [cited by applicant]
US 20240206989A1 · Luengo Muntion et al. · 2024 [cited by applicant]
CA 3107582A1 · 2020 [cited by applicant]
CN 101696943A · 2010 [cited by applicant]
CN 103142313B · 2015 [cited by applicant]
CN 108024833A · 2018 [cited by applicant]
CN 111261252A · 2020 [cited by applicant]
JP 2018029961A · 2018 [cited by applicant]
WO 2017098504A · 2017 [cited by applicant]
WO 2018005842A1 · 2018 [cited by applicant]
WO 2018217433A1 · 2018 [cited by applicant]
WO 2018217444A2 · 2018 [cited by applicant]
WO 2018237187A2 · 2018 [cited by applicant]
WO 2019036006A1 · 2019 [cited by applicant]
WO 2019036007A1 · 2019 [cited by applicant]
WO 2019050729A1 · 2019 [cited by applicant]
WO 2019050886A1 · 2019 [cited by applicant]
WO 2019139949A1 · 2019 [cited by applicant]
WO 2019217893A1 · 2019 [cited by applicant]
WO 2020009830A1 · 2020 [cited by applicant]
WO 2020102584A2 · 2020 [cited by applicant]
WO 2021058294A1 · 2021 [cited by applicant]
Colleoni et al., “Deep Learning Based Robotic Tool Detection and Articulation Estimation with Spatio-Temporal ayers”, IEEE Robotics and Automation Letters, Jul. 1, 2019, 7 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/GR2022/000015; International Filing Date Mar. 18, 2022; Date of Mailing: Jun. 14, 2022; 11 pages. [cited by applicant]
Islam et al.,“AP-MTL: Attention Pruned Multi-task Learning Model for Real-time Instrument Detection and Segmentation in Robot-assisted Surgery”, 2020 IEEE International Conference on Robotics and Automation, May 31, 202… [cited by applicant]
Jin et al.,“Multi-task recurrent convolutional network with correlation loss for surgical video analysis”, Medical Image Analysis, Oct. 2019, 31 pages. [cited by applicant]
Sarikaya et al., “Detection and Localization of Robotic Tools in Robot-Assisted Surgery Videos Using Deep Neural Networks for Region Proposal and Detection”, Cornell University Library, Jul. 2020, 8 pages. [cited by applicant]
Yang et el., Image-based laparoscopic tool detection and tracking using convolutional neural networks: a review of the literature, Computer Assisted Surgery, Jan. 2020, pp. 15-28. [cited by applicant]
Extended European Search Report for Application No. 19220289.3-1207 issued Nov. 24, 2020, 12 pages. [cited by applicant]
Hashimoto, et al., Artificial Intelligence in Surgery: Promises and Perils, Jul. 2018, 15 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/GR2022/000013; International Filing Date Mar. 18, 2022; Date of Mailing: Jun. 22, 2022; 9 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/GR2022/000014; International Filing Date Mar. 18, 2022; Date of Mailing: Jun. 20, 2022; 8 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/GR2022/000016; International Filing Date Mar. 18, 2022; Date of Mailing: Jun. 22, 2022; 14 pages. [cited by applicant]
Twinanda et al, “EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos”, IEEE Transactions, Jan. 2017; 11 pages. [cited by applicant]
Yidan et al., “daVinciNet: Joint Prediction of Motion and Surgical State in Robot-Assisted Surgery” 2020 IEEE International Conference; Oct. 2020; 8 pages. [cited by applicant]
Yidan et al., “Temporal Segmentation of Surgical Sub-tasks through Deep Learning with Multiple Data Sources”, 2020 EEE International Conference, May 2020, 12 pages. [cited by applicant]
Qin et al. “daVinciNet: Joint Prediction of Motion and Surgical Sate in Robot-Assisted Surgery” 2020, arxiv.org, Cornell University Library, XP081771349 (16 Pages). [cited by applicant]
European Examination Report, Dated: Aug. 5, 2025, Application No. 22714529.9-1218, Filed: Mar. 18, 2022, 9 pages. [cited by applicant]
Volkov et al., “Machine Learning and Coresets for Automated Real-Time Video Segmentation of Laproscopic and Robot-Assisted Surgery”, IEEE International Conference on Robotics and Automation, 2017, 6 pgs. [cited by applicant]