IP Library Granted Patent US 12,283,196
Granted Patent B2
US 12,283,196 · App. 17/552,886 · Granted Apr 22, 2025

Surgical simulator providing labeled data

Inventors: Xing Jin (San Jose, CA); Joëlle K. Barral (Mountain View, CA); Lin Yang (Sunnyvale, CA); Martin Habbecke (Palo Alto, CA); Gianni Campion (Mountain View, CA)
Assignee: Verily Life Sciences LLC
G09B23/28G06N3/08G06T7/0012G06V10/454G06V10/751G06V10/764G06V10/82G06V20/20G06V20/41G06T2207/20081G06T2207/20084G06T2207/30004
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,283,196
App. No.
17/552,886
Granted
Apr 22, 2025
Kind
B2
Abstract

A surgical simulator for simulating a surgical scenario comprises a display system, a user interface, and a controller. The controller includes one or more processors coupled to memory that stores instructions that when executed cause the system to perform operations. The operations include generating simulated surgical videos, each representative of the surgical scenario. The operations further include associating simulated ground truth data from the simulation with the simulated surgical videos. The ground truth data corresponds to context information of at least one of a simulated surgical instrument, a simulated anatomical region, a simulated surgical task, or a simulated action. The operations further include annotating features of the simulated surgical videos based, at least in part, on the simulated ground truth data for training a machine learning model.

Claims (46)

1. A method comprising:

generating, with a surgical simulator, simulated surgical videos each representative of a simulation of a surgical scenario;

associating simulated ground truth data from the simulation with the simulated surgical videos, wherein the simulated ground truth data corresponds to context information of at least one of a simulated surgical instrument, a simulated anatomical region, or a simulated surgical task simulated by the surgical simulator;

annotating features of the simulated surgical videos based, at least in part, on the simulated ground truth data to generate annotated training data for training a machine learning model;

pre-training the machine learning model with the annotated training data to probabilistically identify the features:

providing, to a refiner neural network, simulated images from the simulated surgical videos before the pre-training;

refining the simulated surgical videos with the refiner neural network, wherein the refiner neural network adjusts the simulated images until a discriminator neural network determines the simulated images are comparable to unlabeled real images within a first threshold, wherein the features annotated from the simulated surgical videos are included after the refining, and wherein the unlabeled real images are representative of the surgical scenario in a real environment;

receiving unlabeled real videos from a surgical video database, each of the unlabeled real videos representative of the surgical scenario;

identifying the features from the unlabeled real videos with the machine learning model trained to probabilistically identify the features from the unlabeled real videos, wherein the features include at least one of a separation distance between a surgical instrument and an anatomical region, temporal boundaries of a surgical complication, or spatial boundaries of the surgical complication; and

annotating the unlabeled real videos with the machine learning model by labeling the features of the unlabeled real videos identified by the machine learning model.

2. The method of claim 1 , further comprising:

receiving, from the surgical video database, labeled real surgical videos, each representative of the surgical scenario in the real environment, and wherein the labeled real surgical videos include the features annotated for the machine learning model; and

after pre-training the machine learning model with the simulated surgical videos, training the machine learning model with the labeled real surgical videos to train the machine learning model to probabilistically identify the features from the unlabeled real videos.

3. The method of claim 1 , wherein the context information for each of the simulated surgical videos includes at least one of three-dimensional spatial boundaries of the simulated surgical instrument, three-dimensional spatial boundaries of the simulated anatomical region, motion of the simulated surgical instrument, separation distance between the simulated surgical instrument and the simulated anatomical region, orientation of the simulated surgical instrument, temporal boundaries of one or more surgical steps of the simulated surgical task from the surgical scenario, temporal boundaries of a simulated surgical complication, or spatial boundaries of the simulated surgical complication.

4. The method of claim 1 , further comprising:

rendering markers or identifiers superimposed on top of a corresponding one of the unlabeled real videos to generate annotated surgical videos.

5. The method of claim 1 , further comprising:

displaying annotated surgical videos corresponding to the unlabeled real videos with the features annotated by the machine learning model superimposed on a corresponding one of the unlabeled real videos.

6. The method of claim 1 , wherein the machine learning model includes at least one of a radial basis function neural network, a Kohonen self-organizing neural network, a recurrent neural network, a convolution neural network, or a modular neural network.

7. A non-transitory computer-readable storage medium having stored thereon instructions which, when executed by one or more processing units, cause the one or more processing units to perform operations comprising:

generating, with a surgical simulator, simulated surgical videos each representative of a simulation of a surgical scenario;

associating simulated ground truth data from the simulation with the simulated surgical videos, wherein the simulated ground truth data corresponds to context information of at least one of a simulated surgical instrument, a simulated anatomical region, or a simulated surgical task simulated by the surgical simulator;

annotating features of the simulated surgical videos based, at least in part, on the simulated ground truth data to generate annotated training data for training a machine learning model;

pre-training the machine learning model with the annotated training data to probabilistically identify the features;

receiving, from a surgical video database, labeled real surgical videos, each representative of the surgical scenario in a real environment, and wherein the labeled real surgical videos include the features annotated for the machine learning model; and

after the pre-training the machine learning model with the simulated surgical videos, training the machine learning model with the labeled real surgical videos to train the machine learning model to probabilistically identify the features from unlabeled real videos;

receiving the unlabeled real videos from the surgical video database, each of the unlabeled real videos representative of the surgical scenario;

identifying the features from the unlabeled real videos with the machine learning model trained to probabilistically identify the features from the unlabeled real videos, wherein the features include at least one of a separation distance between a surgical instrument and an anatomical region, temporal boundaries of a surgical complication, or spatial boundaries of the surgical complication; and

annotating the unlabeled real videos with the machine learning model by labeling the features of the unlabeled real videos identified by the machine learning model.

8. The non-transitory computer-readable medium of claim 7 , wherein the instructions, which when executed by the one or more processing units, cause the one or more processing units to perform further operations comprising:

providing, to a refiner neural network, simulated images from the simulated surgical videos before the pre-training; and

refining the simulated surgical videos with the refiner neural network, wherein the refiner neural network adjusts the simulated images until a discriminator neural network determines the simulated images are comparable to unlabeled real images within a first threshold, wherein the features annotated from the simulated surgical videos are included after the refining, and wherein the unlabeled real images are representative of the surgical scenario in the real environment.

9. The non-transitory computer-readable medium of claim 7 , wherein the context information for each of the simulated surgical videos includes at least one of three-dimensional spatial boundaries of the simulated surgical instrument, three-dimensional spatial boundaries of the simulated anatomical region, motion of the simulated surgical instrument, separation distance between the simulated surgical instrument and the simulated anatomical region, orientation of the simulated surgical instrument, temporal boundaries of one or more surgical steps of the simulated surgical task from the surgical scenario, temporal boundaries of a simulated surgical complication, or spatial boundaries of the simulated surgical complication.

10. The non-transitory computer-readable medium of claim 7 , wherein the instructions, which when executed by the one or more processing units, cause the one or more processing units to perform further operations comprising rendering markers or identifiers superimposed on top of a corresponding one of the unlabeled real videos to generate annotated surgical videos.

11. The non-transitory computer-readable medium of claim 7 , wherein the instructions, which when executed by the one or more processing units, cause the one or more processing units to perform further operations comprising displaying annotated surgical videos corresponding to the unlabeled real videos with the features annotated by the machine learning model superimposed on a corresponding one of the unlabeled real videos.

12. The non-transitory computer-readable medium of claim 7 , wherein the machine learning model includes at least one of a radial basis function neural network, a Kohonen self-organizing neural network, a recurrent neural network, a convolution neural network, or a modular neural network.

13. A surgical simulator for simulating a surgical scenario, the surgical simulator comprising:

a display system adapted to show simulated surgical videos to a user of the surgical simulator;

a user interface adapted to correlate a physical action of the user with a simulated action of the surgical simulator;

a controller including one or more processors coupled to memory, the display system, and the user interface, wherein the memory stores instructions that when executed by the one or more processors cause the surgical simulator to perform operations including:

generating the simulated surgical videos, each representative of a simulation of the surgical scenario;

associating simulated ground truth data from the simulation with the simulated surgical videos, wherein the simulated ground truth data corresponds to context information of at least one of a surgical instrument, a anatomical region, a surgical task, or a surgical action simulated by the surgical simulator;

annotating features of the simulated surgical videos based, at least in part, on the simulated ground truth data to generate annotated training data for training a machine learning model to probabilistically identify the features from at least one of unlabeled simulated videos or unlabeled real videos corresponding to the surgical scenario, wherein the features include at least one of a separation distance between the surgical instrument and the anatomical region, temporal boundaries of a surgical complication, or spatial boundaries of the surgical complication;

pre-training the machine learning model with the annotated training data;

receiving, from a surgical video database, real surgical videos representative of the surgical scenario, wherein the real surgical videos include the features annotated for the machine learning model; and

after the pre-training, training the machine learning model with the real surgical videos, wherein the training configures the machine learning model to probabilistically identify the features from the unlabeled real videos corresponding to the surgical scenario.

Assignments (1)
CHANGE OF NAME Recorded Apr 1, 2026
From: VERILY LIFE SCIENCES LLC
To: VERILY HEALTH INC.
Reel/Frame 075367/0775 →
Continuity (3)
Continuation 16373261 · Apr 2, 2019
Provisional Application 62660726 · Apr 20, 2018
Related Publication 20220108450A1 · Apr 7, 2022
References Cited (23)
US 9700292B2 · Nawana et al. · 2017 [cited by applicant]
US 10679046B1 · Black · 2020 [cited by examiner]
US 11232556B2 · Jin · 2022 [cited by examiner]
US 20090263775A1 · Ullrich · 2009 [cited by applicant]
US 20180357514A1 · Zisimopoulos · 2018 [cited by examiner]
WO 2017083768A1 · 2017 [cited by applicant]
Shrivastava, Ashish, et al. “Learning from Simulated and Unsupervised Images through Adversarial Training.” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2017 (Year: 2017). [cited by examiner]
Twinanda, Andru P., et al. “Endonet: a deep architecture for recognition tasks on laparoscopic videos.” IEEE transactions on medical imaging 36.1 (2016): 86-97 (Year: 2016). [cited by examiner]
Gillan, Stewart N., and George M. Saleh. “Ophthalmic surgical simulation: a new era.” JAMA ophthalmology 131.12 (2013): 1623-1624 (Year: 2013). [cited by examiner]
Bousmalis, K., et al., “Unsupervised Pixel-Level Domain Adaptation with Generative Adversarial Networks,” Dec. 16, 2016, Arxiv.org, Cornell University Library, Cornell University, Ithaca, New York, 16 pages. [cited by applicant]
International Search Report and Written Opinion mailed Jul. 9, 2019, issued in corresponding International Application No. PCT/US2019/028389, filed Apr. 19, 2019, 12 pages. [cited by applicant]
Johnson-Roberson, M., et al., “Driving in the Matrix: Can Virtual Worlds Replace Human-Generated Annotations for Real World Tasks?” Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA)… [cited by applicant]
Shrivastava, A., et al., “Learning From Simulated and Unsupervised Images Through Adversarial Training,” Dec. 22, 2016, Arxiv.org, Cornell University Library, Cornell University, Ithaca, New York, 16 pages. [cited by applicant]
“Surgery simulator,” Wikipedia, The Free Dictionary, Mar. 24, 2018, <https://en.wikipedia.org/w/index.php?title=Surgery_simulator&oldid=832275506> [retrieved Jun. 17, 2019], 3 pages. [cited by applicant]
Aerial Informatics and Robotics Platform—Microsoft Research, Retrieved from Internet <https://www.microsoft.com/en-us/research/project/aerial-informatics-robotics-platform/> Mar. 1, 2018, 7 pages. [cited by applicant]
Intuitive Surgical—da Vinci Surgical System—Skills Simulator, Retrieved from Internet <https://www.intuitive.com/en-us/products-and-services/da-vinci/education> Mar. 1, 2018, 3 pages. [cited by applicant]
Shrivastava, A. et al., “Learning from Simulated and Unsupervised Images through Adversarial Training”, arXiv:1612.07828 [cs.CV], Dec. 22, 2016, 16 pages. [cited by applicant]
MSim, Realistic and Life-Like Simulation for all Core Robotic Skills, Retrieved from Internet <https://mimicsimulation.com/products/technology/msim/> Mar. 1, 2018, 8 pages. [cited by applicant]
Bartyzal, R., “Multi-Label Image Classification with Inception Net”, Retrieved from Internet <https://towardsdatascience.com/multi-label-image-classification-with-inception-net-cbb2ee538e30> Mar. 9, 2019, 11 pages. [cited by applicant]
“Minimally Invasive Surgery Simulation”, SenseGraphics, Retrieved from Internet <https://sensegraphics.com/simulation/minimally-invasive-surgery-simulation/> Mar. 1, 2018, 4 pages. [cited by applicant]
“Simulators”, 3D Systems, Retrieved from Internet <https://simbionix.com/simulators/> Mar. 1, 2018, 2 pages. [cited by applicant]
“Surgical Simulation Training Solutions”, CAE Healthcare, Retrieved from Internet <https://caehealthcare.com/surgical-simulation/> Mar. 1, 2018, 8 pages. [cited by applicant]
“Virtual Simulator for Robots”, NVIDIA, <https://www.nvidia.com/en-us/deep-learning-ai/industries/robotics/> Mar. 1, 2018, 6 pages. [cited by applicant]