IP Library Granted Patent US 12,562,255
Granted Patent B2
US 12,562,255 · App. 18/596,298 · Granted Feb 24, 2026

Using multiple modalities of surgical data for comprehensive data analytics of a surgical procedure

Inventors: Jagadish Venkataraman (Menlo Park, CA); Pablo Garcia Kilroy (Menlo Park, CA)
Assignee: Verb Surgical Inc.
G16H20/40G06F16/45G16H30/40G16H40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,255
App. No.
18/596,298
Granted
Feb 24, 2026
Kind
B2
Abstract

This patent disclosure provides various embodiments of combining multiple modalities of non-text surgical data in forms of videos, images, and audios in a meaningful manner so that the combined data can be used to perform comprehensive data analytics for a surgical procedure. In some embodiments, the disclosed system can begin by receiving two or more modalities of surgical data during the surgical procedure. The system then time-synchronizes the two or more modalities of surgical data to generate two or more modalities of time-synchronized surgical data. Next, the system converts each modality of the time-synchronized surgical data into a corresponding array of values of a common format. The system then combines the two or more arrays of values to generate a combined set of values. The system subsequently performs comprehensive data analytics on the combined set of values to generate a surgical decision for the surgical procedure.

Claims (49)

1 . A computer-implemented method for performing comprehensive data analytics for a surgical procedure, the computer-implemented method comprising:

time-synchronizing multiple modalities of surgical data during a surgical procedure, the multiple modalities of surgical data including a first modality of surgical data and a second modality of surgical data;

converting, by a first segmentation engine, the first modality of surgical data into a first set of alphanumeric values during the surgical procedure, the first modality of surgical data including an audio feed being captured during the surgical procedure;

converting, by a second segmentation engine, the second modality of surgical data into a second set of alphanumeric values during the surgical procedure, the second modality of surgical data including a video feed being captured during the surgical procedure;

combining the first set of alphanumeric values and the second set of alphanumeric values to generate a prediction during the surgical procedure, wherein generating the prediction includes translating an identifiable location of an event detected in the audio feed to a corresponding location of the event within the video feed; and

based on the prediction, highlighting the corresponding location of the event within the video feed.

2 . The computer-implemented method of claim 1 , wherein the audio feed includes a surgeon narrating or discussing the surgical procedure in real-time.

3 . The computer-implemented method of claim 1 , wherein generating the prediction further includes determining a size of an anomaly, an occurrence of a smoking event, or an occurrence of a bleeding event.

4 . The computer-implemented method of claim 1 , wherein the video feed includes real-time video captured by one or more cameras inside a patient's body or near a surgical site, the one or more cameras including an endoscope camera.

5 . The computer-implemented method of claim 1 , further comprising:

time-synchronizing the multiple modalities of surgical data during the surgical procedure, the multiple modalities of surgical data including a third modality of surgical data;

converting, by a third segmentation engine, the third modality of surgical data into a third set of alphanumeric values during the surgical procedure; and

combining the first set of alphanumeric values, the second set of alphanumeric values, and the third set of alphanumeric values to generate the prediction during the surgical procedure.

6 . The computer-implemented method of claim 5 , wherein the third modality is image data associated with the surgical procedure, the image data being generated from pre-op imaging, intra-op real-time imaging, and post-op imaging.

7 . The computer-implemented method of claim 5 further comprising:

combining text data, the first set of alphanumeric values, the second set of alphanumeric values, and the third set of alphanumeric values to generate the prediction related to the event during the surgical procedure, wherein the text data includes patient vitals, patient information, surgeon information, hospital statistics, treatment plans, medication data, test results and progress notes.

8 . A non-transitory computer-readable medium to store instructions that, when executed by a processor of a computer system, cause the computer system to:

time-synchronize multiple modalities of surgical data during a surgical procedure, the multiple modalities of surgical data including a first modality of surgical data and a second modality of surgical data;

convert, by a first segmentation engine, the first modality of surgical data into a first set of alphanumeric values during the surgical procedure, the first modality of surgical data including an audio feed being captured during the surgical procedure;

convert, by a second segmentation engine, the second modality of surgical data into a second set of alphanumeric values during the surgical procedure, the second modality of surgical data including a video feed being captured during the surgical procedure;

combine the first set of alphanumeric values and the second set of alphanumeric values to generate a prediction during the surgical procedure, wherein generating the prediction includes translating an identifiable location of an event detected in the audio feed to a corresponding location of the event within the video feed; and

based on the prediction, highlight the corresponding location of the event within the video feed.

9 . The non-transitory computer-readable medium of claim 8 , wherein, the audio feed includes a surgeon narrating or discussing the surgical procedure in real-time.

10 . The non-transitory computer-readable medium of claim 8 , wherein generating the prediction further includes determining a size of an anomaly, an occurrence of a smoking event, or an occurrence of a bleeding event.

11 . The non-transitory computer-readable medium of claim 8 , wherein the video feed includes real-time video captured by one or more cameras inside a patient's body or near a surgical site, the one or more cameras including an endoscope camera.

12 . The non-transitory computer-readable medium of claim 8 , further comprising instructions that, when executed by the processor of the computer system, cause the computer system to:

time-synchronize the multiple modalities of surgical data during the surgical procedure, the multiple modalities of surgical data including a third modality of surgical data;

convert, by a third segmentation engine, the third modality of surgical data into a third set of alphanumeric values during the surgical procedure; and

combine the first set of alphanumeric values, the second set of alphanumeric values, and the third set of alphanumeric values to generate the prediction during the surgical procedure.

13 . The non-transitory computer-readable medium of claim 12 , wherein the third modality of is image data associated with a surgical procedure, the image data being generated from pre-op imaging, intra-op real-time imaging, and post-op imaging.

14 . The non-transitory computer-readable medium of claim 12 further comprising instructions that, when executed by the processor of the computer system, cause the computer system to:

combine text data, the first set of alphanumeric values, the second set of alphanumeric values, and the third set of alphanumeric values to generate the prediction related to the event during the surgical procedure, wherein the text data includes patient vitals, patient information, surgeon information, hospital statistics, treatment plans, medication data, test results and progress notes.

15 . A system for performing comprehensive data analytics for a surgical procedure, the system comprising:

one or more processors; and

a memory coupled to the one or more processors, wherein the memory stores a set of instructions that, when executed by the one or more processors, cause the system to:

time-synchronize multiple modalities of surgical data during a surgical procedure, the multiple modalities of surgical data including a first modality of surgical data and a second modality of surgical data;

convert, by a first segmentation engine, the first modality of surgical data into a first set of alphanumeric values during the surgical procedure, the first modality of surgical data including an audio feed being captured during the surgical procedure;

convert, by a second segmentation engine, the second modality of surgical data into a second set of alphanumeric values during a surgical procedure, the second modality of surgical data including a video feed being captured during the surgical procedure;

combine the first set of alphanumeric values and the second set of alphanumeric values to generate a prediction during the surgical procedure, wherein generating the prediction includes translating an identifiable location of an event detected in the audio feed to a corresponding location of the event within the video feed; and

based on the prediction, highlight the corresponding location of the event within the video feed.

16 . The system of claim 15 , wherein the audio feed includes a surgeon narrating or discussing the surgical procedure in real-time.

17 . The system of claim 15 , wherein generating the prediction further includes determining a size of an anomaly, an occurrence of a smoking event, or an occurrence of a bleeding event.

18 . The system of claim 15 , wherein the video feed includes real-time video captured by one or more cameras inside a patient's body or near a surgical site, the one or more cameras including an endoscope camera.

19 . The system of claim 15 , further comprising instructions that, when executed by the one or more processors, cause the system to:

time-synchronize the multiple modalities of surgical data during the surgical procedure, the multiple modalities of surgical data including a third modality of surgical data;

convert, by a third segmentation engine, the third modality of surgical data into a third set of alphanumeric values during a surgical procedure; and

combine the first set of alphanumeric values, the second set of alphanumeric values, and the third set of alphanumeric values to generate the prediction during the surgical procedure.

20 . The system of claim 19 , further comprising instructions that, when executed by the one or more processors, cause the system to:

combine text data, the first set of alphanumeric values, the second set of alphanumeric values, and the third set of alphanumeric values to generate the prediction related to the event during the surgical procedure, wherein the third modality of is image data associated with the surgical procedure, the image data being generated from pre-op imaging, intra-op real-time imaging, and post-op imaging, and text data includes patient vitals, patient information, surgeon information, hospital statistics, treatment plans, medication data, test results and progress notes.

Assignments (1)
MERGER Recorded Jan 27, 2026
From: VERB SURGICAL INC.
To: AURIS HEALTH, INC.
Reel/Frame 073601/0790 →
Continuity (3)
Continuation 17493589 · Oct 4, 2021
Continuation 16418790 · May 21, 2019
Related Publication 20240290459A1 · Aug 29, 2024
References Cited (28)
US 9836654B1 · Alvi et al. · 2017 [cited by applicant]
US 20080062280A1 · Wang et al. · 2008 [cited by applicant]
US 20110273309A1 · Zhang et al. · 2011 [cited by applicant]
US 20130041685A1 · Yegnanarayanan · 2013 [cited by applicant]
US 20150164436A1 · Maron et al. · 2015 [cited by applicant]
US 20160342744A1 · Joao · 2016 [cited by applicant]
US 20170019529A1 · Bostick et al. · 2017 [cited by applicant]
US 20180122506A1 · Grantcharov · 2018 [cited by examiner]
US 20190206562A1 · Shelton, IV et al. · 2019 [cited by applicant]
US 20200273581A1 · Wolf · 2020 [cited by examiner]
CN 105992996A · 2016 [cited by applicant]
JP 2011167301A · 2011 [cited by applicant]
KR 1020110092350A · 2011 [cited by applicant]
KR 101881862B1 · 2018 [cited by applicant]
WO WO2017075541A1 · 2017 [cited by examiner]
WO 2017220788A1 · 2017 [cited by applicant]
Volkov, Mikhail, et al. “Machine learning and coresets for automated real-time video segmentation of laparoscopic and robot-assisted surgery.” 2017 IEEE international conference on robotics and automation (ICRA). IEEE, … [cited by examiner]
Office Action received for Korean Patent Application No. 10-2021-7040698, mailed on Mar. 20, 2024, 18 pages (6 pages of English Translation and 12 pages of Original Document). [cited by applicant]
Office Action received for Chinese Patent Application No. 201980096658.8, mailed on Mar. 12, 2025, 20 pages (11 Pages of Original Document and 9 Pages of English Translation). [cited by applicant]
Second Office Action received for Chinese Patent Application No. 201980096658.8, mailed on Jun. 19, 2025, 19 pages (11 Pages of Original Document and 8 Pages of English Translation). [cited by applicant]
Song et al., “Principles and Application Technology of Sensors,” Southwest Jiaotong University Press, pp. 277-278 (2 Pages of Original Document and 3 Pages of English Translation) pp. 277-278. [cited by applicant]
First Office Action received for Chinese Patent Application No. 201980096658.8, mailed Mar. 12, 2025, 22 pages (11 Pages of Original Document and 11 Pages of English Translation). [cited by applicant]
Non-Final Office Action received for U.S. Appl. No. 16/418,790, mailed on Dec. 30, 2020, 24 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2019/034063 mailed Feb. 21, 2020, 10 pages. [cited by applicant]
International Preliminary Report on Patentability for International Application No. PCT/US2019/034063 mailed Dec. 2, 2021, 6 pages. [cited by applicant]
Extended European Search Report for European Application No. 19929554.4 mailed Dec. 21, 2022, 7 pages. [cited by applicant]
Merloz, Phillippe, et al; “Basic Concept in Computer Assisted Surgery”; Chinese Journal of Reparative and Reconstructive Surgery, vol. 20 (3), 2006, 99-276-278. [cited by applicant]
Notice of Allowance received for Korean Application No. 10-2021-7040698, mailed on Jun. 17, 2024, 3 pages (2 pages of original office action and 1 page of English Translation). [cited by applicant]