IP Library Granted Patent US 12,548,159
Granted Patent B2
US 12,548,159 · App. 17/776,073 · Granted Feb 10, 2026

Scene perception systems and methods

Inventors: Omid Mohareri (San Francisco, CA); Simon P. DiMaio (San Carlos, CA); Zhaoshuo Li (Baltimore, MD); Amirreza Shaban (Atlanta, GA); Jean-Gabriel Simard (Montreal, CA)
Assignee: Intuitive Surgical Operations, Inc.
G06T7/11A61B34/20G06T1/0014G06T7/215G06T7/292G06T7/80A61B2034/2065A61B2090/364G06T2207/20081G06T2207/20221G06T2207/30244
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,159
App. No.
17/776,073
Granted
Feb 10, 2026
Kind
B2
Abstract

Scene perception systems and methods are described herein. In certain illustrative examples, a system combines data sets associated with imaging devices included in a dynamic multi-device architecture and uses the combined data sets to perceive a scene (e.g., a surgical scene) imaged by the imaging devices. To illustrate, the system may access tracking data for imaging devices capturing images of a scene and fuse, based on the tracking data, data sets respectively associated with the imaging devices to generate fused sets of data for the scene. The tracking data may represent a change in a pose of at least one of the image devices that occurs while the imaging devices capture images of the scene. The fused sets of data may represent or be used to generate perceptions of the scene. In certain illustrative examples, scene perception is dynamically optimized using a feedback control loop.

Claims (64)

1 . An apparatus comprising:

a memory storing instructions; and

a processor communicatively coupled to the memory and configured to execute the instructions to:

access first tracking data for imaging devices capturing images from different viewpoints of a scene at a first point in time, the first tracking data indicating first poses of the imaging devices, the imaging devices of a same imaging modality;

fuse, based on the first tracking data, first data sets respectively associated with the imaging devices to generate a first fused set of data for the scene, the first fused set of data corresponding to the first point in time;

access second tracking data for the imaging devices capturing images from different viewpoints of the scene at a second point in time, the second tracking data indicating second poses of the imaging devices and representing a change in a pose of at least one of the image devices that occurs while the imaging devices capture images of the scene;

fuse, based on the second tracking data, second data sets respectively associated with the imaging devices to generate a second fused set of data for the scene, the second fused set of data corresponding to the second point in time; and

generate, based on the first fused set of data, a perception of the scene.

2 . The apparatus of claim 1 , wherein the imaging devices are visible light imaging devices.

3 . The apparatus of claim 1 , wherein the imaging devices have same intrinsic parameters.

4 . The apparatus of claim 1 , wherein:

the first and second data sets respectively associated with the imaging devices comprise first and second segmentation data sets respectively associated with the imaging devices;

the fusing of the first data sets to generate the first fused set of data comprises fusing the first segmentation data sets to form a first fused set of segmentation data; and

the fusing of the second data sets to generate the second fused set of data comprises fusing the second segmentation data sets to form a second fused set of segmentation data.

5 . The apparatus of claim 1 , wherein:

the first and second data sets respectively associated with the imaging devices comprise first and second image data sets respectively associated with the imaging devices;

the fusing of the first data sets to generate the first fused set of data comprises stitching a first set of images of the scene together along non-overlapping boundaries; and

the fusing of the second data sets to generate the second fused set of data comprises stitching a second set of images of the scene together along non-overlapping boundaries.

6 . The apparatus of claim 1 , wherein at least one of the imaging devices is mounted to an articulating component of a robotic system.

7 . The apparatus of claim 1 , wherein at least one of the imaging devices is mounted to an articulating support structure in a surgical facility.

8 . The apparatus of claim 1 , wherein:

the instructions comprise a machine learned algorithm; and

the processor is configured to apply the machine learned algorithm to perform the fusing of the first data sets to generate the first fused set of data and the fusing of the second data sets to generate the second fused set of data.

9 . The apparatus of claim 1 , wherein the processor is further configured to execute the instructions to:

determine a potential to improve the perception of the scene; and

provide output indicating an operation to be performed to improve the perception of the scene.

10 . The apparatus of claim 9 , wherein the processor provides the output to a robotic system to instruct the robotic system to change the pose of at least one of the imaging devices.

11 . The apparatus of claim 1 , wherein the fusing of the first data sets comprises:

segmenting each of the first data sets; and

combining confidences of the segmentations of the first data sets to generate the first fused set of data for the scene.

12 . The apparatus of claim 11 , wherein the combining confidences of the segmentations comprises projecting confidences of one of the first data sets onto another of the first data sets.

13 . A system comprising:

a first imaging device;

a second imaging device of a same imaging modality as the first imaging device and having a dynamic relationship with the first imaging device based at least on the second imaging device being dynamically moveable relative to the first imaging device during imaging of a scene by the first and second imaging devices; and

a processing system communicatively coupled to the imaging devices and configured to:

access first tracking data for the first and second imaging devices during the imaging of the scene by the first and second imaging devices from different viewpoints of the scene at a first point in time, the first tracking data indicating first poses of the first and second imaging devices;

fuse, based on the first tracking data, first data sets respectively associated with the first and second imaging devices to generate a first fused set of data for the scene, the first fused set of data corresponding to a first point in time;

access second tracking data for the first and second imaging devices during the imaging of the scene by the first and second imaging devices from different viewpoints of the scene at a second point in time, the second tracking data indicating second poses of the first and second imaging devices and representing a change in a pose of the second image device that occurs during the imaging of the scene by the first and second imaging devices;

fuse, based on the second tracking data, second data sets respectively associated with the first and second imaging devices to generate a second fused set of data for the scene, the second fused set of data corresponding to the second point in time; and

generate, based on the first fused set of data, a perception of the scene.

14 . The system of claim 13 , wherein:

the scene comprises a surgical scene proximate a robotic surgical system;

the first imaging device is mounted on a first component of the robotic surgical system; and

the second imaging device is mounted on a second component of the robotic surgical system, the second component configured to articulate.

15 . The system of claim 13 , wherein:

the scene comprises a surgical scene at a surgical facility;

the first imaging device is mounted on a first component at the surgical facility; and

the second imaging device is mounted on a second component at the surgical facility, the second component configured to articulate.

16 . The system of claim 13 , wherein:

the first imaging device is mounted to a first robotic system; and

the second imaging device is mounted to a second robotic system that is separate from the first robotic system.

17 . A method comprising:

accessing, by a processing system, first tracking data for imaging devices capturing images from different viewpoints of a scene at a first point in time, the first tracking data indicating first poses of the imaging devices, the imaging devices of a same imaging modality;

fusing, by the processing system based on the first tracking data, first data sets respectively associated with the imaging devices to generate a first fused set of data for the scene, the first fused set of data corresponding to a first point in time;

accessing, by the processing system, second tracking data for the imaging devices capturing images from different viewpoints of the scene at a second point in time, the second tracking data indicating second poses of the imaging devices and representing a change in a pose of at least one of the image devices that occurs while the imaging devices capture images of the scene;

fusing, by the processing system based on the second tracking data, second data sets respectively associated with the imaging devices to generate a second fused set of data for the scene, the second fused set of data corresponding to the second point in time; and

generating, based on the first fused set of data, a perception of the scene.

18 . The method of claim 17 , wherein the imaging devices are visible light imaging devices.

19 . The method of claim 17 , wherein the fusing of the first data sets comprises:

segmenting each of the first data sets; and

combining confidences of the segmentations of the first data sets to generate the first fused set of data for the scene.

20 . The method of claim 17 , further comprising:

determining, by the processing system, a potential to improve the perception of the scene; and

providing, by the processing system, output indicating an operation to be performed to improve the perception of the scene.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2022
From: MOHARERI, OMID; DIMAIO, SIMON P.; LI, ZHAOSHUO; SHABAN, AMIRREZA; SIMARD, JEAN-GABRIEL
To: INTUITIVE SURGICAL OPERATIONS, INC.
Reel/Frame 059893/0643 →
Continuity (3)
Provisional Application 63017506 · Apr 29, 2020
Provisional Application 62936343 · Nov 15, 2019
Related Publication 20220392084A1 · Dec 8, 2022
References Cited (13)
US 20140303491A1 · Shekhar et al. · 2014 [cited by applicant]
US 20180064316A1 · Charles et al. · 2018 [cited by applicant]
US 20190184582A1 · Namiki · 2019 [cited by examiner]
US 20200175669A1 · Bian · 2020 [cited by examiner]
US 20200289222A1 · Denlinger · 2020 [cited by examiner]
CN 106937531A · 2017 [cited by applicant]
CN 107456278A · 2017 [cited by applicant]
CN 108348143A · 2018 [cited by applicant]
TW I639136B · 2018 [cited by examiner]
Golodetz, Stuart et.al., “Collaborative Large-Scale Dense 3D Reconstruction with Online Inter-Agent Pose Optimisation,” IEEE Transactions on Visualization and Computer Graphics, Jan. 2018, pp. 1-14. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2020/060567,mailed Feb. 26, 2021, 11 pages. [cited by applicant]
Vertut, J, and Coiffet, P., “Robot Technology: Teleoperation and Robotics Evolution and Development,” English translation, Prentice-Hall, Inc., Inglewood Cliffs, NJ, USA 1986, vol. 3A, 332 pages. [cited by applicant]
International Preliminary Report on Patentability for Application No. PCT/US2020/060567, mailed May 27, 2022, 8 pages. [cited by applicant]