IP Library Granted Patent US 12,186,139
Granted Patent B2
US 12,186,139 · App. 17/663,034 · Granted Jan 7, 2025

Systems and methods for scene-adaptive image quality in surgical video

Inventors: Nishant Verma (Burlingame, CA); Ryan Robertson (Sunnyvale, CA); Thomas Teisseyre (Montara, CA)
Assignee: VERILY LIFE SCIENCES LLC
A61B90/361A61B1/000096A61B1/0661A61B90/37G06T7/0012G06T2207/10068
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,186,139
App. No.
17/663,034
Granted
Jan 7, 2025
Kind
B2
Abstract

One example method for scene-adaptive image quality in surgical video includes receiving a first video frame from an endoscope, the first video frame generated from a first raw image captured by an image sensor of the endoscope and processed by an image signal processing (“ISP”) pipeline having a plurality of ISP parameters; recognizing, using a trained machine learning (“ML”) model, a first scene type or a first scene feature type based on the first video frame; determining a first set of ISP parameters based on the first scene type or the first scene feature type; applying the first set of ISP parameters to the ISP pipeline; and receiving a second video frame from the endoscope, the second video frame generated from a second raw image captured by the image sensor and processed by the ISP pipeline using the first set of ISP parameters.

Claims (76)

1. A method comprising:

receiving a first video frame from an endoscope, the first video frame generated from a first raw image captured by an image sensor of the endoscope and processed by an image signal processing (“ISP”) pipeline having a plurality of ISP parameters;

recognizing, using a trained machine learning (“ML”) model, a first scene type or a first scene feature type based on the first video frame;

determining a first set of ISP parameters based on the first scene type or the first scene feature type;

applying the first set of ISP parameters to the ISP pipeline; and

receiving a second video frame from the endoscope, the second video frame generated from a second raw image captured by the image sensor and processed by the ISP pipeline using the first set of ISP parameters.

2. The method of claim 1 , wherein recognizing the first scene type or the first scene feature type comprises obtaining, using the trained ML model, a plurality of probabilities, each probability corresponding to a different scene type of a plurality of scene types or a different scene feature type of a plurality of scene feature types, each probability indicating a likelihood that the first video frame is of the corresponding scene type or scene feature type.

3. The method of claim 2 , further comprising determining a subset of the plurality of scene types or a subset of the plurality of scene feature types based on respective probabilities satisfying a threshold.

4. The method of claim 2 , wherein determining a first set of ISP parameters comprises:

obtaining a plurality of sets of ISP parameters, each set of ISP parameters of the plurality of sets of ISP parameters corresponding to a scene type of the plurality of scene types or to a scene feature type of the plurality of scene feature types; and

generating the first set of ISP parameters based on interpolating between the plurality of sets of ISP parameters.

5. The method of claim 4 , wherein the plurality of sets of ISP parameters comprises a second set of ISP parameters and a third set of ISP parameters, and wherein each of the second and third sets of ISP parameters comprises values for a first ISP parameter and a second ISP parameter, wherein:

the value for the first ISP parameter of the second set of ISP parameters is different than the value for the first ISP parameter of the third set of ISP parameters; and

the value for the second ISP parameter of the second set of ISP parameters is different than the value for the second ISP parameter of the third set of ISP parameters; and

wherein generating the first set of ISP parameters comprises:

interpolating a first interpolated parameter value based on the value for the first ISP parameter of the second set of ISP parameters and the value for the first ISP parameter of the third set of ISP parameters according to a first interpolation technique; and

interpolating a second interpolated parameter value based on the value for the second ISP parameter of the second set of ISP parameters and the value for the second ISP parameter of the third set of ISP parameters according to a second interpolation technique.

6. The method of claim 5 , wherein the first interpolation technique is different from the second interpolation technique.

7. The method of claim 1 , further comprising:

identifying, using a second trained ML model, a scene feature type based on the first video frame;

determining a scene feature set of ISP parameters based on the first scene feature;

combining the scene feature set of ISP parameters with the first set of ISP parameters; and

wherein applying the first set of ISP parameters to the ISP pipeline comprises applying the combination of the scene feature set of ISP parameters with the first set of ISP parameters to the ISP pipeline.

8. The method of claim 7 , wherein the first scene feature comprises a tool or an anatomical feature.

9. A system comprising:

a non-transitory computer-readable medium; and

one or more processors communicatively coupled to the non-transitory computer-readable medium and configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

receive a first video frame from an endoscope, the first video frame generated from a first raw image captured by an image sensor of the endoscope and processed by an image signal processing (“ISP”) pipeline having a plurality of ISP parameters;

recognize, using a trained machine learning (“ML”) model, a first scene type or a first scene feature type based on the first video frame;

determine a first set of ISP parameters based on the first scene type or the first scene feature type;

apply the first set of ISP parameters to the ISP pipeline; and

receive a second video frame from the endoscope, the second video frame generated from a second raw image captured by the image sensor and processed by the ISP pipeline using the first set of ISP parameters.

10. The system of claim 9 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to obtain, using the trained ML model, a plurality of probabilities, each probability corresponding to a different scene type of a plurality of scene types or a different scene feature type of a plurality of scene feature types, each probability indicating a likelihood that the first video frame is of the corresponding scene type or scene feature type.

11. The system of claim 10 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to determine a subset of the plurality of scene types or a subset of the plurality of scene feature types based on respective probabilities satisfying a threshold.

12. The system of claim 10 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

obtain a plurality of sets of ISP parameters, each set of ISP parameters of the plurality of sets of ISP parameters corresponding to a scene type of the plurality of scene types or to a scene feature type of the plurality of scene feature types; and

generate the first set of ISP parameters based on interpolating between the plurality of sets of ISP parameters.

13. The system of claim 12 , wherein the plurality of sets of ISP parameters comprises a second set of ISP parameters and a third set of ISP parameters, and wherein each of the second and third sets of ISP parameters comprises values for a first ISP parameter and a second ISP parameter, wherein:

the value for the first ISP parameter of the second set of ISP parameters is different than the value for the first ISP parameter of the third set of ISP parameters; and

the value for the second ISP parameter of the second set of ISP parameters is different than the value for the second ISP parameter of the third set of ISP parameters; and

wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

interpolate a first interpolated parameter value based on the value for the first ISP parameter of the second set of ISP parameters and the value for the first ISP parameter of the third set of ISP parameters according to a first interpolation technique;

interpolate a first interpolated parameter value based on the value for the second ISP parameter of the second set of ISP parameters and the value for the second ISP parameter of the third set of ISP parameters according to a second interpolation technique; and

generate the first set of ISP parameters based on the first and second interpolated parameter values.

14. The system of claim 13 , wherein the first interpolation technique is different from the second interpolation technique.

15. The system of claim 9 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

identify, using a second trained ML model, a scene feature type based on the first video frame;

determine a scene feature set of ISP parameters based on the first scene feature;

combine the scene feature set of ISP parameters with the first set of ISP parameters; and

apply the combination of the scene feature set of ISP parameters with the first set of ISP parameters to the ISP pipeline.

16. The system of claim 15 , wherein the first scene feature comprises a tool or an anatomical feature.

17. A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:

receive a first video frame from an endoscope, the first video frame generated from a first raw image captured by an image sensor of the endoscope and processed by an image signal processing (“ISP”) pipeline having a plurality of ISP parameters;

recognize, using a trained machine learning (“ML”) model, a first scene type or a first scene feature type based on the first video frame;

determine a first set of ISP parameters based on the first scene type or the first scene feature type;

apply the first set of ISP parameters to the ISP pipeline; and

receive a second video frame from the endoscope, the second video frame generated from a second raw image captured by the image sensor and processed by the ISP pipeline using the first set of ISP parameters.

18. The non-transitory computer-readable medium of claim 17 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to obtain, using the trained ML model, a plurality of probabilities, each probability corresponding to a different scene type of a plurality of scene types or a different scene feature type of a plurality of scene feature types, each probability indicating a likelihood that the first video frame is of the corresponding scene type or scene feature type.

19. The non-transitory computer-readable medium of claim 18 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to determine a subset of the plurality of scene types or a subset of the plurality of scene feature types based on respective probabilities satisfying a threshold.

20. The non-transitory computer-readable medium of any of claim 18 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

obtain a plurality of sets of ISP parameters, each set of ISP parameters of the plurality of sets of ISP parameters corresponding to a scene type of the plurality of scene types or to a scene feature type of the plurality of scene feature types; and

generate the first set of ISP parameters based on interpolating between the plurality of sets of ISP parameters.

21. The non-transitory computer-readable medium of claim 20 , wherein the plurality of sets of ISP parameters comprises a second set of ISP parameters and a third set of ISP parameters, and wherein each of the second and third sets of ISP parameters comprises values for a first ISP parameter and a second ISP parameter, wherein:

the value for the first ISP parameter of the second set of ISP parameters is different than the value for the first ISP parameter of the third set of ISP parameters; and

the value for the second ISP parameter of the second set of ISP parameters is different than the value for the second ISP parameter of the third set of ISP parameters; and

wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

interpolate a first interpolated parameter value based on the value for the first ISP parameter of the second set of ISP parameters and the value for the first ISP parameter of the third set of ISP parameters according to a first interpolation technique;

interpolate a first interpolated parameter value based on the value for the second ISP parameter of the second set of ISP parameters and the value for the second ISP parameter of the third set of ISP parameters according to a second interpolation technique; and

generate the first set of ISP parameters based on the first and second interpolated parameter values.

22. The non-transitory computer-readable medium of claim 21 , wherein the first interpolation technique is different from the second interpolation technique.

23. The non-transitory computer-readable medium of claim 17 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

identify, using a second trained ML model, a scene feature type based on the first video frame;

determine a scene feature set of ISP parameters based on the first scene feature;

combine the scene feature set of ISP parameters with the first set of ISP parameters; and

apply the combination of the scene feature set of ISP parameters with the first set of ISP parameters to the ISP pipeline.

24. The non-transitory computer-readable medium of claim 23 , wherein the first scene feature comprises a tool or an anatomical feature.

Assignments (2)
CHANGE OF ADDRESS Recorded Nov 19, 2024
From: VERILY LIFE SCIENCES LLC
To: VERILY LIFE SCIENCES LLC
Reel/Frame 069390/0656 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2022
From: VERMA, NISHANT; ROBERTSON, RYAN; TEISSEYRE, THOMAS
To: VERILY LIFE SCIENCES LLC
Reel/Frame 060195/0501 →
Continuity (2)
Provisional Application 63191481 · May 21, 2021
Related Publication 20220370167A1 · Nov 24, 2022
References Cited (16)
US 8428308B2 · Jasinski et al. · 2013 [cited by applicant]
US 20030159141A1 · Zacharias · 2003 [cited by applicant]
US 20070273759A1 · Krupnick et al. · 2007 [cited by applicant]
US 20130142418A1 · Van Zwol et al. · 2013 [cited by applicant]
US 20160344996A1 · Olilla · 2016 [cited by applicant]
US 20170270508A1 · Roach et al. · 2017 [cited by applicant]
US 20190110856A1 · Barral · 2019 [cited by examiner]
US 20190182421A1 · Piponi · 2019 [cited by examiner]
US 20200304650A1 · Roach et al. · 2020 [cited by applicant]
US 20220301123A1 · Mosleh · 2022 [cited by examiner]
US 20230255443A1 · He · 2023 [cited by examiner]
EP 3135028 · 2019 [cited by applicant]
JP 5981053 · 2016 [cited by applicant]
Huang et al., “Feature Extraction of Video Data for Automatic Visual Tool Tracking in Robot Assisted Surgery”, Proceedings of the 2019 4th International Conference on Robotics, Control and Automation, Available Online a… [cited by applicant]
Kim et al., “Designing a New Endoscope for Panoramic-View with Focus-Area 3D-Vision in Minimally Invasive Surgery”, Journal of Medical and Biological Engineering, vol. 40, Available Online at URL: https://link.springer.… [cited by applicant]
International Application No. PCT/US2022/072322 , “International Search Report and Written Opinion”, Jul. 29, 2022, 11 pages. [cited by applicant]