IP Library Granted Patent US 12,652,409
Granted Patent B2
US 12,652,409 · App. 18/240,783 · Granted Jun 9, 2026

Combined compression and feature extraction models for storing and analyzing medical videos

Inventor: Joel Shor (Somerville, MA)
Assignee: Verily Health Inc.
H04N19/50G06T7/73G06V10/774G06V10/82G06V20/46H04N19/136H04N19/147H04N19/184H04N19/436G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30032
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,652,409
App. No.
18/240,783
Granted
Jun 9, 2026
Kind
B2
Abstract

A method of compressing and detecting target features of a medical video is presented herein. In some embodiments, the method may include receiving an uncompressed medical video comprising at least one target feature, compressing the uncompressed medical video to generate a compressed medical video based on a predicted location of the at least one target feature using a first pretrained machine learning model, and detecting the location of the at least one target feature of the compressed medical video using a second pretrained machine learning model. In some embodiments, the first pretrained machine learning model and the second pretrained machine learning model may be trained in tandem using domain-specific medical videos.

Claims (36)

1 . A method comprising:

receiving an uncompressed medical video comprising at least one target feature;

compressing the uncompressed medical video to generate a compressed medical video based on a predicted location of the at least one target feature using a first pretrained machine learning model; and,

detecting a location of the at least one target feature of the compressed medical video using a second pretrained machine learning model,

wherein the first pretrained machine learning model and the second pretrained machine learning model are trained in tandem using domain-specific medical videos in a manner that, in a first iteration, a training medical video is both compressed and analyzed to detect a location of at least one training target feature before a start of a second iteration after the first iteration, and

wherein the first pretrained machine learning model is trained to compress a first video frame differently from a second video frame, the first video frame comprising the location of at least one training target feature and the second video frame not comprising the location of at least one training target feature.

2 . The method of claim 1 , wherein the uncompressed medical video is obtained during a colonoscopy.

3 . The method of claim 2 , wherein one or more target features of the at least one target feature is a polyp.

4 . The method of claim 1 , wherein the first pretrained machine learning model comprises a video compression transformer.

5 . The method of claim 1 , wherein the second pretrained machine learning model comprises a RetinaNet detector.

6 . The method of claim 1 , wherein the first pretrained machine learning model is configured to allocate a larger proportion of bits to the predicted location of the at least one target feature than to other parts of the compressed medical video.

7 . The method of claim 1 , wherein the second pretrained machine learning model is designed to optimize an accuracy of detection of the at least one target feature.

8 . The method of claim 1 , wherein the first pretrained machine learning model is designed to optimize an accuracy of detection of the at least one target feature and to minimize a compression loss of the compressed medical video.

9 . A system comprising:

a memory configured to store a plurality of processor-executable instructions, the memory including:

a video compressor based on a first pretrained machine learning model;

a target feature detector based on a second pretrained machine learning model; and,

a processor configured to execute the plurality of processor-executable instructions instruction to perform operations including:

receiving an uncompressed medical video comprising at least one target feature;

compressing the uncompressed medical video to generate a compressed medical video based on a predicted location of the at least one target feature; and,

detecting a location of the at least one target feature of the compressed medical video,

wherein the first pretrained machine learning model and the second pretrained machine learning model are trained in tandem using domain-specific medical videos in a manner that, in a first iteration, a training medical video is both compressed and analyzed to detect a location of at least one training target feature before a start of a second iteration after the first iteration, and

wherein the first pretrained machine learning model is trained to compress a first video frame differently from a second video frame, the first video frame comprising the location of at least one training target feature and the second video frame not comprising the location of at least one training target feature.

10 . The system of claim 9 , wherein the first pretrained machine learning model comprises a video compression transformer.

11 . The system of claim 9 , wherein the second pretrained machine learning model comprises a RetinaNet detector.

12 . The system of claim 9 , wherein the second pretrained machine learning model is designed to optimize an accuracy of detection of the at least one target feature.

13 . The system of claim 9 , wherein the first pretrained machine learning model is designed to optimize an accuracy of detection of the at least one target feature and to minimize a compression loss of the compressed medical video.

14 . A non-transitory processor-readable storage medium storing a plurality of processor-executable instructions, the plurality of processor-executable instructions being executed by a processor to perform operations comprising:

receiving an uncompressed medical video comprising at least one target feature;

compressing the uncompressed medical video to generate a compressed medical video based on a predicted location of the at least one target feature using a first pretrained machine learning model; and,

detecting a location of the at least one target feature of the compressed medical video using a second pretrained machine learning model,

wherein the first pretrained machine learning model and the second pretrained machine learning model are trained in tandem using domain-specific medical videos in a manner that, in a first iteration, a training medical video is both compressed and analyzed to detect a location of at least one training target feature before a start of a second iteration after the first iteration, and

wherein the first pretrained machine learning model is trained to compress a first video frame differently from a second video frame, the first video frame comprising the location of at least one training target feature and the second video frame not comprising the location of at least one training target feature.

15 . The non-transitory processor-readable storage medium of claim 14 , wherein the first pretrained machine learning model comprises a video compression transformer.

16 . The non-transitory processor-readable storage medium of claim 14 , wherein the second pretrained machine learning model comprises a RetinaNet detector.

17 . The non-transitory processor-readable storage medium of claim 14 , wherein the first pretrained machine learning model is designed to optimize an accuracy of detection of the at least one target feature and to minimize a compression loss of the compressed medical video and the second pretrained machine learning model is designed to optimize a second accuracy of detection of the at least one target feature.

Assignments (2)
CHANGE OF NAME Recorded May 4, 2026
From: VERILY LIFE SCIENCES LLC
To: VERILY HEALTH INC.
Reel/Frame 075501/0627 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2023
From: SHOR, JOEL
To: VERILY LIFE SCIENCES LLC
Reel/Frame 064871/0024 →
Continuity (2)
Provisional Application 63410522 · Sep 27, 2022
Related Publication 20240129515A1 · Apr 18, 2024
References Cited (20)
US 10860929B2 · Rippel · 2020 [cited by examiner]
US 11082720B2 · Tsai et al. · 2021 [cited by applicant]
US 11521377B1 · Wang · 2022 [cited by examiner]
US 12026874B2 · Protsenko · 2024 [cited by examiner]
US 20210314629A1 · Tsai et al. · 2021 [cited by applicant]
US 20220230311A1 · Zhang · 2022 [cited by examiner]
US 20240257497A1 · Goldenberg · 2024 [cited by examiner]
CN 114298978A · 2022 [cited by examiner]
CN 114913164A · 2022 [cited by examiner]
WO WO2020210734A1 · 2020 [cited by examiner]
WO WO2021159774A1 · 2021 [cited by examiner]
Citation for cited IDS NPL Item #5 showing Jun. 2022 Publication Date for prior art date used. (Year: 2022). [cited by examiner]
Espacenet provided translation of Lu (Year: 2022). [cited by examiner]
STIC provided translation of Nuo (Year: 2022). [cited by examiner]
STIC provided translation of Zhou (Year: 2021). [cited by examiner]
“Learning Binary Residual Representations for Domain-Specific Video Streaming.” University of California, The Thirty-Second AAAI Conference on Artificial Intelligence (AAAI-18). (8 pgs). [cited by applicant]
“Scale-space flow for end-to-end optimized video compression.” (2020) IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (10 pgs). [cited by applicant]
“Relevance-Based Compression of Cataract Surgery Videos Using Convolutional Neural Networks.” MM '20, Oct. 12-16, 2020, Seattle, WA, USA. (10 pgs). [cited by applicant]
“Focal Loss for Dense Object Detection.” arXiv:1708.02002v2 [cs.CV] Feb. 7, 2018 (10 pgs). [cited by applicant]
“VCT: A Video Compression Transformer.” arXiv:2206.07307v1 [cs.CV] Jun. 15, 2022. (16 pgs). [cited by applicant]