IP Library Granted Patent US 12,488,577
Granted Patent B2
US 12,488,577 · App. 17/762,906 · Granted Dec 2, 2025

System and method for analyzing medical images based on spatio-temporal data

Inventors: John Galeotti (Pittsburgh, PA); Tejas Sudharshan Mathai (Seattle, WA)
Assignee: Carnegie Mellon University
G06V10/82G06T7/0016G06T7/20G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30004
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,577
App. No.
17/762,906
Granted
Dec 2, 2025
Kind
B2
Abstract

Provided is a system, method, and computer program product for analyzing spatio-temporal medical images using an artificial neural network. The method includes capturing a series of medical images of a patient, the series of medical images comprising visual movement of at least one entity, tracking time-varying spatial data associated with the at least one entity based on the visual movement, generating spatio-temporal data by correlating the time-varying spatial data with the series of medical images, and analyzing the series of medical images based on an artificial neural network comprising a plurality of layers, one or more layers of the plurality of layers each combining features from at least three different scales, at least one layer of the plurality of layers of the artificial neural network configured to learn spatio-temporal relationships based on the spatio-temporal data.

Claims (36)

1 . A method for analyzing spatio-temporal medical images using an artificial neural network, comprising:

capturing a series of medical images of a patient with an imaging device, the series of medical images comprising visual movement of at least one entity comprising at least a portion of at least one of the patient and an object;

tracking, with a computing device, time-varying spatial data associated with the at least one entity based on the visual movement of the at least one entity in one or more images of the series of medical images;

generating, with a computing device, spatio-temporal data by correlating the time-varying spatial data with the series of medical images;

inputting the spatio-temporal data into an artificial neural network; and

analyzing, with a computing device, the series of medical images based on the spatio-temporal data by using the artificial neural network comprising a plurality of layers, one or more layers of the plurality of layers each combining features from at least three different scales, wherein a layer of the plurality of layers of the artificial neural network is configured to learn multi-scale spatio-temporal relationships of features from at least two different scales of the at least three different scales, the features from the layer and a preceding layer, based on the spatio-temporal data, and wherein the one or more layers that combine features from the at least three different scales comprise at least one of the following: dilated convolutions of different scales, dense and/or residual connections between at least a subset of layers of the plurality of layers wherein the at least the subset of layers comprises features from at least three different scales, or any combination thereof.

2 . The method of claim 1 , wherein the one or more layers that combine features from the at least three different scales comprise dilated convolutions of different scales.

3 . The method of claim 1 , wherein the one or more layers that combine features from the at least three different scales comprise dense and/or residual connections between at least a subset of layers of the plurality of layers, the at least the subset of layers comprising features from at least three different scales.

4 . The method of claim 1 , wherein the one or more layers that combine features from the at least three different scales comprise convolutions of at least two different scales and connections to a subset of layers of the plurality of layers comprising features from at least two different scales, resulting in features of at least three different scales.

5 . The method of claim 1 , wherein the at least one entity comprises at least one of the following: an instrument, the imaging device, a physical artifact, a manifested artifact, or any combination thereof.

6 . The method of claim 1 , wherein tracking the time-varying spatial data comprises tracking at least one of the following: translational/rotational positions of the at least one entity, a velocity of the at least one entity, an acceleration of the at least one entity, an inertial measurement of the at least one entity, or any combination thereof.

7 . The method of claim 1 , wherein tracking the time-varying spatial data is based on at least one of the following: an inertial measurement unit, a tracking system, a position sensor, robotic kinematics, inverse kinematics, or any combination thereof.

8 . The method of claim 1 , wherein the spatio-temporal data comprises at least one of the following: data representing an internal motion within the patient's body, data representing an external motion of the patient's body, data representing a motion of an instrument, data representing an angle of the instrument, data representing a deforming motion of the patient's body, or any combination thereof.

9 . The method of claim 1 , wherein the artificial neural network comprises an encoder and a decoder, and wherein at least one of the decoder and the encoder is configured to utilize the spatio-temporal data as input.

10 . The method of claim 1 , wherein the artificial neural network comprises at least one of the following: Long-Short Term Memory (LSTM) units, Gated Recurrent Units (GRUs), temporal convolutional networks, or any combination thereof.

11 . A system for analyzing spatio-temporal medical images using an artificial neural network, comprising a computing device programmed or configured to:

capture a series of medical images of a patient with an imaging device, the series of medical images comprising visual movement of at least one entity comprising at least a portion of at least one of the patient and an object;

track time-varying spatial data associated with the at least one entity based on the visual movement of the at least one entity in one or more images of the series of medical images;

generate spatio-temporal data by correlating the time-varying spatial data with the series of medical images;

input the spatio-temporal data into an artificial neural network; and

analyze the series of medical images based on the spatio-temporal data by using the artificial neural network comprising a plurality of layers, one or more layers of the plurality of layers each combining features from at least three different scales, wherein a layer of the plurality of layers of the artificial neural network is configured to learn multi-scale spatio-temporal relationships of features from at least two different scales of the at least three different scales, the features from the layer and a preceding layer based on the spatio-temporal data, and wherein the one or more layers that combine features from the at least three different scales comprise at least one of the following: dilated convolutions of different scales, dense and/or residual connections between at least a subset of layers of the plurality of layers wherein the at least the subset of layers comprises features from at least three different scales, or any combination thereof.

12 . The system of claim 11 , wherein the one or more layers that combine features from the at least three different scales comprise dilated convolutions of different scales.

13 . The system of claim 11 , wherein the one or more layers that combine features from the at least three different scales comprise dense and/or residual connections between at least a subset of layers of the plurality of layers, the at least the subset of layers comprising features from at least three different scales.

14 . The system of claim 11 , wherein the one or more layers that combine features from the at least three different scales comprise convolutions of at least two different scales and connections to a subset of layers of the plurality of layers comprising features from at least two different scales, resulting in features of at least three different scales.

15 . The system of claim 11 , wherein the at least one entity comprises at least one of the following: an instrument, the imaging device, a physical artifact, a manifested artifact, or any combination thereof.

16 . The system of claim 11 , wherein tracking the time-varying spatial data comprises tracking at least one of the following: translational/rotational positions of the at least one entity, a velocity of the at least one entity, an acceleration of the at least one entity, an inertial measurement of the at least one entity, or any combination thereof.

17 . The system of claim 11 , wherein tracking the time-varying spatial data is based on at least one of the following: an inertial measurement unit, a tracking system, a position sensor, robotic kinematics, inverse kinematics, or any combination thereof.

18 . The system of claim 11 , wherein the spatio-temporal data comprises at least one of the following: data representing an internal motion within the patient's body, data representing an external motion of the patient's body, data representing a motion of an instrument, data representing an angle of the instrument, data representing a deforming motion of the patient's body, or any combination thereof.

19 . The system of claim 11 , wherein the artificial neural network comprises an encoder and a decoder, and wherein at least one of the decoder and the encoder is configured to utilize the spatio-temporal data as input.

20 . The system of claim 11 , wherein the artificial neural network comprises at least one of the following: Long-Short Term Memory (LSTM) units, Gated Recurrent Units (GRUs), temporal convolutional networks, or any combination thereof.

21 . A computer program product for analyzing medical images using a neural network, comprising at least one non-transitory computer-readable medium including instructions that, when executed by a computing device, cause the computing device to:

capture a series of medical images of a patient with an imaging device, the series of medical images comprising visual movement of at least one entity comprising at least a portion of at least one of the patient and an object;

track time-varying spatial data associated with the at least one entity based on the visual movement of the at least one entity in one or more images of the series of medical images;

generate spatio-temporal data by correlating the time-varying spatial data with the series of medical images;

input the spatio-temporal data into an artificial neural network; and

analyze the series of medical images based on the spatio-temporal data by using the artificial neural network comprising a plurality of layers, one or more layers of the plurality of layers each combining features from at least three different scales, wherein a layer of the plurality of layers of the artificial neural network is configured to learn multi-scale spatio-temporal relationships of features from at least two different scales of the at least three different scales, the features from the layer and a preceding layer, based on the spatio-temporal data, and wherein the one or more layers that combine features from the at least three different scales comprise at least one of the following: dilated convolutions of different scales, dense and/or residual connections between at least a subset of layers of the plurality of layers wherein the at least the subset of layers comprises features from at least three different scales, or any combination thereof.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2022
From: GALEOTTI, JOHN; MATHAI, TEJAS SUDHARSHAN
To: CARNEGIE MELLON UNIVERSITY
Reel/Frame 059376/0336 →
Continuity (2)
Provisional Application 62904728 · Sep 24, 2019
Related Publication 20220383500A1 · Dec 1, 2022
References Cited (36)
US 20180218502A1 · Golden et al. · 2018 [cited by applicant]
US 20190122360A1 · Zhang et al. · 2019 [cited by applicant]
US 20190223725A1 · Lu et al. · 2019 [cited by applicant]
US 20190279361A1 · Meyer et al. · 2019 [cited by applicant]
CA 2534701A1 · 2007 [cited by applicant]
CN 109427058A · 2019 [cited by applicant]
CN 109598727A · 2019 [cited by applicant]
CN 109690554A · 2019 [cited by applicant]
JP H1103121A · 1989 [cited by applicant]
JP 201937692A · 2019 [cited by applicant]
WO 2019173237A1 · 2019 [cited by applicant]
Khened, M., Kollerathu, V.A. and Krishnamurthi, G., 2019. Fully convolutional multi-scale residual DenseNets for cardiac segmentation and automated cardiac diagnosis using ensemble of classifiers. Medical image analysis… [cited by examiner]
Yu, F. and Koltun, V., 2015. Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122. [cited by examiner]
Ilea, D.E., Duffy, C., Kavanagh, L., Stanton, A. and Whelan, P.F., 2012. Fully automated segmentation and tracking of the intima media thickness in ultrasound video sequences of the common carotid artery. IEEE transacti… [cited by examiner]
L Zhang, L Lu, X Wang, RM Zhu, M Bagheri . . .—arXiv preprint arXiv . . . , 2019—arxiv.org. [cited by examiner]
Apostolopoulos et al, “Pathological OCT Retinal Layer Segmentation using Branch Residual U-shape Networks”, Medical Image Computing and Computer Assisted Intervention MICCAI 2017, Lecture Notes in Computer Science vol. … [cited by applicant]
Arbelle et al., “Microscopy Cell Segmentation via Convolutional LSTM Networks”, IEEE ISBI, 2019, pp. 1008-1012. [cited by applicant]
Basty et al., “Super Resolution of Cardiac Cine MRI Sequences Using Deep Learning”, Image Analysis for Moving Organ, Breast, and Thoracic Images, RAMBO 2018, Lecture Notes in Computer Science, vol. 11040, 10 pages. [cited by applicant]
Chaniot et al., “Vessel Segmentation in High-Frequency 2D/3D Ultrasound Images”, IEEE Int Ultrasonics Symp., 2016, 4 pages. [cited by applicant]
Gao et al., “Fully Convolutional Structured LSTM Networks for Joint 4D Medical Image Segmentation”, IEEE ISBI, 2018, pp. 1104-1108. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, IEEE CVPR, 2016, pp. 770-778. [cited by applicant]
Jaeger et al., “Two public chest X-ray datasets for computer-aided screening of pulmonary diseases”, Quant Imaging Med Surg, 2014 4(6), pp. 475-477. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization”, ICLR, 2015, pp. 1-15. [cited by applicant]
Mathai et al., “Fast Vessel Segmentation and Tracking in Ultra High-Frequency Ultrasound Images”, Medical Image Computing and Computer Assisted Intervention, MICCAI 2018, Lecture Notes in Computer Science, vol. 11073, 1… [cited by applicant]
Mathai et al., “Learning to Segment Corneal Tissue Interfaces in OCT Images”, IEEE ISBI, 2019, pp. 1432-1436. [cited by applicant]
Menchón-Lara et al., “Fully automatic segmentation of ultrasound common carotid artery images based on machine learning”, Neurocomputing 151, 2015, pp. 161-167. [cited by applicant]
Milletari et al., “CFCM: Segmentation via Coarse to Fine Context Memory”, Medical Image Computing and Computer Assisted Intervention, MICCAI 2018, Lecture Notes in Computer Science, vol. 11073, 9 pages. [cited by applicant]
Mohler III et al., “High Frequency Ultrasound for Evaluation of Intimal Thickness”, J Am Soc Echocardiogr. 22(10), Oct. 2009, pp. 1129-1133. [cited by applicant]
Ronneberger et al., “U-Net: Convolutional Networks for Biomedical Image Segmentation”, Medical Image Computer and Computer-Assisted Intervention, MICCAI 2018, Lecture Notes in Computer Science, vol. 9351, 8 pages. [cited by applicant]
Shin et al., “Automating Carotid Intima-Media Thickness Video Interpretation with Convolutional Neural Networks”, CVPR, 2016, pp. 2526-2535. [cited by applicant]
Yu et al., “Multi-Scale Context Aggregation by Dilated Convolutions”, ICLR, 2016, pp. 1-13. [cited by applicant]
Zhang et al., “A Multi-Level Convolutional LSTM Model for the Segmentation of Left Ventricle Myocardium in Infarcted Porcine Cine MR Images”, IEEE ISBI, 2018, pp. 470-473. [cited by applicant]
Zhao et al., “Predicting Tongue Motion in Unlabeled Ultrasound Videos Using Convolutional LSTM Neural Networks”, IEEE ICASSP, 2019, pp. 5926-5930. [cited by applicant]
Lu et al., “A 3D Convolutional Neural Network for Volumetric Image Semantic Segmentation”, Procedia Manufacturing, 2019, pp. 422-428. [cited by applicant]
Milletari et al., “CFCM:Segmentation via Coarse to Fine Context Memory”, arXiv:1806.01413v1, 2018, 10 pages. [cited by applicant]
Wang et al., “Comparative study between dual source CT coronary angiography and conventional coronary angiography: initial experience”, Chin J Med Imaging Technol, 2008, pp. 881-884, vol. 24, No. 6 (English language abs… [cited by applicant]