Multi arm machine learning models with attention for lesion segmentation
Embodiments disclosed herein generally relate to multi-arm machine learning models for lesion detection. Particularly, aspects of the present disclosure are directed to accessing a three-dimensional magnetic resonance imaging (MRI) images. Each of the three-dimensional MRI images depict a same volume of a brain of a subject. The volume of the brain includes at least part of one or more lesions. Each three-dimensional MRI image of the three-dimensional MRI images is processed using one or more corresponding encoder arms of a machine-learning model to generate an encoding of the three-dimensional MRI image. The encodings of the three-dimensional MRI images are concatenated to generate a concatenated representation. The concatenated representation is processed using a decoder arm of the machine-learning model to generate a prediction that identifies one or more portions of the volume of the brain predicted to depict at least part of a lesion.
1 . A computer-implemented method comprising:
accessing a plurality of three-dimensional magnetic resonance imaging (MRI) images, wherein each of the plurality of three-dimensional MRI images depicts a same volume of a brain of a subject; and a first three-dimensional MRI image was generated using a first type of MRI sequence that is different from a second type of MRI sequence used to generate a second three-dimensional MRI image;
processing, for each three-dimensional MRI image of the plurality of three-dimensional MRI images, the three-dimensional MRI image using one or more corresponding encoder arms of a machine-learning model to generate an encoding of the three-dimensional MRI image, wherein the one or more corresponding encoder arms comprise a plurality of encoder models, wherein the plurality of encoder models comprises a first encoder model trained for the first type of MRI sequence and a second encoder model trained for the second type of MRI sequence, and wherein the first three-dimensional MRI image is processed by the first encoder model and the second three-dimensional MRI image is processed by the second encoder model;
concatenating the encodings of the plurality of three-dimensional MRI images to generate a concatenated representation; and
processing the concatenated representation using a decoder arm of the machine-learning model to generate a prediction that identifies one or more portions of the volume of the brain predicted to depict at least part of a lesion.
2 . The computer-implemented method of claim 1 , further comprising:
generating, for each three-dimensional MRI image of the plurality of three-dimensional MRI images using the corresponding encoder model, a downsampled encoding having a resolution that is lower than a resolution of the encoding of the three-dimensional MRI image;
processing, for each three-dimensional MRI image of the plurality of three-dimensional MRI images, the downsampled encoding using one or more layers of the one or more corresponding encoder arms; and
concatenating the downsampled encodings to generate another concatenated representation, wherein the prediction is further based on processing of the another concatenated representation using the decoder arm of the machine-learning model.
3 . The computer-implemented method of claim 2 , wherein the machine-learning model includes one or more skip attention modules, each of the one or more skip attention modules connecting an encoding block of the one or more corresponding encoder arms of the machine-learning model to a decoder block of the decoder arm at a same resolution.
4 . The computer-implemented method of claim 3 , wherein each skip attention module of the one or more skip attention modules receives an input of the concatenated representation and an upsampled encoding of the another concatenated representation at the resolution of the three-dimensional MRI image, and wherein the prediction is further based on processing an output of skip-feature encodings from the skip attention modules using the decoder arm of the machine-learning model.
5 . The computer-implemented method of claim 4 , wherein the one or more skip attention modules include a residual connection between the input and the output of the skip attention module to facilitate skipping the skip attention module if relevant high-dimensional features are unavailable.
6 . The computer-implemented method of claim 1 , wherein the machine-learning model was trained using a weighted binary cross entropy loss and/or a Tversky loss.
7 . The computer-implemented method of claim 1 , wherein the machine-learning model was trained using loss calculated at each of multiple depths of the machine-learning model.
8 . The computer-implemented method of claim 1 , wherein the first type of MRI sequence includes a sequence from a sequence set of T1, T2 and fluid-attenuated inversion recovery (FLAIR), and the second type of MRI sequence includes another sequence from the sequence set.
9 . The computer-implemented method of claim 1 , further comprising:
determining a number of lesions using the prediction.
10 . The computer-implemented method of claim 1 , further comprising:
determining one or more lesion sizes or a lesion load using the prediction.
11 . The computer-implemented method of claim 1 , further comprising:
accessing data corresponding to a previous MRI;
determining a change in a quantity, a size or cumulative size of one or more lesions using the prediction and the data; and
generating an output that represents the change.
12 . The computer-implemented method of claim 1 , further comprising:
recommending changing a treatment strategy based on the prediction.
13 . The computer-implemented method of claim 1 , further comprising:
providing an output corresponding to a possible or confirmed diagnosis of the subject of multiple sclerosis based at least in part on the prediction.
14 . A system comprising:
one or more data processors; and
a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform a set of actions including:
accessing a plurality of three-dimensional magnetic resonance imaging (MRI) images, wherein each of the plurality of three-dimensional MRI images depicts a same volume of a brain of a subject; and a first three-dimensional MRI image was generated using a first type of MRI sequence that is different from a second type of MRI sequence used to generate a second three-dimensional MRI image;
processing, for each three-dimensional MRI image of the plurality of three-dimensional MRI images, the three-dimensional MRI image using one or more corresponding encoder arms of a machine-learning model to generate an encoding of the three-dimensional MRI image, wherein the one or more corresponding encoder arms comprise a plurality of encoder models, wherein the plurality of encoder models comprises a first encoder model trained for the first type of MRI sequence and a second encoder model trained for the second type of MRI sequence, and wherein the first three-dimensional MRI image is processed by the first encoder model and the second three-dimensional MRI image is processed by the second encoder model;
concatenating the encodings of the plurality of three-dimensional MRI images to generate a concatenated representation; and
processing the concatenated representation using a decoder arm of the machine-learning model to generate a prediction that identifies one or more portions of the volume of the brain predicted to depict at least part of a lesion.
15 . The system of claim 14 , wherein the set of actions further includes:
generating, for each three-dimensional MRI image of the plurality of three-dimensional MRI images, a downsampled encoding having a resolution that is lower than a resolution of the encoding of the three-dimensional MRI image;
processing, for each three-dimensional MRI image of the plurality of three-dimensional MRI images, the downsampled encoding using one or more layers of the one or more corresponding encoder arms; and
concatenating the downsampled encodings to generate another concatenated representation, wherein the prediction is further based on processing of the another concatenated representation using the decoder arm of the machine-learning model.
16 . The system of claim 14 , wherein the machine-learning model includes one or more skip attention modules, each of the one or more skip attention modules connecting an encoding block of the one or more corresponding encoder arms of the machine-learning model to a decoder block of the decoder arm at a same resolution.
17 . A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform a set of actions including:
accessing a plurality of three-dimensional magnetic resonance imaging (MRI) images, wherein each of the plurality of three-dimensional MRI images depicts a same volume of a brain of a subject; and a first three-dimensional MRI image was generated using a first type of MRI sequence that is different from a second type of MRI sequence used to generate a second three-dimensional MRI image;
processing, for each three-dimensional MRI image of the plurality of three-dimensional MRI images, the three-dimensional MRI image using one or more corresponding encoder arms of a machine-learning model to generate an encoding of the three-dimensional MRI image, wherein the one or more corresponding encoder arms comprise a plurality of encoder models, wherein the plurality of encoder models comprises a first encoder model trained for the first type of MRI sequence and a second encoder model trained for the second type of MRI sequence, and wherein the first three-dimensional MRI image is processed by the first encoder model and the second three-dimensional MRI image is processed by the second encoder model;
concatenating the encodings of the plurality of three-dimensional MRI images to generate a concatenated representation; and
processing the concatenated representation using a decoder arm of the machine-learning model to generate a prediction that identifies one or more portions of the volume of the brain predicted to depict at least part of a lesion.
18 . The computer-implemented method of claim 1 , wherein the first encoder model is trained using MRI images of the first type of MRI sequence, wherein the MRI images of the first type were preprocessed by performing normalization and/or resizing prior to training the first encoder model, and wherein the second encoder model is trained using MRI images of the second type of MRI sequence, wherein the MRI images of the second type are preprocessed by performing intensity rescaling and z-scoring prior to training the second encoder model.
19 . The computer-implemented method of claim 2 , wherein the first encoder model is trained to learn feature representations unique to the first type of MRI sequence, and the second encoder model is trained to learn feature representations unique to the second type of MRI sequence.
20 . The computer-implemented method of claim 19 , wherein the machine-learning model is trained by minimizing a loss function to simultaneously optimize weights of the plurality of encoder models.