SYSTEMS AND METHODS FOR INDEXING AND RETRIEVING ELECTRONIC INFORMATION
Described herein are systems and methods for indexing multimodal datasets including digital image data, structured textual data, unstructured textual data, keyword textual data, or any combination thereof. Further described herein are systems and methods for querying and searching indexed multimodal datasets to retrieve original source data relevant to the query. The systems and methods described herein are agnostic as to mode of data and as to mode of query.
1 . A method for indexing multimodal dataset, the method comprising:
receiving a multimodal dataset comprising a plurality of digital image inputs and a plurality of textual inputs;
generating, using a group level aggregator model, a first plurality of group level vectors based on the plurality of digital image inputs;
generating, using the group level aggregator model, a second plurality of group level vectors based on the plurality of textual inputs; and
storing the first plurality of group level vectors and the second plurality of group level vectors in an index.
2 . The method of claim 1 , wherein generating the first plurality of group level vectors comprises:
generating, using a trained foundation model, a plurality of tile level vectors based on each digital image of the plurality of digital image inputs; and
aggregating, using the group level aggregator model, each plurality of tile level vectors to generate the first plurality of group level vectors.
3 . The method of claim 2 , wherein each group level vector represents one or more features extracted from individual regions within a corresponding digital image of the plurality of digital image inputs.
4 . The method of claim 1 , wherein the plurality of digital image inputs comprises digital medical images.
5 . The method of claim 4 , wherein the digital medical images comprise an image of a cytology specimen, an image of histopathology specimen, a whole slide image, a multiplex immunofluorescent image, a multiplex immunohistochemistry image, a magnetic resonance imaging (MRI) image, a computed tomography (CT) image, an X-ray image, a nuclear medicine imaging (NMI) image, an ultrasound image, a mammography image, an endoscopic image, an angiography image, a confocal microscopy image, a fluorescence in situ hybridization image, an optical coherence tomography image, a bone scan image, a thermography image, an electron microscopy image, and/or other images supporting detailed visualization and characterization of tissue specimens to evaluate disease mechanisms, progression, and therapeutic response across diverse clinical contexts.
6 . The method of claim 1 , wherein the plurality of textual inputs comprises unstructured text, keyword text, structured text, or a combination thereof.
7 . The method of claim 6 , wherein the unstructured text comprises tabular medical data, diagnosis information, notes regarding sample retrieval and/or preparation, histological details, clinical context involving patient history and other modalities of tests, information specific to staining and markers, morphological observations, transcripts from auditory comments or opinions, references to other tests, references to treatment data, or a combination thereof.
8 . The method of claim 6 , wherein the keyword text comprises a medical term, a diagnostic code, or a morphological descriptor.
9 . The method of claim 6 , wherein the structured text comprises genetic sequencing data, genomic data, molecular data, proteomic data, standardized diagnostic reports, coded medical records, or templated clinical forms.
10 . The method of claim 1 , wherein generating the first plurality of group level vectors and/or the second plurality of group level vectors comprises aggregation.
11 . The method of claim 1 , wherein generating the first plurality of group level vectors and/or the second plurality of group level vectors comprises aggregation and compression.
12 . A system for indexing a multimodal dataset, the system comprising:
at least one memory storing instructions; and
at least one processor configured to execute the instructions to perform operations comprising:
receiving a multimodal dataset comprising a plurality of digital image inputs and a plurality of textual inputs;
generating, using a group level aggregator model, a first plurality of group level vectors based on the plurality of digital image inputs;
generating, using the group level aggregator model, a second plurality of group level vectors based on the plurality of textual inputs; and
storing the first plurality of group level vectors and the second plurality of group level vectors in an index.
13 . The system of claim 12 , wherein generating the first plurality of group level vectors comprises:
generating, using a trained foundation model, a plurality of tile level vectors based on each digital image of the plurality of digital image inputs; and
aggregating, using the group level aggregator model, each plurality of tile level vectors to generate the first plurality of group level vectors.
14 . The system of claim 13 , wherein each tile level vector represents one or more features extracted from individual regions within a corresponding digital image of the plurality of digital image inputs.
15 . The system of claim 12 , wherein the plurality of digital image inputs comprises digital medical images.
16 . The system of claim 12 , wherein the plurality of textual inputs comprises unstructured text, keyword text, structured text, or a combination thereof.
17 . The system of claim 16 , wherein:
the unstructured text comprises tabular medical data, diagnosis information, notes regarding sample retrieval and/or preparation, histological details, clinical context involving patient history and other modalities of tests, information specific to staining and markers, morphological observations, transcripts from auditory comments or opinions, references to other tests, references to treatment data, or a combination thereof;
the keyword text comprises a medical term, a diagnostic code, or a morphological descriptor; and/or
the structured text comprises genetic sequencing data, genomic data, molecular data, proteomic data, standardized diagnostic reports, coded medical records, or templated clinical forms.
18 . The system of claim 16 , wherein generating the first plurality of group level vectors and/or the second plurality of group level vectors comprises aggregation.
19 . The system of claim 16 , wherein generating the first plurality of group level vectors and/or the second plurality of group level vectors comprises aggregation and compression.
20 . A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform a method for indexing a multimodal dataset, the method comprising:
receiving a multimodal dataset comprising a plurality of digital image inputs and a plurality of textual inputs;
generating, using a group level aggregator model, a first plurality of group level vectors based on the plurality of digital image inputs;
generating, using the group level aggregator model, a second plurality of group level vectors based on the plurality of textual inputs; and
storing the first plurality of group level vectors and the second plurality of group level vectors in an index.