Object characterization using one or more neural networks
Apparatuses, systems, and techniques are presented to detect one or more objects in one or more images. In at least one embodiment, one or more neural networks can be used to detect one or more objects in one or more images based, at least in part, on textual descriptions of the one or more objects.
1 . One or more processors, comprising circuitry to use one or more neural networks to:
determine, from a report comprising text and one or more images:
one or more image feature embeddings from at least one image of the one or more images, at least one portion of the at least one image depicting at least a portion of at least one object of one or more objects; and
one or more text feature embeddings from at least a portion of the text, wherein the at least a portion of the text describes at least one object of the one or more objects;
determine, based on a measure of similarity between the one or more image feature embeddings and the one or more text feature embeddings, that the one or more text features-feature embeddings and the one or more image feature embeddings correspond to a same object of the at least one object, wherein the measure of similarity is a loss calculation between the one or more image feature embeddings and the one or more text feature embeddings; and
identify the at least one portion of the at least one image with the one or more text feature embeddings.
2 . The one or more processors of claim 1 , wherein the circuitry is further to extract attributes, from the portion of text, corresponding to a set of attributes for a type of the at least one object.
3 . The one or more processors of claim 2 , wherein the circuitry is further to identify one or more regions of interest in the one or more images that have at least a minimum probability of containing at least one object of the one or more objects, wherein the circuitry is further to weight the one or more regions of interest based at least in part upon respective probability values.
4 . The one or more processors of claim 3 , wherein circuitry is further to determine a similarity between the weighted regions of interest and the extracted attributes to use to detect the at least one object; and
determine the similarity surpasses one or more thresholds.
5 . The one or more processors of claim 2 , wherein the circuitry is further to provide one or more locations and one or more characterizations of the one or more objects in the one or more images based at least in part upon the image features-feature embeddings and text feature embeddings.
6 . The one or more processors of claim 1 , wherein the at least one object are instances of a disease, and wherein the one or more images and one or more textual descriptions are contained in one or more medical reports.
7 . The one or more processors of claim 1 , wherein the circuitry is to use the one or more neural networks to detect the at least one object in one or more images based, at least in part, on a similarity between one or more first feature embeddings of one or more textual descriptions of the at least one object and one or more feature embeddings of the at least one object identified by the one or more neural networks.
8 . A system comprising:
one or more processors to use one or more neural networks to:
determine, from a report comprising text and one or more images:
one or more modified image feature embeddings from at least one image of the one or more images, at least one portion of the at least one image depicting at least a portion of at least one object of one or more objects; and
one or more text feature embeddings from at least a portion of the text, wherein the at least a portion of the text describes at least one object of the one or more objects;
determine, based on a measure of similarity between the one or more image feature embeddings and the one or more text feature embeddings, that the one or more text feature embeddings and the one or more image feature embeddings correspond to a same object of the at least one object, wherein the measure of similarity is a loss calculation between the one or more image feature embeddings and one or more text feature embeddings; and
identify the at least one portion of the at least one image with the one or more text feature embeddings.
9 . The system of claim 8 , wherein the one or more processors are further to extract attributes, from the portion of text, corresponding to a set of attributes for a type of the at least one object.
10 . The system of claim 9 , wherein the one or more processors are further to identify one or more regions of interest in the one or more images that have at least a minimum probability of containing at least one object of the one or more objects, wherein the one or more processors are further to weight the one or more regions of interest based at least in part upon respective probability values.
11 . The system of claim 10 , wherein the one or more processors are further to determine a similarity between the weighted regions of interest and the extracted attributes to use to detect the at least one object; and
determine the similarity surpasses one or more thresholds.
12 . The system of claim 9 , wherein the one or more processors are further to provide one or more locations and one or more characterizations of the one or more objects in the one or more images based at least in part upon the image feature embeddings and text feature embeddings.
13 . The system of claim 8 , wherein the at least one object includes instances of a disease, and wherein the one or more images and one or more textual descriptions are contained in one or more medical reports.
14 . A method comprising:
using one or more neural networks to:
determine, from a report comprising text and one or more images:
one or more modified image feature embeddings from at least one image of the one or more images, at least one portion of the at least one image depicting at least a portion of at least one object of one or more objects; and
one or more text feature embeddings from at least a portion of the text, the at least a portion of the text describing at least one object of the one or more objects;
determine, based on a measure of similarity between the one or more image feature embeddings and the one or more text feature embeddings, that the one or more text feature embeddings and the one or more image feature embeddings correspond to a same object of the at least one object, wherein the measure of similarity is a loss calculation between the one or more image feature embeddings and one or more text feature embeddings; and
identify the at least one portion of the at least one image with the one or more text feature embeddings.
15 . The method of claim 14 , further comprising:
extracting one or more attributes, from the portion of text, corresponding to a set of attributes for a type of the at least one object.
16 . The method of claim 15 , further comprising:
identifying one or more regions of interest in the one or more images that have at least a minimum probability of containing at least one object of the one or more objects; and
weighting the one or more regions of interest based at least in part upon respective probability values.
17 . The method of claim 16 , further comprising:
determining a similarity between the weighted regions of interest and the extracted one or more attributes to use to detect the at least one object; and
determining the similarity surpasses one or more thresholds.
18 . The method of claim 14 , further comprising:
providing one or more locations and one or more characterizations of the at least one object in the one or more images based at least in part upon the image feature embeddings and text feature embeddings.
19 . The method of claim 14 , wherein the at least one object includes one or more instances of a disease, and wherein the one or more images and one or more textual descriptions are contained in one or more medical reports.
20 . A machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:
use one or more neural networks to:
determine, from a report comprising text and one or more images:
one or more modified image feature embeddings from at least one image of the one or more images, at least one portion of the at least one image depicting at least a portion of at least one object of one or more objects; and
one or more text feature embeddings from at least a portion of the text, wherein the at least a portion of the text describes at least one object of the one or more objects;
determine, based on a measure of similarity between the one or more image feature embeddings and the one or more text feature embeddings, that the one or more text feature embeddings and the one or more image feature embeddings correspond to a same object of the at least one object, wherein the measure of similarity is a loss calculation between the one or more image feature embeddings and one or more text feature embeddings; and
identify the at least one portion of the at least one image with the one or more text feature embeddings.
21 . The machine-readable medium of claim 20 , wherein the instructions if performed further cause the one or more processors to:
extract ones or more attributes, from the portion of text, corresponding to a set of attributes for a type of the at least one object.
22 . The machine-readable medium of claim 21 , wherein the instructions if performed further cause the one or more processors to:
identify one or more regions of interest in the one or more images that have at least a minimum probability of containing at least one object of the one or more objects; and
weight the one or more regions of interest based at least in part upon respective probability values.
23 . The machine-readable medium of claim 22 , wherein the instructions if performed further cause the one or more processors to:
determine a similarity between the weighted regions of interest and the extracted one or more attributes to use to detect the at least one object; and
determine the similarity surpasses one or more thresholds.
24 . The machine-readable medium of claim 21 , wherein the instructions if performed further cause the one or more processors to:
provide one or more locations and one or more characterizations of the one or more objects in the one or more images based at least in part upon the image feature embeddings and text feature embeddings.
25 . The machine-readable medium of claim 20 , wherein the at least one object includes one or more instances of a disease, and wherein the one or more images and one or more textual descriptions are contained in one or more medical reports.
26 . An object detection system, comprising:
one or more processors to use one or more neural networks to:
determine, from a report comprising text and one or more images:
one or more modified image feature embeddings from at least one image of the one or more images, at least one portion of the at least one image depicting at least a portion of at least one object of one or more objects; and
one or more text feature embeddings from at least a portion of the text, wherein the at least a portion of the text describes at least one object of the one or more objects;
determine, based on a measure of similarity between the one or more image feature embeddings and the one or more text feature embeddings, that the one or more text feature embeddings and the one or more image feature embeddings correspond to a same object of the at least one object, wherein the measure of similarity is a loss calculation between the one or more image feature embeddings and one or more text feature embeddings; and
identify the at least one portion of the at least one image with the one or more text feature embeddings;
and memory for storing network parameters for the one or more neural networks.
27 . The object detection system of claim 26 , wherein the one or more processors are further to extract attributes, from the portion of text, corresponding to a set of attributes for a type of the at least one object.
28 . The object detection system of claim 27 , wherein the one or more processors are further to identify one or more regions of interest in the one or more images that have at least a minimum probability of containing at least one object of the one or more objects, wherein the one or more processors are further to weight the one or more regions of interest based at least in part upon respective probability values.
29 . The object detection system of claim 28 , wherein the one or more processors are further to determine a similarity between the weighted regions of interest and the extracted attributes to use to detect the one or more objects; and
determining the similarity surpasses one or more thresholds.
30 . The object detection system of claim 27 , wherein the one or more processors are further to provide one or more locations and one or more characterizations of the one or more objects in the one or more images based at least in part upon the image feature embeddings and text feature embeddings.
31 . The object detection system of claim 26 , wherein the at least one object includes one or more instances of a disease, and wherein the one or more images and one or more textual descriptions are contained in one or more medical reports.