System, method, and computer program product for streaming data mining with text-image joint embeddings
View Patent ↗Provided are systems, methods, and computer program products for streaming data mining with text-image joint embeddings by obtaining a roadway dataset, having a condition that each record in the roadway dataset includes an image logged by an autonomous vehicle (AV) in a roadway, generating based on a text-image model, an embedding dataset comprising an image embedding for each record in the roadway dataset, each image embedding identifying a point in a high dimension space, determining, by the one or more processors, a parametric description of a high probability field which surrounds a query input text in the high dimension space, determining, by the one or more processors, at least one input image of an input stream of images that is in the high probability field, and automatically labeling, by the one or more processors, the at least one input image based on the query input text.
1 . A computer-implemented method for mining streaming data, comprising:
obtaining, by one or more processors, a roadway dataset comprising a plurality of records, wherein each record in the roadway dataset includes an image logged by an autonomous vehicle (AV) in a roadway;
generating, by the one or more processors based on a text-image model, an embedding dataset comprising an image embedding for each record in the roadway dataset, each image embedding identifying a point in a high dimension space;
receiving, by the one or more processors, a query input text;
generating, by the one or more processors, a text query embedding to identify a query input point by encoding the query input text;
determining, by the one or more processors, a parametric description of a high probability field which surrounds a query input text in the high dimension space;
determining, by the one or more processors, a distance between the query input point and each image embedding of the embedding dataset;
generating, by the one or more processors, a normal distribution;
determining, by the one or more processors, at least one input image of an input stream of images that is in the high probability field;
automatically labeling, by the one or more processors, the at least one input image based on the query input text; and
determining to not label a second image of the input stream of images in response to determining, by the one or more processors, that the second image of the input stream of images is not in the high probability field,
wherein each input image in the input stream of images that resembles the query input text comprises a point in the high dimension space at a position within a threshold distance of a point in the high probability field representing the query input text, and
wherein each input image in the input stream of images that does not resemble the query input text comprises a point in the high dimension space at a position outside the threshold distance of the point in the high probability field representing the query input text that is outside the threshold distance.
2 . The method of claim 1 , further comprising:
encoding an image to provide the image embedding for a corresponding image logged by the AV in the roadway dataset, wherein the image embedding represents an image embedding point in the high dimension space; and
matching an embedding of each input image in the input stream of images, with the parametric description of the high probability field to determine that each input image in the input stream of images is in the high probability field,
wherein the parametric description for the high probability field which surrounds the query input point associated with a query input text, and
wherein it is probable based on parameters in each image embedding of the embedding dataset that each input image fitting the parametric description of the high probability field is positioned in the high probability field of the query input point.
3 . The method of claim 1 , wherein the input query text comprises a textual description and a probability parameter,
wherein the input query text is received from a remote computing device or a remote application,
wherein the input stream of images is obtained or detected by the AV, and
wherein the embedding of the input image in the input stream of images is compared by the AV to the parametric description of the high probability field to determine if the image is in the high probability field.
4 . The method of claim 1 , wherein the high probability field comprises at least one of a spread, a top percentage of probable records, or a range representing an area of the high probability field.
5 . The method of claim 1 , wherein each record in the embedding dataset includes at least one image embedding and a link to a specific image which visualizes the image embedding.
6 . The method of claim 1 , further comprising:
slicing at least one image of the roadway dataset into a plurality of image slices; and
encoding each of the plurality of image slices,
wherein each of the plurality of image slices includes a metadata label.
7 . The method of claim 1 , wherein a content of an input image of the input stream of images is determined by comparing an embedding of the image in the input stream with at least one image embedding in the roadway dataset, without comparing pixel values.
8 . A system, comprising:
a memory; and
at least one processor coupled to the memory and configured to:
obtain a roadway dataset, wherein each record in the roadway dataset includes an image logged by an autonomous vehicle (AV) in a roadway;
generate, based on a text-image model, an embedding dataset comprising an image embedding for each record in the roadway dataset, each image embedding identifying a point in a high dimension space;
receive a query input text;
generate a text query embedding to identify a query input point by encoding the query input text;
determine a parametric description of a high probability field which surrounds a query input text in the high dimension space;
determine a distance between the query input point and each image embedding of the embedding dataset;
generate a normal distribution;
determine at least one input image of an input stream of images that is in the high probability field; and
automatically label the at least one input image based on the query input text;
determining to not label a second image of the input stream of images in response to determining that the second image of the input stream of images is not in the high probability field,
wherein each input image in the input stream of images that resembles the query input text comprises a point in the high dimension space at a position within a threshold distance of a point in the high probability field representing the query input text, and
wherein each input image in the input stream of images that does not resemble the query input text comprises a point in the high dimension space at a position outside the threshold distance of the point in the high probability field representing the query input text that is outside the threshold distance.
9 . The system of claim 8 , wherein the at least one processor coupled to the memory is further configured to:
encode an image to provide the image embedding for a corresponding image logged by the AV in the roadway dataset, wherein the image embedding represents an image embedding point in the high dimension space; and
match an embedding of each input image in the input stream of images, with the parametric description of the high probability field to determine that each input image in the input stream of images is in the high probability field,
wherein the parametric description for the high probability field which surrounds the query input point associated with a query input text, and
wherein it is probable based on parameters in each image embedding of the embedding dataset that each input image fitting the parametric description of the high probability field is positioned in the high probability field of the query input point.
10 . The system of claim 9 ,
wherein the input query text comprises a textual description and a probability parameter,
wherein the input query text is received from a remote computing device or a remote application,
wherein the input stream of images is obtained or detected by the AV, and
wherein the embedding of the input image in the input stream of images is compared by the AV to the parametric description of the high probability field to determine if the image is in the high probability field.
11 . The system of claim 8 , wherein the high probability field comprises at least one of a spread, a top percentage of probable records, or a range representing an area of the high probability field.
12 . The system of claim 8 , wherein each record in the embedding dataset includes at least one image embedding and a link to a specific image which visualizes the image embedding.
13 . The system of claim 8 , wherein the at least one processor coupled to the memory is further configured to:
slice at least one image of the roadway dataset into a plurality of image slices; and
encode each of the plurality of image slices,
wherein each of the plurality of image slices includes a metadata label.
14 . The system of claim 8 , wherein a content of an input image of the input stream of images is determined by comparing an embedding of the image in the input stream with at least one image embedding in the roadway dataset, without comparing pixel values.
15 . A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device comprising one or more processors, cause the at least one computing device to:
obtain a roadway dataset, wherein each record in the roadway dataset includes an image logged by an autonomous vehicle (AV) in a roadway;
generate, based on a text-image model, an embedding dataset comprising an image embedding for each record in the roadway dataset, each image embedding identifying a point in a high dimension space;
receive a query input text;
generate a text query embedding to identify a query input point by encoding the query input text;
determine a parametric description of a high probability field which surrounds a query input text in the high dimension space;
determining a distance between the query input point and each image embedding of the embedding dataset;
generating a normal distribution;
determine at least one input image of an input stream of images that is in the high probability field;
automatically label the at least one input image based on the query input text; and
determine to not label a second image of the input stream of images in response to determining that the second image of the input stream of images is not in the high probability field,
wherein each input image in the input stream of images that resembles the query input text comprises a point in the high dimension space at a position within a threshold distance of a point in the high probability field representing the query input text, and
wherein each input image in the input stream of images that does not resemble the query input text comprises a point in the high dimension space at a position outside the threshold distance of the point in the high probability field representing the query input text that is outside the threshold distance.
16 . The non-transitory computer-readable medium of claim 15 , wherein the instructions stored thereon cause the at least one computing device to:
encode an image to provide the image embedding for a corresponding image logged by the AV in the roadway dataset, wherein the image embedding represents an image embedding point in the high dimension space; and
match an embedding of each input image in the input stream of images, with the parametric description of the high probability field to determine that each input image in the input stream of images is in the high probability field,
wherein the parametric description for the high probability field which surrounds the query input point associated with a query input text, and
wherein it is probable based on parameters in each image embedding of the embedding dataset that each input image fitting the parametric description of the high probability field is positioned in the high probability field of the query input point.