DIGITAL IMAGE ANNOTATION AND RETRIEVAL SYSTEMS AND METHODS
In a digital image annotation and retrieval system, a machine learning model identifies an image feature in an image and generates a plurality of question prompts for the feature. For a particular feature, a feature annotation is generated, which can include capturing a narrative, determining a plurality of narrative units, and mapping a particular narrative unit to the identified image feature. An enriched image is generated using the generated feature annotation. The enriched image includes searchable metadata comprising the feature annotation and the plurality of question prompts.
1 . A computing system comprising at least one processor, at least one memory, and computer-executable instructions stored in the at least one memory, the computer-executable instructions structured, when executed, to cause the at least one processor to perform operations comprising:
generating, by a trained machine learning model, a set of question prompts for one or more image features identified from at least one image of a set of images;
for at least one question prompt of the set of question prompts, generating a feature annotation, wherein generating the feature annotation comprises:
capturing a narrative, via one or more information capturing devices, the narrative comprising at least one of a video file, an audio file, or text;
based on the captured narrative, generating a transcript comprising text corresponding to the captured narrative;
using the transcript to identify a set of narrative units; and
mapping a particular narrative unit from the set of narrative units to the identified image feature to generate the feature annotation;
using the generated feature annotation and based on the at least one image of the set of images, generating an enriched image,
wherein the enriched image comprises rich metadata, the rich metadata comprising the generated feature annotation and the at least one question prompt of the set of question prompts; and
in response to detecting a user input,
identifying a set of search terms using the user input; and
searching the set of images, the rich metadata, and the plurality of question prompts using the search term to identify an item that is responsive to at least one search term in the set of search terms.