Methods and apparatus for identifying video-derived data
A method for generating and selecting images of objects based on video data and text data includes receiving, at a processor of a video camera system, a video stream including a series of video frames depicting at least one object. A set of at least one classification for the object is generated. Additionally, an image that depicts the object and that includes a cropped portion of a video frame from the series of video frames is generated. A set of at least one index key is generated based on the set of at least one classification, and the image is stored based on the set of at least one index key. The processor receives a signal representing a text input from a user, and the processor performs at least one of (1) retrieval of the image or (2) generation of an alert.
1 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:
receive a plurality of temporally arranged images, each image from the plurality of temporally arranged images including a depiction of an object;
generate a set of at least one identification for the object based on at least one image from the plurality of temporally arranged images;
generate a set of at least one cropped image, each cropped image from the set of at least one cropped image including a region of an image from the plurality of temporally arranged images and from a plurality of regions of that image, the image being different from each remaining image from the plurality of temporally arranged images, the region being smaller than an entirety of the image and depicting the object;
cause storage, in a memory operably coupled to the processor, of at least one cropped image from the set of at least one cropped image based on an index that is (1) configured for use with a search operation and (2) associated with the set of at least one identification;
receive, from a compute device of a user, a signal that encodes at least one identification from the set of at least one identification;
retrieve, using the search operation and the index, the at least one cropped image from the memory based on the signal; and
cause transmission of the at least one cropped image to a remote compute device for display.
2 . The non-transitory, processor-readable medium of claim 1 , wherein the plurality of temporally arranged images includes images generated by at least one sensor of a video camera.
3 . The non-transitory, processor-readable medium of claim 1 , further storing instructions to cause the processor to cause transmission of an alert to the compute device of the user in response to the signal.
4 . The non-transitory, processor-readable medium of claim 1 , wherein:
the set of at least one identification is associated with first embedded data;
the signal is associated with second embedded data; and
the search operation retrieves the at least one cropped image based on a comparison between the first embedded data and the second embedded data.
5 . The non-transitory, processor-readable medium of claim 1 , wherein the set of at least one identification is generated based on a motion associated with the object within at least two images from the plurality of temporally arranged images.
6 . The non-transitory, processor-readable medium of claim 1 , wherein:
the set of at least one identification is generated using a machine learning model and a plurality of images from the plurality of temporally arranged images; and
the set of at least one identification indicates an action.
7 . The non-transitory, processor-readable medium of claim 1 , further storing instructions to cause the processor to:
cause metadata associated with the plurality of temporally arranged images to be stored at the memory based on the index, the instructions to retrieve the at least one cropped image including instructions to retrieve the at least one cropped image from the memory based on the metadata.
8 . The non-transitory, processor-readable medium of claim 7 , wherein:
the metadata indicates a first time period; and
the signal encodes a second time period, the instructions to retrieve the at least one cropped image including instructions to retrieve the at least one cropped image from the memory based on the second time period including the first time period.
9 . The non-transitory, processor-readable medium of claim 1 , wherein:
the signal is a first signal;
the non-transitory, processor-readable medium further stores instructions to cause the processor to receive a second signal associated with a time period; and
the instructions to retrieve the at least one cropped image include instructions to retrieve the at least one cropped image from the memory based on the first signal and the second signal.
10 . The non-transitory, processor-readable medium of claim 1 , wherein the plurality of temporally arranged images is a first plurality of temporally arranged images, and the non-transitory, processor-readable medium further stores instructions to cause the processor to:
generate a second plurality of temporally arranged images based on the first plurality of temporally arranged images, the second plurality of temporally arranged images (1) having a resolution lower than a resolution of the first plurality of temporally arranged images and (2) used to generate the set of at least one identification.
11 . The non-transitory, processor-readable medium of claim 10 , wherein the instructions to generate the set of at least one identification include instructions to provide the second plurality of temporally arranged images as input to a machine learning model.
12 . An apparatus, comprising:
a processor; and
a memory operably coupled to the processor, the memory storing instructions to cause the processor to:
receive a plurality of temporally arranged images, each image from the plurality of temporally arranged images including a depiction of an object,
generate a plurality of temporally arranged compressed images based on the plurality of temporally arranged images,
generate an identification for the object based on at least one compressed image from the plurality of temporally arranged compressed images,
generate a set of at least one cropped image based on the identification, each cropped image from the set of at least one cropped image including a region of an image from the plurality of temporally arranged images, the region of the image being from a plurality of regions of that image, the image being different from each remaining image from the plurality of temporally arranged images, the region of the image being smaller than an entirety of the image and depicting the object, and
cause at least one cropped image from the set of at least one cropped image to be indexed in a searchable data structure based on the identification.
13 . The apparatus of claim 12 , wherein the memory further stores instructions to cause the processor to:
use the identification as a key to perform a search the searchable data structure for the at least one cropped image; and
retrieve the at least one cropped image from the searchable data structure in response to identifying, based on the search, the at least one cropped image indexed in the searchable data structure.
14 . The apparatus of claim 12 , further comprising a video camera operably coupled to the processor, the video camera configured to generate the plurality of temporally arranged images.
15 . The apparatus of claim 12 , wherein the memory further stores instructions to cause the processor to detect a motion of the object by providing the plurality of temporally arranged compressed images as input to a Kalman filter, the identification for the object being generated based on the motion.
16 . The apparatus of claim 12 , wherein:
the instructions to generate the identification include instructions to provide at least two temporally arranged compressed images from the plurality of temporally arranged compressed images as input to a machine learning model to generate the identification; and
the identification indicates an activity performed by the object.
17 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:
receive a plurality of temporally arranged images, each image from the plurality of temporally arranged images including a depiction of an object;
generate a plurality of temporally arranged compressed images based on the plurality of temporally arranged images;
generate detection data based on the plurality of temporally arranged compressed images;
generate a plurality of closeup images based on the plurality of temporally arranged images, each closeup image from the plurality of closeup images depicting an enlarged view of the object;
cause the plurality of closeup images and the detection data to be stored in association with each other at a database, the plurality of closeup images being indexed at the database based on the detection data;
retrieve, from the database and based on a text prompt that indicates a portion of the detection data, at least one closeup image that is from the plurality of closeup images and that is associated with the portion of the detection data; and
cause transmission of at least one signal to cause display of the at least one closeup image via a user interface of a remote compute device.
18 . The non-transitory, processor-readable medium of claim 17 , wherein:
the detection data includes a feature vector associated with at least one feature of the object.
19 . The non-transitory, processor-readable medium of claim 17 , further storing instructions to cause the processor to generate embedded data by providing the text prompt as input to a machine learning model, the instructions to retrieve the at least one closeup image including instructions to retrieve the at least one closeup image based on the embedded data.
20 . The non-transitory, processor-readable medium of claim 17 , wherein:
the text prompt is a first text prompt;
the at least one closeup image is a subset of selected closeup images from the plurality of closeup images;
the at least one signal is at least one first signal; and
the non-transitory, processor-readable medium further stores instructions to cause the processor to:
receive a second text prompt that indicates a subset of the portion of the detection data,
select at least one specific closeup image that is from the subset of selected closeup images and that is associated with the subset of the portion of the detection data, and
cause transmission of at least one second signal to cause display of the at least one specific closeup image via the user interface of the remote compute device.