Two-stage camera-based video search using a combination of draft model and full vision language model (VLM)
An apparatus comprising an interface and a processor. The interface may receive pixel data and a user input. The processor may upload encoded video to a remote device, implement a first AI, and upload image match data to the remote device in response to the user input and video frames. The first AI may parse a natural text description of the user input to determine search parameters and search the video frames for a match between the search parameters and video data of the video frames. The remote device may store the encoded video, implement a second AI and generate decoded video in response to the encoded video and the image match data. The second AI may implement a full AI model and the first AI may implement a draft model version of the full AI model. The second AI may generate a plain language description of the decoded video.
1 . An apparatus comprising:
an interface configured to receive (i) pixel data of an environment and (ii) a user input comprising a natural text description; and
a processor configured to (i) process said pixel data arranged as raw video frames, (ii) perform encoding operations to generate encoded video frames in response to said raw video frames, (iii) implement a first AI module configured to operate on said raw video frames and (iv) upload to a remote computing device (A) image match data in response to said user input and said raw video frames and (B) said encoded video frames, wherein
(a) said first AI module is configured to (i) parse said natural text description of said user input to determine search parameters, (ii) search said raw video frames for a match between said search parameters and video data of said raw video frames and (iii) enable said upload of said image match data in response to said match,
(b) said remote computing device is configured to (i) store each of said encoded video frames, (ii) implement a second AI module and (iii) extract a decoded video sequence from only a portion of said encoded video frames identified in said image match data,
(c) said second AI module implements a full AI model and said first AI module implements a draft model version of said full AI model, and
(d) said second AI module is configured to generate a plain language description of said decoded video sequence in response to receiving said upload of said image match data.
2 . The apparatus according to claim 1 , wherein (i) said draft model version of said full AI model is a Contrastive Language-Image Pre-Training (CLIP) model and (ii) said full AI model is a Vision Language Model (VLM).
3 . The apparatus according to claim 1 , wherein (i) said processor is implemented as a system-on-chip (SoC) with a device memory, (ii) said remote computing device comprises a CPU, a GPU and a system memory and (iii) said CPU, said GPU and said system memory provide more Artificial Intelligence (AI) computation resources than said processor and said device memory.
4 . The apparatus according to claim 1 , wherein said image match data comprises a frame number corresponding to said match between said search parameters and video data of said raw video frames.
5 . The apparatus according to claim 1 , wherein said apparatus comprises a security camera.
6 . The apparatus according to claim 1 , wherein said remote computing device is an AI server located on-premise with said apparatus.
7 . The apparatus according to claim 1 , wherein said remote computing device is implemented using a cloud computing network.
8 . The apparatus according to claim 1 , wherein said natural text description of said user input and said plain language description enable a video question answer configuration for searching said raw video frames.
9 . The apparatus according to claim 1 , wherein said draft model version of said full AI model is a zero shot model configured to enable a prompt of said draft model version of said full AI model to be tuned to said environment without training specific to said environment.
10 . The apparatus according to claim 1 , wherein (A) said remote computing device is configured to (i) extract said decoded video sequence from said encoded video frames and (ii) present said decoded video sequence to said second AI module, each in response to receiving said image match data and (B) limiting an input to said second AI module to said decoded video sequence when said image match data is received prevents said second AI module from continuously consuming power to perform AI operations.
11 . The apparatus according to claim 1 , wherein said decoded video sequence comprises a first number of said encoded video frames before a frame number of said image match data and a second number of said encoded video frames after said frame number of said image match data.
12 . The apparatus according to claim 1 , wherein said plain language description of said decoded video sequence is generated by said second AI module in response to performing AI operations to determine time-based and behavioral information of video data in said decoded video sequence.
13 . The apparatus according to claim 1 , wherein providing said decoded video sequence to said second AI module enables said full AI model to perform AI operations at a faster rate than a frame rate of said raw video frames.
14 . The apparatus according to claim 1 , wherein providing said decoded video sequence to said second AI module enables said full AI model to perform AI operations on demand.
15 . The apparatus according to claim 1 , wherein said draft model version of said full AI model and said full AI model are jointly trained.
16 . The apparatus according to claim 1 , wherein said draft model version of said full AI model is trained based on contrastive learning by computing a similarity between text and image pairs.
17 . The apparatus according to claim 1 , wherein said processor is configured to perform said encoding operations to generate said encoded video frames in response to said raw video frames in parallel with said first AI module performing said search of said raw video frames for said match between said search parameters and said video data of said raw video frames in real time.
18 . A system comprising:
a plurality of capture devices each configured to (i) capture pixel data, (ii) process said pixel data arranged as raw video frames, (iii) generate encoded video frames in response to encoding operations performed on said raw video frames, (iv) implement a first AI module configured to operate on said raw video frames, and (v) determine image match data in response to a user input and said raw video frames; and
a remote computing device configured to (i) store each of said encoded video frames uploaded from said plurality of capture devices, (ii) implement a second AI module and (iii) generate a decoded video sequence from only a portion of said encoded video frames in said image match data from one of said plurality of capture devices, wherein
(a) said first AI module is configured to (i) parse a natural text description of said user input to determine search parameters, (ii) search said raw video frames for a match between said search parameters and video data of said raw video frames and (iii) enable an upload of said image match data in response to said match,
(b) said second AI module implements a full AI model and said first AI module implements a draft model version of said full AI model, and
(c) said second AI module is configured to generate a plain language description of said decoded video sequence in response to receiving said upload of said image match data.
19 . The system according to claim 18 , wherein (i) said user input is provided to a companion app by a user inputting said natural text description and (ii) said plain language description is presented as an output to said companion app.
20 . The system according to claim 19 , wherein said companion app is implemented on a smartphone.