Intelligent notifications for truck bed camera enabled by smart metadata using image to text models
An apparatus comprising an interface and a processor. The interface may be configured to receive pixel data of an environment near a vehicle. The processor may be configured to process the pixel data arranged as video frames, perform computer vision operations on the video frames to detect objects, store an inventory of items in response to generating a text description of each of the objects, determine criteria for a notification rule for the objects of the inventory of items in response to a user input, determine whether the criteria for the notification rule for the objects has been met, and generate a notification according to the notification rule in response to detecting that the criteria has been met. The processor may comprise an AI module configured to perform video to text analysis of the video frames to generate the text description and determine the notification rule.
1 . An apparatus comprising:
an interface configured to receive pixel data of an environment near a vehicle; and
a processor configured to (i) process said pixel data arranged as video frames, (ii) perform a video to text analysis on said video frames to generate a text description of said video frames comprising plain language that fully describes visual content captured in said video frames, (iii) generate an inventory of items comprising each object described in said environment from said text description, (iv) determine criteria for a notification rule for one of said objects of said inventory of items in response to a user input, (v) compare said text description of said video frames to said criteria for said notification rule to determine whether said criteria for said notification rule for said one of said objects has been met, and (vi) generate a notification according to said notification rule in response to detecting that said criteria has been met, wherein
said processor comprises an Artificial Intelligence (AI) module configured to (a) perform said video to text analysis of said video frames to generate said text description of said video frames comprising said plain language that fully describes said visual content of said video frames and (b) determine said notification rule in response to a conversational interaction for receiving said user input.
2 . The apparatus according to claim 1 , wherein (i) a camera is installed on a rear window of said vehicle, (ii) said environment comprises a truck bed of said vehicle, and (iii) said camera is configured to capture said pixel data of said truck bed through said rear window of said vehicle.
3 . The apparatus according to claim 2 , wherein said inventory of items corresponds to said objects in said truck bed.
4 . The apparatus according to claim 2 , wherein (i) said camera comprises an adhesive on a side of said camera with a lens and (ii) said adhesive enables said camera to be installed on said rear window.
5 . The apparatus according to claim 2 , wherein said criteria for said notification rule comprises at least one of (i) detecting that said one of said objects is not in said truck bed and (ii) detecting that anyone has touched said one of said objects in said truck bed.
6 . The apparatus according to claim 2 , wherein said text description of each of said objects comprises an identification of said objects and a location of said objects in said truck bed.
7 . The apparatus according to claim 6 , wherein said location of said objects in said truck bed is updated as said objects change said location in said truck bed over time.
8 . The apparatus according to claim 1 , wherein (i) a camera is integrated as part of said vehicle, and (ii) said camera is configured to capture said pixel data of said environment near said vehicle when said vehicle is parked.
9 . The apparatus according to claim 8 , wherein said camera is a backup camera of said vehicle.
10 . The apparatus according to claim 1 , wherein (i) said video to text analysis is configured to perform facial recognition operations to identify a person and (ii) said notification rule comprises one or more approved people for accessing said one of said objects.
11 . The apparatus according to claim 10 , wherein (i) an app for a smartphone is implemented to enable receiving said user input and presenting said notification, and (ii) said app is configured to receive photos captured by a camera of said smartphone to use as reference images for said facial recognition operations.
12 . The apparatus according to claim 1 , wherein said user input comprises a natural language text description for said notification rule and said AI module is a large language model configured to determine said criteria for said notification rule in response to said natural language text description.
13 . The apparatus according to claim 1 , wherein said inventory of items correspond to supplies for a tailgate party.
14 . The apparatus according to claim 1 , wherein (i) a remote server is configured to store a database of items, (ii) said database of items comprises (a) reference images for said inventory of items and (b) an item description of said inventory of items, and (iii) said apparatus is further configured to communicate with said remote server to compare said objects detected to said database of items to determine said inventory of items.
15 . The apparatus according to claim 1 , wherein said AI module comprises (i) a first AI model configured to perform said video to text analysis of said video frames to generate said text description of said objects in said environment and (ii) a second AI model configured to determine said notification rule in response to said user input.
16 . The apparatus according to claim 15 , wherein a third AI model is configured to compare said criteria with said text description to determine whether to generate said notification.
17 . The apparatus according to claim 16 , wherein (i) said text description is stored as smart metadata generated by a transformer network implemented by said AI module, (ii) said smart metadata comprises a natural language text description of said video frames and (iii) said third AI model is configured to compare said criteria to said natural language text description of said video frames.
18 . The apparatus according to claim 1 , further comprising a sensor fusion module, wherein
(i) said interface is further configured to receive data from one or more sensors of said vehicle,
(ii) said sensor fusion module is configured to (a) receive said data and (b) make inferences in response to an analysis of said data and said video frames, and
(iii) said AI module is configured to generate said text description in response to said inferences.
19 . The apparatus according to claim 18 , wherein said sensors comprise one or more of (i) a lidar, (ii) a high resolution radar, and (iii) a thermal camera.
20 . The apparatus according to claim 1 , wherein said processor is further configured to (i) perform computer vision operations on said video frames to detect said objects in said environment, (ii) generate a frame number in response to a detection threshold for said objects, (iii) extract a subset of said video frames in response to said frame number, (iv) present said subset of said video frames to said AI module and (v) perform said video to text analysis on said subset of said video frames after said computer vision operations are performed.