Machine-learned model based event detection
Example embodiments described herein therefore relate to an object-model based event detection system that comprises a plurality of sensor devices, to perform operations that include: generating sensor data at the plurality of sensor devices; accessing the sensor data generated by the plurality of sensor devices; detecting an event, or precursor to an event, based on the sensor data, wherein the detected event corresponds to an event category; accessing an object model associated with the event type in response to detecting the event, wherein the object model defines a procedure to be applied by the event detection system to the sensor data; and streaming at least a portion of a plurality of data streams generated by the plurality of sensor devices to a server system based on the procedure, wherein the server system may perform further analysis or visualization based on the portion of the plurality of data streams.
1. A system comprising:
at least one sensor device to generate sensor data comprising a plurality of data streams;
a memory; and
at least one hardware processor to perform operations comprising:
receiving sensor data at a sensor device that includes a dashcam mounted at a vehicle, the sensor data comprising monocular image data;
applying a stereoscopic inference model to the monocular image data, the stereoscopic inference model trained to generate a 3-dimensional (3D) depth model based on the monocular image data;
constructing the #D depth model based on the monocular image data generated by the dashcam and the stereoscopic inference model;
detecting an event based on the 3D depth model;
accessing at least a portion of the plurality of data streams in response to the detecting the event; and
causing display of a presentation of the portion of the plurality of data streams at a client device, the presentation of the portion of the plurality of data streams including an identifier associated with the sensor device.
2. The system of claim 1 , wherein the detecting the event based on the sensor data includes:
performing a comparison of the 3D depth model against one or more threshold values; and
detecting the event based on the comparison.
3. The system of claim 1 , wherein the sensor data includes video data, and the detecting the event includes:
extracting a set of features from the video data;
applying the set of features from the video data to a machine learned model; and
detecting the event based on an output of the machine learned model.
4. The system of claim 1 , wherein the sensor data includes image data that comprises a set of image features, and the detecting the event based on the sensor data includes:
determining a point of gaze based on the image features; and
detecting the event based on the point of gaze.
5. The system of claim 1 , wherein the at least one hardware processor performs operations further comprising:
applying a first portion of the plurality of data streams to a machine learned model at the sensor device;
detecting a precursor to the event at the sensor device based on an output of the machine learned model;
accessing a second portion of the plurality of data streams in response to the detecting the precursor to the event at the sensor device; and
wherein the detecting the event includes detecting the event based on the second portion of the plurality of data streams.
6. The system of claim 1 , wherein the at least one hardware processor performs operations further comprising:
applying the sensor data from a first portion of the plurality of data streams to a first machine learned model;
detecting a precursor to the event based on a first output of the first machine learned model;
applying the sensor data from a second portion of the plurality of data streams to a second machine learned model in response to the detecting the precursor to the event; and
wherein the detecting the event includes detecting the event based on a second output of the second machine learned model.
7. A method comprising:
receiving sensor data at a sensor device that includes a dashcam mounted at a vehicle, the sensor data monocular image data from one or more of a plurality of data streams;
applying a stereoscopic inference model to the monocular image data, the stereoscopic inference model trained to generate a 3-dimensional (3D) depth model based on the monocular image data;
constructing the #D depth model based on the monocular image data generated by the dashcam and the stereoscopic inference model;
detecting an event based on the 3D depth model;
accessing at least a portion of the plurality of data streams in response to the detecting the event; and
causing display of a presentation of the portion of the plurality of data streams at a client device, the presentation of the portion of the plurality of data streams including an identifier associated with the sensor device.
8. The method of claim 7 , wherein the detecting the event based on the sensor data includes:
performing a comparison of the 3D depth model against one or more threshold values; and
detecting the event based on the comparison.
9. The method of claim 7 , wherein the sensor data includes video data, and the detecting the event includes:
extracting a set of features from the video data;
applying the set of features from the video data to a machine learned model; and
detecting the event based on an output of the machine learned model.
10. The method of claim 7 , wherein the sensor data includes image data that comprises a set of image features, and the detecting the event based on the sensor data includes:
determining a point of gaze based on the image features; and
detecting the event based on the point of gaze.
11. The method of claim 7 , wherein the method further comprises:
applying a first portion of the plurality of data streams to a machine learned model at the sensor device;
detecting a precursor to the event at the sensor device based on an output of the machine learned model;
accessing a second portion of the plurality of data streams in response to the detecting the precursor to the event at the sensor device; and
wherein the detecting the event includes detecting the event based on the second portion of the plurality of data streams.
12. The method of claim 7 , wherein the method further comprises:
applying the sensor data from a first portion of the plurality of data streams to a first machine learned model;
detecting a precursor to the event based on a first output of the first machine learned model;
applying the sensor data from a second portion of the plurality of data streams to a second machine learned model in response to the detecting the precursor to the event; and
wherein the detecting the event includes detecting the event based on a second output of the second machine learned model.
13. A non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations comprising:
receiving sensor data at a sensor device that includes a dashcam mounted at a vehicle, the sensor data monocular image data from one or more of a plurality of data streams;
applying a stereoscopic inference model to the monocular image data, the stereoscopic inference model trained to generate a 3-dimensional (3D) depth model based on the monocular image data;
to constructing the #D depth model on the monocular image data generated by the dashcam and the stereoscopic inference model;
detecting an event based on the 3D depth model;
accessing at least a portion of the plurality of data streams in response to the detecting the event; and
causing display of a presentation of the portion of the plurality of data streams at a client device, the presentation of the portion of the plurality of data streams including an identifier associated with the sensor device.
14. The non-transitory machine-readable storage medium of claim 13 , wherein the detecting the event based on the sensor data includes:
performing a comparison of the 3D depth model against one or more threshold values; and
detecting the event based on the comparison.
15. The non-transitory machine-readable storage medium of claim 13 , wherein the sensor data includes video data, and the detecting the event includes:
extracting a set of features from the video data;
applying the set of features from the video data to a machine learned model; and
detecting the event based on an output of the machine learned model.
16. The non-transitory machine-readable storage medium of claim 13 , wherein the sensor data includes image data that comprises a set of image features, and the detecting the event based on the sensor data includes:
determining a point of gaze based on the image features; and
detecting the event based on the point of gaze.
17. The non-transitory machine-readable storage medium of claim 13 , wherein the instructions cause the machine to perform operations further comprising:
applying a first portion of the plurality of data streams to a machine learned model at the sensor device;
detecting a precursor to the event at the sensor device based on an output of the machine learned model;
accessing a second portion of the plurality of data streams in response to the detecting the precursor to the event at the sensor device; and
wherein the detecting the event includes detecting the event based on the second portion of the plurality of data streams.