Continuous training of an object detection and classification model for varying environmental conditions
Apparatuses, systems, and techniques for continuous training of an object detection/classification model are provided. A first image depicting an environment is identified. Data for an object detected in the first image is obtained based on outputs of a machine learning (ML) model. A determination is made of whether a level of confidence that the object corresponds to an object class satisfies a confidence criterion. If so, first conditions corresponding to the environment depicted in the first image are identified. Noise characteristics associated with a second image depicting the environment are determined based on a difference between the first conditions and second conditions of the second image. The first image is augmented based on the noise characteristics to generate a third image depicting the environment based on the second conditions and the detected object. Training data associated with the third image is provided to train the ML model.
1 . A method comprising:
obtaining, using a processing device and based on one or more outputs of a machine learning model, object data associated with an object detected in a first image comprising a depiction of an environment associated with a first set of conditions, wherein the object data comprises an indication of a region of the first image that includes the object, an indication of an object class, and a level of confidence that the object corresponds to the object class;
responsive to determining that the level of confidence satisfies a level of confidence criterion pertaining to the first set of conditions associated with the environment, selecting, using the processing device, the first image as a baseline image for generating training data associated with re-training the machine learning model to detect objects in images depicting the environment;
determining, using the processing device, one or more noise characteristics associated with a second image comprising a depiction of the environment associated with a second set of conditions that is distinct from the first set of conditions, wherein the one or more noise characteristics are determined in view of a difference between the first set of conditions and the second set of conditions;
based on the selection of the first image as the baseline image, augmenting, using the processing device, the first image based on the one or more determined noise characteristics to generate a third image, the third image reflecting the depiction of the environment based on the second set of conditions and a depiction of the object detected in the first image; and
providing, using the processing device, training data associated with the third image to re-train the machine learning model, the training data comprising the third image, an indication of a region of the third image that includes the object, and the indication of the object class.
2 . The method of claim 1 , wherein determining the one or more noise characteristics associated with the second image comprises:
obtaining a first image characteristic associated with the first image and a second image characteristic associated with the second image, wherein the first image characteristic corresponds to a first amount of noise of the first image in view of the first set of conditions of the environment depicted in the first image and the second image characteristic corresponds to a second amount of noise of the second image in view of the second set of conditions of the environment depicted in the second image; and
calculating a difference between the first image characteristic and the second image characteristic, wherein the one or more noise characteristics associated with the second image correspond to the calculated difference.
3 . The method of claim 1 , wherein determining the one or more noise characteristics associated with the second image comprises:
determining at least one of a structural similarity index based on the first image and the second image or a peak signal-to-noise ratio (PSNR) based on the first image and the second image.
4 . The method of claim 1 , wherein the first set of conditions comprises at least a first distinct environmental condition associated with the environment depicted in the first image, and wherein the second set of conditions comprises at least a second distinct environmental condition associated with the environment depicted in the second image.
5 . The method of claim 1 , wherein augmenting the first image based on the one or more determined noise characteristics comprises:
applying one or more transformations associated with the one or more determined noise characteristics to the first image to generate the third image.
6 . The method of claim 1 , wherein the training data associated with the third image further comprises an indication of at least one of the first set of conditions, the second set of conditions, or a difference between a first condition of the first set of conditions and a corresponding second condition of the second set of conditions.
7 . The method of claim 1 , wherein the processing device is configured to execute instructions of a first process associated with detecting objects depicted in images generated for the environment and instructions of a second process associated with retraining the machine learning model, and wherein the method further comprises:
identifying an empty time slot of a processing schedule associated with the processing device; and
scheduling execution of the instructions of the second process during the identified empty time slot, wherein the instructions of the second process correspond to retraining the machine learning model based on the training data associated with the third image.
8 . The method of claim 7 , further comprising:
receiving an alert indicating a detection of motion within the environment during the execution of the instructions of the second process, wherein the alert corresponds to a fourth image comprising a depiction of the environment;
generating processing state data associated with the execution of the instructions of the second process; and
executing the instructions associated with the first process to detect one or more additional objects depicted in the fourth image.
9 . The method of claim 1 , wherein the level of confidence criterion further pertains to at least one of an operating condition associated with a camera that generated the first image or a setting associated with the camera.
10 . A system comprising:
a memory device; and
a processing device coupled to the memory device, wherein the processing device is to perform operations comprising:
obtaining, based on one or more outputs of a machine learning model, object data associated with an object detected in a first image comprising a depiction of an environment associated with a first set of conditions, wherein the object data comprises an indication of a region of the first image that includes the object, an indication of an object class, and a level of confidence that the object corresponds to the object class;
responsive to determining that the level of confidence satisfies a level of confidence criterion pertaining to the first set of conditions associated with the environment, selecting the first image as a baseline image for generating training data associated with re-training the machine learning model to detect objects in images depicting the environment;
determining one or more noise characteristics associated with a second image comprising a depiction of the environment associated with a second set of conditions that is distinct from the first set of conditions, wherein the one or more noise characteristics are determined in view of a difference between the first set of conditions and the second set of conditions;
based on the selection of the first image as the baseline image, augmenting the first image based on the one or more determined noise characteristics to generate a third image, the third image reflecting the depiction of the environment based on the second set of conditions and a depiction of the object detected in the first image; and
providing training data associated with the third image to re-train the machine learning model, the training data comprising the third image, an indication of a region of the third image that includes the object, and the indication of the object class.
11 . The system of claim 10 , wherein determining the one or more noise characteristics associated with the second image comprises:
obtaining a first image characteristic associated with the first image and a second image characteristic associated with the second image, wherein the first image characteristic corresponds to a first amount of noise of the first image in view of the first set of conditions of the environment depicted in the first image and the second image characteristic corresponds to a second amount of noise of the second image in view of the second set of conditions of the environment depicted in the second image; and
calculating a difference between the first image characteristic and the second image characteristic, wherein the one or more noise characteristics associated with the second image correspond to the calculated difference.
12 . The system of claim 10 , wherein determining the one or more noise characteristics associated with the second image comprises:
determining at least one of a structural similarity index based on the first image and the second image or a peak signal-to-noise ratio (PSNR) based on the first image and the second image.
13 . The system of claim 10 , wherein the first set of conditions comprises at least a first distinct environmental condition associated with the environment depicted in the first image and wherein the second set of conditions comprises at least a second distinct environmental condition associated with the environment depicted in the second image.
14 . The system of claim 10 , wherein augmenting the first image based on the one or more determined noise characteristics comprises:
applying one or more transformations associated with the one or more determined noise characteristics to the first image to generate the third image.
15 . The system of claim 10 , wherein the training data associated with the third image further comprises an indication of at least one of the first set of conditions, the second set of conditions, or a difference between a first condition of the first set of conditions and a corresponding second condition of the second set of conditions.
16 . The system of claim 10 , wherein the processing device is configured to execute instructions of a first process associated with detecting objects depicted in images generated for the environment and instructions of a second process associated with retraining the machine learning model, and wherein the operations further comprise:
identifying an empty time slot of a processing schedule associated with the processing device; and
scheduling execution of the instructions of the second process during the identified empty time slot, wherein the instructions of the second process correspond to retraining the machine learning model based on the training data associated with the third image.
17 . A non-transitory computer readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations comprising:
obtaining, based on one or more outputs of a machine learning model, object data associated with an object detected in a first image comprising a depiction of an environment associated with a first set of conditions, wherein the object data comprises an indication of a region of the first image that includes the object, an indication of an object class, and a level of confidence that the object corresponds to the object class;
responsive to determining that the level of confidence satisfies a level of confidence criterion pertaining to the first set of conditions associated with the environment, selecting the first image as a baseline image for generating training data associated with re-training the machine learning model to detect objects in images depicting the environment;
determining one or more noise characteristics associated with a second image comprising a depiction of the environment associated with a second set of conditions that is distinct from the first set of conditions, wherein the one or more noise characteristics are determined in view of a difference between the first set of conditions and the second set of conditions;
based on the selection of the first image as the baseline image, augmenting the first image based on the one or more determined noise characteristics to generate a third image, the third image reflecting the depiction of the environment based on the second set of conditions and a depiction of the object detected in the first image; and
providing training data associated with the third image to re-train the machine learning model, the training data comprising the third image, an indication of a region of the third image that includes the object, and the indication of the object class.
18 . The non-transitory computer readable storage medium of claim 17 , wherein determining the one or more noise characteristics associated with the second image comprises:
obtaining a first image characteristic associated with the first image and a second image characteristic associated with the second image, wherein the first image characteristic corresponds to a first amount of noise in the first image in view of the first set of conditions of the environment depicted in the first image and the second image characteristic corresponds to a second amount of noise of the second image in view of the second set of conditions of the environment depicted in the second image; and
calculating a difference between the first image characteristic and the second image characteristic, wherein the one or more noise characteristics associated with the second image correspond to the calculated difference.
19 . The non-transitory computer readable storage medium of claim 17 , wherein determining the one or more noise characteristics associated with the second image comprises:
determining at least one of a structural similarity index based on the first image and the second image or a peak signal-to-noise ratio (PSNR) based on the first image and the second image.
20 . The non-transitory computer readable storage medium of claim 17 , wherein the first set of conditions comprises at least a first distinct environmental condition associated with the environment depicted in the first image, and wherein the second set of conditions comprises at least a second distinct environmental condition associated with the environment depicted in the second image.