Video training data for machine learning detection/recognition of products
Described herein are systems, apparatus, methods and computer program products configured for image detection/recognition of products. The disclosed systems and techniques utilize video data to provide the necessary number of images and view angles needed to train a machine learning product detection/recognition system to recognize a specific product within later provided images. In various embodiments, a user may provide video data and the video data may be transformed in a manner that may aid in training of the machine learning system.
1. A computer-implemented method implemented in a machine learning environment, the method comprising:
receiving first video data, the first video data comprising a plurality of frames of a first video;
identifying a first object within the plurality of frames;
determining that the plurality of frames depict the first object in a number of different positions greater than a threshold number of positions;
creating training data by modifying the first video data to highlight the first object within the plurality of frames, the modifying comprising:
identifying the first object within a first plurality of frames;
determining non-first object pixels within the first plurality of frames;
creating first modified video data by replacing, independent of the pixels of the first object, the non-first object pixels with first background pixels in the first plurality of frames; and
creating second modified video data by replacing, independent of the pixels of the first object, the first background pixels with second background pixels in the first plurality of frames; and
providing the training data to a machine learning program, wherein the training data is configured to train the machine learning program to identify the first object within an image.
2. The method of claim 1 , wherein the creating the training data further comprises creating a plurality of sets of second modified video data, each set of the second modified video data based on the first modified video data, wherein each of the sets of the second modified video data highlights the first object within each of the first plurality of frames.
3. The method of claim 2 , wherein each of the sets of the second modified video data is modified to reflect different product placement and lighting conditions.
4. The method of claim 1 , wherein the modifying further comprises:
creating third modified video data by constructing a bounding box around the first object, wherein the bounding box is constructed around the first object in each of the first plurality of frames.
5. The method of claim 1 , further comprising:
annotating the video data to further highlight the first object within each of the first plurality of frames.
6. The method of claim 5 , wherein the annotating the video data comprises providing annotations in a secondary file.
7. The method of claim 1 , wherein the first object within the first video data is a three-dimensional computer aided design (CAD) object.
8. The method of claim 1 , further comprising:
transmitting instructions for obtaining the first video data to a user device.
9. A computer program product comprising computer-readable program code capable of being executed by one or more processors in an object detection/recognition environment when retrieved from a non-transitory computer-readable medium, the program code comprising instructions configurable to cause operations comprising:
receiving first video data, the first video data comprising a plurality of frames of a first video;
identifying a first object within the plurality of frames;
determining that the plurality of frames depict the first object in a number of different positions greater than a threshold number of positions;
creating training data by modifying the first video data to highlight the first object within the plurality of frames, the modifying comprising:
identifying the first object within a first plurality of frames;
determining non-first object pixels within the first plurality of frames;
creating first modified video data by replacing, independent of the pixels of the first object, the non-first object pixels with first background pixels in the first plurality of frames; and
creating second modified video data by replacing, independent of the pixels of the first object, the first background pixels with second background pixels in the first plurality of frames; and
providing the training data to a machine learning program, wherein the training data is configured to train the machine learning program to identify the first object within an image.
10. The computer program product of claim 9 , wherein the creating the training data further comprises creating a plurality of sets of second modified video data each set of the second modified video data based on the first modified video data, wherein each of the sets of the second modified video data highlights the first object within each of the first plurality of frames.
11. The computer program product of claim 10 , wherein each of the sets of the second modified video data is modified to reflect different product placement and lighting conditions.
12. The computer program product of claim 9 , wherein the modifying further comprises:
creating third modified video data by constructing a bounding box around the first object in each of the first plurality of frames.
13. The computer program product of claim 9 , wherein the operations further comprise:
annotating the video data to further highlight the first object within each of the first plurality of frames.
14. The computer program product of claim 13 , wherein the annotating the video data comprises providing annotations in a secondary file.
15. The computer program product of claim 9 , wherein the first object within the first video data is a three-dimensional computer aided design (CAD) object.
16. The computer program product of claim 9 , wherein the operations further comprise:
transmitting instructions for obtaining the first video data to a user device.
17. A computing system implemented using a server system implemented in an object detection/recognition environment, the computer system configurable to cause execution of operations comprising:
receiving first video data, the first video data comprising a plurality of frames of a first video;
identifying a first object within the plurality of frames;
determining that the plurality of frames depict the first object in a number of different positions greater than a threshold number of positions;
creating training data by modifying the first video data to highlight the first object within the plurality of frames, the modifying comprising:
identifying the first object within a first plurality of frames:
determining non-first object pixels within the first plurality of frames;
creating first modified video data by replacing, independent of the pixels of the first object, the non-first object pixels with first background pixels in the first plurality of frames; and
creating second modified video data by replacing, independent of the pixels of the first object, the first background pixels with second background pixels in the first plurality, of frames; and
providing the training data to a machine learning program, wherein the training data is configured to train the machine learning program to identify the first object within an image.
18. The computing system of claim 17 , wherein the creating the training data further comprises creating a plurality of sets of second modified video data, each set of the second modified video data based on the first modified video data, wherein each of the sets of the second modified video data highlights the first object within each of the first plurality of frames, and wherein each of the sets of the second modified video data is modified to reflect different product placement and lighting conditions.
19. The computing system of claim 17 , wherein the modifying further comprises:
creating third modified video data by constructing a bounding box around the first object, wherein the bounding box is constructed around the first object in each of the plurality of frames.
20. The computing system of claim 17 , wherein the operations further comprise:
annotating the video data to further highlight the first object within each of the first plurality of frames, wherein the annotating the video data comprises providing annotations in a secondary file.