Interactive system and method for improved collection and annotation of machine learning datasets
View Patent ↗A method for improved collection and annotation of training datasets for training machine learning models is described. An image is captured using an image capture device and established as a reference image frame. The object of interest in the reference frame is annotated using a geometrical shape and established as seed data along with a pre-defined quantity target. The dimension of the annotated object is scaled to establish an annotation guideline. The annotation guideline is shifted in position in subsequent live images and displayed along with a live image of the object of interest. The user is prompted to adjust the live image so that the object of interest is fully encompassed by annotation guideline and capture a second live image. The captured image is classified in real time as accepted based on a benchmarking threshold and stored. The process is repeated until the pre-defined quantity target is satisfied.
1 . A method for curation of training data for training machine learning models comprising:
capturing, using an image capture device, a reference image frame comprising an object of interest;
obtaining a reference datapoint for a label or a set of labels from the reference image frame; wherein the label or set of labels is associated with the object of interest;
annotating the reference datapoint to ascertain an expected annotation for each label or the set of labels; wherein the expected annotation comprises a geometrical shape;
establishing the expected annotations and a pre-defined quantity target as seed data for each label or the set of labels;
scaling the geometrical shape of the expected annotation based on a predetermined scaling factor;
obtaining a position of the scaled geometric shape as an annotation guideline;
shifting the position of the annotation guideline by varying the position in a two-dimensional coordinate system and displaying the shifted annotation guideline and a live image on the screen of the image capture device; wherein the live image comprises an image including the object of interest;
prompting a user to adjust the live image frame such that the object of interest is tightly encompassed inside the scaled and shifted annotation guideline displayed on the live image screen;
capturing a second live image frame and comparing the objects of interest in the captured second live image frame with the objects of interest in the reference image frame to determine a similarity score;
classifying, in real-time, the captured second live image frame as accepted, if the similarity score is above a pre-defined threshold;
storing the accepted image frames as training dataset for machine learning models; and
repeating, for newly captured live image frame, the steps of shifting, prompting, capturing, classifying and storing until the pre-defined quantity target criteria is satisfied, based on a comparison between the newly captured live image frame and the reference image frame.
2 . The method of claim 1 , wherein the geometrical shape comprises a rectangular bounding box, a square bounding box or a higher dimension polygonal bounding box.
3 . The method of claim 1 , wherein the reference image frame comprises one or more objects.
4 . The method of claim 3 , wherein each of the one or more objects is associated with a label from the set of labels.
5 . The method of claim 1 , wherein the pre-defined quantity target is a threshold number of captured image frames classified as accepted for each of the labels or the set of labels.
6 . The method of claim 1 , wherein predetermined scaling factor is a value by which the size of the geometrical shape of the expected annotation is scaled-up and/or scaled-down to obtain the size of the annotation guideline.
7 . The method of claim 1 , wherein shifting the position of annotation guideline comprises changing the two-dimensional co-ordinates of the annotation guideline using a shifting formula.
8 . The method of claim 1 , wherein prompting the user comprises providing a visual or textual indication.
9 . The method of claim 8 , wherein prompting the user further comprises checking for motion or blur, improper framing of the live image screen.
10 . The method of claim 1 , wherein similarity score is determined based on one or more of a blur detection, a noise detection, an improper framing detection and environmental conditions.
11 . The method of claim 10 , wherein classifying in real-time the captured images comprises, comparing the similarity score of the captured images to one or more of a benchmark criteria.
12 . The method of claim 11 , wherein the benchmark criteria is associated with a label or set of labels.
13 . The method of claim 1 , wherein the training dataset comprises audio dataset, image dataset, video dataset and/or textual dataset.
14 . The method of claim 13 , further comprising: obtaining a reference textual dataset pertaining to a textual object of interest and associating with a textual pre-defined quantity target; wherein the reference textual dataset comprises questions/phrases related to an object of interest; and wherein the object of interest comprises masking words representative of the reference textual dataset;
guiding a user to associate variations of the questions/phrases with the reference textual dataset and determining a textual similarity score; wherein the variations of the questions/phrases comprises grammatical variations, active/passive voice variations and/or contextual variations;
classifying in real time, the variations of the questions/phrases as accepted, if the similarity score is below a second pre-defined threshold;
storing the accepted variations of the questions/phrases as training dataset for training a machine learning model; and
repeating the steps of guiding, classifying and storing until the textual pre-defined quantity target criteria is satisfied.
15 . An image curation system for curation of training dataset for training machine learning models comprises:
an image capture device configured to capture images of a scene;
a hardware device comprising one or more storage devices and one or more processing units;
wherein the storage device comprises instructions stored therein that when executed by the one or more processing unit causes the image curation system to:
obtain a reference image frame comprising an object of interest captured using the image capture device;
obtain a reference datapoint for a label or a set of labels from the reference image; wherein the label or set of labels is associated with the object of interest;
annotate the reference datapoint to ascertain an expected annotation for each label or a set of labels; wherein the expected annotation comprises a geometrical shape;
establish the expected annotations and a pre-defined quantity target as seed data for each label or the set of labels;
scale the geometrical shape of the expected annotation based on a predetermined scaling factor;
obtain a position of the scaled geometric shape as an annotation guideline;
shift the position of the annotation guideline by varying the position in a two-dimensional coordinate system and displaying the shifted annotation guideline and a live image on the screen of the image capture device; wherein the live image comprises an image including the object of interest; prompt a user to adjust the live image frame such that the object of interest is tightly encompassed inside the scaled and shifted annotation guideline displayed on the live image screen;
capture a second live image frame and compare the captured second live image frame with the reference image frame to determine a similarity score;
classify, in real time, the captured second live image frame as accepted, if the similarity score is above a pre-defined threshold; store the accepted image frames in the one or more storage device as training dataset for machine learning models;
repeat, for newly captured live image frame, the shift, prompt, capture, classify and store functions until a pre-defined quantity target is satisfied, based on a comparison between the newly captured live image frame and the reference image frame.
16 . The system of claim 15 , wherein comparing the captured second live image frame further comprises comparing the object of interest in the reference frame and captured second live image frame.
17 . The system of claim 15 , wherein the training dataset further comprises audio dataset, image dataset, video dataset and/or textual dataset.
18 . The system of claim 15 , wherein the storage device comprises further instructions causing the image curation system to:
obtain a reference textual dataset pertaining to a textual object of interest and associating with a textual pre-defined quantity target; wherein the reference textual dataset comprises questions/phrases pertaining to an object of interest;
guide a user to associate variations of the questions/phrases with the reference textual dataset and determining a textual similarity score; wherein the variations of the questions/phrases comprises grammatical variations, active/passive voice variations and/or contextual variations;
classify in real time, the variations of the questions/phrases as accepted, if the textual similarity score is below a second pre-defined threshold;
store the accepted variations of the questions/phrases as training dataset for training a machine learning model; and repeat the guide, classify, and store functions until the textual pre-defined quantity target criteria is satisfied.