IP Library Patent Application 16441170
Patent Application
App. No. 16/441,170

SELECTING IMAGES FOR MANUAL ANNOTATION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/441,170
Abstract

Systems and methods for selecting images for manual annotation are provided. For example, a first image corresponding to a first point in time, a second image corresponding to a second point in time and at least two more images, each corresponds to a point in time, may be accessed. The first image and the second image may be presented. First annotation data associated with the first image and second annotation data associated with the second image may be received and compared. A third image of the at least two more images may be selected based on a result of the comparison, the first point in time, the second point in time, and the points in time corresponding to at least some of the at least two more images. The selected third image may be presented, and third annotation data associated with the third image may be received.

Claims (96)

1 . A method for selecting images for manual annotation, the method comprising:

accessing a group of images, the group of images comprises a first image corresponding to a first point in time, a second image corresponding to a second point in time and at least two more images, each image of the at least two more images corresponds to a point in time;

presenting to a user at least part of the first image;

receiving from the user first annotation data associated with the first image;

presenting to the user at least part of the second image;

receiving from the user second annotation data associated with the second image;

comparing the first annotation data and the second annotation data;

based on a result of the comparison of the first annotation data and the second annotation data, the first point in time, the second point in time, and the points in time corresponding to at least some of the at least two more images, selecting a third image of the at least two more images;

presenting to the user at least part of the selected third image; and

receiving from the user third annotation data associated with the third image.

2 . The method of claim 1 , further comprising:

in response to a first result of the comparison of the first annotation data and the second annotation data, selecting a first quantity;

in response to a second result of the comparison of the first annotation data and the second annotation data, selecting a second quantity, the second quantity differs from the first quantity;

selecting a set of images of the at least two more images, such that a quantity of images in the selected set of images matches the selected quantity;

presenting to the user the selected set of images; and

receiving from the user annotation data associated with each of the selected set of images.

3 . The method of claim 1 , further comprising:

determining an embedding in a mathematical space of at least the first image, the second image and the at least two more images; and

using the first annotation data, the second annotation data and the determined embedding of the images in the mathematical space to select the third image of the at least two more images.

4 . The method of claim 3 , wherein the determined embedding in the mathematical space is based on pixel data of the first image, pixel data of the second image, pixel data of at least some of the at least two more images, the first point in time, the second point in time, and the points in time corresponding to at least some of the at least two more images.

5 . The method of claim 1 , wherein the first image is a first frame of a video, the second image is a second frame of the video, and the at least two more images are at least two more frames of the video.

6 . The method of claim 5 , further comprising:

in response to a first result of the comparison of the first annotation data and the second annotation data, selecting a first frame rate;

in response to a second result of the comparison of the first annotation data and the second annotation data, selecting a second frame rate, the second frame rate differs from the first frame rate;

sampling frames of at least part of the video at the selected frame rate;

presenting the sampled frames to the user; and

receiving from the user annotation data associated with the sampled frames.

7 . The method of claim 1 , further comprising:

in response to a first result of the comparison of the first annotation data and the second annotation data, selecting the third image of the at least two more images to correspond to a point in time between the first point in time and the second point in time; and

in response to a second result of the comparison of the first annotation data and the second annotation data, selecting the third image of the at least two more images to correspond to a third point in time such that the second point in time is between the first point in time and the third point in time.

8 . The method of claim 1 , further comprising:

in response to a first result of the comparison of the first annotation data and the second annotation data, selecting the third image of the at least two more images to correspond to a point in time closer to the first point in time than to the second point in time; and

in response to a second result of the comparison of the first annotation data and the second annotation data, selecting the third image of the at least two more images to correspond to a point in time closer to the second point in time than to the first point in time.

9 . The method of claim 1 , further comprising:

in response to a first result of the comparison of the first annotation data and the second annotation data, selecting a first image resolution;

in response to a second result of the comparison of the first annotation data and the second annotation data, selecting a second image resolution, the second image resolution differs from the first image resolution;

presenting to the user at least part of the selected third image in the selected resolution; and

receiving from the user third annotation data associated with the third image in response to the presentation of the at least part of the selected third image in the selected resolution.

10 . The method of claim 1 , further comprising:

in response to a first result of the comparison of the first annotation data and the second annotation data, selecting a first image augmentation;

in response to a second result of the comparison of the first annotation data and the second annotation data, selecting a second image augmentation, the second image augmentation differs from the first image augmentation;

applying the selected image augmentation to at least part of the third image to obtain an augmented image;

presenting to the user at least part of the obtained augmented image; and

receiving from the user third annotation data associated with the third image in response to the presentation of the at least part of the obtained augmented image.

11 . The method of claim 1 , further comprising updating a dataset of images and annotations according to the selected third image and the third annotation data received from the user.

12 . The method of claim 1 , further comprising:

using the first image, the second image, the first annotation data and the second annotation data to train a machine learning algorithm and obtain an inference model;

using the obtained inference model to obtain at least one inferred result for each image of the at least two more images; and

based on a result of the comparison of the first annotation data and the second annotation data, the first point in time, the second point in time, the points in time corresponding to at least some of the at least two more images and the obtained inferred results associated with at least some of the at least two more images, selecting the third image of the at least two more images.

13 . The method of claim 1 , further comprising:

using the first image, the second image, the first annotation data and the second annotation data to train a machine learning algorithm and obtain an inference model;

generating at least two augmented images for each image of the at least two more images;

using the obtained inference model to obtain a plurality of inferred results for each image of the at least two more images, wherein the plurality of inferred results corresponding to any particular image of the at least two more images comprises at least one inferred result for each augmented image of the generated at least two augmented images corresponding to the particular image;

analyzing the plurality of inferred results of each image of the at least two more images to obtain a stability assessment for each image of the at least two more images; and

using the stability assessments associated with at least some of the at least two more images to select the third image of the at least two more images.

14 . A system for selecting images for manual annotation, the system comprising:

at least one processor configured to:

access a group of images, the group of images comprises a first image corresponding to a first point in time, a second image corresponding to a second point in time and at least two more images, each image of the at least two more images corresponds to a point in time;

present to a user at least part of the first image;

receive from the user first annotation data associated with the first image;

present to the user at least part of the second image;

receive from the user second annotation data associated with the second image;

compare the first annotation data and the second annotation data;

based on a result of the comparison of the first annotation data and the second annotation data, the first point in time, the second point in time, and the points in time corresponding to at least some of the at least two more images, select a third image of the at least two more images;

present to the user at least part of the selected third image; and

receive from the user third annotation data associated with the third image.

15 . The system of claim 14 , wherein the at least one processor is further configured to:

determine an embedding in a mathematical space of at least the first image, the second image and the at least two more images; and

use the first annotation data, the second annotation data and the determined embedding of the images in the mathematical space to select the third image of the at least two more images.

16 . The system of claim 14 , wherein the first image is a first frame of a video, the second image is a second frame of the video, and the at least two more images are at least two more frames of the video.

17 . The system of claim 14 , wherein the at least one processor is further configured to:

in response to a first result of the comparison of the first annotation data and the second annotation data, select a first image augmentation;

in response to a second result of the comparison of the first annotation data and the second annotation data, select a second image augmentation, the second image augmentation differs from the first image augmentation;

apply the selected image augmentation to at least part of the third image to obtain an augmented image;

present to the user at least part of the obtained augmented image; and

receive from the user third annotation data associated with the third image in response to the presentation of the at least part of the obtained augmented image.

18 . The system of claim 14 , wherein the at least one processor is further configured to:

use the first image, the second image, the first annotation data and the second annotation data to train a machine learning algorithm and obtain an inference model;

use the obtained inference model to obtain at least one inferred result for each image of the at least two more images; and

based on a result of the comparison of the first annotation data and the second annotation data, the first point in time, the second point in time, the points in time corresponding to at least some of the at least two more images and the obtained inferred results associated with at least some of the at least two more images, select the third image of the at least two more images.

19 . The system of claim 14 , wherein the at least one processor is further configured to:

use the first image, the second image, the first annotation data and the second annotation data to train a machine learning algorithm and obtain an inference model;

generate at least two augmented images for each image of the at least two more images;

use the obtained inference model to obtain a plurality of inferred results for each image of the at least two more images, wherein the plurality of inferred results corresponding to any particular image of the at least two more images comprises at least one inferred result for each augmented image of the generated at least two augmented images corresponding to the particular image;

analyze the plurality of inferred results of each image of the at least two more images to obtain a stability assessment for each image of the at least two more images; and

use the stability assessments associated with at least some of the at least two more images to select the third image of the at least two more images.

20 . A non-transitory computer readable medium storing data and computer implementable instructions for carrying out a method for selecting images for manual annotation, the method comprising:

accessing a group of images, the group of images comprises a first image corresponding to a first point in time, a second image corresponding to a second point in time and at least two more images, each image of the at least two more images corresponds to a point in time;

presenting to a user at least part of the first image;

receiving from the user first annotation data associated with the first image;

presenting to the user at least part of the second image;

receiving from the user second annotation data associated with the second image;

comparing the first annotation data and the second annotation data;

based on a result of the comparison of the first annotation data and the second annotation data, the first point in time, the second point in time, and the points in time corresponding to at least some of the at least two more images, selecting a third image of the at least two more images;

presenting to the user at least part of the selected third image; and

receiving from the user third annotation data associated with the third image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2019
From: GUTTMANN, MOSHE
To: ALLEGRO ARTIFICIAL INTELLIGENCE LTD
Reel/Frame 049894/0820 →