IP Library Granted Patent US 11,003,928
Granted Patent B2
US 11,003,928 · App. 16/535,335 · Granted May 11, 2021

Using captured video data to identify active turn signals on a vehicle

Inventors: Rotem Littman (Hod Hasharon, IL); Gilad Saban (Rehovot, IL); Noam Presman (Ramat Gan, IL); Dana Berman (Tel Aviv, IL); Asaf Kagan (Herzliya, IL)
Assignee: Argo AI, LLC
G06K9/00825G05D1/0088G05D1/0231G06K9/00718G06K9/6212G06K9/6267G06T7/20G06T7/70G06T7/90G06T7/97G06T11/20G05D2201/0213G06T2207/10016G06T2207/20084G06T2207/30244G06T2207/30252G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,003,928
App. No.
16/535,335
Granted
May 11, 2021
Kind
B2
Abstract

A system uses video of a vehicle or other object to detect and classify an active turn sign on the object. The system generates an image stack by scaling and shifting a set of digital image frames from the video to a fixed scale, yielding a sequence of images over a time period. The system processes the image stack with a classifier to determine a pose of the object, as well as the state and class of each visible turn signals on the object. When the system determines that a turn signal is active, the system will predict an action that the object will take based on the class of that signal.

Claims (48)

1. A computer-implemented method of detecting and classifying a turn signal on an object captured in a video sequence, the method comprising:

receiving a video sequence comprising a plurality of digital image frames that contain an image of an object;

generating an image stack by scaling and shifting a set of the digital image frames to a fixed scale and yielding a sequence of images of the object over a time period;

processing the image stack with a classifier to determine a state and a class of each turn signal that appears on the object in the video sequence, wherein the state is one of a group of candidate states that comprise active and inactive; and

when the classifying determines that the state of one of the turn signals is active, identifying the class of the active turn signal as one of a group of candidate classes that comprise left turn signal and right turn signal.

2. The method of claim 1 , wherein:

processing the image stack with the classifier also determines a pose of the object; and

determining the class of each turn signal comprises determining the class of each turn signal based on the pose.

3. The method of claim 1 further comprising, before generating the image stack, cropping each of the digital image frames in the set to eliminate information outside of the bounding boxes of each frame.

4. The method of claim 1 , wherein:

the classifier comprises a convolutional neural network (CNN); and

the method further comprises, before processing the image stack with the classifier, training the CNN on a plurality of training image stack sets that include, for each training image stack, labels indicative of turn signal state and turn signal class.

5. The method of claim 1 , wherein processing the image stack with the classifier further comprises determining a class of the object, wherein candidate classes include vehicle and bicyclist.

6. The method of claim 1 further comprising predicting a direction of movement of the object based on the state and class of the turn signal.

7. The method of claim 6 , wherein:

receiving the video sequence is performed by a camera of an autonomous vehicle (AV);

determining the state of the turn signal, determining the class of the turn signal and predicting the action of the object are performed by an on-board processor of the AV; and

the method further comprises, by the on-board processor, causing the AV to take an action responsive to the predicted direction of movement of the object.

8. The method of claim 1 further comprising, before generating the image stack:

processing the digital image frames to detect the object in the digital image frames by adding bounding boxes to the digital image frames; and

performing registration on the set of digital image frames to cause the bounding boxes of each frame in the set to share in a common location and scale within each digital image frame.

9. The method of claim 8 , wherein processing the digital image frames to detect the object in the digital image frames comprises applying Mask R-CNN to the digital image frames.

10. The method of claim 8 further comprising, before performing registration, tracking the object across the digital image frames to eliminate frames that are less likely to contain the object, yielding the set on which registration will be performed.

11. The method of claim 10 , wherein tracking the object across the digital image frames comprises:

performing Intersection over Union matching between a plurality of pairs of the digital image frames; or

performing color histogram matching between a plurality of pairs of the digital image frames.

12. A vehicle having an on-board system for detecting and classifying turn signals on other objects observed in a video sequence, the vehicle comprising:

a video camera;

an on-board processor; and

a computer-readable memory containing programming instructions that are configured to cause the on-board processor to:

receive a video sequence comprising a plurality of digital image frames that contain an image of an object,

generate an image stack by scaling and shifting a set of the digital image frames to a fixed scale and yielding a sequence of images of the object over a time period,

process the image stack with a classifier to determine a state and a class of each turn signal that appears on the object in the video sequence, wherein the state is one of a group of candidate states that comprise active and inactive, and

when the classifying determines that the state of one of the turn signals is active, identifying the class of the active turn signal as one of a group of candidate classes that comprise left turn signal and right turn signal.

13. The system of claim 12 , wherein:

the instructions to process the image stack with the classifier also comprise instructions to determine a pose of the object; and

the instructions to determine the class of each turn signal comprise instructions to determine the class of each turn signal based on the pose.

14. The system of claim 12 further comprising instructions configured to cause the on-board processor to, before generating the image stack, crop each of the digital image frames in the set to eliminate information outside of the bounding boxes of each frame.

15. The system of claim 12 , further comprising programming instructions configured to cause the on-board processor to predict a direction of movement of the object based on the state and class of the turn signal.

16. The system of claim 15 , further comprising programming instructions configured to cause the on-board processor to instruct a vehicle operational system to take an action in response to the predicted direction of movement.

17. The system of claim 12 further comprising additional programming instructions configured to cause the on-board processor to, before generating the image stack:

process the digital image frames to detect the object in the digital image frames by adding bounding boxes to the digital image frames; and

perform registration on the set of digital image frames to cause the bounding boxes of each frame in the set to share in a common location and scale within each digital image frame.

18. The system of claim 17 , wherein the instructions to process the digital image frames to detect the object in the digital image frames comprise instructions to apply Mask R-CNN to the digital image frames.

19. The system of claim 17 further comprising instructions configured to cause the on-board processor to, before performing registration, track the object across the digital image frames to eliminate frames that are less likely to contain the object, yielding the set on which registration will be performed.

20. The system of claim 19 , wherein the instructions to track the object across the digital image frames comprise instructions to:

perform Intersection over Union matching between a plurality of pairs of the digital image frames; or

perform color histogram matching between a plurality of pairs of the digital image frames.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2024
From: ARGO AI, LLC
To: VOLKSWAGEN GROUP OF AMERICA INVESTMENTS, LLC
Reel/Frame 069177/0099 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2019
From: LITTMAN, ROTEM; SABAN, GILAD; PRESMAN, NOAM; BERMAN, DANA; KAGAN, ASAF
To: ARGO AI, LLC
Reel/Frame 050322/0462 →
Continuity (1)
Related Publication 20210042542A1 · Feb 11, 2021
Cited By (1)
US 12,608,955