IP Library › Granted Patent US 11,410,549
Granted Patent B2
US 11,410,549 · App. 16/609,729 · Granted Aug 9, 2022

Method, device, readable medium and electronic device for identifying traffic light signal

Inventor: Jinglin Yang (Beijing, CN)
Assignee: BOE Technology Group Co., Ltd.
G08G1/095G06K9/6232G06K9/6288G06N3/0454G06N3/08G06T3/4007G06T3/4046G06T7/11G06V20/582G08G1/0116G08G1/0145
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,410,549
App. No.
16/609,729
Granted
Aug 9, 2022
Kind
B2
Abstract

The present disclosure relates to a method, device, computer readable media, and electronic devices for identifying a traffic light signal from an image. The method for identifying a traffic light signal from an image includes extracting, based on a deep neural network, multiple layers of first feature maps corresponding to different layers of the deep neural network from the image. The method includes selecting at least two layers of the first feature maps having different scales from the multiple layers of the first feature maps. The method includes inputting the at least two layers of the first feature maps to a convolution layer having a convolution kernel matching a shape of a traffic light to obtain a second feature map. The method includes obtaining a detection result of the traffic light signal based on the second feature map.

Claims (29)

1. A method for identifying an object from an image, comprising:

extracting, based on a deep neural network, multiple layers of first feature maps corresponding to different layers of the deep neural network from the image;

selecting at least two layers of the first feature maps having different scales from the multiple layers of the first feature maps; inputting the at least two layers of the first feature maps to a convolution layer having a convolution kernel matching a shape of the object to obtain a second feature map; and

obtaining a detection result of the object based on the second feature map, wherein the object is a traffic light, and wherein the deep neural network is a VGG network, and selecting at least two layers of the first feature maps having different scales from the multiple layers of the first feature maps comprises: selecting feature maps output from seventh, tenth, and thirteenth layers of the VGG network for improving accuracy of matching the shape of the traffic light.

2. The method according to claim 1 , wherein obtaining a detection result of the object based on the second feature map comprises:

performing a pooling operation on the second feature map to obtain a fusion feature with uniform size; and

obtaining the detection result of a traffic light signal from the fusion feature through a fully connected layer and an activation function layer.

3. The method according to claim 1 , further comprising:

filtering the detection result with a positional confidence,

wherein the positional confidence indicates a probability that the traffic light appears in a lateral or longitudinal position in the image.

4. The method according to claim 1 , further comprising:

extracting a first recommended region from an R channel of the image and a second recommended region from a G channel of the image,

wherein the obtaining a detection result of the object based on the second feature map comprises:

fusing the second feature map, the first recommended region and the second recommended region to obtain a fusion feature; and

obtaining the detection result of a traffic light signal based on the fusion feature.

5. The method according to claim 4 , wherein fusing the second feature map, the first recommended region, and the second recommended region to obtain a fusion feature comprises:

performing a pooling operation on the second feature map to obtain a third feature map with a uniform size;

mapping the first recommended region and the second recommended region to the third feature map with the uniform size to obtain a first mapping region and a second mapping region; and

performing the pooling operation on the first mapping region and the second mapping region to obtain the fusion feature with the uniform size.

6. The method according to claim 4 , wherein obtaining the detection result of the object based on the fusion feature comprises:

obtaining the detection result of the traffic light signal from the fusion feature through a fully connected layer and an activation function layer.

7. The method according to claim 4 , wherein after extracting a first recommended region from an R channel of the image and extracting a second recommended region from a G channel of the image, the method further comprises:

determining positional confidences of the first recommended region and the second recommended region; and

excluding the first recommended region and the second recommended region having the positional confidence less than a predetermined threshold.

8. The method according to claim 4 , wherein extracting a first recommended region from an R channel of the image and a second recommended region from a G channel of the image comprises:

extracting, by using an RPN network, the first recommended region and the second recommended region from any layer of the first feature map of the multiple layers of the first feature maps of the R channel and the G channel of the image respectively.

9. The method according to claim 4 , wherein extracting a first recommended region from an R channel of the image and a second recommended region from a G channel of the image comprises:

obtaining the first recommended region and the second recommended region from the R channel and the G channel respectively, via selective search or morphological filtering.

10. A non-transitory computer readable medium, having a computer program stored thereon, wherein the computer program comprises executable instructions, and when the executable instructions are executed by a processor, a method comprising: extracting, based on a deep neural network, multiple layers of first feature maps corresponding to different layers of the deep neural network from the image; selecting at least two layers of the first feature maps having different scales from the multiple layers of the first feature maps; inputting the at least two layers of the first feature maps to a convolution layer having a convolution kernel matching a shape of a traffic light to obtain a second feature map; and obtaining a detection result of the traffic light signal based on the second feature map, wherein the object is a traffic light, and wherein the deep neural network is a VGG network, and selecting at least two layers of the first feature maps having different scales from the multiple layers of the first feature maps comprises: selecting feature maps output from seventh, tenth, and thirteenth layers of the VGG network for improving accuracy of matching the shape of the traffic light.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2019
From: YANG, JINGLIN
To: BOE TECHNOLOGY GROUP CO., LTD.
Reel/Frame 050896/0130 →
Priority Claims (1)
CN 201810554025.4 · May 31, 2018 · national
Continuity (1)
Related Publication 20210158699A1 · May 27, 2021
Cited By (1)
US 12,307,789