IP Library › Granted Patent US 10,503,966
Granted Patent B1
US 10,503,966 · App. 16/157,173 · Granted Dec 10, 2019

Binocular pedestrian detection system having dual-stream deep learning neural network and the methods of using the same

Inventor: Zixiao Pan (Shanghai, CN)
Assignee: TINDEI NETWORK TECHNOLOGY (SHANGHAI) CO., LTD.
G06K9/00342G06K9/6256G06K9/6262G06N3/04G06N3/08G06T7/593H04N13/239H04N13/271G06T2207/10012G06T2207/10021G06T2207/20081G06T2207/20084G06T2207/30196H04N2013/0081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,503,966
App. No.
16/157,173
Granted
Dec 10, 2019
Kind
B1
Abstract

Aspects of present disclosure relates to a binocular pedestrian detection system (BPDS). BPDS includes: a binocular camera to capture certain binocular images of pedestrians passing through a predetermined area, an image/video processing ASIC to process binocular images captured, and a binocular pedestrian detection system controller having a processor, a network interface, and a memory storing computer executable instructions. When executed by processor, computer executable instructions cause processor to perform: capturing, by binocular camera, binocular images of pedestrians, binocularly rectifying binocular images, calculating disparity maps of binocular images rectified, training a dual-stream deep learning neural network, and detecting pedestrians passing through predetermined area using dual-stream deep learning neural network trained. Dual-stream deep learning neural network includes a neural network for extracting disparity features from disparity maps of binocular images, and a neural network for learning and fusing features from rectified left images and disparity maps of binocular images.

Claims (68)

1. A binocular pedestrian detection system, comprising:

a binocular camera having a first camera, and a second camera to capture a plurality of binocular images of a plurality of pedestrians passing through a predetermined area over a predetermined period of time, wherein each binocular image comprises a left image and a right image;

an image/video processing application specific integrated circuit (ASIC) to preprocess the plurality of binocular images of the pedestrians captured by the binocular camera; and

a binocular pedestrian detection system controller, having a processor, a network interface, and a memory storing an operating system, and computer executable instructions, wherein when executed in the processor, the computer executable instructions cause the processor to perform one or more of following:

capturing, by the binocular camera, the plurality of binocular images of the pedestrians passing through the predetermined area over the predetermined period of time;

binocularly rectifying, by the image/video processing ASIC, the plurality of binocular images of the pedestrians captured to obtain a plurality of rectified binocular images, wherein each of the plurality of rectified binocular images comprises a rectified left image and a rectified right image;

calculating, by the image/video processing ASIC, disparity maps of the plurality of rectified binocular images;

training a dual-stream deep learning neural network, wherein the dual-stream deep learning neural network comprises a neural network for extracting disparity features from the disparity maps of the plurality of binocular images of the pedestrians, and a neural network for learning and fusing features from the plurality of rectified left images and disparity maps of the plurality of binocular images of the pedestrians; and

detecting, via the dual-stream deep learning neural network trained, the plurality of pedestrians passing through the predetermined area over the predetermined period of time.

2. The binocular pedestrian detection system of claim 1 , wherein the first camera comprises a first lens, and a first CMOS sensor for capturing the left image through the first lens, the second camera comprises a second lens, and a second CMOS sensor for capturing the right image through the second lens, and the left image and the right image form a binocular image.

3. The binocular pedestrian detection system of claim 1 , wherein the image/video processing ASIC is programmed to perform one or more of following operations:

performing calibration of the binocular camera;

binocularly rectifying the plurality of binocular images of the pedestrians, wherein the plurality of binocular images of the pedestrians comprises a plurality of training binocular images of the pedestrians and a plurality of real-time binocular images of the pedestrians captured by the binocular camera for pedestrian detection;

calculating the disparity maps of the plurality of training binocular images of the pedestrians during a training phase; and

calculating the disparity maps of the plurality of real-time binocular images of the pedestrians during an application phase.

4. The binocular pedestrian detection system of claim 3 , wherein the training phase comprises:

training of the neural network for extracting disparity features using the disparity maps of the plurality of binocular images of the pedestrians;

training of the neural network for learning and fusing the features of RGB and disparity using the plurality of left images of pedestrians; and

stacking the trained neural network for extracting disparity features and the neural network for learning and fusing features to form the dual-stream deep learning neural network.

5. The binocular pedestrian detection system of claim 4 , wherein the application phase comprises:

capturing, by the binocular camera, the plurality of real-time binocular images of the pedestrians;

binocularly rectifying, by the image/video processing ASIC, the plurality of real-time binocular images of the pedestrians captured;

calculating, by the image/video processing ASIC, disparity maps of the plurality of real-time binocular images of the pedestrians binocularly rectified;

detecting the plurality of the pedestrians from the plurality of rectified left images and disparity maps of the plurality of real-time binocular images of the pedestrians using the dual-stream deep learning neural network formed during the training phase; and

performing non-maximum suppression to detection results to obtain final pedestrian detection results.

6. The binocular pedestrian detection system of claim 5 , wherein the detecting the plurality of the pedestrians using the dual-stream deep learning neural network comprises:

extracting disparity features from the disparity maps of the plurality of real-time binocular images of the pedestrians using the neural network for extracting disparity features;

learning RGB features from the plurality of rectified left images of pedestrians using the first N layers of the neural network for learning and fusing the features of RGB and disparity, wherein N is a positive integer;

stacking the disparity features extracted and the RGB features learned through a plurality of channels;

fusing features of disparity and RGB using the last M−N layers of the neural network for learning and fusing features to obtain the final pedestrian detection results, wherein M is a positive integer greater than N and is a total number of layers of the neural network for learning and fusing features.

7. The binocular pedestrian detection system of claim 6 , wherein N is 7, and M is 15.

8. The binocular pedestrian detection system of claim 1 , wherein the binocular pedestrian detection system is installed over a doorway having binocular camera facing down to the plurality of pedestrians passing through the doorway.

9. The binocular pedestrian detection system of claim 1 , wherein the binocular pedestrian detection system is installed over a doorway having binocular camera facing the plurality of pedestrians passing through the doorway in a predetermined angle.

10. The binocular pedestrian detection system of claim 1 , wherein the network interface comprises a power-on-ethernet (POE) network interface, wherein the power supplying to the binocular pedestrian detection system is provided by the POE network interface, and the final pedestrian detection results are transmitted to a server that collects the final pedestrian detection results through the network interface and a communication network.

11. A method of detecting pedestrians using a binocular pedestrian detection system, comprising:

capturing, using a binocular camera of the binocular pedestrian detection system, a plurality of binocular images of a plurality of pedestrians passing through a predetermined area over a predetermined period of time, wherein each of the plurality of binocular images comprises a left image and a right image;

binocularly rectifying, via an image/video processing application specific integrated circuit (ASIC) of the binocular pedestrian detection system, the plurality of binocular images of the pedestrians captured to obtain a plurality of rectified binocular images, wherein each of the plurality of rectified binocular images comprises a rectified left image and a rectified right image;

calculating, via the image/video processing ASIC, disparity maps of the plurality of rectified binocular images;

training a dual-stream deep learning neural network using the plurality of rectified left images and disparity maps of the plurality of binocular images of the pedestrians calculated, wherein the dual-stream deep learning neural network comprises a neural network for extracting disparity features from the disparity maps of the plurality of binocular images of the pedestrians, and a neural network for learning and fusing features from the plurality of rectified left images and disparity maps of the plurality of binocular images of the pedestrians; and

detecting, via the dual-stream deep learning neural network trained, the plurality of pedestrians passing through the predetermined area over the predetermined period of time.

12. The method of claim 11 , wherein the binocular pedestrian detection system comprises:

the binocular camera having a first camera, and a second camera to capture a plurality of binocular images of the plurality of pedestrians passing through the predetermined area over a predetermined period of time, wherein each of the plurality of binocular images comprises a left image captured by the first camera, and a right image captured by the second camera;

the image/video processing ASIC to preprocess the plurality of binocular images of the pedestrians captured by the binocular camera; and

a binocular pedestrian detection system controller, having a processor, a network interface, and a memory storing an operating system, and computer executable instructions, wherein when executed in the processor, the computer executable instructions cause the processor to perform one or more operations of the method.

13. The method of claim 12 , wherein the network interface comprises a power-on-ethernet (POE) network interface, wherein the power supply to the binocular pedestrian detection system is provided by the POE network interface, and the final pedestrian detection results are transmitted to a server that collects the final pedestrian detection results through the network interface and a communication network.

14. The method of claim 11 , wherein the image/video processing ASIC is programmed to perform one or more of following operations:

performing calibration of the binocular camera;

binocularly rectifying the plurality of binocular images of the pedestrians, wherein the plurality of binocular images of the pedestrians comprises a plurality of training binocular images of the pedestrians and a plurality of real-time binocular images of the pedestrians captured by the binocular camera for pedestrian detection;

calculating the disparity maps of the plurality of training binocular images of the pedestrians during a training phase; and

calculating the disparity maps of the plurality of real-time binocular images of the pedestrians during an application phase.

15. The method of claim 14 , wherein the training phase comprises:

training of the neural network for extracting disparity features using the disparity maps of the plurality of binocular images of the pedestrians;

training of the neural network for learning and fusing the features of RGB and disparity using the plurality of left images of pedestrians; and

stacking the trained neural network for extracting disparity features and the neural network for learning and fusing features to form the dual-stream deep learning neural network.

16. The method of claim 15 , wherein the application phase comprises:

capturing, by the binocular camera, the plurality of real-time binocular images of the pedestrians;

binocularly rectifying, by the image/video processing ASIC, the plurality of real-time binocular images of the pedestrians captured;

calculating, by the image/video processing ASIC, disparity maps of the plurality of real-time binocular images of the pedestrians binocularly rectified;

detecting the plurality of the pedestrians from the plurality of rectified left images and disparity maps of the plurality of real-time binocular images of the pedestrians using the dual-stream deep learning neural network formed during the training phase; and

performing non-maximum suppression to detection results to obtain final pedestrian detection results.

17. The method of claim 16 , wherein the detecting the plurality of the pedestrians using the dual-stream deep learning neural network comprises:

extracting disparity features from the disparity maps of the plurality of real-time binocular images of the pedestrians using the neural network for extracting disparity features;

learning RGB features from the plurality of rectified left images of pedestrians using the first N layers of the neural network for learning and fusing the features of RGB and disparity, wherein N is a positive integer;

stacking the disparity features extracted and the RGB features learned through a plurality of channels;

fusing features of disparity and RGB using the last M−N layers of the neural network for learning and fusing features to obtain the final pedestrian detection results, wherein M is a positive integer greater than N and is a total number of layers of the neural network for learning and fusing features.

18. The method of claim 17 , wherein N is 7, and M is 15.

19. The method of claim 11 , wherein the binocular pedestrian detection system is installed over a doorway having binocular camera facing down to the plurality of pedestrians passing through the doorway.

20. The method of claim 11 , wherein the binocular pedestrian detection system is installed over a doorway having binocular camera facing the plurality of pedestrians passing through the doorway in a predetermined angle.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2021
From: TINDEI NETWORK TECHNOLOGY (SHANGHAI) CO., LTD.
To: TD INTELLIGENCE (GUANGZHOU) CO., LTD
Reel/Frame 055177/0348 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2018
From: PAN, ZIXIAO
To: TINDEI NETWORK TECHNOLOGY (SHANGHAI) CO., LTD.
Reel/Frame 047130/0195 →
Cited By (1)
US 12,333,693