IP Library Granted Patent US 10,380,741
Granted Patent B2
US 10,380,741 · App. 15/478,947 · Granted Aug 13, 2019

System and method for a deep learning machine for object detection

Inventors: Arvind Yedla (La Jolla, CA); Marcel Nassar (San Diego, CA); Mostafa El-Khamy (San Diego, CA); Jungwon Lee (San Diego, CA)
Assignee: Samsung Electronics Co., Ltd
G06T7/11G06K9/00369G06K9/3241G06K9/4628G06K9/627G06N3/0454G06T7/194
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,380,741
App. No.
15/478,947
Filed
Apr 4, 2017
Granted
Aug 13, 2019
Kind
B2
Art Unit
2667
USPC
382/103
Abstract

Apparatuses and methods of manufacturing same, systems, and methods for object detection using a region-based deep learning model are described. In one aspect, a method is provided, in which a region proposal network (RPN) is used to identify regions of interest (RoI) in an image by assigning a confidence levels, the assigned confidence levels of the RoIs are used to boost the background score assigned by the downstream classifier to each RoI, and the background scores are used in a softmax function to calculate the final class probabilities for each object class.

Claims (89)

1. A method of object detection in an image using a region-based deep learning model, the method comprising:

identifying, using a region proposal network (RPN), regions of interest (RoI) in the image and assigning a confidence levels to each identified RoI;

boosting a background score assigned by a downstream classifier to each RoI, using the confidence level assigned to the ROI and optical flow magnitude;

using the boosted background scores in a softmax function to calculate final class probabilities; and

identifying each RoI as including an object in a foreground of the image or as a part of the background of the image, based on the final class probabilities.

2. The method of claim 1 , wherein the object includes a pedestrian.

3. The method of claim 1 , wherein the region-based deep learning model is a faster region-based convolutional neural network (R-CNN).

4. The method of claim 1 , wherein the region-based deep learning model is a region-based fully convolutional network (R-FCN).

5. The method of claim 1 , wherein the confidence levels comprise P B , which is a probability of the RoI being the background, and P F , which is a probability of the RoI including the object in the foreground.

6. The method of claim 5 , wherein s 0 is the background score assigned by the downstream classifier to a RoI boosted according to the formula:

s

0

=

{

s

0

if

P

B

<

P

F

P

B

·

s

0

P

F

otherwise

.

7. The method of claim 1 , wherein using the assigned confidence levels of the RoIs to boost the background score assigned by the downstream classifier to each RoI comprises:

iteratively refining the boosted background scores.

8. The method of claim 1 , wherein semantic segmentation masks are also used to boost the background score assigned by the downstream classifier to each RoI.

9. An apparatus capable of object detection using a region-based deep learning model, comprising:

one or more non-transitory computer-readable media; and

at least one processor which, when executing instructions stored on one or more non-transitory computer readable media, performs the steps of:

identifying, using a region proposal network (RPN), regions of interest (RoI) in the image and assigning a confidence levels to each identified RoI;

boosting a background score assigned by a downstream classifier to each RoI, using the confidence level assigned to the ROI and optical flow magnitude;

using the boosted background scores in a softmax function to calculate final class probabilities; and

identifying each RoI as including an object in a foreground of the image or as a part of the background of the image, based on the final class probabilities.

10. The apparatus of claim 9 , where the object includes a pedestrian.

11. The apparatus of claim 9 , wherein the region-based deep learning model is a faster region-based convolutional neural network (R-CNN).

12. The apparatus of claim 9 , wherein the region-based deep learning model is a region-based fully convolutional network (R-FCN).

13. The apparatus of claim 9 , wherein the confidence levels comprise P B , which is a probability of the RoI being the background, and P F , which is a probability of the RoI including the object in the foreground.

14. The apparatus of claim 13 , wherein so is the background score assigned by the downstream classifier to a RoI boosted according to the formula:

s

0

=

{

s

0

if

P

B

<

P

F

P

B

·

s

0

P

F

otherwise

.

15. The apparatus of claim 9 , wherein using the assigned confidence levels of the RoIs to boost the background score assigned by the downstream classifier to each RoI comprises:

iteratively refining the boosted background scores.

16. The apparatus of claim 9 , wherein semantic segmentation masks are also used to boost the background score assigned by the downstream classifier to each RoI.

17. A method, comprising:

manufacturing a chipset comprising:

at least one processor which, when executing instructions stored on one or more non-transitory computer readable media, performs the steps of:

identifying, using a region proposal network (RPN), regions of interest (RoI) in the image and assigning a confidence levels to each identified RoI;

boosting a background score assigned by a downstream classifier to each RoI, using the confidence level assigned to the ROI and optical flow magnitude;

using the boosted background scores in a softmax function to calculate final class probabilities; and

identifying each RoI as including an object in a foreground of the image or as a part of the background of the image, based on the final class probabilities, and

the one or more non-transitory computer-readable media which store the instructions.

18. A method of testing an apparatus, comprising:

testing whether the apparatus has at least one processor which, when executing instructions stored on one or more non-transitory computer readable media, performs the steps of:

identifying, using a region proposal network (RPN), regions of interest (RoI) in the image and assigning a confidence levels to each identified RoI;

boosting a background score assigned by a downstream classifier to each RoI, using the confidence level assigned to the ROI and optical flow magnitude;

using the boosted background scores in a softmax function to calculate final class probabilities; and

identifying each RoI as including an object in a foreground of the image or as a part of the background of the image, based on the final class probabilities, and

testing whether the apparatus has the one or more non-transitory computer-readable media which store the instructions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2019
From: YEDLA, ARVIND; NASSAR, MARCEL; EL-KHAMY, MOSTAFA; LEE, JUNGWON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 049642/0340 →
Continuity (2)
Provisional Application 62431086 · Dec 7, 2016
Related Publication 20180158189A1 · Jun 7, 2018
Cited By (16)
US 12,198,396 US 12,216,610 US 12,223,428 US 12,236,689 US 12,307,350 US 12,346,816 US 12,367,405 US 12,455,739 US 12,462,575 US 12,522,243 US 12,536,131 US 12,554,467 US 12,591,240 US 12,618,976 US 12,623,691 US 12,709,294