IP Library › Granted Patent US 10,296,815
Granted Patent B2
US 10,296,815 · App. 15/749,693 · Granted May 21, 2019

Cascaded convolutional neural network

Inventors: Lior Wolf (Herzliya, IL); Assaf Mushinsky (Tel Aviv, IL)
Assignee: Ramot at Tel Aviv University Ltd.
G06K9/6274G06K9/00228G06K9/46G06K9/4628G06K9/6202G06K9/6268G06K9/6281G06K9/66G06N3/0454G06N5/046G06T7/70G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,296,815
App. No.
15/749,693
Granted
May 21, 2019
Kind
B2
Abstract

A convolutional neural network system for detecting at least one object in at least one image. The system includes a plurality of object detectors, corresponding to a predetermined image window size in the at least one image. Each object detector is associated with a respective down-sampling ratio with respect to the at least one image. Each object detector includes a respective convolutional neural network and an object classifier coupled with the convolutional neural network. The respective convolutional neural network includes a plurality of convolution layers. The object classifier classifies objects in the image according to the results from the convolutional neural network. Object detectors associated with the same respective down-sampling ratio define at least one group of object detectors. Object detectors in a group of object detectors being associated with common convolution layers.

Claims (47)

1. A convolutional neural network system for detecting at least one object in at least one image, the system comprising:

a plurality of object detectors, each object detector corresponding to a predetermined image window size in said at least one image, each object detector being associated with a respective down sampling ratio with respect to said at least one image, each object detector including:

a respective convolutional neural network, said convolutional neural network including a plurality of convolution layers; and

an object classifier, coupled with said convolutional neural network, for classifying objects in said image according to the results from said convolutional neural network,

wherein, object detectors associated with a same respective down sampling ratio define at least one group of object detectors, object detectors in a group of object detectors being associated with common convolution layers,

wherein said object classifier is a convolution classifier, convolving at least one classification filter with a features map provided by said respective convolutional neural network

wherein said respective convolutional neural network producing the features map including a plurality of features, each entry represents the features intensities within an image window associated with said entry, said image window exhibiting said respective image window size

wherein said object classifier provides a classification vector which includes a probability that said object is located at each of the image windows associated with said features

wherein said classification vector further includes image window correction factors for each image window associated with said features map, said image window correction factors include corrections to a width and a height of each image window and corrections to a location of each image window and correction to an orientation of each image window.

2. A convolutional neural network method comprising the procedures of:

down sampling an image according to a plurality of down sampling ratios, to produce a plurality of down sampled images, each down sampled image being associated with a respective down sampling ratio;

for each down sampled image, detecting by a corresponding convolutional neural network, objects at a predetermined image window size with respect to at least one image; and

classifying objects in said image,

wherein convolutional neural networks, detecting objects in respective down sampled images associated with a same respective down sampling ratio, define at least one group of convolutional neural networks, convolutional neural networks in a group of convolutional neural networks being associated with common convolution layers,

further including, prior to said procedure of down-sampling said image, the procedures of:

producing augmented training samples from an initial training set; and

training the convolutional neural networks to have the common layers.

3. The convolutional neural network method according to claim 2 , wherein training the convolutional neural networks to have common layers includes averaging weights and parameters, of all the groups of layers with identical characteristics of the object detectors.

4. The convolutional neural network method according to claim 2 , wherein, training the convolutional neural networks to have common layers includes training a single training scale detector by employing said augmented training samples and deploying duplicates of said training scale detector, each duplicate being associated with a respective scaled version of said at least one image, the duplicates of the training scale detector defining a convolutional neural network system.

5. The convolutional neural network method according to claim 2 , wherein said procedure of producing augmented training samples include the sub procedures of:

determining a location of each object key point within a respective training sample bounding box;

for object key point type, determining a respective key point reference location according to the average location of the object key points of a same type, the average being determined according to the object key point locations of all objects in the initial training set;

registering all the training samples in the initial training set with the features reference locations; and

randomly perturbing each of the aligned training samples from this reference location.

6. The convolutional neural network system according to claim 1 , further including a plurality of down samplers each associated with a respective down sampling ratio, said down samplers being configured to produce said scaled versions of said image, each scaled version being associated with a respective down sampling ratio.

7. The convolutional neural network system according to claim 6 , wherein down samplers, and object detectors associated with a same respective image window size with respect to said image, define a scale detector, each scale detector being associated with a respective scaled version of said image.

8. The convolutional neural network system according to claim 7 , wherein, a single training scale detector is trained when the scale detectors exhibit a same configuration of object detectors, and when the convolutional neural networks in the object detectors exhibit groups layers with identical characteristics.

9. The convolutional neural network system according to claim 8 , wherein prior to training said training scale detector a number of training samples in a training set is increased beyond the initial number of training samples by:

determining a location of each object key point within a respective training sample bounding box;

for object key point type, determining a respective feature reference location according to an average location of the object key points of a same type, the average being determined according to the object key point locations of all objects in an initial training set;

registering all the training samples in the initial training set with the features reference locations; and randomly perturbing each of the aligned training samples from this reference location.

10. A convolutional neural network system for detecting at least one object in at least one image, the system comprising:

a plurality of object detectors, each object detector corresponding to a predetermined image window size in said at least one image, each object detector being associated with a respective down sampling ratio with respect to said at least one image, each object detector including:

a respective convolutional neural network, said convolutional neural network including a plurality of convolution layers; and

an object classifier, coupled with said convolutional neural network, for classifying objects in said image according to the results from said convolutional neural network;

a plurality of down samplers each associated with a respective down sampling ratio, said down samplers being configured to produce said scaled versions of said image, each scaled version being associated with a respective down sampling ratio,

wherein, object detectors associated with a same respective down sampling ratio define at least one group of object detectors, object detectors in a group of object detectors being associated with common convolution layers,

wherein down samplers, and object detectors associated with a same respective image window size with respect to said image, define a scale detector, each scale detector being associated with a respective scaled version of said image, and

wherein, a single training scale detector is trained when the scale detectors exhibit a same configuration of object detectors, and when the convolutional neural networks in the object detectors exhibit groups layers with identical characteristics.

11. The convolutional neural network system according to claim 10 , wherein said object classifier is a convolution classifier, convolving at least one classification filter with a features map provided by said respective convolutional neural network.

12. The convolutional neural network system according to claim 11 , wherein said respective convolutional neural network producing a features map including a plurality of features, each entry represents the features intensities within an image window associated with said entry, said image window exhibiting said respective image window size.

13. The convolutional neural network system according to claim 12 , where said object classifier provides a probability that said object is located at each of the image windows associated said features.

14. The convolutional neural network system according to claim 13 , wherein said classification vector further including image window correction factors for each image window associated with said features map, said image window correction factors include corrections to a width and a height of each image window and corrections to a location of each image window and correction to an orientation of each image window.

15. The convolutional neural network system according to claim 14 , wherein prior to training said training scale detector a number of training samples in a training set is increased beyond an initial number of training samples by:

determining a location of each object key point within a respective training sample bounding box;

for object key point type, determining a respective feature reference location according to an average location of the object key points of a same type, the average being determined according to the object key point locations of all objects in the initial training set;

registering all the training samples in the initial training set with the features reference locations; and randomly perturbing each of the aligned training samples from this reference location.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2018
From: WOLF, LIOR; MUSHINSKY, ASSAF
To: RAMOT AT TEL AVIV UNIVERSITY LTD.
Reel/Frame 044803/0274 →
Priority Claims (1)
GB 1614009.7 · Aug 16, 2016 · national
Continuity (5)
Provisional Application 62486997 · Apr 19, 2017
Provisional Application 62325562 · Apr 21, 2016
Provisional Application 62325551 · Apr 21, 2016
Provisional Application 62325553 · Apr 21, 2016
Related Publication 20190042892A1 · Feb 7, 2019