IP Library Granted Patent US 11,113,583
Granted Patent B2
US 11,113,583 · App. 16/544,532 · Granted Sep 7, 2021

Object detection apparatus, object detection method, computer program product, and moving object

Inventor: Daisuke Kobayashi (Kunitachi Tokyo, JP)
Assignee: Kabushiki Kaisha Toshiba
G06K9/629G06K9/6232G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,113,583
App. No.
16/544,532
Granted
Sep 7, 2021
Kind
B2
Abstract

An object detection apparatus includes a calculation section, a first generation section, and a second generation section. The calculation section calculates a plurality of first feature maps from an input image. The first generation section generates a spatial attention map for which a higher first weighted value is defined for an element having a higher relation in terms of a first space on the basis of the first feature maps. The second generation section generates a plurality of second feature maps by performing weighting on each of the first feature maps in accordance with the first weighted value indicated for the spatial attention map. A detection section detects an object included in an input image by using the second feature maps.

Claims (40)

1. An object detection apparatus comprising:

a hardware processor configured to:

calculate a plurality of first feature maps from an image, at least one feature map having a different amount of an element included in an input image than at least one other feature map from the plurality of first feature maps;

generates a plurality of first combination maps by subjecting each element group of the first feature maps to a linear embedding process;

generate a spatial attention map for which a higher first weighted value is defined for a spatial attention map element comprising a higher direction relation in terms of a first space defined by a positional direction in the first feature maps and a relational direction between the first feature maps than a weighted value defined for another spatial attention map element, based at least in part on the first feature maps and the plurality of first combination maps;

generate a plurality of second feature maps by performing weighting on each of the first feature maps in accordance with a first weighted value indicated for the spatial attention map; and

detect an object included in the input image by using the second feature maps.

2. The object detection apparatus according to claim 1 , wherein:

the hardware processor calculates the first feature maps that differ in at least one of a resolution and a scale, and

the relational direction is an increase or decrease direction of the resolution or a magnification or reduction direction of the scale.

3. The object detection apparatus according to claim 1 , wherein the hardware processor generates the spatial attention map for which an inner product result of a vector sequence of the feature amounts along each of the relational direction and the positional direction in each element group of corresponding elements of the first feature maps is defined for each element as the first weighted value.

4. The object detection apparatus according to claim 1 , wherein

the plurality of first combination maps have different weight values at the linear embedding, and the hardware processor:

generates the spatial attention map for which an inner product result of a vector sequence of the feature amounts along each of the relational direction and the positional direction in each element corresponding to each other between the first combination maps is defined for each element as the first weighted value.

5. The object detection apparatus according to claim 4 , wherein the hardware processor:

generates a second combination map in which respective feature amounts of the elements included in each element group are linearly embedded for each element group of the corresponding elements of the first feature maps, the second combination map having a different weight value at the linear embedding from the weight value of the first combination map, and

generates the second feature maps by performing weighing on each of the feature amounts of each element included in the second combination map in accordance with the first weighted value indicated for the spatial attention map.

6. The object detection apparatus according to claim 5 , wherein the hardware processor generates the second feature maps in which the feature amounts of each element of a third combination map that is weighted in accordance with the first weighted value indicated for the spatial attention map and the feature amounts of each element of the first feature maps are added to the respective feature amounts of each element included in the second combination map, for each corresponding element.

7. The object detection apparatus according to claim 1 , further comprising:

a storage unit configured to store therein the second feature maps, wherein

the hardware processor is further configured to:

generate a temporal attention map for which a higher second weighted value is defined for a temporal attention map element comprising a higher relation in a temporal direction between a first group of the second feature maps generated this time and a second group of the second feature maps generated in the past on the basis of the first group and the second group,

generate third feature maps by performing weighting on each of the second feature maps included in the first group or the second group in accordance with a second weighted value indicated for the temporal attention map, and

detect the object included in the input image by using a plurality of the third feature maps generated from the second feature maps.

8. The object detection apparatus according to claim 1 , wherein the hardware processor calculates the first feature maps from the input image by using a convolutional neural network.

9. A system comprising:

the object detection apparatus according to claim 1 ; and

a controller configured to control an operation processor configured to operate the system based at least in part on information indicating a detection result of an object.

10. An object detection method performed by a computer, the method comprising:

calculating a plurality of first feature maps from an image, at least one feature map having a different amount of an element included in an input image than at least one other feature map from the plurality of first feature maps;

generating a plurality of first combination maps by subjecting each element group of the first feature maps to a linear embedding process;

generating a spatial attention map for which a higher first weighted value is defined for a spatial attention map element comprising a higher relation in terms of a first space defined by a positional direction in the first feature maps and a relational direction between the first feature maps than a weighted value defined for another spatial attention map element, based at least in part on the first feature maps and the plurality of first combination maps;

generating a plurality of second feature maps by performing weighting on each of the first feature maps in accordance with a first weighted value indicated for the spatial attention map; and

detecting an object included in the input image by using the second feature maps.

11. A computer program product having a non-transitory computer readable medium comprising instructions, wherein the instructions, when executed by a computer, cause the computer to perform:

calculating a plurality of first feature maps from an image, at least one feature map having a different amount of an element included in an input image than at least one other feature map from the plurality of first feature maps;

generating a plurality of first combination maps by subjecting each element group of the first feature maps to a linear embedding process;

generating a spatial attention map for which a higher first weighted value is defined for a spatial attention map element comprising a higher relation in terms of a first space defined by a positional direction in the first feature maps and a relational direction between the first feature maps than a weighted value defined for another spatial attention map element, based at least in part on the first feature maps and the plurality of first combination maps;

generating a plurality of second feature maps by performing weighting on each of the first feature maps in accordance with a first weighted value indicated for the spatial attention map; and

detecting an object included in the input image by using the second feature maps.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2019
From: KOBAYASHI, DAISUKE
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 050628/0224 →
Priority Claims (1)
JP JP2019-050503 · Mar 18, 2019 · national
Continuity (1)
Related Publication 20200302222A1 · Sep 24, 2020
Cited By (2)
US 12,350,835 US 12,400,338