IP Library Granted Patent US 12,694,471
Granted Patent B2
US 12,694,471 · App. 18/030,989 · Granted Jul 28, 2026

Image recognition with a wide area recognition target network that feeds a narrow area recognition target network

Inventor: Toshihide Horii (Osaka, JP)
Assignee: Panasonic Intellectual Property Management Co., Ltd.
G06T3/4046G06T5/50G06T2207/20212
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,694,471
App. No.
18/030,989
Filed
Apr 8, 2023
Granted
Jul 28, 2026
Kind
B2
Art Unit
2663
USPC
382/100
Abstract

Aspects relate to a processing device for increasing the accuracy of image recognition in a neural network that does not include a fully connected layer. A first processor generates a first feature map by executing processing of a first neural network on a target image. An enlarger enlarges the first feature map. A combiner combines the first feature map and the target image and generates a combined image. A second processor generates a second feature map by executing processing of a second neural network on the combined image.

Claims (21)

1 . A processing device adapted to execute inference processing after learning processing, comprising:

a computer-readable non-transitory recording medium storing executable instructions that, in response to execution, causes the processing device to perform operations comprising:

processing, in the inference processing after learning, of a first neural network on a target image to be processed, thereby generating a first feature map having a smaller size than the target image,

up-sampling, in the inference processing after learning, the first feature map to have the same size as the target image,

concatenating, in the inference processing after learning, the first feature map that was up-sampled and the target image, thereby generating a combined image, and

processing, in the inference processing after learning, of a second neural network on the combined image, thereby generating a second feature map having a smaller size than the target image and a larger size than the first feature map,

wherein the first neural network does not include a fully connected layer, and the second neural network does not include a fully connected layer,

in the learning processing, first-stage learning is performed only on the first neural network, and

in the learning processing, a coefficient of a spatial filter derived by the first-stage learning is set to each convolution layer included in the first neural network, and

second-stage learning is performed on the second neural network after the first-stage learning has been performed.

2 . The processing device according to claim 1 , wherein the concatenating combines two inputs as different channels.

3 . A processing method performed by a processing device adapted to execute recognition processing after learning processing, the method including:

a computer-readable non-transitory recording medium storing executable instructions that, in response to execution, causes the processing device to perform the processing method, the processing method comprising:

a step of executing processing of a first neural network on a target image to be processed and generating a first feature map having a smaller size than the target image;

a step of up-sampling the generated first feature map to have the same size as the target image;

a step of concatenating the first feature map that was up-sampled and the target image to generate a combined image; and

a step of executing processing of a second neural network on the generated combined image and generating a second feature map having a smaller size than the target image and a larger size than the first feature map,

wherein the first neural network does not include a fully connected layer, and the second neural network does not include a fully connected layer,

in the learning processing, first-stage learning is performed only on the first neural network, and

in the learning processing, a coefficient of a spatial filter derived by the first-stage learning is set to each convolution layer included in the first neural network, and

second-stage learning is performed on the second neural network after the first-stage learning has been performed on the first neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2023
From: HORII, TOSHIHIDE
To: PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO., LTD.
Reel/Frame 064785/0271 →
Priority Claims (1)
JP 2020-170752 · Oct 8, 2020 · national
Continuity (1)
Related Publication 20230377094A1 · Nov 23, 2023
References Cited (13)
US 10783640B2 · Chen · 2020 [cited by examiner]
US 10796184B2 · Senay · 2020 [cited by examiner]
US 11315235B2 · Horii · 2022 [cited by examiner]
US 20200380665A1 · Horii · 2020 [cited by applicant]
US 20210248761A1 · Liu · 2021 [cited by examiner]
CN 111712853A · 2020 [cited by applicant]
WO 2019159419A1 · 2019 [cited by applicant]
Peng, Chengli, and Jiayi Ma. “Semantic segmentation using stride spatial pyramid pooling and dual attention decoder.” Pattern Recognition 107 (2020): 107498 (retrieved from https://www.sciencedirect.com/science/article/… [cited by examiner]
Chen, Ying-Nong, et al. “Facial/license plate detection using a two-level cascade classifier and a single convolutional feature map.” International Journal of Advanced Robotic Systems 12.12 (2015): 183. (Year: 2015). [cited by examiner]
International Search Report for corresponding Application No. PCT/JP2021/024225, mailed Sep. 21, 2021. [cited by applicant]
Extended European Search Report for corresponding Application No. 21877182.2 dated Mar. 11, 2024. [cited by applicant]
Chen, Wuyang et al., Collaborative Global-Local Networks for Memory-Efficient Segmentation of Ultra-High Resolution Images, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 15, 2019… [cited by applicant]
Chen at al., “Collaborative Global-Local Networks for Memory-Efficient Segmentation of Ultra-High Resolution Images”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2019, pp. 8924-8933. [cited by applicant]