IP Library Granted Patent US 12,561,964
Granted Patent B2
US 12,561,964 · App. 18/023,506 · Granted Feb 24, 2026

Method and system for detecting and classifying objects of image

Inventors: Changdong Yoo (Daejeon, KR); Thang Vu (Daejeon, KR); Xuan Trung Pham (Daejeon, KR); Hyunjun Jang (Daejeon, KR)
Assignee: KOREA ADVANCED INSTITUTE OF SCIENCE AND TECHNOLOGY
G06V10/82G06V10/776
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,964
App. No.
18/023,506
Granted
Feb 24, 2026
Kind
B2
Abstract

The present invention provides a method comprising the steps of: generating a second anchor on the second convolutional feature map by scaling and shifting a first anchor in the ground-truth box; generating a third convolutional feature map by convolving the second convolutional feature map by means of a second convolution; determining whether the overlap ratio between the ground-truth box and a second single anchor is greater than or equal to a reference value; generating a third anchor by scaling and shifting the second anchor having an overlap ratio that is greater than or equal to the reference value; assigning an objectivity score to the third anchor; and presenting the third anchor, which has an objectivity score that is equal to or greater than a reference value, as a proposal on the third convolutional feature map.

Claims (61)

1 . A method comprising:

receiving an input image and generating a first convolutional feature map;

converting the first convolutional feature map into a second convolutional feature map by a first convolution;

generating a first anchor for each point of the input image;

determining whether the first anchor is within a ground truth box;

generating a second anchor on the second convolutional feature map by scaling and shifting the first anchor within the ground truth box;

generating a third convolutional feature map by convolving the second convolutional feature map by a second convolution;

determining whether an overlap ratio between the ground truth box and the second single anchor is greater than or equal to a reference value;

generating a third anchor by scaling and shifting the second anchor having an overlap ratio greater than or equal to the reference value;

assigning an objectness score to the third anchor;

proposing the third anchor having an objectness score equal to or greater than a reference value as a proposal on the third convolutional feature map;

generating a fourth convolutional feature map by convolving the third convolutional feature map by a third convolution;

determining whether an overlapping ratio between the ground truth box and the third single anchor is greater than or equal to a reference value;

generating a fourth anchor by scaling and shifting the third anchor having the overlapping ratio greater than or equal to the reference value;

assigning an objectness score to the fourth anchor; and

proposing the fourth anchor having the objectness score equal to or greater than a reference value as a proposal on the fourth convolutional feature map.

2 . The method of claim 1 , further comprising identifying a candidate object at the point based on the objectness score.

3 . The method of claim 2 , further comprising:

determining a category of the candidate object; and

assigning a confidence score to the category of the candidate object.

4 . The method of claim 1 , wherein the first convolution is a dilated convolution.

5 . The method of claim 1 , wherein the second convolution is an adaptive convolution.

6 . A system comprising:

a processor; and

a computer-readable medium including instructions for executing an object detection and classification network by the processor,

wherein the object detection and classification network includes an initial processing module configured to input an image and generate a convolutional feature map, and an object proposal module configured to generate a proposal corresponding to a candidate object in the image, and

wherein the object proposal module performs:

converting a first convolutional feature map into a second convolutional feature map by a first convolution;

generating a first anchor for each point of the input image;

determining whether the first anchor is within a ground truth box;

generating a second anchor on the second convolutional feature map by scaling and shifting the first anchor within the ground truth box;

generating a third convolutional feature map by convolving the second convolutional feature map by a second convolution;

determining whether an overlap ratio between the ground truth box and the second single anchor is greater than or equal to a reference value;

generating a third anchor by scaling and shifting the second anchor having an overlap ratio greater than or equal to the reference value;

assigning an objectness score to the third anchor;

proposing the third anchor having an objectness score equal to or greater than a reference value as a proposal on the third convolutional feature map;

generating a fourth convolutional feature map by convolving the third convolutional feature map by a third convolution;

determining whether an overlapping ratio between the ground truth box and the third single anchor is greater than or equal to a reference value;

generating a fourth anchor by scaling and shifting the third anchor having the overlapping ratio greater than or equal to the reference value;

assigning an objectness score to the fourth anchor; and

proposing the fourth anchor having the objectness score equal to or greater than a reference value as a proposal on the fourth convolutional feature map.

7 . The system of claim 6 , wherein the first convolution is a dilated convolution.

8 . The system of claim 6 , wherein the second convolution is an adaptive convolution.

9 . The system of claim 6 , wherein the object proposal module performs identifying the candidate object at the point based on the objectness score.

10 . The system of claim 6 , further comprising a proposal classifier for determining a category of the candidate object and assigning a confidence score to the category of the candidate object.

11 . A system comprising:

a processor; and

a computer-readable medium including instructions for executing an object detection and classification network by the processor,

wherein the object detection and classification network includes an initial processing module configured to input an image and generate a convolutional feature map, and an object proposal module configured to generate a proposal corresponding to a candidate object in the image, and

wherein the object proposal module performs:

converting a first convolutional feature map into a second convolutional feature map by a first convolution;

generating a first anchor for each point of the input image;

determining whether the first anchor is within a ground truth box;

generating a second anchor on the second convolutional feature map by scaling and shifting the first anchor within the ground truth box;

generating a third convolutional feature map by convolving the second convolutional feature map by a second convolution;

determining whether an overlap ratio between the ground truth box and the second single anchor is greater than or equal to a reference value;

generating a third anchor by scaling and shifting the second anchor having an overlap ratio greater than or equal to the reference value;

assigning an objectness score to the third anchor; and

proposing the third anchor having an objectness score equal to or greater than a reference value as a proposal on the third convolutional feature map,

further comprising a proposal classifier for determining a category of the candidate object and assigning a confidence score to the category of the candidate object,

further comprising a machine learning module for training at least one parameter of the initial processing module and the object proposal module to generate at least one proposal on a training image, and training at least one parameter of the proposal classifier module to assign a category to each of the at least one proposal on the training image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2023
From: YOO, CHANGDONG; VU, THANG; PHAM, XUAN TRUNG; JANG, HYUNJUN
To: KOREA ADVANCED INSTITUTE OF SCIENCE AND TECHNOLOGY
Reel/Frame 063841/0414 →
Priority Claims (1)
KR 10-2020-0107359 · Aug 25, 2020 · national
Continuity (1)
Related Publication 20230316737A1 · Oct 5, 2023
References Cited (26)
US 10482603B1 · Fu · 2019 [cited by examiner]
US 10579897B2 · Redmon · 2020 [cited by examiner]
US 10713794B1 · He · 2020 [cited by examiner]
US 11354577B2 · Ren · 2022 [cited by examiner]
US 20170206431A1 · Sun · 2017 [cited by examiner]
US 20190012802A1 · Liu · 2019 [cited by examiner]
US 20190228215A1 · Najafirad · 2019 [cited by examiner]
US 20190258878A1 · Koivisto · 2019 [cited by examiner]
US 20200302297A1 · Jaganathan · 2020 [cited by examiner]
US 20210056351A1 · Peng · 2021 [cited by examiner]
US 20210142097A1 · Zheng · 2021 [cited by examiner]
US 20210326656A1 · Lee · 2021 [cited by examiner]
US 20220147753A1 · Fang · 2022 [cited by examiner]
US 20220156554A1 · Fu · 2022 [cited by examiner]
KR 20180065856A · 2018 [cited by applicant]
Yu Peng Chen et al. , “An Enhanced Region Proposal Network for object detection using deep learning method,” Sep. 20, 2018, PLoS ONE, 13(9): e0203897, pp. 1-15. [cited by examiner]
Fukoeng Wong et al.,“Adaptive learning feature pyramid for object detection ,” Dec. 6, 2019, IET Computer Vision, vol. 13, Issue 8,pp. 742-746. [cited by examiner]
Thang Vu et al.,“Cascade RPN: Delving into High-Quality Region Proposal Network with Adaptive Convolution,” Dec. 4, 2019, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada,pp. 1-… [cited by examiner]
Yuelei Xu et al.,“End-to-End Airport Detection in Remote Sensing Images Combining Cascade Region Proposal Networks and Multi-Threshold Detection Networks,” Sep. 21, 2018, Remote Sens. 2018, 10, 1516,pp. 1-15. [cited by examiner]
Wei Guo et al.,“Extended Feature Pyramid Network with Adaptive Scale Training Strategy and Anchors for Object Detection in Aerial Images,” Mar. 1, 2020, Remote Sens. 2020, 12, 784,pp. 1-20. [cited by examiner]
Yanjie Wang et al.,“Multi-scale dilated convolution of convolutional neural network for crowd counting,” Oct. 17, 2019, Multimedia Tools and Applications (2020) 79,pp. 1-7. [cited by examiner]
Xiaotong Zhao et al.,“Aggregated Residual Dilation-Based Feature Pyramid Network for Object Detection,” Sep. 27, 2019, IEEE Access vol. 7,2019, pp. 134014-134024. [cited by examiner]
Fan, H., et al., “Siamese Cascaded Region Proposal Networks for Real-Time Visual Tracking”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jan. 9, 2020, pp. 7952-7961. [cited by applicant]
Xu, Y., et al., “End-to-End Airport Detection in Remote Sensing Images Combining Cascade Region Proposal Networks and Multi-Threshold Detection Networks”, Remote Sens. 2018, Received Aug. 8, 2018, Accepted Sep. 12, 2018… [cited by applicant]
Vu, T., et al., “Cascade RPN: Delving into High-Quality Region Proposal Network with Adaptive Convolution”, arXiv:1909.06720v1 [cs.CV], Sep. 15, 2019, 12 pages. [cited by applicant]
International Search Report dated May 13, 2021 issued in PCT/KR2020/012450, 4 pages. [cited by applicant]