IP Library Granted Patent US 12,475,676
Granted Patent B2
US 12,475,676 · App. 18/000,338 · Granted Nov 18, 2025

Object detection method, object detection device, and program

Inventor: Taiki Sekii (Tokyo, JP)
Assignee: KONICA MINOLTA, INC.
G06V10/476G06T7/66G06T7/75G06V10/764G06V10/82G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,676
App. No.
18/000,338
Granted
Nov 18, 2025
Kind
B2
Abstract

An object detection method includes a key point estimation step of estimating key point candidates for each object in an image; and a detection step of detecting key points for each object based on the estimated key point candidates. Considering an object model that models shape of an object, the key points are points that satisfy a defined condition among points indicating a boundary of the object model that are projected onto defined coordinate axes. The defined coordinate axes have an origin at a geometric center of the object model and each forms a defined angle relative to a polar axis in a polar coordinate system set for the object model.

Claims (46)

1 . An object detection method for detecting each object in an image containing one or more objects of one or more defined categories, comprising:

a key point estimation step of estimating key point candidates for each object in the image; and

a detection step of detecting key points for each object based on the estimated key point candidates, wherein

considering an object model that models shape of an object, the key points are points that satisfy a defined condition among points indicating a boundary of the object model that are projected onto defined coordinate axes,

the defined coordinate axes have an origin at a geometric center of the object model and each forms a defined angle relative to a polar axis in a polar coordinate system set for the object model, and

two key points of the key points are defined for each of the coordinate axes, the two key points for each of the coordinate axes are a point on the boundary of the object having a maximum value and a point on the boundary of the object having a minimum value in a positive range of the each of the coordinate axes.

2 . The object detection method of claim 1 , wherein

the key points are selected as points on the object having a local maximum or local minimum value, and the two key points for each of the coordinate axes are determined from the points on the object having a local maximum or local minimum value.

3 . The object detection method of claim 1 , further comprising

a center position estimation step of estimating a center candidate for each object in the image and a confidence indicating likelihood of accurate estimation, wherein

the detection step uses the confidence to detect a center position of each object from the center candidates, and uses each detected center position in detection of the key points of each object from the key point candidates.

4 . The object detection method of claim 3 , wherein

the key point estimation step and the center position estimation step are executed by a machine-learning model trained to detect each object.

5 . The object detection method of claim 1 , wherein

the key point estimation step estimates the key point candidates as areas each having a size based on a corresponding object of the one or more objects.

6 . The object detection method of claim 1 , wherein

the key point estimation step is executed by a machine-learning model trained to detect each object.

7 . The object detection method of claim 6 , wherein

the machine-learning model is a convolutional neural network, and

parameters of the convolutional neural network are defined by machine-learning based on a training image including a detection target object, a true value of a center position of the detection target object in the training image, and a true value of a key point of the detection target object in the training image.

8 . The object detection method of claim 1 , wherein each of the key points that satisfies the defined condition is a point on the surface of the object model that is a point of a local maximum coordinate value or a local minimum coordinate value among points on the surface of the object model such that the each key point protrudes from other portions of the object or is recessed from other portions of the object.

9 . The object detection method of claim 8 , further comprising

a center position estimation step of estimating a center candidate for each object in the image and a confidence indicating likelihood of accurate estimation, wherein

the detection step uses the confidence to detect a center position of each object from the center candidates, and uses each detected center position in detection of the key points of each object from the key point candidates.

10 . The object detection method of claim 9 , wherein

the key point estimation step and the center position estimation step are executed by a machine-learning model trained to detect each object.

11 . The object detection method of claim 8 , wherein

the key point estimation step estimates the key point candidates as areas each having a size based on a corresponding object of the one or more objects.

12 . The object detection method of claim 8 , wherein

the key point estimation step is executed by a machine-learning model trained to detect each object.

13 . The object detection method of claim 12 , wherein

the machine-learning model is a convolutional neural network, and

parameters of the convolutional neural network are defined by machine-learning based on a training image including a detection target object, a true value of a center position of the detection target object in the training image, and a true value of a key point of the detection target object in the training image.

14 . An object detection device for detecting each object in an image containing one or more objects of one or more defined categories, comprising:

a machine-learning model trained to detect each object, which executes a key point estimation process of estimating key point candidates for each object in the image; and

a detection unit that detects key points for each object based on the estimated key point candidates, wherein

considering an object model that models shape of an object, the key points are points that satisfy a defined condition among points indicative of a boundary of the object model that are projected onto a defined coordinate axis, and

the defined coordinate axes have an origin at a geometric center of the object model and each forms a defined angle relative to a polar axis in a polar coordinate system set for the object model, and

two key points of the key points are defined for each of the coordinate axes, the two key points for each of the coordinate axes are a point on the boundary of the object having a maximum value and a point on the boundary of the object having a minimum value in a positive range of the each of the coordinate axes.

15 . A non-transitory computer readable medium storing a program causing a computer to execute object detection processing for detecting each object in an image containing one or more objects of one or more defined categories, wherein

the object detection processing comprises:

a key point estimation step of estimating key point candidates for each object in the image; and

a detection step of detecting key points for each object based on the estimated key point candidates, wherein

considering an object model that models shape of an object, the key points are points that satisfy a defined condition among points indicative of a boundary of the object model that are projected onto a defined coordinate axis, and

the defined coordinate axes have an origin at a geometric center of the object model and each forms a defined angle relative to a polar axis in a polar coordinate system set for the object model, and

two key points of the key points are defined for each of the coordinate axes, the two key points for each of the coordinate axes are a point on the boundary of the object having a maximum value and a point on the boundary of the object having a minimum value in a positive range of the each of the coordinate axes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2022
From: SEKII, TAIKI
To: KONICA MINOLTA, INC.
Reel/Frame 061930/0435 →
Priority Claims (1)
JP 2020-098325 · Jun 5, 2020 · national
Continuity (1)
Related Publication 20240029394A1 · Jan 25, 2024
References Cited (17)
US 11348269B1 · Ebrahimi Afrouzi · 2022 [cited by examiner]
US 20120206438A1 · Porikli · 2012 [cited by examiner]
US 20210012106A1 · Markhasin · 2021 [cited by examiner]
US 20210157998A1 · Rodriguez · 2021 [cited by examiner]
US 20210201083A1 · Wang · 2021 [cited by examiner]
US 20210295606A1 · Kim · 2021 [cited by examiner]
US 20220148153A1 · Zhu · 2022 [cited by examiner]
US 20220331841A1 · Filler · 2022 [cited by examiner]
US 20230351573A1 · Zhang · 2023 [cited by examiner]
US 20250037492A1 · Skoryukina · 2025 [cited by examiner]
JP 2006202135A · 2006 [cited by applicant]
JP 2014109555A · 2014 [cited by applicant]
Zhang et al., “Image retrieval using the extended salient region,” Information Sciences 399 (2017) 154-182 (Year: 2017). [cited by examiner]
Abdel-Kader et al., “A boundary-based approach to shape orientability using particle swarm optimization,” SIViP (2014) 8:779-788 DOI 10.1007/s11760-013-0598-z (Year: 2014). [cited by examiner]
Xingyi Zhou, Jiacheng Zhuo, Philipp Krahenbuhl, “Bottom-up Object Detection by Grouping Extreme and Center Points”, Computer Vision and Pattern Recognition (CVPR) 2019; 10 pages. [cited by applicant]
Joseph Redmon, Santosh Divvala, Ross Girshick, Ali Farhadi, “You Only Look Once: Unified, Real-Time Object Detection”, Computer Vision and Pattern Recognition (CVPR) 2016; 10 pages. [cited by applicant]
International Search Report and Written Opinion for the corresponding patent application No. PCT/JP2021/019555 dated Aug. 10, 2021, with English translation. [cited by applicant]