IP Library Granted Patent US 12,450,889
Granted Patent B2
US 12,450,889 · App. 18/636,615 · Granted Oct 21, 2025

Method and apparatus with object classification

Inventors: Sangil Jung (Yongin-si, KR); Seungin Park (Yongin-si, KR); Byung In Yoo (Seoul, KR)
Assignee: Samsung Electronics Co., Ltd.
G06V10/806G06V10/40G06V10/764G06V10/7715G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,889
App. No.
18/636,615
Granted
Oct 21, 2025
Kind
B2
Abstract

An object classification method and apparatus are disclosed. The object classification method includes receiving an input image, storing first feature data extracted by a first feature extraction layer of a neural network configured to extract features of the input image, receiving second feature data from a second feature extraction layer which is an upper layer of the first feature extraction layer, generating merged feature data by merging the first feature data and the second feature data, and classifying an object in the input image based on the merged feature data.

Claims (46)

1. A processor-implemented method, comprising:

receiving an input image;

generating first feature data corresponding to global information of the input image by extracting the first feature data using a first feature extraction layer of a neural network with the input image as input;

generating second feature data corresponding to local information of the input image by extracting the second feature data using a second feature extraction layer of the neural network, the second feature extraction layer being an upper layer of the first feature extraction layer;

generating sequentially merged feature data by merging the first feature data corresponding to the global information and sets of second feature data corresponding to local information of the input image output from a plurality of second feature extraction layers of the neural network; and

classifying an object in the input image based on the sequentially merged feature data.

2. The method of claim 1 , wherein the generating of the sequentially merged feature data comprising:

storing the sets of the second feature data corresponding to local features output from the plurality of second feature extraction layers.

3. The method of claim 1 , wherein the first feature data comprises one of a feature vector, a feature map, activation data and an activation map, and

the second feature data comprises one of a feature vector, a feature map, activation data and an activation map.

4. The method of claim 3 , wherein the generating of the sequentially merged feature data is based on an inner product between a feature map corresponding to the first feature data and a feature vector corresponding to the second feature data.

5. The method of claim 1 , wherein the second feature extraction layer is an uppermost feature extraction layer among a plurality of feature extraction layers comprised in the neural network.

6. The method of claim 1 , wherein the generating of the sequentially merged feature data comprises:

determining a weight to be applied to the first feature data;

applying the determined weight to the first feature data and determining first feature data to which the weight is applied; and

generating the sequentially merged feature data by merging the second feature data and the first feature data to which the weight is applied.

7. The method of claim 1 , wherein the generating of the sequentially merged feature data comprises:

converting the second feature data such that a dimension of the second feature data corresponds to a dimension of the first feature data; and

generating the sequentially merged feature data by merging the converted second feature data and the first feature data.

8. The method of claim 7 , wherein the generating of the sequentially merged feature data further comprises:

converting the sequentially merged feature data such that a dimension of the sequentially merged feature data corresponds to the dimension of the second feature data.

9. The method of claim 1 , wherein the first feature data comprises a local feature of the input image, and the second feature data comprises a global feature of the input image.

10. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 .

11. An electronic apparatus, comprising:

one or more processors, configured to:

receive an input image;

generate first feature data corresponding to global information of the input image by extracting the first feature data using a first feature extraction layer of a neural network with the input Image as input;

generate second feature data corresponding to local information of the input image by extracting the second feature data using a second feature extraction layer of the neural network, the second feature extraction layer being an upper layer of the first feature extraction layer;

generate sequentially merged feature data by merging the first feature data corresponding to the global information and sets of second feature data corresponding to local features-information of the input image output from a plurality of second feature extraction layers of the neural network; and

classify an object in the input image based on the sequentially merged feature data.

12. The electronic apparatus of claim 11 , wherein the processors are further configured to:

store the sets of the second feature data corresponding to local features output from the plurality of second feature extraction layers.

13. The electronic apparatus of claim 11 , wherein the first feature data comprises one of a feature vector, a feature map, activation data and an activation map, and

the second feature data comprises one of a feature vector, a feature map, activation data and an activation map.

14. The electronic apparatus of claim 13 , wherein the sequentially

merged feature data is based on an inner product between a feature map corresponding to the first feature data and a feature vector corresponding to the second feature data.

15. The electronic apparatus of claim 11 , wherein the second feature extraction layer is an uppermost feature extraction layer among a plurality of feature extraction layers comprised in the neural network.

16. The electronic apparatus of claim 11 , wherein the processors are further configured to:

determine a weight to be applied to the first feature data;

apply the determined weight to the first feature data and determining first feature data to which the weight is applied; and

generate the sequentially merged feature data by merging the second feature data and the first feature data to which the weight is applied.

17. The electronic apparatus of claim 11 , wherein the processors are further configured to:

convert the second feature data such that a dimension of the second feature data corresponds to a dimension of the first feature data; and

generate the sequentially merged feature data by merging the converted second feature data and the first feature data.

18. The electronic apparatus of claim 17 , wherein the processors are further configured to:

convert the sequentially merged feature data such that a dimension of the sequentially merged feature data corresponds to the dimension of the second feature data.

Priority Claims (1)
KR 10-2021-0126062 · Sep 24, 2021 · national
Continuity (2)
Continuation 17697160 · Mar 17, 2022
Related Publication 20240273884A1 · Aug 15, 2024
References Cited (34)
US 7599554B2 · Agnihotri et al. · 2009 [cited by applicant]
US 8711248B2 · Jandhyala et al. · 2014 [cited by applicant]
US 8763114B2 · Alperovitch et al. · 2014 [cited by applicant]
US 9536293B2 · Lin et al. · 2017 [cited by applicant]
US 10002313B2 · Vaca Castano et al. · 2018 [cited by applicant]
US 10304233B2 · Oh · 2019 [cited by applicant]
US 20170061328A1 · Majumdar · 2017 [cited by examiner]
US 20170169315A1 · Vaca Castano et al. · 2017 [cited by applicant]
US 20210089823A1 · Iio · 2021 [cited by examiner]
US 20220092351A1 · Huang · 2022 [cited by examiner]
US 20220172378A1 · Sharma · 2022 [cited by examiner]
US 20230095716A1 · Jung · 2023 [cited by examiner]
KR 100924690B1 · 2009 [cited by applicant]
KR 1020130060274A · 2013 [cited by applicant]
KR 1020130073852A · 2013 [cited by applicant]
KR 101479225B1 · 2015 [cited by applicant]
KR 1020160142760A · 2016 [cited by applicant]
KR 1020200101514A · 2020 [cited by applicant]
KR 102249663B1 · 2021 [cited by applicant]
WO WO2020104542A1 · 2020 [cited by applicant]
Deformation and Refined Features Based Lesion Detection on Chest X-Ray. Ce Li 1, Dong Zhang 1 , Shaoyi Du 2. Jan. 24, 2020. (Year: 2020). [cited by examiner]
Dosovitskiy, Alexey, et al. “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale.” arXiv:2010.11929v2 [cs.CV] Jun. 3, 2021. pp. 1-22 (22 pages in English). [cited by applicant]
He, Kaiming, et al. “Deep Residual Learning for Image Recognition.” [cited by applicant]
Heo, Byeongho, et al. “Rethinking spatial dimensions of vision transformers.” Proceedings of the IEEE/CVF international conference on computer vision. arXiv:2103.16302v2 [cs.CV] Aug. 18, 2021. (10 pages in English). [cited by applicant]
Howard, Andrew G., et al. “Mobilenets: Efficient Convolutional Neural Networks for Mobile Vision Applications.” arXiv:1704.04861vl [cs.CV] Apr. 17, 2017. (9 pages in English). [cited by applicant]
Krizhevsky, Alex, et al. “ImageNet Classification with Deep Convolutional Neural Networks.” Advances in neural information processing systems vol. 25 (2012). pp. 1-9 (9 pages in English). [cited by applicant]
Li Ce et al., “Deformation and Refined Features based Lesion Detection on Chest X-Ray,” IEEE Access, vol. 8, Jan. 2, 2020, pp. 14675-14689. [cited by applicant]
Simonyan, Karen, et al. “Very Deep Convolutional Networks for Large-Scale Image Recognition.” arXiv:1409.1556 ver. 6 [cs.CV] Apr. 10, 2015. pp. 1-14 (14 pages in English). [cited by applicant]
Extended European Search Report issued on Oct. 5, 2022, in European Patent Application No. 22169643.8 (60 Pages). [cited by applicant]
Korean Office Action issued on Aug. 8, 2025, in counterpart Korean Application No. 10-2021-0126062(4 pages in English, 9 pages in Korean). [cited by applicant]
Lin, Tsung-Yi, et al. “Feature pyramid networks for object detection.” [cited by applicant]
Cao, Yue, et al. “Gcnet: Non-local networks meet squeeze-excitation networks and beyond.” [cited by applicant]
Touvron, Hugo, et al. “Going deeper with image transformers.” [cited by applicant]
Frome, Andrea, et al. “Devise: A deep visual-semantic embedding model.” [cited by applicant]