IP Library › Granted Patent US 12,749,296
Granted Patent B2
US 12,749,296 · App. 17/970,907 · Granted Sep 29, 2026

Method and apparatus with object recognition

Inventors: Insoo Kim (Seongnam-si, KR); Kikyung Kim (Hwaseong-si, KR); Seungju Han (Seoul, KR); Jiwon Baek (Hwaseong-si, KR); Jaejoon Han (Seoul, KR)
Assignee: Samsung Electronics Co., Ltd.
G06V10/806G06V10/7715G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,749,296
App. No.
17/970,907
Granted
Sep 29, 2026
Kind
B2
Abstract

A method and apparatus for object recognition are provided. A processor-implemented method includes extracting feature maps including local feature representations from an input image, generating a global feature representation corresponding to the input image by fusing the local feature representations, and performing a recognition task on the input image based on the local feature representations and the global feature representation.

Claims (70)

1 . A processor-implemented method, the method comprising:

extracting feature maps comprising local feature representations from an input image;

generating a global feature representation corresponding to the input image by fusing the local feature representations;

estimating a first recognition result corresponding to the local feature representations, using a first recognition model, and

estimating a second recognition result corresponding to the global feature representation, using a second recognition model, wherein the second recognition model comprises a classification model configured to estimate a classification result from the global feature representation; and

performing a recognition task on the input image based on the first recognition result corresponding to the local feature representations and the second recognition result corresponding to the global feature representation.

2 . The method of claim 1 , wherein the generating of the global feature representation comprises:

fusing pooling results corresponding to the local feature representations.

3 . The method of claim 2 , wherein the pooling comprises global average pooling.

4 . The method of claim 1 , wherein the generating of the global feature representation comprises:

performing an attention mechanism using query data pre-trained in association with the recognition task.

5 . The method of claim 4 , wherein the generating of the global feature representation by performing the attention mechanism comprises:

determining key data and value data corresponding to the local feature representations;

determining a weighted sum of the value data based on similarity between the key data and the query data; and

determining the global feature representation based on the weighted sum.

6 . The method of claim 1 , wherein the first recognition model comprises:

an object detection model configured to estimate a detection result from the local feature representations.

7 . The method of claim 6 , wherein the detection result comprises:

one or more of bounding box information, objectness information, or class information, and

wherein the classification result comprises:

one or more of multi-class classification information, context classification information, or object count information.

8 . The method of claim 6 , further comprising training the first model and training the second model, wherein the training the first recognition model affects the training of the second recognition model, and the training of the second recognition model affects the training of the first recognition model.

9 . The method of claim 8 , wherein the training of the first model and the training of the second model comprises:

using an in-training feature extraction model, or a trained feature extraction model, to extract training feature maps comprising in-training local feature representations from a training input image;

using an in-training feature fusion model, or a trained fusion model, to determine a training global feature representation corresponding to the training input image by fusing the training local feature representations;

using an in-training first recognition model to estimate a training first recognition result corresponding to the training local feature representations;

using an in-training second recognition model to estimate a training second recognition result corresponding to the training global feature representation; and

generating the first model and the second model by training the in-training first recognition model and the in-training second recognition model together based on the training first recognition result and the training second recognition result.

10 . The method of claim 6 , wherein the performing of the recognition task further comprises:

determining a task result recognized by the recognition task by fusing the first recognition result and the second recognition result.

11 . The method of claim 1 , wherein the recognition task corresponds to:

one of plural task candidates, the task candidates having respectively associated pre-trained query data items, and

wherein the generating of the global feature representation comprises:

selecting, from among the pre-trained query data items, a query data item associated with the recognition task; and

determining the global feature representation by performing an attention mechanism based on the selected query data item.

12 . The method of claim 1 , further comprising capturing the input image using a camera.

13 . A processor-implemented method, the method comprising:

using a feature extraction model to extract feature maps comprising local feature representations from an input image;

using a feature fusion model to determine a global feature representation corresponding to the input image by fusing the local feature representations;

using a first recognition model to estimate a first recognition result corresponding to the local feature representations;

using a second recognition model to estimate a second recognition result corresponding to the global feature representation, wherein the second recognition model comprises a plurality of classification models respectively corresponding to task candidates; and

based on the first recognition result corresponding to the local feature representations and the second recognition result corresponding to the global feature representation, training one or more of the feature extraction model, the feature fusion model, the first recognition model, or the second recognition model.

14 . The method of claim 13 , further comprising determining a training loss based on the first recognition result and the second recognition result, wherein the training is based on the training loss.

15 . The training method of claim 13 , wherein the first recognition model and the second recognition model are trained as an integrated model to influence the training of both the first recognition model and the second recognition model.

16 . The method of claim 13 , wherein the feature fusion model is configured to:

determine the global feature representation by performing an attention mechanism based on query data corresponding to a current task candidate among the task candidates, and

wherein the determining of the training loss comprises:

determining the training loss by applying a classification result of a classification model corresponding to the current task candidate among the task candidates as the second recognition result.

17 . The method of claim 13 , wherein training the first recognition model affects training the second recognition model and training the second recognition model affects training the first recognition model.

18 . A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the method of claim 1 .

19 . An electronic apparatus, comprising:

one or more processors comprising processing circuitry; and

a memory comprising one or more storage media storing instructions that, when executed individually or collectively by the one or more processors, cause the electronic device to: extract feature maps comprising respective local feature representations from an input image;

determine a global feature representation corresponding to the input image by fusing the local feature representations;

use a first recognition model to estimate a first recognition result corresponding to the local feature representations;

use a second recognition model to estimate a second recognition result corresponding to the global feature representation, wherein the second recognition model comprises a classification model configured to estimate a classification result from the global feature representation; and

perform a recognition task on the input image based on the first recognition result corresponding to the local feature representations and the second recognition result corresponding to the global feature representation.

20 . The electronic apparatus according to claim 19 , further comprising a camera configured to generate the input image.

21 . The electronic apparatus of claim 19 , wherein the processor is further configured to:

determine the global feature representation by fusing pooling results corresponding to the local feature representations.

22 . The electronic apparatus of claim 19 , wherein the processor is further configured to:

determine the global feature representation by performing an attention mechanism using query data pre-trained in response to the recognition task.

23 . The electronic apparatus of claim 22 , wherein the attention mechanism comprises a vision transformer model that performs the fusing based on similarity of keys and values of the query data.

24 . The electronic apparatus of claim 19 , wherein the processor is further configured to select between the first recognition model and the second recognition model based on the recognition task.

25 . The electronic apparatus of claim 19 , wherein the first recognition model comprises:

an object detection model configured to estimate a detection result corresponding to each of the local feature representations, and

wherein the second recognition model comprises:

a classification model configured to estimate a classification result corresponding to the global feature representation.

26 . The electronic apparatus of claim 25 , wherein the processor is further configured to:

determine a task result of the recognition task by fusing the first recognition result and the second recognition result.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 21, 2022
From: KIM, INSOO; KIM, KIKYUNG; HAN, SEUNGJU; BAEK, JIWON; HAN, JAEJOON
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061497/0532 →
Priority Claims (1)
KR 10-2022-0003434 · Jan 10, 2022 · national
Continuity (1)
Related Publication 20230222781A1 · Jul 13, 2023
References Cited (48)
US 6563950B1 · Wiskott et al. · 2003 [cited by applicant]
US 7646909B2 · Jiang et al. · 2010 [cited by applicant]
US 8121357B2 · Imaoka · 2012 [cited by applicant]
US 8571332B2 · Kumar et al. · 2013 [cited by applicant]
US 8856541B1 · Chaudhury et al. · 2014 [cited by applicant]
US 9171217B2 · Pawlicki et al. · 2015 [cited by applicant]
US 9547790B2 · Kim et al. · 2017 [cited by applicant]
US 9798871B2 · Suh et al. · 2017 [cited by applicant]
US 9928410B2 · Yoo et al. · 2018 [cited by applicant]
US 10083343B2 · Suh et al. · 2018 [cited by applicant]
US 10121059B2 · Yoo et al. · 2018 [cited by applicant]
US 10210379B2 · Suh et al. · 2019 [cited by applicant]
US 10248845B2 · Han et al. · 2019 [cited by applicant]
US 10275684B2 · Han et al. · 2019 [cited by applicant]
US 10331941B2 · Rhee et al. · 2019 [cited by applicant]
US 10878592B2 · Croxford et al. · 2020 [cited by applicant]
US 10891514B2 · Liu · 2021 [cited by examiner]
US 10990658B2 · Suh et al. · 2021 [cited by applicant]
US 12249138B2 · Kilickaya · 2025 [cited by examiner]
US 20120288167A1 · Sun et al. · 2012 [cited by applicant]
US 20130322705A1 · Wong · 2013 [cited by applicant]
US 20140044359A1 · Rousson · 2014 [cited by applicant]
US 20140169643A1 · Todoroki · 2014 [cited by applicant]
US 20150125049A1 · Taigman et al. · 2015 [cited by applicant]
US 20160070952A1 · Kim et al. · 2016 [cited by applicant]
US 20210056344A1 · Zhang et al. · 2021 [cited by applicant]
CN 110619369B · 2020 [cited by examiner]
CN 113011386A · 2021 [cited by examiner]
KR 101731461B1 · 2017 [cited by applicant]
KR 1020190089777A · 2019 [cited by applicant]
KR 1020200007022A · 2020 [cited by applicant]
KR 102102161B1 · 2020 [cited by applicant]
KR 102186913B1 · 2020 [cited by applicant]
KR 1020210011707A · 2021 [cited by applicant]
KR 1020210079922A · 2021 [cited by applicant]
WO WO2021097378A1 · 2021 [cited by applicant]
S. Chen and Y. Tian, “Pyramid of Spatial Relatons for Scene-Level Land Use Classification,” in IEEE Transactions on Geoscience and Remote Sensing, vol. 53, No. 4, pp. 1947-1957, Apr. 2015, doi: 10.1109/TGRS.2014.2351395… [cited by examiner]
S. Guo, Li Liu, W. Wang, Songyang Lao and L. Wang, “An attention model based on spatial transformers for scene recognition,” 2016 23rd International Conference on Pattern Recognition (ICPR), Cancun, Mexico, 2016, pp. 37… [cited by examiner]
Dosovitskiy, Alexey, et al. “An image is worth 16×16 words: Transformers for image recognition at scale.” arXiv preprint arXiv: 2010.11929 (2020). (Year: 2020). [cited by examiner]
Jetley, Saumya, et al. “Learn to pay attention.” arXiv preprint arXiv: 1804.02391 (2018) (Year: 2018). [cited by examiner]
Minaee, Shervin, et al. “Image segmentation using deep learning: A survey.” IEEE transactions on pattern analysis and machine intelligence 44.7 (2021): 3523-3542. (Year: 2021). [cited by examiner]
Jetley, Saumya, et al. “Learn To Pay Attention.” arXiv:1804.02391v2, Apr. 26, 2018, (14 pages in English). [cited by applicant]
Minaee, Shervin, et al. “Image Segmentation Using Deep Learning: A Survey.” IEEE transactions on pattern analysis and machine intelligence 44.7, arXiv:2001.05566v5, Nov. 15, 2020: (22 pages in English). [cited by applicant]
Extended European search report issued on Jun. 14, 2023, in counterpart European Patent Application No. 22210327.7 (12 pages in English). [cited by applicant]
Kathuria, Ayoosh. “What's new in YOLO v3?” Towards Date Science. (2018). pp 1-15. [cited by applicant]
Rozsa, Andras, et al. “Are accuracy and robustness correlated.” [cited by applicant]
Choi, Jun-Ho, et al. “Evaluating robustness of deep image super-resolution against adversarial attacks.” [cited by applicant]
European Office Action issued on Jun. 10, 2025, in corresponding European Patent Application No. 22210327.7. (8pages in English). [cited by applicant]