IP Library › Granted Patent US 12,738,043
Granted Patent B2
US 12,738,043 · App. 18/367,193 · Granted Sep 15, 2026

Electronic apparatus for identifying a region of interest in an image and control method thereof

Inventors: Ilhyun Cho (Suwon-si, KR); Wookhyung Kim (Suwon-si, KR); Jayoon Koo (Suwon-si, KR); Namuk Kim (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06V10/82G06T3/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,043
App. No.
18/367,193
Granted
Sep 15, 2026
Kind
B2
Abstract

An electronic apparatus includes a memory configured to store a neural network model including a first network and a second network. The electronic apparatus also includes at least one processor connected to the memory. The at least one processor is configured to obtain description information corresponding to a first image by inputting the first image to the first network, obtain a second image based on the description information, obtain a third image representing a region of interest of the first image by inputting the first image and the second image to the second network. The neural network model is a model trained based on a plurality of sample images, a plurality of sample description information corresponding to the plurality of sample images, and a sample region of interest of the plurality of sample images.

Claims (52)

1 . An electronic apparatus, comprising:

a memory configured to store a neural network model comprising a first network and a second network, wherein the neural network model comprises weights; and

at least one processor connected to the memory and configured to control the electronic apparatus,

wherein the at least one processor is configured to:

obtain first description information corresponding to a first image by inputting the first image to the first network using the weights,

obtain a second image based on the first description information, and

obtain a third image representing a first region of interest of the first image by inputting the first image and the second image to the second network using the weights,

wherein the weights of the neural network model are trained based on: i) a plurality of sample images, ii) a plurality of sample description information corresponding to the plurality of sample images, and iii) a sample region of interest for each sample image of the plurality of sample images, and

wherein the first description information comprises at least one word and the at least one processor is further configured to obtain the second image by converting each word of the at least one word to a corresponding color.

2 . The electronic apparatus of claim 1 , wherein the at least one processor is further configured to:

obtain the third image by inputting the first image to an input layer of the second network using the weights, and

input the second image to an intermediate layer of the second network using the weights.

3 . The electronic apparatus of claim 1 , wherein the at least one processor is further configured to:

obtain the first image by downscaling an original image to a pre-set resolution, or

downscale the original image to a pre-set scaling rate.

4 . The electronic apparatus of claim 3 , wherein the at least one processor is further configured to upscale the third image to correspond to a resolution of the original image.

5 . The electronic apparatus of claim 1 , wherein the third image depicts the first region of interest of the first image in a first color, and the third image depicts a background region in a second color, wherein the background region is a remaining region excluding the first region of interest of the first image, and

a resolution of the third image is the same as that of the first image.

6 . The electronic apparatus of claim 1 , wherein the first network is configured so as to be trained of a first relationship of the plurality of sample description information for the plurality of sample images through an artificial intelligence algorithm, and

the second network is configured so as to be trained, through the artificial intelligence algorithm, of a second relationship of the plurality of sample images and the sample region of interest for the each sample image of the plurality of sample images, wherein each sample image corresponds to a sample description information of the plurality of sample description information.

7 . The electronic apparatus of claim 6 , wherein the first network and the second network are simultaneously trained.

8 . The electronic apparatus of claim 1 , wherein the first network comprises a convolution network and a plurality of long short-term memory networks (LSTMs), and

the plurality of LSTMs are configured to output the first description information.

9 . The electronic apparatus of claim 1 , wherein the at least one processor is further configured to:

identify a remaining region excluding the first region of interest from the first image as a background region, and

image process the first region of interest and the background region differently.

10 . The electronic apparatus of claim 1 , wherein the at least one processor is further configured to:

read the weights of the neural network model from the memory; and

implement the first network and the second network in the at least one processor using the weights.

11 . The electronic apparatus of claim 1 , wherein the at least one processor is further configured to:

identify a remaining region excluding the first region of interest from the first image as a background region, and

lower a brightness of the background region and maintain the brightness of the first region of interest.

12 . A control method of an electronic apparatus, the control method comprising:

obtaining first description information corresponding to a first image by inputting the first image to a first network comprised in a neural network model;

obtaining a second image based on the first description information; and

obtaining a third image showing a region of interest of the first image by inputting the first image and the second image to a second network comprised in the neural network model,

wherein the neural network model is a model trained based on a plurality of sample images, a plurality of sample description information corresponding to the plurality of sample images, and a sample region of interest for each sample image of the plurality of sample images, and

wherein the first description information comprises at least one word and the obtaining the second image comprises converting each word of the at least one word to a corresponding color.

13 . The control method of claim 12 , wherein the obtaining the third image comprises:

obtaining the third image by inputting the first image in an input layer of the second network, and

inputting the second image to an intermediate layer of the second network.

14 . The control method of claim 12 , further comprising:

obtaining the first image by downscaling an original image to a pre-set resolution, or

downscaling the original image to a pre-set scaling rate.

15 . The control method of claim 14 , further comprising upscaling the third image to correspond to a resolution of the original image.

16 . The control method of claim 15 , wherein the resolution of the first image is a first resolution of 320 by 240 after the downscaling, and the third image has a full high definition (FHD) resolution after the upscaling, wherein the original image has an FHD resolution.

17 . The control method of claim 14 , wherein the pre-set resolution is 320 by 240.

18 . The control method of claim 17 , wherein the downscaling is configured to reduce a power consumption by using the first image, wherein the first image is of low resolution.

19 . The control method of claim 14 , wherein the downscaling maintains a horizontal width of the original image and a vertical height of the original image.

20 . The control method of claim 12 , further comprising:

identifying a remaining region excluding the first region of interest from the first image as a background region, and

lowering a brightness of the background region and maintaining the brightness of the first region of interest.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 12, 2023
From: CHO, ILHYUN; KIM, WOOKHYUNG; KOO, JAYOON; KIM, NAMUK
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 064881/0377 →
Priority Claims (1)
KR 10-2022-0136326 · Oct 21, 2022 · national
Continuity (3)
Continuation PCTKR2023011051 · Jul 28, 2023
Related Publication 20240135697A1 · Apr 25, 2024
Related Publication 20240233356A9 · Jul 11, 2024
References Cited (58)
US 8019175B2 · Lee et al. · 2011 [cited by applicant]
US 10217195B1 · Agrawal · 2019 [cited by examiner]
US 10679063B2 · Cheng et al. · 2020 [cited by applicant]
US 11017255B2 · Lee · 2021 [cited by applicant]
US 11388465B2 · Park et al. · 2022 [cited by applicant]
US 11551433B2 · Lee · 2023 [cited by applicant]
US 11657230B2 · Lee · 2023 [cited by examiner]
US 11825033B2 · Park et al. · 2023 [cited by applicant]
US 20030137522A1 · Kaasila · 2003 [cited by examiner]
US 20060215753A1 · Lee et al. · 2006 [cited by applicant]
US 20070097268A1 · Relan · 2007 [cited by examiner]
US 20080201282A1 · Garcia et al. · 2008 [cited by applicant]
US 20130028514A1 · Kihira · 2013 [cited by examiner]
US 20150254538A1 · Fukazawa · 2015 [cited by examiner]
US 20160225053A1 · Romley · 2016 [cited by examiner]
US 20170295372A1 · Lawrence · 2017 [cited by examiner]
US 20180211438A1 · Gonzalez Aguirre · 2018 [cited by examiner]
US 20180247405A1 · Kisilev et al. · 2018 [cited by applicant]
US 20190042574A1 · Kim et al. · 2019 [cited by applicant]
US 20200053408A1 · Park et al. · 2020 [cited by applicant]
US 20200226410A1 · Liu · 2020 [cited by examiner]
US 20210019890A1 · Chen · 2021 [cited by examiner]
US 20210034905A1 · Lee · 2021 [cited by applicant]
US 20210166013A1 · Tensmeyer · 2021 [cited by examiner]
US 20210166345A1 · Kim · 2021 [cited by examiner]
US 20210241432A1 · Li et al. · 2021 [cited by applicant]
US 20210256365A1 · Wang et al. · 2021 [cited by applicant]
US 20210279495A1 · Lee · 2021 [cited by applicant]
US 20220007946A1 · Ziegle · 2022 [cited by examiner]
US 20220030291A1 · Park et al. · 2022 [cited by applicant]
US 20220058332A1 · Ke et al. · 2022 [cited by applicant]
US 20230096682A1 · Chou · 2023 [cited by examiner]
US 20230230198A1 · Zhang · 2023 [cited by examiner]
US 20230274453A1 · Tang · 2023 [cited by examiner]
US 20240040179A1 · Park et al. · 2024 [cited by applicant]
US 20240080462A1 · Park · 2024 [cited by examiner]
US 20240153258A1 · Mangla · 2024 [cited by examiner]
US 20240389961A1 · Lee · 2024 [cited by examiner]
US 20250254329A1 · Zhang · 2025 [cited by examiner]
JP 2008536211A · 2008 [cited by applicant]
JP 2022505115A · 2022 [cited by applicant]
KR 100946813B1 · 2010 [cited by applicant]
KR 101930940B1 · 2018 [cited by applicant]
KR 1020190013427A · 2019 [cited by applicant]
KR 1020190030151A · 2019 [cited by applicant]
KR 102022648B1 · 2019 [cited by applicant]
KR 1020200062891A · 2020 [cited by applicant]
KR 102199094B1 · 2021 [cited by applicant]
KR 1020220037617A · 2022 [cited by applicant]
Zhang et al, CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection, 2019 (Year: 2019). [cited by examiner]
Communication issued on Apr. 22, 2025 by the European Patent Office in European Patent Application No. 23879964.7. [cited by applicant]
Lu Zhang et al., “CapSal: Leveraging Captioning to Boost Semantics for Salient Object Detection”, 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 6017-6026, DOI: 10.1109/CVPR.2… [cited by applicant]
Tengfei Xing et al., “Text to Region: Visual-Word Guided Saliency Detection”, Advances in Multimedia Information Processing—PCM 2018, Sep. 2018, pp. 740-749, DOI: 10.1007/978-3-030-00764-5_68, XP047686275. [cited by applicant]
Vasili Ramanishka et al., “Top-Down Visual Saliency Guided by Captions”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, pp. 3135-3144, DOI: 10.1109/CVPR.2017.334, XP033249660. [cited by applicant]
International Search Report and Written Opinion (PCT/ISA/210 and PCT/ISA/237) dated Nov. 27, 2023, issued by International Searching Authority in International Application No. PCT/KR2023/011051. [cited by applicant]
Shen Yang et al., “A Dilated Inception Network for Visual Saliency Prediction”, IEEE Transactions on Multimedia, arXiv:1904.03571v2 [cs.CV], May 9, 2019, 14 pages. [cited by applicant]
Xiongkuo Min et al., “A Multimodal Saliency Model for Videos With High Audio-Visual Correspondence”, IEEE Transactions on Image Processing, vol. 29, 2020, 15 pages. [cited by applicant]
Extended European Search Report dated Apr. 1, 2026, issued by the European Patent Office in European Application No. 23879964.7. [cited by applicant]