IP Library › Granted Patent US 12,632,662
Granted Patent B2
US 12,632,662 · App. 17/947,758 · Granted May 19, 2026

Electronic apparatus and controlling method thereof

Inventors: Hyungtak Choi (Suwon-si, KR); Munjo Kim (Suwon-si, KR); Seonghan Ryu (Suwon-si, KR); Sejin Kwak (Suwon-si, KR); Lohith Ravuru (Suwon-si, KR); Haehun Yang (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F40/30G06F40/42G06V40/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,662
App. No.
17/947,758
Granted
May 19, 2026
Kind
B2
Abstract

The disclosure refers to electronic apparatuses and controlling methods thereof. In an embodiment, an electronic apparatus includes an input interface, an output interface, and a processor that is communicatively coupled to the input interface and the output interface. The processor is configured to control the input interface to receive conversation data including one or more texts and one or more images. The processor is further configured to extract a first text and an image from the conversation data. The processor is further configured to identify a meaning of the conversation data based on at least one of the first text and the image. The processor is further configured to control the output interface to output the meaning of the conversation data.

Claims (68)

1 . An electronic apparatus, comprising:

a transceiver;

an output interface; and

a processor communicatively coupled to the transceiver and the output interface,

wherein the processor is configured to:

control the transceiver to receive conversation data comprising one or more texts and one or more images;

extract a first text and an image from the one or more texts and the one or more images of the conversation data, the image comprising an uniform record locator (URL);

connect to a web server to acquire main information and a screen of the URL;

identify a meaning of the conversation data based on at least one of the first text, the image, and the main information of the URL; and

control the output interface to output the meaning of the conversation data and the screen of the URL,

wherein to identify the meaning of the conversation data comprises to:

perform a first attempt to identify the meaning of the conversation data based on the first text excluding the image;

determine, by an artificial intelligence neural network model and using a language model, an importance of the image based on the meaning of the conversation data;

based on at least one of the meaning of the conversation data not being identified by the first attempt or the importance of the image exceeding a predetermined threshold:

extract coordinates for an area of the image;

replace the area of the image with a predetermined second text based on the coordinates, the predetermined second text being determined, by the artificial intelligence neural network model, using the language model and based on a previous text of the conversation data; and

perform a second attempt to identify the meaning of the conversation data based on the first text and the predetermined second text;

based on the meaning of the conversation data not being identified by the first attempt and the second attempt:

recognize a third text included in the image by using an optical character recognition (OCR) method; and

perform a third attempt to identify the meaning of the conversation data based on the first text and the third text.

2 . The electronic apparatus of claim 1 , wherein the processor is further configured to:

identify the meaning of the conversation data based on the first text and the image.

3 . The electronic apparatus of claim 1 ,

wherein the predetermined second text comprises at least one of a positive response text, a negative response text, and a text related to the first text.

4 . The electronic apparatus of claim 1 , wherein the processor is further configured to:

based on the first text being a first language and the third text being a second language different from the first language, translate the third text into the first language.

5 . The electronic apparatus of claim 1 , wherein the processor is further configured to:

based on the meaning of the conversation data not being identified by the first attempt, the second attempt, and the third attempt:

extract caption information from the image; and

perform a fourth attempt to identify the meaning of the conversation data based on the first text and the caption information.

6 . The electronic apparatus of claim 5 , wherein the processor is further configured to:

based on the meaning of the conversation data not being identified by the first attempt, the second attempt, the third attempt, and the fourth attempt:

recognize a facial expression included in the image;

identify an emotion corresponding to the facial expression; and

identify the meaning of the conversation data based on the first text and the emotion.

7 . The electronic apparatus of claim 1 , wherein to control the output interface to output the meaning of the conversation data comprises to:

control the output interface to output the meaning of the conversation data as an output image.

8 . The electronic apparatus of claim 1 ,

wherein the image includes at least one of an emoticon, a thumbnail image, and a meme.

9 . A controlling method of an electronic apparatus, the controlling method comprising:

receiving, by a transceiver of the electronic apparatus, conversation data comprising one or more texts and one or more images;

extracting a first text and an image from the one or more texts and the one or more images of the conversation data, the image comprising an uniform record locator (URL);

connecting to a web server and acquiring main information and a screen of the URL;

identifying a meaning of the conversation data based on at least one of the first text, the image, and the main information of the URL; and

outputting, to an output interface of the electronic apparatus, the meaning of the conversation data and the screen of the URL,

wherein the identifying of the meaning of the conversation data comprises:

performing a first attempt to identify the meaning of the conversation data based on the first text excluding the image;

determining, by an artificial intelligence neural network model and using a language model, an importance of the image based on the meaning of the conversation data;

based on at least one of the meaning of the conversation data not being identified by the first attempt or the importance of the image exceeding a predetermined threshold:

extracting coordinates for an area of the image;

replacing the area of the image with a predetermined second text based on the coordinates, the predetermined second text being determined, by the artificial intelligence neural network model, using the language model and based on a previous text of the conversation data; and

performing a second attempt to identify the meaning of the conversation data based on the first text and the predetermined second text;

based on the meaning of the conversation data not being identified by the first attempt and the second attempt:

recognizing a third text included in the image by using an optical character recognition (OCR) method; and

performing a third attempt to identify the meaning of the conversation data based on the first text and the third text.

10 . The controlling method of claim 9 , wherein the identifying of the meaning of the conversation data comprises:

identifying the meaning of the conversation data based on the first text and the image.

11 . The controlling method of an electronic apparatus of claim 9 ,

wherein the predetermined second text comprises at least one of a positive response text, a negative response text, and a text related to the first text.

12 . The controlling method of claim 9 , wherein the identifying of the meaning of the conversation data further comprises:

based on the meaning of the conversation data not being identified by the first attempt, the second attempt, and the third attempt:

extracting caption information from the image; and

performing a fourth attempt to identify the meaning of the conversation data based on the first text and the caption information.

13 . The controlling method of claim 12 , wherein the identifying of the meaning of the conversation data further comprises:

based on the meaning of the conversation data not being identified by the first attempt, the second attempt, the third attempt, and the fourth attempt:

recognizing a facial expression included in the image;

identifying an emotion corresponding to the facial expression; and

identifying the meaning of the conversation data based on the first text and the emotion.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2022
From: CHOI, HYUNGTAK; KIM, MUNJO; RYU, SEONGHAN; KWAK, SEJIN; RAVURU, LOHITH; YANG, HAEHUN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061139/0839 →
Priority Claims (1)
KR 10-2021-0138349 · Oct 18, 2021 · national
Continuity (2)
Continuation PCTKR2022010639 · Jul 20, 2022
Related Publication 20230020143A1 · Jan 19, 2023
References Cited (27)
US 9740677B2 · Kim et al. · 2017 [cited by applicant]
US 10986046B2 · Kim · 2021 [cited by applicant]
US 20140067397A1 · Radebaugh · 2014 [cited by applicant]
US 20140163957A1 · Tesch · 2014 [cited by examiner]
US 20140245115A1 · Zhang · 2014 [cited by examiner]
US 20160210962A1 · Kim et al. · 2016 [cited by applicant]
US 20190012376A1 · Fujimoto · 2019 [cited by examiner]
US 20190197102A1 · Lerner et al. · 2019 [cited by applicant]
US 20190386937A1 · Kim · 2019 [cited by applicant]
US 20200026766A1 · Ji · 2020 [cited by examiner]
US 20200036831A1 · Kim · 2020 [cited by examiner]
US 20210103610A1 · Lee · 2021 [cited by examiner]
US 20220182343A1 · Lee · 2022 [cited by examiner]
US 20240106769A1 · Lee · 2024 [cited by examiner]
CN 110209897A · 2019 [cited by applicant]
KR 1020160010746A · 2016 [cited by applicant]
KR 1020170061647A · 2017 [cited by applicant]
KR 101964514B1 · 2019 [cited by applicant]
KR 1020190096304A · 2019 [cited by applicant]
KR 1020190123093A · 2019 [cited by applicant]
KR 1020200025177A · 2020 [cited by applicant]
KR 1020200048716A · 2020 [cited by applicant]
KR 1020210039610A · 2021 [cited by applicant]
KR 1020210039618A · 2021 [cited by applicant]
Communication dated Oct. 28, 2022, issued by the International Searching Authority in counterpart International Application No. PCT/KR2022/010639 (PCT/ISA/210). [cited by applicant]
Communication dated Oct. 28, 2022, issued by the International Searching Authority in counterpart International Application No. PCT/KR2022/010639 (PCT/ISA/237). [cited by applicant]
Communication dated Sep. 12, 2024, issued by the European Patent Office for European Patent Application No. 22883725.8. [cited by applicant]