Electronic apparatus and controlling method thereof
The disclosure refers to electronic apparatuses and controlling methods thereof. In an embodiment, an electronic apparatus includes an input interface, an output interface, and a processor that is communicatively coupled to the input interface and the output interface. The processor is configured to control the input interface to receive conversation data including one or more texts and one or more images. The processor is further configured to extract a first text and an image from the conversation data. The processor is further configured to identify a meaning of the conversation data based on at least one of the first text and the image. The processor is further configured to control the output interface to output the meaning of the conversation data.
1 . An electronic apparatus, comprising:
a transceiver;
an output interface; and
a processor communicatively coupled to the transceiver and the output interface,
wherein the processor is configured to:
control the transceiver to receive conversation data comprising one or more texts and one or more images;
extract a first text and an image from the one or more texts and the one or more images of the conversation data, the image comprising an uniform record locator (URL);
connect to a web server to acquire main information and a screen of the URL;
identify a meaning of the conversation data based on at least one of the first text, the image, and the main information of the URL; and
control the output interface to output the meaning of the conversation data and the screen of the URL,
wherein to identify the meaning of the conversation data comprises to:
perform a first attempt to identify the meaning of the conversation data based on the first text excluding the image;
determine, by an artificial intelligence neural network model and using a language model, an importance of the image based on the meaning of the conversation data;
based on at least one of the meaning of the conversation data not being identified by the first attempt or the importance of the image exceeding a predetermined threshold:
extract coordinates for an area of the image;
replace the area of the image with a predetermined second text based on the coordinates, the predetermined second text being determined, by the artificial intelligence neural network model, using the language model and based on a previous text of the conversation data; and
perform a second attempt to identify the meaning of the conversation data based on the first text and the predetermined second text;
based on the meaning of the conversation data not being identified by the first attempt and the second attempt:
recognize a third text included in the image by using an optical character recognition (OCR) method; and
perform a third attempt to identify the meaning of the conversation data based on the first text and the third text.
2 . The electronic apparatus of claim 1 , wherein the processor is further configured to:
identify the meaning of the conversation data based on the first text and the image.
3 . The electronic apparatus of claim 1 ,
wherein the predetermined second text comprises at least one of a positive response text, a negative response text, and a text related to the first text.
4 . The electronic apparatus of claim 1 , wherein the processor is further configured to:
based on the first text being a first language and the third text being a second language different from the first language, translate the third text into the first language.
5 . The electronic apparatus of claim 1 , wherein the processor is further configured to:
based on the meaning of the conversation data not being identified by the first attempt, the second attempt, and the third attempt:
extract caption information from the image; and
perform a fourth attempt to identify the meaning of the conversation data based on the first text and the caption information.
6 . The electronic apparatus of claim 5 , wherein the processor is further configured to:
based on the meaning of the conversation data not being identified by the first attempt, the second attempt, the third attempt, and the fourth attempt:
recognize a facial expression included in the image;
identify an emotion corresponding to the facial expression; and
identify the meaning of the conversation data based on the first text and the emotion.
7 . The electronic apparatus of claim 1 , wherein to control the output interface to output the meaning of the conversation data comprises to:
control the output interface to output the meaning of the conversation data as an output image.
8 . The electronic apparatus of claim 1 ,
wherein the image includes at least one of an emoticon, a thumbnail image, and a meme.
9 . A controlling method of an electronic apparatus, the controlling method comprising:
receiving, by a transceiver of the electronic apparatus, conversation data comprising one or more texts and one or more images;
extracting a first text and an image from the one or more texts and the one or more images of the conversation data, the image comprising an uniform record locator (URL);
connecting to a web server and acquiring main information and a screen of the URL;
identifying a meaning of the conversation data based on at least one of the first text, the image, and the main information of the URL; and
outputting, to an output interface of the electronic apparatus, the meaning of the conversation data and the screen of the URL,
wherein the identifying of the meaning of the conversation data comprises:
performing a first attempt to identify the meaning of the conversation data based on the first text excluding the image;
determining, by an artificial intelligence neural network model and using a language model, an importance of the image based on the meaning of the conversation data;
based on at least one of the meaning of the conversation data not being identified by the first attempt or the importance of the image exceeding a predetermined threshold:
extracting coordinates for an area of the image;
replacing the area of the image with a predetermined second text based on the coordinates, the predetermined second text being determined, by the artificial intelligence neural network model, using the language model and based on a previous text of the conversation data; and
performing a second attempt to identify the meaning of the conversation data based on the first text and the predetermined second text;
based on the meaning of the conversation data not being identified by the first attempt and the second attempt:
recognizing a third text included in the image by using an optical character recognition (OCR) method; and
performing a third attempt to identify the meaning of the conversation data based on the first text and the third text.
10 . The controlling method of claim 9 , wherein the identifying of the meaning of the conversation data comprises:
identifying the meaning of the conversation data based on the first text and the image.
11 . The controlling method of an electronic apparatus of claim 9 ,
wherein the predetermined second text comprises at least one of a positive response text, a negative response text, and a text related to the first text.
12 . The controlling method of claim 9 , wherein the identifying of the meaning of the conversation data further comprises:
based on the meaning of the conversation data not being identified by the first attempt, the second attempt, and the third attempt:
extracting caption information from the image; and
performing a fourth attempt to identify the meaning of the conversation data based on the first text and the caption information.
13 . The controlling method of claim 12 , wherein the identifying of the meaning of the conversation data further comprises:
based on the meaning of the conversation data not being identified by the first attempt, the second attempt, the third attempt, and the fourth attempt:
recognizing a facial expression included in the image;
identifying an emotion corresponding to the facial expression; and
identifying the meaning of the conversation data based on the first text and the emotion.