Method and electronic device for recognizing text in image
A method and an electronic device for recognizing text are provided. The method includes detecting positions of pieces of text included in the text in the image, generating cropped images by cropping areas corresponding to the pieces of text in the image, recognizing characters of the pieces of text based on the cropped images, generating a sentence by inputting the positions of the pieces of text and the characters of the pieces of text to a multimodal language model, wherein the multimodal language model is an artificial intelligence (AI) model for inferring an original sentence of the text, and displaying the sentence.
1 . A method, performed by an electronic device, of recognizing a text in an image, the method comprising:
detecting spatial positions of pieces of text included in the text in the image;
generating cropped images by cropping an area for each of the pieces of text in the image based on the detected spatial positions;
recognizing characters of the pieces of text for each of the cropped images;
generating a sentence by inputting the characters and respective spatial positions of each piece of text into a multimodal language model ; and
displaying the sentence,
wherein the multimodal language model is an artificial intelligence (AI) model for inferring an original sentence of the text,
wherein the multimodal language model includes a plurality of sub-networks having different layers appropriate for processing different input modality data, the sub-networks including a first sub-network configured to calculate first input modality data comprising positions of text and a second sub-network configured to calculate second input modality data comprising characters of text, and
wherein the multimodal language model applies a first weight to the first sub-network and a second weight to the second sub-network, and obtains the sentence by applying the first weight to the positions of the text and the second weight to the characters of the text.
2 . The method of claim 1 , wherein the multimodal language model has been trained based on a training data set including positions of a sentence and words in the sentence.
3 . The method of claim 2 , wherein the detecting of the positions of the pieces of text comprises obtaining data indicating the positions of the pieces of text by applying the image to a text detection model.
4 . The method of claim 3 , wherein the recognizing of the characters of the pieces of text comprises obtaining the characters of the pieces of text corresponding to each of the cropped images, respectively, by applying each of the cropped images to a text recognition model.
5 . The method of claim 1 , further comprising:
generating a text-position set by matching a character of a first piece of text of the text with a position of the first piece of text and matching a character of a second piece of text of the text with a position of the second piece of text,
wherein the generating of the sentence comprises inputting the text-position set to the multimodal language model.
6 . The method of claim 5 , further comprising:
indexing the text-position set,
wherein the inputting of the text-position set to the multimodal language model comprises further inputting an index of the text-position set to the multimodal language model.
7 . The method of claim 1 , wherein the generating of the sentence comprises applying a different weight to each of the positions of the pieces of text and the characters of the pieces of text.
8 . The method of claim 1 ,
wherein the displaying of the sentence comprises separately displaying elements of the sentence, and
wherein the elements of the sentence comprising at least one of a subject, an object, or a verb.
9 . The method of claim 8 , wherein the displaying of the sentence further comprises displaying a recommended word for replacing a word in the sentence in order to modify a grammar or spelling error of the sentence.
10 . An electronic device for recognizing a text in an image, the electronic device comprising:
a display;
a memory storing one or more instructions; and
at least one processor configured to execute the one or more instructions stored in the memory to:
detect spatial positions of pieces of text included in the text in the image,
generate cropped images by cropping an area for each of the pieces of text in the image based on the detected spatial positions,
recognize characters of the pieces of text for each of the cropped images,
generate a sentence by inputting the characters and respective spatial positions of each piece of text into a multimodal language model, and
control the display to display the sentence,
wherein the multimodal language model is an artificial intelligence (AI) model for inferring an original sentence of the text,
wherein the multimodal language model includes a plurality of sub-networks having different layers appropriate for processing different input modality data, the sub-networks including a first sub-network configured to calculate first input modality data comprising positions of text and a second sub-network configured to calculate second input modality data comprising characters of text, and
wherein the multimodal language model applies a first weight to the first sub-network and a second weight to the second sub-network, and obtains the sentence by applying the first weight to the positions of the text and the second weight to the characters of the text.
11 . The electronic device of claim 10 , wherein the multimodal language model has been trained based on a training data set including positions of a sentence and words in the sentence.
12 . The electronic device of claim 11 , wherein the at least one processor is further configured to execute the one or more instructions to obtain data indicating the positions of the pieces of text by applying the image to a text detection model.
13 . The electronic device of claim 12 , wherein the at least one processor is further configured to execute the one or more instructions to obtain the characters of the pieces of text corresponding to each of the cropped images, respectively, by applying each of the cropped images to a text recognition model.
14 . The electronic device of claim 10 , wherein the at least one processor is further configured to execute the one or more instructions to:
generate a text-position set by matching a character of a first piece of text of the text with a position of the first piece of text and matching a character of a second piece of text of the text with a position of the second piece of text, and
input the text-position set to the multimodal language model.
15 . The electronic device of claim 14 , wherein the at least one processor is further configured to execute the one or more instructions to:
index the text-position set, and
input an index of the text-position set to the multimodal language model.
16 . The electronic device of claim 10 ,
wherein the at least one processor is further configured to execute the one or more instructions to separately display elements of the sentence, and
wherein the elements of the sentence comprising at least one of a subject, an object, or a verb.
17 . The electronic device of claim 16 , wherein the at least one processor is further configured to execute the one or more instructions to control the display to display a recommended word for replacing a word in the sentence in order to modify a grammar or spelling error of the sentence.
18 . One or more non-transitory computer-readable storage media storing one or more computer programs including computer-executable instructions that, when executed by one or more processors of an electronic device individually or collectively, cause the electronic device to perform operations for recognizing a text in an image, the operations comprising:
detecting spatial positions of pieces of text included in the text in the image;
generating cropped images by cropping an area for each of the pieces of text in the image based on the detected spatial positions;
recognizing characters of the pieces of text for each of the cropped images;
generating a sentence by inputting the characters and respective spatial positions of each piece of text into a multimodal language model; and
displaying the sentence,
wherein the multimodal language model is an artificial intelligence (AI) model for inferring an original sentence of the text,
wherein the multimodal language model includes a plurality of sub-networks having different layers appropriate for processing different input modality data, the sub-networks including a first sub-network configured to calculate first input modality data comprising positions of text and a second sub-network configured to calculate second input modality data comprising characters of text, and
wherein the multimodal language model applies a first weight to the first sub-network and a second weight to the second sub-network, and obtains the sentence by applying the first weight to the positions of the text and the second weight to the characters of the text.