IP Library Granted Patent US 12,154,361
Granted Patent B2
US 12,154,361 · App. 17/313,755 · Granted Nov 26, 2024

Method and apparatus of image-to-document conversion based on OCR, device, and readable storage medium

Inventors: Xingyao Chen (Shenzhen Guangdong, CN); Canlu Huang (Shenzhen Guangdong, CN); Wencan Hu (Shenzhen Guangdong, CN); Yidong Chen (Guangdong, CN); Hanquan Lin (Guangdong, CN); Fei Huang (Guangdong, CN); Geyang Ke (Guangdong, CN); Zhiquan Yang (Guangdong, CN)
Assignee: Tencent Technology (Shenzhen) Company Limited
G06V30/414G06F18/214G06F40/106G06V10/44G06V30/1463G06V30/15G06V30/10G06V30/146
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,154,361
App. No.
17/313,755
Granted
Nov 26, 2024
Kind
B2
Abstract

A method of image-to-document conversion based on optical character recognition (OCR) includes obtaining an image to be converted into a target document, and performing layout segmentation on the image according to image content of the image, to obtain n image layouts, each of the n image layouts corresponding to a content type, and n being a positive integer. The method also includes, for each of the n image layouts, processing image content in the respective image layout according to the content type corresponding to the respective image layout, to obtain converted content corresponding to the respective image layout. The method further includes adding the converted content corresponding to the n image layouts to an electronic document, to obtain the target document.

Claims (75)

1. A method of image-to-document conversion based on optical character recognition (OCR), the method comprising:

obtaining an image to be converted into a target document;

classifying regions of the image according to image content of the image, to obtain n image regions, each of the n image regions being classified to a corresponding content type, n being an integer greater than or equal to 3, and at least 3 of the n image regions being classified into different respective content types, wherein the n image regions are obtained by performing (i) combination processing including combining consecutive regions of the image belonging to a same content type, (ii) generating a binary tree having the n image regions as nodes, (iii) performing depth traversing of the binary tree to obtain a reading sequence, and (iv) performing intersection processing including adjusting positions of regions that intersect each other based on the reading sequence;

for each of the n image regions, processing image content in the respective image region, by processing circuitry of a server, according to the content type to which the respective image region was classified, to obtain converted content corresponding to the respective image region; and

adding the converted content corresponding to the n image regions to an electronic document, to obtain the target document.

2. The method according to claim 1 , wherein the classifying comprises:

encoding the image by using an encoder, to obtain encoded data;

decoding the encoded data by using a decoder, to obtain a mask image; and

obtaining the n image regions according to regions in the mask image.

3. The method according to claim 2 , wherein the obtaining the n image regions according to the regions in the mask image comprises:

performing correction processing for the regions in the mask image, to obtain the n image regions, the correction processing comprising denoising processing, the combination processing, and the intersection processing to generate a corrected mask image, and

the denoising processing comprising filtering out regions in the mask image whose areas are smaller than a preset area.

4. The method according to claim 3 , wherein the mask image further comprises a single-column splitter bar; and

the performing the correction processing for the regions in the mask image comprises:

splitting the corrected mask image according to the single-column splitter bar, to obtain at least two split mask images;

correcting regions in each of the at least two split mask images; and

generating rectangular boxes corresponding to the corrected regions in the split mask images as the n image regions.

5. The method according to claim 1 , wherein the content type of one of the n image regions comprises text; and

the processing image content in the respective image region comprises:

performing text recognition on the image content in the respective image region, to obtain a text recognition result of segmentation based on text lines;

determining a paragraph formation result of the text lines according to line-direction features of the text lines, the paragraph formation result representing a segmentation manner for the text recognition result, and the line-direction features comprising at least one of a line height and a line spacing; and

re-segmenting the text recognition result according to the paragraph formation result, to obtain a text conversion result corresponding to the respective image region.

6. The method according to claim 5 , wherein the determining comprises:

generating a histogram according to the line-direction features of the text lines;

setting a threshold corresponding to the line-direction features according to a distribution of the line-direction features in the histogram; and

determining a text line as a paragraph formation line in response to a determination that a line-direction feature of the text line reaches the threshold, the determination that the text line is the paragraph formation line representing that the text line is a beginning or an end of a paragraph.

7. The method according to claim 1 , wherein the content type of one of the n image regions comprises a table; and

the processing the image content in the respective image region comprises:

obtaining cells of a target table according to borders in the respective image region;

performing calculation for the image content in the respective image region, to obtain character coordinates; and

obtaining the target table as a table conversion result corresponding to the respective image region according to the character coordinates and the cells.

8. The method according to claim 7 , wherein the obtaining the cells of the target table comprises:

determining horizontal borders and vertical borders according to the borders in the respective image region, and determining intersections between the horizontal borders and the vertical borders; and

obtaining the cells of the target table according to the horizontal borders, the vertical borders, and the intersections between the horizontal borders and the vertical borders.

9. The method according to claim 8 , wherein the determining the horizontal borders and the vertical borders comprises:

recognizing the borders in the respective image region; and

obtaining the horizontal borders and the vertical borders by correcting the recognized borders in the respective image region to a horizontal direction or a vertical direction.

10. The method according to claim 9 , wherein the method further comprises:

correcting the respective image region by correcting the recognized borders in the respective image region to the horizontal direction or the vertical direction; and

the performing the calculation for the image content in the respective image region, to obtain the character coordinates comprises performing calculation for the image content in the corrected image region, to obtain the character coordinates.

11. The method according claim 1 , wherein the content type of one of the n image regions comprises a picture; and

the processing the image content in the respective image region comprises:

performing picture cropping for the image content in the respective image region, and using a picture obtained through the picture cropping as converted picture content corresponding to the respective image region.

12. The method according to claim 1 , wherein the content type of one of the n image regions comprises a formula; and

the processing the image content in the respective image region comprises:

performing picture cropping for the image content in the respective image region, and using a picture obtained through the picture cropping as converted formula content corresponding to the respective image region.

13. The method according to claim 1 , wherein the obtaining the image comprises:

obtaining a to-be-rectified image; and

inputting the to-be-rectified image to a rectification neural network configured to output the image, the rectification neural network being a network obtained through training with a simulation dataset, simulation data in the simulation dataset being data obtained after distortion processing is performed on a sample image.

14. A method of image-to-document conversion based on optical character recognition (OCR), the method comprising:

displaying a conversion interface, the conversion interface comprising a conversion control and an image selection region;

selecting an image in the image selection region, the image to be converted into a target document;

triggering a conversion function corresponding to the conversion control in response to triggering of the conversion control, the conversion function converting the image into a document format; and

displaying, by processing circuitry of a terminal device, a target document display interface, the target document display interface comprising the target document obtained after the image is converted, a typesetting manner of the target document corresponding to a typesetting manner of the image, the target document being obtained by the conversion function in a manner comprising

classifying regions of the image according to image content of the image, to obtain n image regions, each of the image regions being classified to a corresponding content type, n being an integer greater than or equal to 3, and at least 3 of the n image regions being classified into different respective content types, wherein the n image regions are obtained by performing (i) combination processing including combining consecutive regions of the image belonging to a same content type, (ii) generating a binary tree having the n image regions as nodes, (iii) performing depth traversing of the binary tree to obtain a reading sequence, and (iv) performing intersection processing including adjusting positions of regions that intersect each other based on the reading sequence;

for each of the n image regions, processing the image content in the respective image region according to the content type to which the respective image region was classified, to obtain converted content corresponding to the respective image region; and

adding the converted content corresponding to the n image regions to an electronic document, to obtain the target document.

15. The method according to claim 14 , wherein a content type in a first target region of the target document is consistent with a content type in a second target region of the image, and a position of the first target region in the target document corresponds to a position of the second target region in the image,

the content type comprising at least one of text, a picture, a table, and a formula.

16. An apparatus of image-to-document conversion based on optical character recognition (OCR), comprising:

processing circuitry configured to

obtain an image to be converted into a target document;

classifying regions of the image according to image content of the image, to obtain n image regions, each of the image regions being classified to a corresponding content type, n being an positive integer greater than or equal to 3, and at least 3 of the n image regions being classified into different respective content types, wherein the n image regions are obtained by performing (i) combination processing including combining consecutive regions of the image belonging to a same content type, (ii) generating a binary tree having the n image regions as nodes, (iii) performing depth traversing of the binary tree to obtain a reading sequence, and (iv) performing intersection processing including adjusting positions of regions that intersect each other based on the reading sequence;

for each of the image regions, process image content in the respective image region according to the content type to which the respective image region was classified, to obtain converted content corresponding to the respective image region; and

add the converted content corresponding to the n image regions to an electronic document, to obtain the target document.

17. A computer device, comprising:

a processor; and

a memory connected to the processor,

the memory storing machine-readable instructions, the machine-readable instructions being loaded and executed by the processor to implement the method of image-to-document conversion based on optical character recognition (OCR) according to claim 1 .

18. A non-transitory computer-readable storage medium, storing machine-readable instructions, the machine-readable instructions being loaded and executed by a processor to implement the method of image-to-document conversion based on optical character recognition (OCR) according to claim 1 .

19. A computer device, comprising:

a processor; and

a memory connected to the processor,

the memory storing machine-readable instructions, the machine-readable instructions being loaded and executed by the processor to implement the method of image-to-document conversion based on optical character recognition (OCR) according to claim 14 .

20. A non-transitory computer-readable storage medium, storing machine-readable instructions, the machine-readable instructions being loaded and executed by a processor to implement the method of image-to-document conversion based on optical character recognition (OCR) according to claim 14 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2021
From: CHEN, XINGYAO; HUANG, CANLU; HU, WENCAN; CHEN, YIDONG; LIN, HANQUAN; HUANG, FEI; KE, GEYANG; YANG, ZHIQUAN
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 056216/0252 →
Priority Claims (1)
CN 201910224228.1 · Mar 22, 2019 · national
Continuity (2)
Continuation PCTCN2020078181 · Mar 6, 2020
Related Publication 20210256253A1 · Aug 19, 2021
Cited By (1)
US 12,346,649