IP Library Granted Patent US 11,854,285
Granted Patent B2
US 11,854,285 · App. 17/316,318 · Granted Dec 26, 2023

Neural network architecture for extracting information from documents

Inventor: Dasaprakash Krishnamurthy (Chennai, IN)
Assignee: UST Global (Singapore) Pte. Ltd.
G06V30/412G06N3/08G06V30/413G06V30/43
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,854,285
App. No.
17/316,318
Granted
Dec 26, 2023
Kind
B2
Abstract

A system to extract data from regions of interest on a document is provided. The system includes a storage device storing an image derived from a document having text information. The system includes a document importer operable to perform optical character recognition to convert image data in the image to machine readable data. The system includes a neural network that identifies at least one region of interest on the image to classify an area of the at least one region of interest as a table. The neural network is operable to take as input the machine readable data and the image and combine both the machine readable data and the image to determine that the classified area is the table.

Claims (48)

1. A system to extract data from regions of interest on a document, the system comprising:

a storage device storing an image derived from a document having text information;

a document importer operable to perform optical character recognition to convert image data in the image to machine readable data; and

a neural network that identifies at least one region of interest on the image to classify an area of the at least one region of interest as a table or column, the neural network taking in as input the machine readable data and the image and combining both the machine readable data and the image to determine that the classified area is the table or column,

wherein the neural network includes:

a first convolutional neural network backbone configured to process the machine readable data,

a second convolutional neural network backbone configured to process the image,

a region proposal network configured to determine the at least one region of interest, and

a region of interest pool layer configured to match dimensionality of the region proposal network and a fully connected layer,

the fully connected layer configured to classify the area as the table or column.

2. The system of claim 1 , wherein the neural network further includes a 1×1 convolution that combines outputs of the first and the second convolutional neural network backbones.

3. The system of claim 1 , wherein the first and the second convolutional neural network backbones include a VGG backbone, an AlexNet backbone, or a ResNet backbone.

4. The system of claim 1 , wherein the first and the second convolutional neural network are initialized with pre-trained weights of a ResNet-101 backbone.

5. The system of claim 1 , wherein the image is obtained from a scanner or a photograph of the document.

6. The system of claim 1 , wherein the neural network further classifies a second area of the at least one region of interest as a column.

7. The system of claim 1 , wherein the neural network predicts a bounding box for the area classified as the table.

8. The system of claim 1 , wherein the fully connected layer is further configured to indicate size of the table or column.

9. The system of claim 1 , further comprising:

an extractor configured to parse text from the machine readable data, the parsed text corresponding to text within the classified area.

10. A method of extracting data from regions of interest from a document, the method comprising:

retrieving from a storage device an image of the document;

converting the image to machine readable data using optical character recognition of a data importer;

combining the machine readable data and the image using a neural network, the neural network including a first convolutional neural network backbone to process the machine readable data and a second convolutional neural network backbone to process the image;

identifying, by a region proposal network of the neural network, at least one region of interest on the image using the combined machine readable data and image; and

classifying, using a fully connected layer of the neural network, an area of the at least one region of interest as a table or a column, the neural network further including a region of interest pool layer for matching dimensionality of the region proposal network and the fully connected layer.

11. The method of claim 10 , wherein the neural network further includes a 1×1 convolution that combines outputs of the first and the second convolutional neural network backbones.

12. The method of claim 10 , wherein the first and the second convolutional neural network backbones include a VGG backbone, an AlexNet backbone, or a ResNet backbone.

13. The method of claim 10 , further comprising:

initializing the first and the second convolutional neural network with pre-trained weights of a ResNet-101 backbone.

14. The method of claim 10 , wherein the image is obtained from a scanner or a photograph of the document.

15. The method of claim 10 , further comprising:

classifying, by the neural network, a second area of the at least one region of interest as a column.

16. The method of claim 10 , further comprising indicating, using the fully connected layer of the neural network, a size of the table or column.

17. The method of claim 10 , further comprising extracting parsed text from the machine readable data, the parsed text corresponding to text within the classified area.

18. A system to extract data from regions of interest on a document, the system comprising:

a storage device storing an image derived from a document having text information;

a document importer operable to

perform optical character recognition to convert image data in the image to machine readable data, and

perform semantic enrichment on the machine readable data to obtain a semantically enriched document, the semantic enrichment involving extracting labels of fields from the machine readable data; and

a neural network including a first convolutional neural network, a second convolutional neural network, and a fully connected layer, the neural network operable to:

preprocess the semantically enriched document using the first convolutional neural network;

preprocess the image using the second convolutional neural network;

concatenate outputs of the first convolutional neural network and the second convolutional neural network;

identify, from the concatenated outputs, at least one region of interest on the image using the combined machine readable data and image; and

classify an area of the at least one region of interest as a table using the fully connected layer of the neural network.

19. The method of claim 18 , wherein the neural network is further operable to indicate size of the table or column.

20. The system of claim 18 , further comprising:

an extractor configured to parse text from the machine readable data, the parsed text corresponding to text within the classified area.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2025
From: UST GLOBAL (SINGAPORE) PTE. LIMITED
To: UST GLOBAL PRIVATE LIMITED
Reel/Frame 072012/0778 →
SECURITY INTEREST Recorded Aug 13, 2025
From: UST GLOBAL PRIVATE LIMITED
To: CITIBANK, N.A., AS AGENT
Reel/Frame 072012/0804 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 1, 2023
From: KRISHNAMURTHY, DASAPRAKASH
To: UST GLOBAL (SINGAPORE) PTE. LTD.
Reel/Frame 064803/0563 →
Priority Claims (1)
IN 202111002935 · Jan 21, 2021 · national
Continuity (1)
Related Publication 20220230013A1 · Jul 21, 2022
Cited By (1)
US 12,675,674