IP Library Granted Patent US 10,628,668
Granted Patent B2
US 10,628,668 · App. 16/058,476 · Granted Apr 21, 2020

Systems and methods for generating and using semantic images in deep learning for classification and data extraction

Inventor: Uwe Ast (Constance, DE)
Assignee: Open Text SA ULC
G06K9/00463G06F17/2785G06K9/00442G06K9/00456G06K9/00469G06K9/6267G06N3/0427G06N3/0454G06N3/08G06N5/046G06N20/00G06K2209/01G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,628,668
App. No.
16/058,476
Granted
Apr 21, 2020
Kind
B2
Abstract

Disclosed is a new document processing solution that combines the powers of machine learning and deep learning and leverages the knowledge of a knowledge base. Textual information in an input image of a document can be converted to semantic information utilizing the knowledge base. A semantic image can then be generated utilizing the semantic information and geometries of the textual information. The semantic information can be coded by semantic type determined utilizing the knowledge base and positioned in the semantic image utilizing the geometries of the textual information. A region-based convolutional neural network (R-CNN) can be trained to extract regions from the semantic image utilizing the coded semantic information and the geometries. The regions can be mapped to the textual information for classification/data extraction. With semantic images, the number of samples and time needed to train the R-CNN for document processing can be significantly reduced.

Claims (47)

1. A system comprising:

a processor;

a non-transitory computer-readable medium; and

stored instructions translatable by the processor to perform:

receiving an input image of a document;

converting textual information in the document to semantic information utilizing a language knowledge base;

generating a semantic image utilizing the semantic information and geometries of the textual information, the semantic information coded in the semantic image by semantic type determined utilizing the language knowledge base, the coded semantic information positioned in the semantic image utilizing the geometries of the textual information;

extracting regions from the semantic image utilizing a machine learning engine with a deep learning feedback loop, the deep learning feedback loop including a field type knowledge base and a region-based convolutional neural network, the field type knowledge base containing field-specific reference data, the region-based convolutional neural network configured for recognizing regions of interest in semantic images utilizing the field-specific reference data in the field type knowledge base;

mapping the regions extracted from the semantic image to the textual information; and

extracting data from the document by the regions.

2. The system of claim 1 , wherein the semantic image contains no text.

3. The system of claim 1 , wherein the stored instructions are further translatable by the processor to perform:

classifying the document based on the regions extracted from the semantic image.

4. The system of claim 1 , wherein the stored instructions are further translatable by the processor to perform:

generating the textual information with the geometries utilizing optical character recognition.

5. The system of claim 1 , wherein the region-based convolutional neural network is configured for grouping pieces of coded semantic information into a region by proximity in position and in semantic type.

6. The system of claim 1 , wherein the semantic image comprises a plurality of semantic layers, wherein each semantic layer of the semantic layers contains a subset of the regions, and wherein each semantic layer represents a unique semantic type of the coded semantic information.

7. The system of claim 1 , wherein the semantic information is coded in the semantic image by colors, shades, textures, intensities, or semantic codes associated with different semantic types.

8. A method comprising:

receiving, by a computer, an input image of a document;

converting, by the computer, textual information in the document to semantic information utilizing a language knowledge base;

generating, by the computer, a semantic image utilizing the semantic information and geometries of the textual information, the semantic information coded in the semantic image by semantic type determined by the computer utilizing the language knowledge base, the coded semantic information positioned in the semantic image utilizing the geometries of the textual information;

extracting, by the computer, regions from the semantic image utilizing a machine learning engine with a deep learning feedback loop, the deep learning feedback loop including a field type knowledge base and a region-based convolutional neural network, the field type knowledge base containing field-specific reference data, the region-based convolutional neural network configured for recognizing regions of interest in semantic images utilizing the field-specific reference data in the field type knowledge base;

mapping, by the computer, the regions extracted from the semantic image to the textual information; and

extracting, by the computer, data from the document by the regions.

9. The method according to claim 8 , wherein the semantic image contains no text.

10. The method according to claim 8 , further comprising:

classifying the document based on the regions extracted from the semantic image.

11. The method according to claim 8 , further comprising:

generating the textual information with the geometries utilizing optical character recognition.

12. The method according to claim 8 , wherein the region-based convolutional neural network is configured for grouping pieces of coded semantic information into a region by proximity in position and in semantic type.

13. The method according to claim 8 , wherein the semantic image comprises a plurality of semantic layers, wherein each semantic layer of the semantic layers contains a subset of the regions, and wherein each semantic layer represents a unique semantic type of the coded semantic information.

14. The method according to claim 8 , wherein the semantic information is coded in the semantic image by colors, shades, textures, intensities, or semantic codes associated with different semantic types.

15. A computer program product comprising a non-transitory computer-readable medium storing instructions translatable by a processor to perform:

receiving an input image of a document;

converting textual information in the document to semantic information utilizing a language knowledge base;

generating a semantic image utilizing the semantic information and geometries of the textual information, the semantic information coded in the semantic image by semantic type determined utilizing the language knowledge base, the coded semantic information positioned in the semantic image utilizing the geometries of the textual information;

extracting regions from the semantic image utilizing a machine learning engine with a deep learning feedback loop, the deep learning feedback loop including a field type knowledge base and a region-based convolutional neural network, the field type knowledge base containing field-specific reference data, the region-based convolutional neural network configured for recognizing regions of interest in semantic images utilizing the field-specific reference data in the field type knowledge base;

mapping the regions extracted from the semantic image to the textual information; and

extracting data from the document by the regions.

16. The computer program product of claim 15 , wherein the semantic image contains no text.

17. The computer program product of claim 15 , wherein the instructions are further translatable by the processor to perform:

classifying the document based on the regions extracted from the semantic image.

18. The computer program product of claim 15 , wherein the instructions are further translatable by the processor to perform:

generating the textual information with the geometries utilizing optical character recognition.

19. The computer program product of claim 15 , wherein the region-based convolutional neural network is configured for grouping pieces of coded semantic information into a region by proximity in position and in semantic type.

20. The computer program product of claim 15 , wherein the semantic image comprises a plurality of semantic layers, wherein each semantic layer of the semantic layers contains a subset of the regions, and wherein each semantic layer represents a unique semantic type of the coded semantic information.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2024
From: OPEN TEXT CORP.
To: CROWDSTRIKE, INC.
Reel/Frame 068122/0068 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2024
From: OPEN TEXT SA ULC
To: OPEN TEXT CORP.
Reel/Frame 067400/0102 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 5, 2018
From: OPEN TEXT SOFTWARE GMBH
To: OPEN TEXT SA ULC
Reel/Frame 046786/0381 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2018
From: AST, UWE
To: OPEN TEXT SOFTWARE GMBH
Reel/Frame 046588/0259 →
Cited By (1)
US 12,367,667