IP Library Granted Patent US 11,250,255
Granted Patent B2
US 11,250,255 · App. 17/081,705 · Granted Feb 15, 2022

Systems and methods for generating and using semantic images in deep learning for classification and data extraction

Inventor: Uwe Ast (Constance, DE)
Assignee: OPEN TEXT SA ULC
G06K9/00463G06F40/30G06K9/00442G06K9/00456G06K9/00469G06K9/6267G06N3/0427G06N3/0454G06N3/08G06N5/046G06N20/00G06K2209/01G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,250,255
App. No.
17/081,705
Granted
Feb 15, 2022
Kind
B2
Abstract

Disclosed is a new document processing solution that combines the powers of machine learning and deep learning and leverages the knowledge of a knowledge base. Textual information in an input image of a document can be converted to semantic information utilizing the knowledge base. A semantic image can then be generated utilizing the semantic information and geometries of the textual information. The semantic information can be coded by semantic type determined utilizing the knowledge base and positioned in the semantic image utilizing the geometries of the textual information. A region-based convolutional neural network (R-CNN) can be trained to extract regions from the semantic image utilizing the coded semantic information and the geometries. The regions can be mapped to the textual information for classification/data extraction. With semantic images, the number of samples and time needed to train the R-CNN for document processing can be significantly reduced.

Claims (55)

1. A method, comprising:

obtaining, by a computer, textual information and geometries of the textual information from a document image;

converting, by the computer, the textual information to semantic information utilizing a language knowledge base;

generating, by the computer utilizing the semantic information, a semantic image, the generating comprising positioning the semantic information in the semantic image utilizing the geometries of the textual information from the document image;

extracting, by the computer, regions from the semantic image, each of the regions representing a portion of the semantic information in the semantic image;

mapping, by the computer, the regions extracted from the semantic image to data blocks on a text layer of the document image; and

extracting, by the computer utilizing the regions extracted from the semantic image and the geometries of the textual information, text data from the data blocks on the text layer of the document image.

2. The method according to claim 1 , further comprising:

determining, based at least in part on the text data extracted from the data blocks on the text layer of the document image utilizing the regions extracted from the semantic image and the geometries of the textual information, a document classification for the document image and a confidence level for the document classification.

3. The method according to claim 2 , further comprising:

determining whether the confidence level for the document classification meets a minimum threshold; and

depending upon whether the confidence level for the document classification meets the minimum threshold, providing the document classification to a computing facility or a machine learning training process.

4. The method according to claim 1 , further comprising:

updating a database to include the text data extracted from the data blocks on the text layer of the document image.

5. The method according to claim 1 , wherein the converting the textual information to the semantic information comprises determining a code for each text string in the textual information, the code corresponding to a semantic type.

6. The method according to claim 1 , further comprising:

training an artificial neural network to recognize semantic types in semantic images.

7. The method according to claim 1 , wherein the extracting the regions from the semantic image comprises providing the semantic image to a region-based convolutional neural network trained to identify regions of interest from semantic images.

8. A system, comprising:

a processor;

a non-transitory computer-readable medium; and

stored instructions translatable by the processor for:

obtaining textual information and geometries of the textual information from a document image;

converting the textual information to semantic information utilizing a language knowledge base;

generating, utilizing the semantic information, a semantic image, the generating comprising positioning the semantic information in the semantic image utilizing the geometries of the textual information from the document image;

extracting regions from the semantic image, each of the regions representing a portion of the semantic information in the semantic image;

mapping the regions extracted from the semantic image to data blocks on a text layer of the document image; and

extracting, utilizing the regions extracted from the semantic image and the geometries of the textual information, text data from the data blocks on the text layer of the document image.

9. The system of claim 8 , wherein the stored instructions are further translatable by the processor for:

determining, based at least in part on the text data extracted from the data blocks on the text layer of the document image utilizing the regions extracted from the semantic image and the geometries of the textual information, a document classification for the document image and a confidence level for the document classification.

10. The system of claim 9 , wherein the stored instructions are further translatable by the processor for:

determining whether the confidence level for the document classification meets a minimum threshold; and

depending upon whether the confidence level for the document classification meets the minimum threshold, providing the document classification to a computing facility or a machine learning training process.

11. The system of claim 8 , wherein the stored instructions are further translatable by the processor for:

updating a database to include the text data extracted from the data blocks on the text layer of the document image.

12. The system of claim 8 , wherein the converting the textual information to the semantic information comprises determining a code for each text string in the textual information, the code corresponding to a semantic type.

13. The system of claim 8 , wherein the stored instructions are further translatable by the processor for:

training an artificial neural network to recognize semantic types in semantic images.

14. The system of claim 8 , wherein the extracting the regions from the semantic image comprises providing the semantic image to a region-based convolutional neural network trained to identify regions of interest from semantic images.

15. A computer program product comprising a non-transitory computer-readable medium storing instructions translatable by a processor for:

obtaining textual information and geometries of the textual information from a document image;

converting the textual information to semantic information utilizing a language knowledge base;

generating, utilizing the semantic information, a semantic image, the generating comprising positioning the semantic information in the semantic image utilizing the geometries of the textual information from the document image;

extracting regions from the semantic image, each of the regions representing a portion of the semantic information in the semantic image;

mapping the regions extracted from the semantic image to data blocks on a text layer of the document image; and

extracting, utilizing the regions extracted from the semantic image and the geometries of the textual information, text data from the data blocks on the text layer of the document image.

16. The computer program product of claim 15 , wherein the instructions are further translatable by the processor for:

determining, based at least in part on the text data extracted from the data blocks on the text layer of the document image utilizing the regions extracted from the semantic image and the geometries of the textual information, a document classification for the document image and a confidence level for the document classification.

17. The computer program product of claim 16 , wherein the instructions are further translatable by the processor for:

determining whether the confidence level for the document classification meets a minimum threshold; and

depending upon whether the confidence level for the document classification meets the minimum threshold, providing the document classification to a computing facility or a machine learning training process.

18. The computer program product of claim 15 , wherein the instructions are further translatable by the processor for:

updating a database to include the text data extracted from the data blocks on the text layer of the document image.

19. The computer program product of claim 15 , wherein the converting the textual information to the semantic information comprises determining a code for each text string in the textual information, the code corresponding to a semantic type.

20. The computer program product of claim 15 , wherein the extracting the regions from the semantic image comprises providing the semantic image to a region-based convolutional neural network trained to identify regions of interest from semantic images.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2024
From: OPEN TEXT CORP.
To: CROWDSTRIKE, INC.
Reel/Frame 068122/0068 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2024
From: OPEN TEXT SA ULC
To: OPEN TEXT CORP.
Reel/Frame 067400/0102 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2020
From: AST, UWE
To: OPEN TEXT SOFTWARE GMBH
Reel/Frame 054539/0021 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 3, 2020
From: OPEN TEXT SOFTWARE GMBH
To: OPEN TEXT SA ULC
Reel/Frame 054539/0035 →
Continuity (4)
Continuation 16842097 · Apr 7, 2020
Continuation 16058476 · Aug 8, 2018
Provisional Application 62543246 · Aug 9, 2017
Related Publication 20210073533A1 · Mar 11, 2021
Cited By (2)
US 12,327,424 US 12,406,516