IP Library Granted Patent US 12,367,667
Granted Patent B2
US 12,367,667 · App. 18/457,498 · Granted Jul 22, 2025

Systems and methods for generating and using semantic images in deep learning for classification and data extraction

Inventor: Uwe Ast (Constance, DE)
Assignee: CrowdStrike, Inc.
G06V10/82G06F18/24G06F40/30G06N3/042G06N3/045G06N3/08G06N5/046G06N20/00G06V30/19173G06V30/40G06V30/413G06V30/414G06V30/416G06N5/022G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,667
App. No.
18/457,498
Granted
Jul 22, 2025
Kind
B2
Abstract

Disclosed is a new document processing solution that combines the powers of machine learning and deep learning and leverages the knowledge of a knowledge base. Textual information in an input image of a document can be converted to semantic information utilizing the knowledge base. A semantic image can then be generated utilizing the semantic information and geometries of the textual information. The semantic information can be coded by semantic type determined utilizing the knowledge base and positioned in the semantic image utilizing the geometries of the textual information. A region-based convolutional neural network (R-CNN) can be trained to extract regions from the semantic image utilizing the coded semantic information and the geometries. The regions can be mapped to the textual information for classification/data extraction. With semantic images, the number of samples and time needed to train the R-CNN for document processing can be significantly reduced.

Claims (55)

1. A method, comprising:

determining, by a computer utilizing a language knowledge base, classification types associated with content in a first image;

generating, by the computer, a second image of the content, the second image having representations of the content in the first image, wherein a representation includes positioning information for a piece of content represented by the representation and a code corresponding to a classification type of the piece of content;

providing, by the computer, the second image as input to a machine learning engine, wherein the machine learning engine implements a neural network trained to:

recognize the representations of the content; and

extract regions from the second image utilizing the representations; and

providing, by the computer, the regions extracted from the second image as input to one or more processors, wherein the one or more processors are operable to extract data from the first image utilizing the regions extracted from the second image.

2. The method according to claim 1 , wherein the neural network is trained to learn from the representations and output a learned result without relying on the content in the first image, the learned result comprising a classification for the first image.

3. The method according to claim 2 , further comprising:

providing, responsive to a confidence level for the learned result not meeting a threshold, the learned result to a training process, wherein the training process updates the language knowledge base utilizing the learned result.

4. The method according to claim 1 , wherein the neural network comprises a region-based convolutional neural network.

5. The method according to claim 4 , further comprising:

providing outputs from the machine learning engine to a deep learning (DL) module, wherein the DL module implements a DL algorithm to analyze the outputs from the machine learning engine;

learning from differences between the outputs from the machine learning engine and actual results containing corrected data from the first image; and

producing a set of parameters that reflect the differences thus learned.

6. The method according to claim 5 , further comprising:

providing the set of parameters to the machine learning engine, wherein the machine learning engine is operable to utilize the set of parameters to improve a performance of the machine learning engine.

7. The method according to claim 1 , wherein the second image comprises a plurality of layers, each of the plurality of layers corresponding to one of the classification types associated with content in the first image.

8. A system, comprising:

a memory; and

one or more processors, operatively couped to the memory, to;

determine, utilizing a language knowledge base, classification types associated with content in a first image;

generate a second image of the content, the second image having representations of the content in the first image, wherein a representation includes positioning information for a piece of content represented by the representation and a code corresponding to a classification type of the piece of content;

provide the second image as input to a machine learning engine, wherein the machine learning engine implements a neural network trained to:

recognize the representations of the content; and

extract regions from the second image utilizing the representations; and

provide the regions extracted from the second image as input to one or more processors, wherein the one or more processors are operable to extract data from the first image utilizing the regions extracted from the second image.

9. The system of claim 8 , wherein the neural network is trained to learn from the representations and output a learned result without relying on the content in the first image, the learned result comprises a classification for the first image.

10. The system of claim 9 , wherein the one or more processors are further to:

provide, responsive to a confidence level for the learned result not meeting a threshold, the learned result to a training process, wherein the training process updates the language knowledge base utilizing the learned result.

11. The system of claim 8 , wherein the neural network comprises a region-based convolutional neural network.

12. The system of claim 11 , wherein the one or more processors are further to:

provide outputs from the machine learning engine to a deep learning (DL) module, wherein the DL module implements a DL algorithm to analyze the outputs from the machine learning engine;

learn from differences between the outputs from the machine learning engine and actual results containing corrected data from the first image; and

produce a set of parameters that reflect the differences thus learned.

13. The system of claim 12 , wherein the one or more processors are to:

provide the set of parameters to the machine learning engine, wherein the machine learning engine is operable to utilize the set of parameters to improve its own performance.

14. The system of claim 8 , wherein the second image comprises a plurality of layers, each of the plurality of layers corresponding to one of the classification types associated with content in the first image.

15. A computer program product comprising a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:

determine, utilizing a language knowledge base, classification types associated with content in a first image;

generate a second image of the content, the second image having representations of the content in the first image, wherein a representation includes positioning information for a piece of content represented by the representation and a code corresponding to a classification type of the piece of content;

provide the second image as input to a machine learning engine, wherein the machine learning engine implements a neural network trained to:

recognize the representations of the content; and

extract regions from the second image utilizing the representations; and

provide the regions extracted from the second image as input to one or more processors, wherein the one or more processors are operable to extract data from the first image utilizing the regions extracted from the second image.

16. The computer program product of claim 15 , wherein the neural network is trained to learn from the representations and output a learned result without relying on the content in the first image, the learned result comprises a classification for the first image.

17. The computer program product of claim 16 , wherein the one or more processors are further to:

provide, responsive to a confidence level for the learned result not meeting a threshold, the learned result to a training process, wherein the training process updates the language knowledge base utilizing the learned result.

18. The computer program product of claim 15 , wherein the neural network comprises a region-based convolutional neural network, and wherein the one or more processors are further to:

provide outputs from the machine learning engine to a deep learning (DL) module, wherein the DL module implements a DL algorithm to analyze the outputs from the machine learning engine and,

learn from differences between the outputs from the machine learning engine and actual results containing corrected data from the first image; and

produce producing a set of parameters that reflect the differences thus learned.

19. The computer program product of claim 18 , wherein the one or more processors are further to:

providing the set of parameters to the machine learning engine, wherein the machine learning engine is operable to utilize the set of parameters to improve its own performance.

20. The computer program product of claim 15 , wherein the second image comprises a plurality of layers, each of the plurality of layers corresponding to one of the classification types associated with content in the first image.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2024
From: OPEN TEXT CORP.
To: CROWDSTRIKE, INC.
Reel/Frame 068122/0068 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2024
From: OPEN TEXT SA ULC
To: OPEN TEXT CORP.
Reel/Frame 067400/0102 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2023
From: AST, UWE
To: OPEN TEXT SOFTWARE GMBH
Reel/Frame 064822/0799 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2023
From: OPEN TEXT SOFTWARE GMBH
To: OPEN TEXT SA ULC
Reel/Frame 064822/0806 →
Continuity (6)
Continuation 17579339 · Jan 19, 2022
Continuation 17081705 · Oct 27, 2020
Continuation 16842097 · Apr 7, 2020
Continuation 16058476 · Aug 8, 2018
Provisional Application 62543246 · Aug 9, 2017
Related Publication 20230401840A1 · Dec 14, 2023
References Cited (9)
US 10628668B2 · Ast · 2020 [cited by applicant]
US 10832048B2 · Ast · 2020 [cited by applicant]
US 20080170786A1 · Tomizawa et al. · 2008 [cited by applicant]
US 20180204360A1 · Bekas et al. · 2018 [cited by applicant]
He, Multi-scale Multi-task FCN for Semantic Page Segmentation and Table Detection, 2017, IEEE 14th IAPR International Conference on Document Analysis and Recognition (Year: 2017). [cited by examiner]
Lucia Noce, “Document Image Classification Combining Textual and Visual Features” PHD Thesis, University of Insubria, (Year: 2016). [cited by applicant]
Yang et al. “Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Network”, arXiv.org, Cornell University (Year: 2017). [cited by applicant]
Ren et al., “Faster R—CNN” Towards Real-Time Object Detection with Region Proposal Networks, arXiv: 1506.01497v3 [cs.CV] Jan. 6, 2016, [retrieved from << https://arxiv.org/pdf/1506.01497.pdf >> on Oct. 2, 2018] 14 pages. [cited by applicant]
Notice of Allowance issued for U.S. Appl. No. 16/058,476, issued Dec. 16, 2019, 10 pages. [cited by applicant]
Cited By (3)
US 12,585,657 US 12,587,546 US 12,694,146