IP Library › Granted Patent US 12,204,860
Granted Patent B2
US 12,204,860 · App. 18/112,969 · Granted Jan 21, 2025

Data-driven structure extraction from text documents

Inventor: Christian Reisswig (Berlin, DE)
Assignee: SAP SE
G06F40/295G06F16/355G06F40/14G06F40/284G06N3/04G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,204,860
App. No.
18/112,969
Granted
Jan 21, 2025
Kind
B2
Abstract

Methods and apparatus are disclosed for extracting structured content, as graphs, from text documents. Graph vertices and edges correspond to document tokens and pairwise relationships between tokens. Undirected peer relationships and directed relationships (e.g. key-value or composition) are supported. Vertices can be identified with predefined fields, and thence mapped to database columns for automated storage of document content in a database. A trained neural network classifier determines relationship classifications for all pairwise combinations of input tokens. The relationship classification can differentiate multiple relationship types. A multi-level classifier extracts multi-level graph structure from a document. Disclosed embodiments support arbitrary graph structures with hierarchical and planar relationships. Relationships are not restricted by spatial proximity or document layout. Composite tokens can be identified interspersed with other content. A single token can belong to multiple higher level structures according to its various relationships. Examples and variations are disclosed.

Claims (51)

1. A system comprising:

one or more hardware processors with memory coupled thereto and one or more network interfaces;

computer-readable media storing instructions which, when executed by the one or more hardware processors, cause the hardware processors to perform operations comprising:

extracting a structure graph of a received text document with a trained machine learning classifier, the structure graph comprising a plurality of vertices having respective values defined by content of the received text document;

wherein the trained machine learning classifier is a multi-level classifier, each level of the multi-level classifier being configured to generate a corresponding level of the structure graph;

wherein the multi-level classifier comprises two successive levels L1, L2 having respective classifiers, and the extracting further comprises deriving an input for the level L2 classifier from:

an input to the level L1 classifier; and

the level of the structure graph generated by the level L1 classifier;

mapping the vertices to columns of a database; and

for each of the vertices, storing the respective value of the each vertex in the database.

2. The system of claim 1 , further comprising a database management system storing the database.

3. The system of claim 1 , wherein a given level of the multi-level classifier comprises a transformer neural network.

4. The system of claim 3 , wherein the transformer neural network is configured to receive inputs representing a number N of tokens of the received text document and to generate outputs representing N·(N−1)/2 pairwise combinations of the N tokens, wherein one or more of the outputs define edges of the structure graph.

5. A computer-implemented method comprising:

extracting a structure graph of a received text document with a trained machine learning classifier, the structure graph comprising a plurality of vertices having respective values defined by content of the received text document;

wherein the trained machine learning classifier is a multi-level classifier, each level of the multi-level classifier being configured to generate a corresponding level of the structure graph;

wherein the multi-level classifier comprises two successive levels L1, L2 having respective classifiers, and the extracting further comprises deriving an input for the level L2 classifier from:

an input to the level L1 classifier; and

the level of the structure graph generated by the level L1 classifier;

mapping the vertices to columns of a database; and

for each of the vertices, storing the respective value of the each vertex in the database.

6. The computer-implemented method of claim 5 , wherein the received text document is a first text document, the respective values are first respective values stored in a first record of the database, and the method further comprises:

receiving a second text document;

preprocessing the second text document;

evaluating the preprocessed second text document with the trained machine learning classifier;

adding a second record for the second text document to the database; and

populating columns of the second record with second respective values of the corresponding vertices in the second text file.

7. The computer-implemented method of claim 6 , wherein the preprocessing comprises:

recognizing at least one field in the second text document as a named entity; and

replacing the at least one field in the second text document with the named entity.

8. The computer-implemented method of claim 5 , wherein a given level of the multi-level classifier comprises a deep neural network.

9. The computer-implemented method of claim 5 , wherein the extracting further comprises:

identifying, at an initial level of the multi-level classifier, relationships between single-character tokens of the received text document; and

identifying words of the received text document based on the identified relationships.

10. The method of claim 5 , wherein the structure graph further comprises a directed first edge, among a plurality of edges, joining a first pair of the vertices representing a first pair of tokens of the received text document; and

wherein the first edge indicates a key-value relationship between the first pair of tokens.

11. The method of claim 5 , wherein the structure graph further comprises an undirected second edge, among a plurality of edges, joining a second pair of the vertices representing a second pair of tokens of the received text document; and

wherein the second edge indicates a peer relationship between the second pair of tokens.

12. One or more computer-readable media storing instructions which, when executed by one or more hardware processors, cause the hardware processors to perform operations comprising:

extracting a structure graph of a received text document with a trained machine learning classifier, the structure graph comprising a plurality of vertices having respective values defined by content of the received text document;

wherein the trained machine learning classifier is a multi-level classifier, each level of the multi-level classifier being configured to generate a corresponding level of the structure graph;

wherein the multi-level classifier comprises two successive levels L1, L2 having respective classifiers, and the extracting further comprises deriving an input for the level L2 classifier from:

an input to the level L1 classifier; and

the level of the structure graph generated by the level L1 classifier;

mapping the vertices to columns of a database; and

for each of the vertices, storing the respective value of the each vertex in the database.

13. The one or more computer-readable media of claim 12 , wherein the extracting further comprises merging a plurality of the generated levels of the structure graph.

14. The one or more computer-readable media of claim 12 , wherein the deriving comprises:

replacing a plurality of tokens in the input to the L1 classifier with a composite token in the input for the level L2 classifier.

15. The one or more computer-readable media of claim 12 , wherein a given level of the multi-level classifier comprises a transformer neural network.

16. The one or more computer-readable media of claim 15 , wherein the transformer neural network is configured to receive inputs representing a number N of tokens of the received text document and to generate outputs representing N·(N−1)/2 pairwise combinations of the N tokens, wherein one or more of the outputs define edges of the structure graph.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2023
From: REISSWIG, CHRISTIAN
To: SAP SE
Reel/Frame 062952/0155 →
Continuity (2)
Division 16891819 · Jun 3, 2020
Related Publication 20230206000A1 · Jun 29, 2023
References Cited (51)
US 6772149B1 · Morelock et al. · 2004 [cited by applicant]
US 8234312B2 · Thomas · 2012 [cited by applicant]
US 10540579B2 · Reisswig et al. · 2020 [cited by applicant]
US 10963501B1 · Ramachandrappa · 2021 [cited by examiner]
US 20030095135A1 · Kaasila et al. · 2003 [cited by applicant]
US 20080059945A1 · Sauer et al. · 2008 [cited by applicant]
US 20130246480A1 · Lemcke et al. · 2013 [cited by applicant]
US 20160179982A1 · Dietrich · 2016 [cited by applicant]
US 20170323425A1 · Masuko et al. · 2017 [cited by applicant]
US 20200356851A1 · Li et al. · 2020 [cited by applicant]
US 20210081717A1 · Creed · 2021 [cited by examiner]
US 20210209472A1 · Chen et al. · 2021 [cited by applicant]
US 20210279422A1 · Mihindukulasooriya · 2021 [cited by examiner]
US 20210357588A1 · Friedrich · 2021 [cited by examiner]
US 20210383067A1 · Reisswig · 2021 [cited by applicant]
Yao et al., KG-BERT: BERT for knowledge graph completion, https://arxiv.org/abs/1909.03193, Sep. 11, 2019, pp. 1-8 (Year: 2019). [cited by examiner]
Li et al., Hierarchical Graph Attention Networks for semi-supervised node classification, Jun. 10, 2020, Springer Science and Business media, LLC, pp. 3441-3451 (Year: 2020). [cited by examiner]
Alammar, “The Illustrated BERT, ELMo, and co. (How NLP Cracked Transfer Learning),” http://jalammar.github.io/illustrated-bert/, 20 pages (Dec. 2018). [cited by applicant]
Alammar, “The Illustrated Transformer,” http://jalammar.github.io/illustrated-transformer/, 22 pages (Jun. 2018). [cited by applicant]
Algorithmia, “Introduction to Loss Functions,” https://algorithmia.com/blog/introduction-to-loss-functions, 8 pages (Apr. 2018). [cited by applicant]
Al-Masri, “How Does Back-Propagation in Artificial Neural Networks Work?” https://towardsdatascience.com/how-does-back-propagation-in-artificial-neural-networks-work-c7cad873ea7, 11 pages (Jan. 2019). [cited by applicant]
Bartz et al., “STN-OCR: A single Neural Network for Text Detection and Text Recognition,” arXiv.org 1707.08831v1, Cornell University Library, 9 pages (Jul. 2017). [cited by applicant]
Benaffane, “Transformer vs RNN and CNN for Translation Task,” https://medium.com/analytics-vidhya/transformer-vs-rnn-and-cnn-18eeefa3602b, 12 pages (Aug. 2019). [cited by applicant]
Bird et al., “Extracting Information from Text,” https://www.nltk.org/book/ch07.html, 19 pages (Sep. 2019). [cited by applicant]
Cheng et al., “Long Short-Term Memory-Networks for Machine Reading,” arXiv.org 1601.06733v7, Cornell University Library, 11 pages (Sep. 2016). [cited by applicant]
Chromiak, “The Transformer—Attention is all you need,” https://mchromiak.github.io/articles/2017/Sep/12/Transformer-Attention-is-all-you-need/#.XqxyzGhKiUk, 19 pages (Sep. 2017). [cited by applicant]
Crnoja, “How Automation is Changing Financial Services—SAP Business Entity Recognition,” https://blogs.sap.com/2019/12/10/how-automation-is-changing-financial-services/, 8 pages (Dec. 2019). [cited by applicant]
Dharmadhikari et al., “Multi Label Text Classification through Label Propagation,” International Journal of Engineering Research and Development ISSN, pp. 9-14 (Jun. 2012). [cited by applicant]
Denk et al., “BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding,” arXiv.org 1909.04948v2, Cornell University Library, 4 pages (Oct. 2019). [cited by applicant]
Extended European Search Report for Application No. 21161736.0, 18 pages, mailed Sep. 3, 2021. [cited by applicant]
European Search Report from EP Patent Application No. 18204940.3, 9 pages (Jun. 24, 2019). [cited by applicant]
Harish et al., “Representation and Classification of Text Documents: A Brief Review,” IJCA Special Issue on RTIPPR, pp. 110-119 (Feb. 2010). [cited by applicant]
Jaderberg et al., “Reading Text in the Wild with Convolutional Neural Networks,” ArXiv.org 1412.1842v1, Cornell University Library (Dec. 2014). [cited by applicant]
Jiang et al., “Text Classification using Graph Mining-based Feature Extraction,” 14 pages, also published as Jiang et al., “Text Classification using Graph Mining-based Feature Extraction,” Research and Development in I… [cited by applicant]
Knoblock et al., “Interactively Mapping Data Sources into the Semantic Web,” pp. 1-12 (purportedly Oct. 2011). [cited by applicant]
Nicholson, “A.I. Wiki A Beginner's Guide to Important Topics in AI, Machine Learning, and Deep Learning,” https://pathmind.com/wiki/attention-mechanism-memory-network, 12 pages (Dec. 2019). [cited by applicant]
Palm et al., “CloudScan—A configuration-free invoice analysis system using recurrent neural networks,” arXiv.org 1708.07403v1, Cornell University Library, 8 pages (Aug. 2017). [cited by applicant]
Ramshaw et al., “Text Chunking using Transformation-Based Learning,” arXiv.org cmp-lg/9505040v1, Cornell University Library, 13 pages (May 1995). [cited by applicant]
Rausch et al., “DocParser: Hierarchical Structure Parsing of Document Renderings,” arXiv.org 1911.01702v1, Cornell University Library, 14 pages (Nov. 2019). [cited by applicant]
SAP, “Document Information Extraction,” https://help.sap.com/doc/cc96d8c853274be4b58b1d57acbff284/SHIP/en-US/ab18043fd70e424dbcbcfd81585ff029.pdf, pp. 1-47 (Mar. 2020). [cited by applicant]
Stewart et al., “Document Image p. Segmentation and Character Recognition as Semantic Segmentation,” Proceedings of the 4th International Workshop on Historical Document Imaging and Processing, pp. 101-106 (Nov. 2017). [cited by applicant]
Vaswani et al., “Attention Is All You Need,” arXiv.org 1706.03762v5, Cornell University Library, 15 pages (Dec. 2017). [cited by applicant]
Wang et al., “Extracting Multiple-Relations in One-Pass with Pre-Trained Transformers,” arXiv.org 1902.01030v2, Cornell University Library, 7 pages (Jun. 2019). [cited by applicant]
Weng, “Attention? Attention!” https://lilianweng.github.io/lil-log/2018/06/24/attention-attention.html, 28 pages (Jun. 2018). [cited by applicant]
Wikipedia, “BERT (language model),” downloaded from https://en.wikipedia.org/wiki/BERT_(language_model), 3 pages (document marked Apr. 2020). [cited by applicant]
Wikipedia, “GloVe (machine learning),” downloaded from https://en.wikipedia.org/wiki/GloVe_(machine_learning), 2 pages (document marked Apr. 2020). [cited by applicant]
Wikipedia, “Softmax function,” downloaded from https://en.wikipedia.org/wiki/Softmax_function, 8 pages (document marked Mar. 2020). [cited by applicant]
Wikipedia, “Transformer (machine learning model),” downloaded from https://en.wikipedia.org/wiki/Transformer_(machine_learning_model), 5 pages (document marked Mar. 2020). [cited by applicant]
Wikipedia, “Word2vec,” downloaded from https://en.wikipedia.org/wiki/Word2vec, 6 pages (document marked May 2020). [cited by applicant]
Wu et al., “Enriching Pre-trained Language Model with Entity Information for Relation Classification,” Proceedings of the 28th ACM International Conference on Information and Knowledge Management, pp. 2361-2364 (Nov. 20… [cited by applicant]
Yang et al., “Learning to Extract Semantic Structure from Documents Using Multimodal Fully Convolutional Neural Networks,” arXiv.org, Cornell University Library, 16 pages (Jun. 2017). [cited by applicant]