IP Library › Granted Patent US 11,461,552
Granted Patent B2
US 11,461,552 · App. 16/920,916 · Granted Oct 4, 2022

Automated document review system combining deterministic and machine learning algorithms for legal document review

Inventors: Jianglei Han (Singapore, SG); Traci Zheng Wen Lim (Singapore, SG); Juanlei Rocco Hu (Singapore, SG); Lijie Quan (Singapore, SG); Wei Jin (Singapore, SG); Lingxiao Liang (Singapore, SG)
Assignee: SAP SE
G06F40/289G06F40/284G06N20/00G06Q50/18G06V30/153
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,461,552
App. No.
16/920,916
Granted
Oct 4, 2022
Kind
B2
Abstract

Methods, systems, and computer-readable storage media for receiving, by an automated review system, a legal document as a computer-readable file, and determining, by the automated review system, that the legal document is of a first type, and in response: converting the legal document to a set of images, extracting text data from one or more images in the set of images, the text data including sub-sets of text data, each sub-set of text data representing text in a respective clause of a set of clauses of the legal document, for each sub-set of text data receiving a prediction from a machine learning (ML) model in a set of ML models, the ML model being specific to a clause in the set of clauses, and outputting a set of predictions and respective prediction values for display in a user interface (UI).

Claims (46)

1. A computer-implemented method for automated review of legal documents using an automated review system, the method being executed by one or more processors and comprising:

receiving, by the automated review system, a legal document as a computer-readable file, the automated review system comprising a customized optical character recognition (OCR) module for processing documents of a first type and an OCR module for processing documents of a second type; and

determining, by the automated review system, that the legal document is of the first type, and in response:

converting, by the customized OCR module, the legal document to a set of images,

extracting, by the customized OCR module, text data from one or more images in the set of images, the text data comprising sub-sets of text data, each sub-set of text data representing text in a respective clause of a set of clauses of the legal document,

for each sub-set of text data, determining a clause from a set of clauses that the sub-set of text data is associated with, selecting a machine learning (ML) model from a set of ML models, the ML model being specific to the clause, and receiving a prediction from the ML model, and

outputting a set of predictions and respective prediction values for display in a user interface (UI).

2. The method of claim 1 , wherein extracting text data from one or more images in the set of images comprises:

for at least one image in the set of images, determining that the at least one image depicts a table; and

extracting text data from the table within the at least one image.

3. The method of claim 1 , wherein at least one ML model in the set of ML models is trained using training data comprising a set of relevant sentences for each clause in the set of clauses, each relevant sentence determined to be relevant to a legal term occurring within a respective clause.

4. The method of claim 3 , wherein the set of relevant sentences is provided at least partially based on, for each clause in the set of clauses, representing the legal term as a hash structure, parsing the legal document into an array of sentences, and identifying matches between tokens of sentences and the hash structure.

5. The method of claim 1 , further comprising receiving user input provided through the UI, the user input changing at least one prediction value for a respective clause, and being used to retrain a respective ML model.

6. The method of claim 1 , wherein the set of ML models comprises a first ML model of a first type that is specific to a first clause, and a second ML model of a second type that is specific to a second clause, the second type being different from the first type.

7. The method of claim 1 , wherein the first type comprises an image-based document.

8. A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for automated review of legal documents using a hybrid system, the operations comprising:

receiving, by the automated review system, a legal document as a computer-readable file the automated review system comprising a customized optical character recognition (OCR) module for processing documents of a first type and an OCR module for processing documents of a second type; and

determining, by the automated review system, that the legal document is of the first type, and in response:

converting, by the customized OCR module, the legal document to a set of images,

extracting, by the customized OCR module, text data from one or more images in the set of images, the text data comprising sub-sets of text data, each sub-set of text data representing text in a respective clause of a set of clauses of the legal document,

for each sub-set of text data, determining a clause from a set of clauses that the sub-set of text data is associated with, selecting a machine learning (ML) model from a set of ML models, the ML model being specific to the clause, and receiving a prediction from the ML model, and

outputting a set of predictions and respective prediction values for display in a user interface (UI).

9. The computer-readable storage medium of claim 8 , wherein extracting text data from one or more images in the set of images comprises:

for at least one image in the set of images, determining that the at least one image depicts a table; and

extracting text data from the table within the at least one image.

10. The computer-readable storage medium of claim 8 , wherein at least one ML model in the set of ML models is trained using training data comprising a set of relevant sentences for each clause in the set of clauses, each relevant sentence determined to be relevant to a legal term occurring within a respective clause.

11. The computer-readable storage medium of claim 10 , wherein the set of relevant sentences is provided at least partially based on, for each clause in the set of clauses, representing the legal term as a hash structure, parsing the legal document into an array of sentences, and identifying matches between tokens of sentences and the hash structure.

12. The computer-readable storage medium of claim 8 , wherein operations further comprise receiving user input provided through the UI, the user input changing at least one prediction value for a respective clause, and being used to retrain a respective ML model.

13. The computer-readable storage medium of claim 8 , wherein the set of ML models comprises a first ML model of a first type that is specific to a first clause, and a second ML model of a second type that is specific to a second clause, the second type being different from the first type.

14. The computer-readable storage medium of claim 8 , wherein the first type comprises an image-based document.

15. A system, comprising:

a computing device; and

a computer-readable storage device coupled to the computing device and having instructions stored thereon which, when executed by the computing device, cause the computing device to perform operations for automated review of legal documents using a hybrid system, the operations comprising:

receiving, by the automated review system, a legal document as a computer-readable file, the automated review system comprising a customized optical character recognition (OCR) module for processing documents of a first type and an OCR module for processing documents of a second type; and

determining, by the automated review system, that the legal document is of the first type, and in response:

converting, by the customized OCR module, the legal document to a set of images,

extracting, by the customized OCR module, text data from one or more images in the set of images, the text data comprising sub-sets of text data, each sub-set of text data representing text in a respective clause of a set of clauses of the legal document,

for each sub-set of text data, determining a clause from a set of clauses that the sub-set of text data is associated with, selecting a machine learning (ML) model from a set of ML models, the ML model being specific to the clause, and receiving a prediction from the ML model, and

outputting a set of predictions and respective prediction values for display in a user interface (UI).

16. The system of claim 15 , wherein extracting text data from one or more images in the set of images comprises:

for at least one image in the set of images, determining that the at least one image depicts a table; and

extracting text data from the table within the at least one image.

17. The system of claim 15 , wherein at least one ML model in the set of ML models is trained using training data comprising a set of relevant sentences for each clause in the set of clauses, each relevant sentence determined to be relevant to a legal term occurring within a respective clause.

18. The system of claim 17 , wherein the set of relevant sentences is provided at least partially based on, for each clause in the set of clauses, representing the legal term as a hash structure, parsing the legal document into an array of sentences, and identifying matches between tokens of sentences and the hash structure.

19. The system of claim 15 , wherein operations further comprise receiving user input provided through the UI, the user input changing at least one prediction value for a respective clause, and being used to retrain a respective ML model.

20. The system of claim 15 , wherein the set of ML models comprises a first ML model of a first type that is specific to a first clause, and a second ML model of a second type that is specific to a second clause, the second type being different from the first type.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 6, 2020
From: HAN, JIANGLEI; LIM, TRACI ZHENG WEN; HU, JUANLEI ROCCO; QUAN, LIJIE; JIN, WEI; LIANG, LINGXIAO
To: SAP SE
Reel/Frame 053123/0853 →
Continuity (1)
Related Publication 20220004713A1 · Jan 6, 2022
Cited By (2)
US 12,265,787 US 12,596,877