IP Library › Granted Patent US 12,596,877
Granted Patent B2
US 12,596,877 · App. 18/463,509 · Granted Apr 7, 2026

Computer-implemented contract risk assessment platform leveraging transformers

Inventors: Jianglei Han (Singapore, SG); Qisheng Hu (Singapore, SG); My Hoa Ha (Singapore, SG); Yue Yang (Singapore, SG)
Assignee: SAP SE
G06F40/284G06F16/34G06F40/169G06F40/40G06N3/0455G06N3/048G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,877
App. No.
18/463,509
Filed
Sep 8, 2023
Granted
Apr 7, 2026
Kind
B2
Examiner
WONG, LINDA
Art Unit
2655
USPC
704/9
Abstract

Methods, systems, and computer-readable storage media for receiving a document provided as a computer-readable file, receiving a set of questions, for each question in the set of questions, generating an inference input including a question, at least a portion of text of the document, and multiple tokens, processing, by a PLM, the inference input to generate a set of text embeddings, processing, by a neural network, the set of text embeddings to provide sets of tokens, each set of tokens being specific to a segment of the document and including a start token and an end token respectively identifying a start position and an end position of the segment, determining, from the sets of tokens, a segment for display, and displaying at least a portion of the document in a UI and an annotation indicating the segment within the at least a portion of the document.

Claims (40)

1 . A computer-implemented method for automatic identification and display of segments of text in computer-readable documents, the method being executed by one or more processors and comprising:

receiving a document provided as a computer-readable file;

receiving a set of questions;

for each question in the set of questions, generating an inference input comprising a question, at least a portion of text of the document, and multiple tokens;

processing, by a pre-trained language model (PLM), the inference input to generate a set of text embeddings;

processing, by a neural network, the set of text embeddings to provide sets of tokens, each set of tokens being specific to a segment of the document and comprising a start token and an end token respectively identifying a start position and an end position of the segment;

determining, from the sets of tokens, a segment for display by determining a score difference based on a null score and a non-null score and selecting the segment; and

displaying at least a portion of the document in a user interface (UI) and an annotation indicating the segment within the at least a portion of the document.

2 . The method of claim 1 , wherein tokens in each set of tokens is associated with a logit comprising a non-normalized value that is normalized by a softmax function.

3 . The method of claim 1 , wherein the segment is selected based on a null prediction represented by the null score in response to the score difference exceeding a threshold.

4 . The method of claim 1 , wherein the document comprises a contract and the segment comprises a portion of a clause of the contract.

5 . The method of claim 1 , wherein the annotation comprises highlighting of the segment in the UI, the highlighting extending from the start token and the end token determined for the segment.

6 . The method of claim 1 , wherein the PLM comprises a Bidirectional Encoder Representations from Transformers (BERT) model.

7 . A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for automatic identification and display of segments of text in computer-readable documents, the operations comprising:

receiving a document provided as a computer-readable file;

receiving a set of questions;

for each question in the set of questions, generating an inference input comprising a question, at least a portion of text of the document, and multiple tokens;

processing, by a pre-trained language model (PLM), the inference input to generate a set of text embeddings;

processing, by a neural network, the set of text embeddings to provide sets of tokens, each set of tokens being specific to a segment of the document and comprising a start token and an end token respectively identifying a start position and an end position of the segment;

determining, from the sets of tokens, a segment for display by determining a score difference based on a null score and a non-null score and selecting the segment; and

displaying at least a portion of the document in a user interface (UI) and an annotation indicating the segment within the at least a portion of the document.

8 . The non-transitory computer-readable storage medium of claim 7 , wherein tokens in each set of tokens is associated with a logit comprising a non-normalized value that is normalized by a softmax function.

9 . The non-transitory computer-readable storage medium of claim 7 , wherein the segment is selected based on a null prediction represented by the null score in response to the score difference exceeding a threshold.

10 . The non-transitory computer-readable storage medium of claim 7 , wherein the document comprises a contract and the segment comprises a portion of a clause of the contract.

11 . The non-transitory computer-readable storage medium of claim 7 , wherein the annotation comprises highlighting of the segment in the UI, the highlighting extending from the start token and the end token determined for the segment.

12 . The non-transitory computer-readable storage medium of claim 7 , wherein the PLM comprises a Bidirectional Encoder Representations from Transformers (BERT) model.

13 . A system, comprising:

a computing device; and

a computer-readable storage device coupled to the computing device and having instructions stored thereon which, when executed by the computing device, cause the computing device to perform operations for automatic identification and display of segments of text in computer-readable documents, the operations comprising:

receiving a document provided as a computer-readable file;

receiving a set of questions;

for each question in the set of questions, generating an inference input comprising a question, at least a portion of text of the document, and multiple tokens;

processing, by a pre-trained language model (PLM), the inference input to generate a set of text embeddings;

processing, by a neural network, the set of text embeddings to provide sets of tokens, each set of tokens being specific to a segment of the document and comprising a start token and an end token respectively identifying a start position and an end position of the segment;

determining, from the sets of tokens, a segment for display by determining a score difference based on a null score and a non-null score and selecting the segment; and

displaying at least a portion of the document in a user interface (UI) and an annotation indicating the segment within the at least a portion of the document.

14 . The system of claim 13 , wherein tokens in each set of tokens is associated with a logit comprising a non-normalized value that is normalized by a softmax function.

15 . The system of claim 13 , wherein the segment is selected based on a null prediction represented by the null score in response to the score difference exceeding a threshold.

16 . The system of claim 13 , wherein the document comprises a contract and the segment comprises a portion of a clause of the contract.

17 . The system of claim 13 , wherein the annotation comprises highlighting of the segment in the UI, the highlighting extending from the start token and the end token determined for the segment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2023
From: HAN, JIANGLEI; HA, MY HOA; YANG, YUE; HU, QISHENG
To: SAP SE
Reel/Frame 064844/0684 →
Continuity (1)
Related Publication 20250086392A1 · Mar 13, 2025
References Cited (6)
US 11461552B2 · Han et al. · 2022 [cited by applicant]
US 20070010992A1 · Hon · 2007 [cited by examiner]
US 20220405336A1 · Lippe · 2022 [cited by examiner]
US 20230061647A1 · Arbelle · 2023 [cited by examiner]
CN 115952803A · 2023 [cited by examiner]
CN 116340467A · 2023 [cited by examiner]