IP Library › Granted Patent US 12,093,300
Granted Patent B1
US 12,093,300 · App. 18/463,519 · Granted Sep 17, 2024

Enhancing accuracy of entity matching inference using large language models

Inventors: Yi Quan Zhou (Singapore, SG); Rajesh Vellore Arumugam (Singapore, SG); Raja Sekhar Juluri (Singapore, SG); Xingce Bao (Singapore, SG); Eshwin Sukhdeve (Singapore, SG)
Assignee: SAP SE
G06F16/35G06F16/3329G06F16/3344G06F40/174G06F40/186
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,093,300
App. No.
18/463,519
Granted
Sep 17, 2024
Kind
B1
Abstract

Methods, systems, and computer-readable storage media for receiving a first document including structured data and unstructured data, providing a first sub-document and a second sub-document, the first sub-document including the structured data of the first document, the second sub-document including the unstructured data of the first document, generating a prompt using the second sub-document and a second document, inputting the prompt to a LLM, receiving a response from the LLM, providing a calibrated first document by merging the response into the first sub-document, and processing the calibrated first document and the second document using a ML model to provide a prediction, the prediction indicating a matching class between the first document and the second document.

Claims (43)

1. A computer-implemented method for improved document matching using machine learning (ML) models in ML-based decision systems, the method being executed by one or more processors and comprising:

receiving a first document comprising structured data and unstructured data;

providing a first sub-document and a second sub-document, the first sub-document comprising the structured data of the first document, the second sub-document comprising the unstructured data of the first document;

generating a prompt using the second sub-document and a second document;

inputting the prompt to a large language model (LLM);

receiving a response from the LLM;

providing a calibrated first document by merging the response into the first sub-document; and

processing the calibrated first document and the second document using a ML model to provide a prediction, the prediction indicating a matching class between the first document and the second document.

2. The method of claim 1 , wherein generating a prompt comprises populating a prompt template using data from each of the second sub-document and the second document.

3. The method of claim 1 , wherein the prompt comprises a few-shot prompt.

4. The method of claim 1 , wherein the matching class comprises one of a match, a multi-match, and no match.

5. The method of claim 1 , wherein the unstructured data comprises a string of characters.

6. The method of claim 1 , wherein the ML model comprises a generic line-item matching (GLIM) model.

7. The method of claim 1 , further comprising automatically executing a task in a ML-based decision system in response to the prediction.

8. A non-transitory computer-readable storage medium coupled to one or more processors and having instructions stored thereon which, when executed by the one or more processors, cause the one or more processors to perform operations for improved document matching using machine learning (ML) models in ML-based decision systems, the operations comprising:

receiving a first document comprising structured data and unstructured data;

providing a first sub-document and a second sub-document, the first sub-document comprising the structured data of the first document, the second sub-document comprising the unstructured data of the first document;

generating a prompt using the second sub-document and a second document;

inputting the prompt to a large language model (LLM);

receiving a response from the LLM;

providing a calibrated first document by merging the response into the first sub-document; and

processing the calibrated first document and the second document using a ML model to provide a prediction, the prediction indicating a matching class between the first document and the second document.

9. The non-transitory computer-readable storage medium of claim 8 , wherein generating a prompt comprises populating a prompt template using data from each of the second sub-document and the second document.

10. The non-transitory computer-readable storage medium of claim 8 , wherein the prompt comprises a few-shot prompt.

11. The non-transitory computer-readable storage medium of claim 8 , wherein the matching class comprises one of a match, a multi-match, and no match.

12. The non-transitory computer-readable storage medium of claim 8 , wherein the unstructured data comprises a string of characters.

13. The non-transitory computer-readable storage medium of claim 8 , wherein the ML model comprises a generic line-item matching (GLIM) model.

14. The non-transitory computer-readable storage medium of claim 8 , wherein operations further comprise automatically executing a task in a ML-based decision system in response to the prediction.

15. A system, comprising:

a computing device; and

a computer-readable storage device coupled to the computing device and having instructions stored thereon which, when executed by the computing device, cause the computing device to perform operations for improved document matching using machine learning (ML) models in ML-based decision systems, the operations comprising:

receiving a first document comprising structured data and unstructured data;

providing a first sub-document and a second sub-document, the first sub-document comprising the structured data of the first document, the second sub-document comprising the unstructured data of the first document;

generating a prompt using the second sub-document and a second document;

inputting the prompt to a large language model (LLM);

receiving a response from the LLM;

providing a calibrated first document by merging the response into the first sub-document; and

processing the calibrated first document and the second document using a ML model to provide a prediction, the prediction indicating a matching class between the first document and the second document.

16. The system of claim 15 , wherein generating a prompt comprises populating a prompt template using data from each of the second sub-document and the second document.

17. The system of claim 15 , wherein the prompt comprises a few-shot prompt.

18. The system of claim 15 , wherein the matching class comprises one of a match, a multi-match, and no match.

19. The system of claim 15 , wherein the unstructured data comprises a string of characters.

20. The system of claim 15 , wherein the ML model comprises a generic line-item matching (GLIM) model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 8, 2023
From: ZHOU, YI QUAN; ARUMUGAM, RAJESH VELLORE; JULURI, RAJA SEKHAR; BAO, XINGCE; SUKHDEVE, ESHWIN
To: SAP SE
Reel/Frame 064844/0401 →
Cited By (3)
US 12,346,649 US 12,461,932 US 12,619,876