IP Library › Granted Patent US 12,001,802
Granted Patent B2
US 12,001,802 · App. 17/337,835 · Granted Jun 4, 2024

Training enrichment system for natural language processing

Inventors: Tassilo Klein (Berlin, DE); Moin Nabi (Berlin, DE)
Assignee: SAP SE
G06F40/30G06F16/951G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,001,802
App. No.
17/337,835
Granted
Jun 4, 2024
Kind
B2
Abstract

Disclosed herein are various embodiments for training and enriching a natural language processing system. An embodiment operates by identifying a natural language processor (NLP) trained on a first set of documents, wherein the NLP is trained to perform a set of functionality based on the first set of documents. An industry, set of words corresponding to the industry, and set of sentences including at least a subset of the set of words in which the NLP is to be configured to perform the set of functionality are identified. A set of sentences that exceed a similarity threshold are identified. The NLP is trained with the subset of the set of sentences that exceed the similarity threshold, wherein the trained NLP with the subset is configured to perform the set of functionality within the industry with a greater accuracy than NLP trained on only the first set of documents.

Claims (52)

1. A method comprising:

identifying a natural language processor (NLP) trained on a first set of documents, wherein the NLP is trained to perform a set of functionality based on the first set of documents;

determining an industry in which the NLP is to be configured to perform the set of functionality;

identifying a set of words corresponding to the industry;

identifying a set of sentences including at least a subset of the set of words corresponding to the industry;

scoring the set of sentences based on a similarity to one another;

identifying a subset of the set of sentences that exceed a similarity threshold; and

training the NLP with the subset of the set of sentences that exceed the similarity threshold, wherein the trained NLP with the subset is configured to perform the set of functionality within the industry with a greater accuracy than an NLP trained on only the first set of documents.

2. The method of claim 1 , wherein the first set of documents are annotated with indications of correct and incorrect input and output.

3. The method of claim 2 , wherein the subset of sentences are augmented indicating different variations of sentences with similar meanings.

4. The method of claim 1 , wherein the identifying a set of sentences comprises:

identifying a second set of documents, different from the first set of documents; and

identifying, from the second set of documents, the set of sentences including at least a subset of the set of words corresponding to the industry.

5. The method of claim 4 , wherein the second set of documents correspond to different webpages available on the Internet associated with the industry.

6. The method of claim 5 , further comprising:

crawling the different webpages publicly available on the Internet to identify the set of sentences.

7. The method of claim 1 , wherein the scoring is based on BERTScore.

8. A system, comprising:

a memory; and

at least one processor coupled to the memory and configured to perform instructions that cause the at least one processor to perform operations comprising:

identifying a natural language processor (NLP) trained on a first set of documents, wherein the NLP is trained to perform a set of functionality based on the first set of documents;

determining an industry in which the NLP is to be configured to perform the set of functionality;

identifying a set of words corresponding to the industry;

identifying a set of sentences including at least a subset of the set of words corresponding to the industry;

scoring the set of sentences based on a similarity to one another;

identifying a subset of the set of sentences that exceed a similarity threshold; and

training the NLP with the subset of the set of sentences that exceed the similarity threshold, wherein the trained NLP with the subset is configured to perform the set of functionality within the industry with a greater accuracy than an NLP trained on only the first set of documents.

9. The system of claim 8 , wherein the first set of documents are annotated with indications of correct and incorrect input and output.

10. The system of claim 9 , wherein the subset of sentences are augmented indicating different variations of sentences with similar meanings.

11. The system of claim 8 , wherein the identifying a set of sentences comprises:

identifying a second set of documents, different from the first set of documents; and

identifying, from the second set of documents, the set of sentences including at least a subset of the set of words corresponding to the industry.

12. The system of claim 11 , wherein the second set of documents correspond to different webpages available on the Internet associated with the industry.

13. The system of claim 12 , the operations further comprising:

crawling the different webpages publicly available on the Internet to identify the set of sentences.

14. The system of claim 8 , wherein the scoring is based on BERTScore.

15. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:

identifying a natural language processor (NLP) trained on a first set of documents, wherein the NLP is trained to perform a set of functionality based on the first set of documents;

determining an industry in which the NLP is to be configured to perform the set of functionality;

identifying a set of words corresponding to the industry;

identifying a set of sentences including at least a subset of the set of words corresponding to the industry;

scoring the set of sentences based on a similarity to one another;

identifying a subset of the set of sentences that exceed a similarity threshold; and

training the NLP with the subset of the set of sentences that exceed the similarity threshold, wherein the trained NLP with the subset is configured to perform the set of functionality within the industry with a greater accuracy than an NLP trained on only the first set of documents.

16. The non-transitory computer-readable medium of claim 15 , wherein the first set of documents are annotated with indications of correct and incorrect input and output.

17. The non-transitory computer-readable medium of claim 16 , wherein the subset of sentences are augmented indicating different variations of sentences with similar meanings.

18. The non-transitory computer-readable medium of claim 15 , wherein the identifying a set of sentences comprises:

identifying a second set of documents, different from the first set of documents; and

identifying, from the second set of documents, the set of sentences including at least a subset of the set of words corresponding to the industry.

19. The non-transitory computer-readable medium of claim 18 , wherein the second set of documents correspond to different webpages available on the Internet associated with the industry.

20. The non-transitory computer-readable medium of claim 19 , the operations further comprising:

crawling the different webpages publicly available on the Internet to identify the set of sentences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2021
From: KLEIN, TASSILO; NABI, MOIN
To: SAP SE
Reel/Frame 058322/0632 →
Continuity (1)
Related Publication 20220391592A1 · Dec 8, 2022