IP Library Granted Patent US 11,720,752
Granted Patent B2
US 11,720,752 · App. 16/922,922 · Granted Aug 8, 2023

Machine learning enabled text analysis with multi-language support

Inventor: Tobias Weller (Buseck, DE)
Assignee: SAP SE
G06F40/295G06F18/2148G06N3/02G06N20/10G06N20/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,720,752
App. No.
16/922,922
Granted
Aug 8, 2023
Kind
B2
Abstract

A language determination model may be applied to select a first machine learning model or a second machine learning model to analyze the input text. The first machine learning model trained to analyze text in a first language, the second machine learning model trained to analyze text in a second language, and the input text may be in a third language. The language determination model may select the first machine learning model based on the first machine learning model having a better performance analyzing text in the third language than the second machine learning model. The language determination model may be updated based on an actual performance of the first machine learning model analyzing the input text. Moreover, the first machine learning model may be subject to additional training if the actual performance of the first machine learning model analyzing the input text is below a threshold value.

Claims (31)

1. A system, comprising:

at least one data processor; and

at least one memory storing instructions which, when executed by the at least one data processor, result in operations comprising:

training, based at least on a training data, a language determination model, the training data including a first performance of a first machine learning model analyzing text in a first language, a second language, and/or a third language, and the training data further including a second performance of a second machine learning model analyzing text in the first language, the second language, and/or the third language;

applying the language determination model to select the first machine learning model instead of the second machine learning model to analyze an input text in the third language, the language determination model selecting the first machine learning model trained to analyze text in the first language based at least on the first machine learning model having a better performance analyzing text in the third language than the second machine learning model trained to analyze text in the second language; and

analyzing the input text by at least applying, to the input text, the first machine leaning model.

2. The system of claim 1 , wherein the training data further includes a family, a branch, a sub-branch, and/or a script associated with each of the first language, the second language, and the third language.

3. The system of claim 1 , wherein each of the first performance and the second performance correspond to a quantity of named entities assigned to a correct category.

4. The system of claim 1 , further comprising:

updating, based at least on a third performance of the first machine learning model analyzing the input text, the language determination model.

5. The system of claim 4 , further comprising:

in response to the third performance of the first machine learning model analyzing the input text being below a threshold value, training, based on additional training data in the first language, the first machine learning model.

6. The system of claim 1 , wherein the language determination model is further trained to select, based at least on a type of the input text, the first machine learning model or the second machine learning model to analyze the input text.

7. The system of claim 1 , wherein the first machine learning model, the second machine learning model, and the language determination model each comprise a support vector machine, a boosted decision tree, a regularized logistic regression model, a neural network, and/or a random forest.

8. The system of claim 1 , wherein the analysis of the input text includes assigning each named entity present in the input text to one or more categories including a person name, an organization, a location, a medical code, a time expression, a quantity, a monetary value, and a percentage.

9. A computer-implemented method, comprising:

training, based at least on a training data, a language determination model, the training data including a first performance of a first machine learning model analyzing text in a first language, a second language, and/or a third language, and the training data further including a second performance of a second machine learning model analyzing text in the first language, the second language, and/or the third language;

applying the language determination model to select the first machine learning model instead of the second machine learning model to analyze an input text in the third language, the language determination model selecting the first machine learning model trained to analyze text in the first language based at least on the first machine learning model having a better performance analyzing text in the third language than the second machine learning model trained to analyze text in the second language; and

analyzing the input text by at least applying, to the input text, the first machine leaning model.

10. The method of claim 9 , wherein the training data further includes a family, a branch, a sub-branch, and/or a script associated with each of the first language, the second language, and the third language.

11. The method of claim 9 , wherein each of the first performance and the second performance correspond to a quantity of named entities assigned to a correct category.

12. The method of claim 9 , further comprising:

updating, based at least on a third performance of the first machine learning model analyzing the input text, the language determination model.

13. The method of claim 12 , further comprising:

in response to the third performance of the first machine learning model analyzing the input text being below a threshold value, training, based on additional training data in the first language, the first machine learning model.

14. The method of claim 9 , wherein the language determination model is further trained to select, based at least on a type of the input text, the first machine learning model or the second machine learning model to analyze the input text.

15. The method of claim 9 , wherein the first machine learning model, the second machine learning model, and the language determination model each comprise a support vector machine, a boosted decision tree, a regularized logistic regression model, a neural network, and/or a random forest.

16. A non-transitory computer readable medium storing instructions, which when executed by at least one data processor, result in operations comprising:

training, based at least on a training data, a language determination model, the training data including a first performance of a first machine learning model analyzing text in a first language, a second language, and/or a third language, and the training data further including a second performance of a second machine learning model analyzing text in the first language, the second language, and/or the third language;

applying the language determination model to select the first machine learning model instead of the second machine learning model to analyze an input text in the third language, the language determination model selecting the first machine learning model trained to analyze text in the first language based at least on the first machine learning model having a better performance analyzing text in the third language than the second machine learning model trained to analyze text in the second language; and

analyzing the input text by at least applying, to the input text, the first machine leaning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2020
From: WELLER, TOBIAS
To: SAP SE
Reel/Frame 053143/0333 →
Continuity (1)
Related Publication 20220012429A1 · Jan 13, 2022