IP Library Granted Patent US 12,423,529
Granted Patent B2
US 12,423,529 · App. 17/789,396 · Granted Sep 23, 2025

Low-resource multilingual machine learning framework

Inventors: Uday Kamath (Ashburn, VA); Wael Emara (Plainsboro, NJ)
Assignee: DIGITAL REASONING SYSTEMS, INC.
G06F40/58G06F40/49
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,529
App. No.
17/789,396
Granted
Sep 23, 2025
Kind
B2
Abstract

Systems and methods for performing machine learning on multilingual text data.

Claims (34)

1. A method for multilingual machine learning on text data, comprising:

receiving English-language labeled training data comprising at least one indication of human conduct that violates a policy, ethical standard, and/or law;

applying a pre-trained cross-lingual model to the received labeled training data;

applying zero-shot machine learning to the received labeled training data to generate supervised learning task specific layers;

producing a zero-shot machine learning model comprising applying the pre-trained cross-lingual model to input data to generate an output and applying the supervised learning task specific layers to the output to generate a classification; and

applying the zero-shot machine learning model to unlabeled unseen text data to perform functions including the classification of unseen text data to detect at least one identification of a potential violation condition.

2. The method of claim 1 , wherein the text data corresponds to one or more electronic communications between persons.

3. The method of claim 2 , wherein the text classification includes identifying the unseen text data as corresponding to conduct of the persons.

4. The method of claim 1 , wherein the cross-lingual model is pre-trained based on one or more spoken languages as represented in labeled training data.

5. The method of claim 1 , wherein the text data comprises at least one of a sentence, paragraph, or document.

6. A method for multilingual machine learning on text data, comprising:

receiving labeled and/or unlabeled English training data comprising at least one indication of human conduct that violates a policy, ethical standard, and/or law;

receiving labeled and/or unlabeled non-English training data;

applying a pre-trained cross-lingual model to the received training data to generate supervised learning task specific layers;

generating a semi-supervised multi-lingual model comprising applying the pre-trained cross-lingual model to input data to generate an output and applying the supervised learning task specific layers to the output to generate a classification; and

applying the semi supervised multi-lingual model to unseen multi-lingual text data to perform functions including the classification of unseen text data to detect at least one identification of a potential violation condition.

7. The method of claim 6 , wherein the text data corresponds to one or more electronic communications between persons.

8. The method of claim 6 , wherein the unseen multi-lingual text data comprises at least one of a sentence, paragraph, or document that includes multiple languages.

9. A method for multilingual machine learning on text data, comprising:

receiving labeled data, unlabeled data, multi-lingual data, and human annotated data;

applying a selection criteria;

based at least in part on the selection criteria, providing text data to one or more annotators to perform one or more functions that include at least one of assigning a label to the text data, evaluating semantic similarity, or translating the text data;

generating a multi-lingual model based at least in part on the functions performed by the one or more annotators; and

applying the multi-lingual model to unseen text data to produce a pre-trained cross-lingual model.

10. The method of claim 9 , further comprising using an annotator allocation resource cost and/or budget to determine the one or more functions performed by the one or more annotators.

11. The method of claim 10 , wherein a number and/or type of the one or more annotators is based at least in part on the resource cost and/or budget.

12. The method of claim 9 , wherein a type of the one or more annotators comprises an English-speaking subject matter expert or a multi-lingual subject matter expert.

13. The method of claim 9 , wherein the selection criteria comprises at least one of classifier confusion, data space coverage, colloquial language, domain-specific language, or idiomatic language.

14. A system for multilingual machine learning on text data, comprising:

a plurality of cross-lingual encoders coupled to a corresponding plurality of neural networks; and

an attention-based ensemble encoder component configured to receive an output of the plurality of neural networks, wherein the system is configured to classify the outputs of the plurality of neural networks by detecting at least one identification of a potential violation condition.

15. The system of claim 14 , further comprising a component configured for performing concatenation, received from the attention-based ensemble encoder.

16. The system of claim 14 , further comprising a fully connected neural network configured to produce a cross-lingual sentence representation.

17. The system of claim 14 , wherein an output of each of the plurality of cross-lingual encoders is input into a corresponding one of the corresponding plurality of neural networks.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2025
From: KAMATH, UDAY; EMARA, WAEL
To: DIGITAL REASONING SYSTEMS, INC.
Reel/Frame 070063/0736 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2022
From: KAMATH, UDAY; EMARA, WAEL
To: DIGITAL REASONING SYSTEMS, INC.
Reel/Frame 060348/0552 →
Continuity (2)
Provisional Application 63067450 · Aug 19, 2020
Related Publication 20230177281A1 · Jun 8, 2023
References Cited (22)
US 9923931B1 · Wagster et al. · 2018 [cited by applicant]
US 10445424B2 · Medlock · 2019 [cited by examiner]
US 10878184B1 · Estes et al. · 2020 [cited by applicant]
US 11797530B1 · Bouyarmane · 2023 [cited by examiner]
US 20140350920A1 · Medlock · 2014 [cited by examiner]
US 20200265356A1 · Lee · 2020 [cited by examiner]
US 20200357391A1 · Ghoshal · 2020 [cited by examiner]
KR 20200092446A · 2019 [cited by examiner]
T. Mensink, E. Gavves and C. G. M. Snoek, “COSTA: Co-Occurrence Statistics for Zero-Shot Classification,” 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 2014, pp. 2441-2448, doi: 10.… [cited by examiner]
Y. Xian, B. Schiele and Z. Akata, “Zero-Shot Learning—The Good, the Bad and the Ugly,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 3077-3086, doi: 10.1109/CVPR.20… [cited by examiner]
Zero-Shot Cross-Lingual Opinion Target Extraction Soufian Jebbara and Philipp Cimiano Semalytix GmbH, Bielefeld, Germany Semantic Computing Group, CITEC—Bielefeld University, Bielefeld, Germany (Year: 2019). [cited by examiner]
T. Mensink, E. Gawes and C. G. M. Snoek, “COSTA: Co-Occurrence Statistics for Zero-Shot Classification,” 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 2014, pp. 2441-2448, doi: 10.1… [cited by examiner]
Y. Xian, B. Schiele and Z. Akata, “Zero-Shot Learning—The Good, the Bad and the Ugly,” 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 3077-3086, doi: 10.1109/CVPR.20… [cited by examiner]
Zero-Shot Cross-Lingual Opinion Target Extraction Soufian Jebbara and Philipp Cimiano Semalytix GmbH, Bielefeld, Germany Semantic Computing Group, CITEC—Bielefeld University, Bielefeld, Germany (Year: 2019) (Year: 2019). [cited by examiner]
Zero-Shot Cross-Lingual Opinion Target Extraction Soufian Jebbara and Philipp Cimiano Semalytix GmbH, Bielefeld, Germany Semantic Computing Group, CITEC—Bielefeld University, Bielefeld, Germany (Year: 2019) (Year: 2019)… [cited by examiner]
T. Mensink, E. Gawes and C. G. M. Snoek, “COSTA: Co-Occurrence Statistics for Zero-Shot Classification,” 2014 IEEE Conference U_ | on Computer Vision and Pattern Recognition, Columbus, OH, USA, 2014, pp. 2441-2448, doi:… [cited by examiner]
Y. Xian, B. Schiele and Z. Akata, “Zero-Shot Learning—The Good, the Bad and the Ugly,” 2017 IEEE Conference on Computer V_ | Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 2017, pp. 3077-3086, doi: 10.1109/CV… [cited by examiner]
Zewen Chi et al: “Can Monolingual Pretrained Models Help Cross-Lingual Classification?”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 10, 2019 (Nov. 10, 2019), pp. 1-… [cited by applicant]
Yichao Lu et al: “A neural interlingua for multilingual machine translation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 23, 2018 (Apr. 23, 2018), pp. 1-9. [cited by applicant]
Dong Xin XD48@Rutgers Edu et al: “Leveraging Adversarial Training in Self-Learning for Cross-Lingual Text Classification”, Proceedings of the 21th ACM International Conference on Intelligent Virtual Agents, ACMPUB27, Ne… [cited by applicant]
Sara Rosenthal et al: “A Large-Scale Semi-Supervised Dataset for Offensive Language Identification”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Apr. 29, 2020 (Apr. 29, 2… [cited by applicant]
PCT Application No. PCT/US2021/046725, International Search Report and Written Opinion, Feb. 11, 2022, 19 pages. [cited by applicant]