IP Library › Granted Patent US 12,242,820
Granted Patent B2
US 12,242,820 · App. 17/651,555 · Granted Mar 4, 2025

Generating synthetic code-switched data for training language models

Inventors: Cesa Salaam (Upper Marlboro, MD); Seunghyun Yoon (San Jose, CA); Trung Huu Bui (San Jose, CA); Franck Dernoncourt (San Jose, CA)
Assignee: Adobe Inc.
G06F40/58G06F40/47G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,820
App. No.
17/651,555
Granted
Mar 4, 2025
Kind
B2
Abstract

Techniques for training a language model for code switching content are disclosed. Such techniques include, in some embodiments, generating a dataset, which includes identifying one or more portions within textual content in a first language, the identified one or more portions each including one or more of offensive content or non-offensive content; translating the identified one or more salient portions to a second language; and reintegrating the translated one or more portions into the textual content to generate code-switched textual content. In some cases, the textual content in the first language includes offensive content and non-offensive content, the identified one or more portions include the offensive content, and the translated one or more portions include a translated version of the offensive content. In some embodiments, the code-switched textual content is at least part of a synthetic dataset usable to train a language model, such as a multilingual classification model.

Claims (33)

1. A method of generating code-switched content for training a language model, the method comprising:

identifying one or more portions within textual content in a first language, the identified one or more portions each comprising one or more of offensive content or non-offensive content, the identifying comprising tagging, based on an output of a first trained language model, the one or more portions with at least one content tag;

translating the tagged one or more portions to a second language using a second trained language model; and

replacing, in the textual content, the tagged one or more portions with the translated one or more portions to generate code-switched textual content;

wherein:

the first trained language model comprises a context-aware language model configured to identify one or more salient portions from the textual content based at least on: one or more simplex or complex noun phrases of the textual content, one or more maximally frequent contiguous portions of the textual content, one or more salient portions of the textual content determined based at least on a training dataset used by the first trained language model, one or more denoised portions of the textual content, or a combination thereof; and

the second trained language model comprises a language model configured to translate between at least the first and second languages.

2. The method of claim 1 , wherein:

the textual content in the first language comprises the offensive content and the non-offensive content;

the tagged one or more portions comprise the offensive content; and

the translated one or more portions in the second language comprise a translated version of the offensive content.

3. The method of claim 1 , further comprising causing a multilingual model to be trained using the generated code-switched textual content.

4. The method of claim 1 , wherein the first trained language model has been trained using at least a real-world abusive speech dataset, and is configured to include the offensive content within the tagged one or more portions.

5. The method of claim 1 , wherein the tagging of the one or more portions comprises inserting one or more of an end-of-sentence token or a beginning-of-sentence token.

6. The method of claim 3 , wherein the multilingual model comprises a Cross-lingual Language Model Robustly Optimized Bidirectional Encoder Representations from Transformers (BERT) Pre-training Approach (XLM-ROBERTa).

7. The method of claim 5 , wherein the inserting of the end-of-sentence token comprises adding the end-of-sentence token to an end of the one or more portions, and the inserting of the beginning-of-sentence token comprises adding the beginning-of-sentence token to a beginning of the one or more portions.

8. The method of claim 7 , further comprising discarding one or more portions of the textual content in the first language which are not tagged.

9. A method of generating code-switched content for training a language model, the method comprising:

identifying one or more portions within textual content in a first language, the identified one or more portions each comprising one or more of offensive content or non-offensive content, the identifying comprising tagging, based on an output of a first trained language model, the one or more portions with at least one content tag;

translating the tagged one or more portions to a second language using a second trained language model;

replacing, in the textual content, the tagged one or more portions with the translated one or more portions to generate code-switched textual content; and

causing a multilingual model comprising a Cross-lingual Language Model Robustly Optimized Bidirectional Encoder Representations from Transformers (BERT) Pre-training Approach (XLM-ROBERTa) to be trained using the generated code-switched textual content.

10. The method of claim 9 , wherein:

the textual content in the first language comprises the offensive content and the non-offensive content;

the tagged one or more portions comprise the offensive content; and

the translated one or more portions in the second language comprise a translated version of the offensive content.

11. The method of claim 9 , wherein:

the first trained language model comprises a context-aware language model configured to identify one or more salient portions from the textual content based at least on: one or more simplex or complex noun phrases of the textual content, one or more maximally frequent contiguous portions of the textual content, one or more salient portions of the textual content determined based at least on a training dataset used by the first trained language model, one or more denoised portions of the textual content, or a combination thereof; and

the second trained language model comprises a language model configured to translate between at least the first and second languages.

12. The method of claim 9 , wherein the first trained language model has been trained using at least a real-world abusive speech dataset, and is configured to include the offensive content within the tagged one or more portions.

13. The method of claim 9 , wherein the tagging of the one or more portions comprises inserting an end-of-sentence token, a beginning-of-sentence token, or both.

14. The method of claim 13 , wherein the inserting of the end-of-sentence token comprises adding the end-of-sentence token to an end of the one or more portions, and the inserting of the beginning-of-sentence token comprises adding the beginning-of-sentence token to a beginning of the one or more portions.

15. The method of claim 14 , further comprising discarding one or more portions of the textual content in the first language which are not tagged.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2022
From: SALAAM, CESA; YOON, SEUNGHYUN; BUI, TRUNG HUU; DERNONCOURT, FRANCK
To: ADOBE INC.
Reel/Frame 059042/0252 →
Continuity (1)
Related Publication 20230259718A1 · Aug 17, 2023
References Cited (39)
US 6782510B1 · Gross · 2004 [cited by examiner]
US 9405741B1 · Schaaf · 2016 [cited by examiner]
US 11010687B2 · Mehdad · 2021 [cited by examiner]
US 11095585B2 · Prabhu · 2021 [cited by examiner]
US 11829717B1 · Chen · 2023 [cited by examiner]
US 11861512B1 · Joshi · 2024 [cited by examiner]
US 20100280828A1 · Fein · 2010 [cited by examiner]
US 20110137948A1 · Davis · 2011 [cited by examiner]
US 20110307241A1 · Waibel · 2011 [cited by examiner]
US 20140337989A1 · Orsini · 2014 [cited by examiner]
US 20150309987A1 · Epstein · 2015 [cited by examiner]
US 20200126533A1 · Doyle · 2020 [cited by examiner]
Aguilar, G. et al. “LinCE: A Centralized Benchmark for Linguistic Code-switching Evaluation”, Proc. of the 12th Conference on Language Resources and Evaluation (LREC 2020), pp. 1803-1813, 2020. [cited by applicant]
Bohra, A. et al. “A Dataset of Hindi-English Code-Mixed Social Media Text for Hate Speech Detection”, Proc. of the Second Workshop on Computational Modeling of People's Opinions, Personality, and Emotions in Social Medi… [cited by applicant]
Chakravarthi, B. R. et al. “Corpus Creation for Sentiment Analysis in Code-Mixed Tamil-English Text”, Proc. of the 1st Joint Spoken Language Technologies for Under-resourced languages (STLU) and Collaboration and Comput… [cited by applicant]
Claesson, C. et al. “Weakly Supervised Deep Learning Classification: Concept Classification from Electronic Health Records without Ground Truth Data”, Master's thesis in Complex Adaptive Systems, Department of Physics, … [cited by applicant]
Conneau, A. et al. “Unsupervised Cross-lingual Representation Learning at Scale”, Proc. of the 58th Meeting of the Association for Computational Linguistics, pp. 8440-8451, 2020. [cited by applicant]
Devlin J., et al., “BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding”, Proc. of NAACL-HLT 2019, Available online at: https://arxiv.org/pdf/1810.04805.pdf, May 24, 2019, 16 pages. [cited by applicant]
Fan, A. et al. “Beyond English-Centric Multilingual Machine Translation”, Journal of Machine Learning Research, 22:1-48, 2021. [cited by applicant]
Gaskins, D. et al. “A Crosslinguistic Study of Child Code-Switching within the Noun Phrase: A Usage-Based Perspective”, Languages 6(1): 29, Feb. 13, 2021, doi: 10.3390/languages6010029. [cited by applicant]
Gu, X. et al. “UCPhrase: Unsupervised Context-aware Quality Phrase Tagging”, KDD '21, Proc. of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining, pp. 478-486, Aug. 2021, doi: 10.1145/3447548.3467397. [cited by applicant]
Gupta, A. et al. “Unsupervised Self-Training for Sentiment Analysis of Code-Switched Data”, Proc. of the Fifth Workshop on Computational Approaches to Linguistic Code-Switching (2021), arXiv:2103.14797v2 [cs.CL], Oct. 1… [cited by applicant]
Hatebase [online]. Hatebase—a collaborative, regionalized repository of multilingual hate speech. URL: https://hatebase.org/. [cited by applicant]
Kapoor, R. et al. “Mind Your Language: Abuse and Offense Detection for Code-Switched Languages”, Proc. of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), vol. 33:9951-9952, Jul. 2019, doi: 10.1609… [cited by applicant]
Mandl, T. et al. “Overview of the HASOC Track at FIRE 202: Hate Speech and Offensive Content Identification in Indo-European Languages”, CEUR Workshop Proceedings: Forum for Information Retrieval Evaluation (FIRE '20), … [cited by applicant]
Mathew, B. et al. “HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection”, Proc. of the Thirty-Fifth AAAI Conference on Artificial Intelligence (AAAI-21), 35(17):14867-14875, May 18, 2021. [cited by applicant]
Mehra, S. “Detection of Offensive Language in Social Media Posts”, Thesis for: MSc Artificial Intelligence, Faculty of Engineering and Science Department of Computer Science, Cork Institute of Technology, May 2020, doi:… [cited by applicant]
Pant, K. et al. “Towards Code-switched Classification Exploiting Constituent Language Resources”, Proc. of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th Int… [cited by applicant]
Parafita Couto, M. C. et al. “Code-switching within the noun phrase: Evidence from three corpora”, International Journal of Bilingualism, 23(2):695-714, 2019, doi: 10.1177/1367006917729543. [cited by applicant]
Patwa, P. et al. “SemEval—2020 Task 9: Overview of Sentiment Analysis of Code-Mixed Tweets”, Proc. of the 14th International Workshop on Semantic Evaluation, pp. 774-790, Dec. 2020, doi: 10.18653/v1/2020.semeval-1.100. [cited by applicant]
Pitsilis, G. et al. “Effective hate-speech detection in Twitter data using recurrent neural networks”, Applied Intelligence, 48:4730-4742, Jul. 26, 2018, doi: 10.1007/s10489-018-1242-y. [cited by applicant]
Pratapa, A. et al. “Language Modeling for Code-Mixing: The Role of Linguistic Theory based Synthetic Data”, Proc. of the 56th Annual Meeting of the Association for Computational Linguistics (Long Papers), pp. 1543-1553,… [cited by applicant]
Qin, L. et al. “CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual NLP”, Proc. of the Twenty-Ninth International Joint Conference on Artificial Intelligence (IJCAI-20), arXiv:2006.06402… [cited by applicant]
Roy, S. G. et al. “Leveraging Multilingual Transformers for Hate Speech Detection”, CEUR Workshop Proceedings: Forum for Information Retrieval Evaluation (FIRE '20), arXiv:2101.03207v1 [cs.CL], Jan. 8, 2021. [cited by applicant]
Saha, K. et al. “Prevalence and Psychological Effects of Hateful Speech in Online College Communities”, WebSci '19: Proc. of the 10th ACM Conference on Web Science, vol. 2019:255-264, 2019, doi: 10.1145/3292522.3326032. [cited by applicant]
Samanta, B. et al. “NEVAE: A Deep Generative Model for Molecular Graphs”, Proc. of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), 33(01):1110-1117, Jul. 17, 2019, doi: 10.1609/aaai.v33i01.3301111… [cited by applicant]
Sanh, V. et al. “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter”, 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing—NeurIPS 2019, arXiv:1910.01108v4 [cs.CL], Mar. 1… [cited by applicant]
Tang, T. et al. “Fine-Tuning BERT for Multi-Label Sentiment Analysis in Unbalanced Code-Switching Text”, IEEE Access, vol. 8:193248-193256, 2020, doi: 10.1109/ACCESS.2020.3030468. [cited by applicant]
Winata, G. I. et al. “Code-Switched Language Models Using Neural Based Synthetic Data from Parallel Sentences”, Proc. of the 23rd Conference on Computational Natural Language Learning (CoNLL), pp. 271-280, Nov. 2019, do… [cited by applicant]
Cited By (1)
US 12,499,309