IP Library › Granted Patent US 12,493,717
Granted Patent B2
US 12,493,717 · App. 18/318,315 · Granted Dec 9, 2025

Multi-lingual natural language generation

Inventors: Praneet Pabolu (Bangalore, IN); Karan Dua (Najibabad, IN); Sriram Chaudhury (Bangalore, IN)
Assignee: Oracle International Corporation
G06F21/6254G06F16/345G06F40/166G06F40/216G06F40/284G06F40/40G06F40/47G06F40/56G06F40/58G06N3/045G06N3/09G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,493,717
App. No.
18/318,315
Granted
Dec 9, 2025
Kind
B2
Abstract

A method includes preparing a base model using an input model pretrained on at least three languages different from each other and a base vocabulary including words corresponding to two languages among the at least three languages, where the preparing the base model includes constraining the input model to the words included in the base vocabulary; training the base model using a first enhanced training dataset generated from public data, to generate a text summarization model; training the base model using a second enhanced training dataset generated from the first enhanced training dataset, to generate a text generation model; and training the base model using a third enhanced training dataset that is generated using the second enhanced training dataset and the text summarization model, to generate a next sentence generation model.

Claims (76)

1 . A computer-implemented method comprising:

preparing a base model using an input model pretrained on at least three languages different from each other and a base vocabulary including words corresponding to two languages among the at least three languages, wherein the preparing the base model comprises constraining the input model to the words included in the base vocabulary;

training the base model using a first enhanced training dataset generated from public data, to generate a text summarization model;

training the base model using a second enhanced training dataset generated from the first enhanced training dataset, to generate a text generation model; and

training the base model using a third enhanced training dataset that is generated using the second enhanced training dataset and the text summarization model, to generate a next sentence generation model,

wherein the generated text generation model is trained to, based on an input of one or more first keywords, output text including at least one first keyword among the one or more first keywords,

wherein the generated text summarization model is trained to, based on an input of the text, output a text summary, and

wherein the generated next sentence generation model is trained to, based on an input of one or more second keywords and the text summary, output a next sentence that is appendable to the text generated by the text generation model and includes at least one second keyword among the one or more second keywords.

2 . The computer-implemented method of claim 1 , wherein each of the text generation model, the text summarization model, and the next sentence generation model is a bilingual model that is trained on the two languages including a target language and English, and configured to output predictions based on an input provided in the target language, English, or a mixed language in which the target language and English are intermixed.

3 . The computer-implemented method of claim 1 , wherein the input model is a transformer-based model.

4 . The computer-implemented method of claim 1 , wherein the input model is an mT5 model.

5 . The computer-implemented method of claim 4 , wherein:

the base model is a modified mT5 model, and

in the base model, a vocabulary of the mT5 model is restricted to a first number of words in a target language and a second number of words in English.

6 . The computer-implemented method of claim 1 , wherein the first enhanced training dataset comprises article-summary pairs, and for each article-summary pair, an article serves as an input training datapoint and a corresponding summary serves as a given output,

the second enhanced training dataset comprises keyword-text pairs, and, for each keyword-text pair, one or more keywords serve as an input training datapoint and a corresponding text serves as a given output, and

the third enhanced training dataset comprises summary-keywords-next sentence triplets, and, for each summary-keywords-next sentence triplet, a summary and keywords serve as an input training datapoint and a corresponding next sentence serves as a given output.

7 . The computer-implemented method of claim 1 , further comprising:

generating a plurality of refined models by training the text generation model and the next sentence generation model using a plurality of refined training datasets generated using the text summarization model and private data.

8 . The computer-implemented method of claim 7 , wherein:

the plurality of refined models comprises a refined text generation model and a refined next sentence generation model,

the generating the plurality of refined models further comprises:

training the text generation model using a first refined training dataset among the plurality of refined training datasets, to generate the refined text generation model, wherein the first refined training dataset is generated based on the private data and includes fake values given to entity values included in the private data, and

training the next sentence generation model using a second refined training dataset among the plurality of refined training datasets, to generate the refined next sentence generation model, wherein the second refined training dataset is generated using the first refined training dataset and the text summarization model,

the refined text generation model is trained to, based on an input of one or more first entity values, output text including at least one first entity value among the one or more first entity values, and

the refined next sentence generation model is trained to, based on an input of one or more second entity values and a primary text summary generated by the text summarization model based on the text output by the refined text generation model, output a next sentence that is appendable to the text output by the refined text generation model and includes at least one second entity value among the one or more second entity values.

9 . A system comprising:

one or more processors;

a memory that is coupled to the one or more processors and stores one or more instructions that, when executed by the one or more processors, cause the one or more processors to perform a method including:

preparing a base model using an input model pretrained on at least three languages different from each other and a base vocabulary including words corresponding to two languages among the at least three languages, wherein the preparing the base model includes constraining the input model to the words included in the base vocabulary;

training the base model using a first enhanced training dataset generated from public data, to generate a text summarization model;

training the base model using a second enhanced training dataset generated from the first enhanced training dataset, to generate a text generation model; and

training the base model using a third enhanced training dataset that is generated using the second enhanced training dataset and the text summarization model, to generate a next sentence generation model,

wherein the generated text generation model is trained to, based on an input of one or more first keywords, output text including at least one first keyword among the one or more first keywords,

wherein the generated text summarization model is trained to, based on an input of the text, output a text summary, and

wherein the generated next sentence generation model is trained to, based on an input of one or more second keywords and the text summary, output a next sentence that is appendable to the text generated by the text generation model and includes at least one second keyword among the one or more second keywords.

10 . The system of claim 9 , wherein each of the text generation model, the text summarization model, and the next sentence generation model is a bilingual model that is trained on the two languages including a target language and English, and configured to output predictions based on an input provided in the target language, English, or a mixed language in which the target language and English are intermixed.

11 . The system of claim 9 , wherein:

the base model is a modified mT5 model, and

in the base model, a vocabulary of an mT5 model is restricted to a first number of words in a target language and a second number of words in English.

12 . The system of claim 9 , wherein the first enhanced training dataset comprises article-summary pairs, and for each article-summary pair, an article serves as an input training datapoint and a corresponding summary serves as a given output,

the second enhanced training dataset comprises keyword-text pairs, and, for each keyword-text pair, one or more keywords serve as an input training datapoint and a corresponding text serves as a given output, and

the third enhanced training dataset comprises summary-keywords-next sentence triplets, and, for each summary-keywords-next sentence triplet, a summary and keywords serve as an input training datapoint and a corresponding next sentence serves as a given output.

13 . The system of claim 9 , wherein the method further includes:

generating a plurality of refined models by training the text generation model and the next sentence generation model using a plurality of refined training datasets generated using the text summarization model and private data.

14 . The system of claim 13 , wherein:

the plurality of refined models comprises a refined text generation model and a refined next sentence generation model,

the generating the plurality of refined models further includes:

training the text generation model using a first refined training dataset among the plurality of refined training datasets, to generate the refined text generation model, wherein the first refined training dataset is generated based on the private data and includes fake values given to entity values included in the private data, and

training the next sentence generation model using a second refined training dataset among the plurality of refined training datasets, to generate the refined next sentence generation model, wherein the second refined training dataset is generated using the first refined training dataset and the text summarization model,

the refined text generation model is trained to, based on an input of one or more first entity values, output text including at least one first entity value among the one or more first entity values, and

the refined next sentence generation model is trained to, based on an input of one or more second entity values and a primary text summary generated by the text summarization model based on the text output by the refined text generation model, output a next sentence that is appendable to the text output by the refined text generation model and includes at least one second entity value among the one or more second entity values.

15 . A non-transitory computer-readable memory storing one or more instructions that, when executed by one or more processors, cause the one or more processors to perform a method including:

preparing a base model using an input model pretrained on at least three languages different from each other and a base vocabulary including words corresponding to two languages among the at least three languages, wherein the preparing the base model includes constraining the input model to the words included in the base vocabulary;

training the base model using a first enhanced training dataset generated from public data, to generate a text summarization model;

training the base model using a second enhanced training dataset generated from the first enhanced training dataset, to generate a text generation model; and

training the base model using a third enhanced training dataset that is generated using the second enhanced training dataset and the text summarization model, to generate a next sentence generation model,

wherein the generated text generation model is trained to, based on an input of one or more first keywords, output text including at least one first keyword among the one or more first keywords,

wherein the generated text summarization model is trained to, based on an input of the text, output a text summary, and

wherein the generated next sentence generation model is trained to, based on an input of one or more second keywords and the text summary, output a next sentence that is appendable to the text generated by the text generation model and includes at least one second keyword among the one or more second keywords.

16 . The non-transitory computer-readable memory of claim 15 , wherein each of the text generation model, the text summarization model, and the next sentence generation model is a bilingual model that is trained on the two languages including a target language and English, and configured to output predictions based on an input provided in the target language, English, or a mixed language in which the target language and English are intermixed.

17 . The non-transitory computer-readable memory of claim 15 , wherein:

the base model is a modified mT5 model, and

in the base model, a vocabulary of an mT5 model is restricted to a first number of words in a target language and a second number of words in English.

18 . The non-transitory computer-readable memory of claim 15 , wherein the first enhanced training dataset comprises article-summary pairs, and for each article-summary pair, an article serves as an input training datapoint and a corresponding summary serves as a given output,

the second enhanced training dataset comprises keyword-text pairs, and, for each keyword-text pair, one or more keywords serve as an input training datapoint and a corresponding text serves as a given output, and

the third enhanced training dataset comprises summary-keywords-next sentence triplets, and, for each summary-keywords-next sentence triplet, a summary and keywords serve as an input training datapoint and a corresponding next sentence serves as a given output.

19 . The non-transitory computer-readable memory of claim 15 , wherein the method further includes:

generating a plurality of refined models by training the text generation model and the next sentence generation model using a plurality of refined training datasets generated using the text summarization model and private data.

20 . The non-transitory computer-readable memory of claim 19 , wherein:

the plurality of refined models comprises a refined text generation model and a refined next sentence generation model,

the generating the plurality of refined models further includes:

training the text generation model using a first refined training dataset among the plurality of refined training datasets, to generate the refined text generation model, wherein the first refined training dataset is generated based on the private data and includes fake values given to entity values included in the private data, and

training the next sentence generation model using a second refined training dataset among the plurality of refined training datasets, to generate the refined next sentence generation model, wherein the second refined training dataset is generated using the first refined training dataset and the text summarization model,

the refined text generation model is trained to, based on an input of one or more first entity values, output text including at least one first entity value among the one or more first entity values, and

the refined next sentence generation model is trained to, based on an input of one or more second entity values and a primary text summary generated by the text summarization model based on the text output by the refined text generation model, output a next sentence that is appendable to the text output by the refined text generation model and includes at least one second entity value among the one or more second entity values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2023
From: PABOLU, PRANEET; DUA, KARAN; CHAUDHURY, SRIRAM
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 063658/0190 →
Continuity (2)
Provisional Application 63416779 · Oct 17, 2022
Related Publication 20240127008A1 · Apr 18, 2024
References Cited (38)
US 10546054B1 · Foroughi et al. · 2020 [cited by applicant]
US 11907672B2 · Nugent · 2024 [cited by examiner]
US 20190347570A1 · Wang · 2019 [cited by examiner]
US 20210004485A1 · Summers et al. · 2021 [cited by applicant]
US 20210097201A1 · Wasicek et al. · 2021 [cited by applicant]
US 20210326652A1 · Hazard et al. · 2021 [cited by applicant]
US 20210390127A1 · Fox et al. · 2021 [cited by applicant]
US 20220059224A1 · Tulley · 2022 [cited by examiner]
US 20220180234A1 · Kamthe et al. · 2022 [cited by applicant]
US 20240054290A1 · Iyer et al. · 2024 [cited by applicant]
US 20240249068A1 · Aldred et al. · 2024 [cited by applicant]
EP 3985540A1 · 2022 [cited by applicant]
Unified Language Model Pre-training for Natural Language understanding and generation, Dong et al, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Oct. 15, 2019, 13 pages (Year: 2019). [cited by examiner]
AdaFT: an efficient domain-adaptive fine-tuning framework for sentiment analysis in Chinese financial texts, Yan et al, Applied Intelligence, Apr. 29, 2025,23 pgs (Year: 2025). [cited by examiner]
U.S. Appl. No. 18/318,327, Non-Final Office Action, Mailed On Jul. 2, 2025, 34 pages. [cited by applicant]
Cao et al., “MultiSumm: Towards a Unified Model for Multi-Lingual Abstractive Summarization”, Proceedings of the Association for the Advancement of Artificial Intelligence Conference on Artificial Intelligence, vol. 34,… [cited by applicant]
Dong et al., “Unified Language Model Pre-training for Natural Language Understanding and Generation”, 33rd Conference on Neural Information Processing Systems., Oct. 15, 2019, 13 pages. [cited by applicant]
Beyond English-Centric Multilingual Machine Translation, Available Online at: https://github.com/facebookresearch/fairseq/tree/main/examples/m2m_100, Sep. 20, 2021, 7 pages. [cited by applicant]
C4, Datasets, Available Online at: https://www.tensorflow.org/datasets/catalog/c4#c4multilingual_nights_stay, Dec. 6, 2022, 15 pages. [cited by applicant]
Dataset Card for XL-Sum, Available Online at: https://huggingface.co/datasets/csebuetnlp/xlsum, 11 pages, retrieved Apr. 21, 2023. [cited by applicant]
Language-agnostic BERT Sentence Embedding (LaBSE), Available Online at: https://github.com/bojone/labse, Aug. 24, 2020, 3 pages. [cited by applicant]
M2M100 418M, Available Online at: https://huggingface.co/facebook/m2m100_418M, 2010, 5 pages, retrieved Apr. 21, 2023. [cited by applicant]
Welcome to the Leipzig Corpora Collection/Deutscher Wortschatz, A project of Leipzig University, the Saxon Academy of Sciences and Humanities in Leipzig and the Institute for Applied Informatics. Available Online at: ht… [cited by applicant]
Bennani-Smires et al., Simple Unsupervised Keyphrase Extraction using Sentence Embeddings, Available Online at: https://arxiv.org/pdf/1801.04470.pdf, Sep. 5, 2018, 9 pages. [cited by applicant]
Carbonell et al., The Use of MMR, Diversity-Based Reranking for Reordering Documents and Producing Summaries, In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Informati… [cited by applicant]
Dorr et al., Testing Code Using Synthetic Data, Technical Disclosure Commons, Oct. 2, 2019, pp. 1-8. [cited by applicant]
Fan et al., Beyond English-Centric Multilingual Machine Translation, Computation and Language, Available Online at: https://arxiv.org/abs/2010.11125, Oct. 21, 2020, pp. 1-38. [cited by applicant]
Feng et al., Language-Agnostic BERT Sentence Embedding, Computation and Language, Available Online at: https://arxiv.org/abs/2007.01852, Mar. 8, 2022, 14 pages. [cited by applicant]
Galloni et al., A Novel Evaluation Metric for Synthetic Data Generation, Intelligent Data Engineering and Automated Learning, Oct. 2020, 2 pages. [cited by applicant]
Grootendorst, Keyword Extraction with BERT, A Minimal Method for Extracting Keywords and Keyphrases, Towards Data Science, Oct. 29, 2020, pp. 1-19. [cited by applicant]
Hendricks, Generate Data Containing Fake Personally Identifiable Information, Available Online at: https://github.com/paulhendricks/generator, Aug. 26, 2015, 8 pages. [cited by applicant]
Hillborn, Anonymizing Datasets at Scale Leveraging Databricks Interoperability, Databricks, Feb. 13, 2017, 8 pages. [cited by applicant]
Kondamari et al., Custom Models and Text Translation Come to OCI Language, Oracle AI & Data Science Blog, Available Online at: https://blogs.oracle.com/ai-and-datascience/post/custom-models-text-translation-oci-language… [cited by applicant]
Ladhak et al., Wiki_lingua, Available Online at: https://gem-benchmark.com/data_cards/wiki_lingua, 2020, 5 pages. [cited by applicant]
Lee, Natural Language Generation for Electronic Health Records, npj Digital Medicine, vol. 1, No. 63, Nov. 19, 2018, 7 pages. [cited by applicant]
Stahlberg et al., C4_200M Synthetic Dataset for Grammatical Error Correction, Available Online at: https://github.com/google-research-datasets/C4_200M-synthetic-dataset-for-grammatical-error-correction, Aug. 10, 2021, p… [cited by applicant]
Xue et al., mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer, Computation and Language, Mar. 11, 2021, 17 pages. [cited by applicant]
U.S. Appl. No. 18/318,327, Notice of Allowance, mailed Oct. 29, 2025, 15 pages. [cited by applicant]
Cited By (1)
US 12,675,653