IP Library › Granted Patent US 12,579,050
Granted Patent B2
US 12,579,050 · App. 18/228,423 · Granted Mar 17, 2026

Large language models for creating a multi-lingual, low-resource code translation dataset

Inventors: Zilu Tang (Cambridge, MA); Mayank Agarwal (Somerville, MA); Jie Chen (Briarcliff Manor, NY); Alexander Gregory Shypula (East Brunswick, NJ); Bailin Wang (Cambridge, MA); Yoon Hyung Kim (Cambridge, MA)
Assignees: International Business Machines Corporation; MASSACHUSETTS INSTITUTE OF TECHNOLOGY
G06F11/3608G06F8/51G06F40/47G06F40/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,050
App. No.
18/228,423
Granted
Mar 17, 2026
Kind
B2
Abstract

One or more unit-test cases are generated from a monolingual code corpus and the generated unit-test cases are filtered to generate a corpus of unit-test cases which have acceptability scores exceeding one or more predefined thresholds. One or more of the code samples of the monolingual code corpus are translated from a source language to a target language using a pretrained Large Language Model and the generated unit-test cases are translated from the source language to the target language. The LLM-translated code samples are validated using the translated unit-test cases and a parallel-data training corpus comprising the LLM-translated code samples that pass the validation is created. The pretrained large language model (LLM) is fine-tuned using the parallel-data training corpus, a given code segment is translated using the fine-tuned large language model (LLM), the translated given code segment is tested and the tested given code segment is deployed.

Claims (57)

1 . A method comprising:

generating, using at least one hardware processor, one or more unit-test cases from a monolingual code corpus;

filtering, using the at least one hardware processor, the generated unit-test cases to generate a corpus of unit-test cases which have acceptability scores exceeding one or more predefined thresholds;

translating, using the at least one hardware processor, one or more code samples of the monolingual code corpus from a source language to a target language using a pretrained large language model (LLM);

translating, using the at least one hardware processor, the generated unit-test cases from the source language to the target language;

validating, using the at least one hardware processor, the LLM-translated code samples using the translated unit-test cases;

creating, using the at least one hardware processor, a parallel-data training corpus comprising the LLM-translated code samples that pass the validation;

fine-tuning, using the at least one hardware processor, the pretrained large language model (LLM) using the parallel-data training corpus;

translating, using the at least one hardware processor, a given code segment using the fine-tuned large language model (LLM);

testing, using the at least one hardware processor, the translated given code segment; and

facilitating, using the at least one hardware processor, deployment of the tested given code segment in the target language.

2 . The method of claim 1 , further comprising:

collecting one or more the one or more of the code samples in one or more programming languages; and

extracting the monolingual code corpus from the collection of code samples.

3 . The method of claim 2 , wherein the filtering of the generated unit-test cases further comprises retaining the generated unit-test cases that have toolkit-metrics exceeding the one or more predefined thresholds.

4 . The method of claim 2 , further comprising performing code translation of the given segment of code in the source language to the target language using the Large Language Model.

5 . The method of claim 2 , wherein the large language model is a multi-billion parameter machine learning model pretrained in an unsupervised manner.

6 . The method of claim 2 , wherein the parallel-data training corpus comprises functionally-equivalent implementation of logic in multiple programming languages.

7 . The method of claim 2 , wherein the generated unit-test cases are expected input-output pairs with assert statements.

8 . The method of claim 2 , further comprising repeating the translating of the one or more of the code samples of the monolingual code corpus, the validating, and the creating operations for all code samples of the monolingual code corpus.

9 . The method of claim 2 , further comprising retraining the large language model using the parallel-data training corpus.

10 . The method of claim 2 , further comprising repeating the translating of the one or more of the code samples of the monolingual code corpus, the validating, and the creating operations for code samples of the monolingual code corpus that failed the validation operation.

11 . The method of claim 2 , further comprising verifying that the code samples of the monolingual code corpus and the translated code are functionally equivalent using the translated unit-test cases.

12 . The method of claim 1 , wherein the acceptability scores are one or more of a coverage score and a mutation score.

13 . The method of claim 1 , further comprising running the deployed tested given code segment.

14 . A computer program product, comprising:

one or more tangible computer-readable storage media and program instructions stored on at least one of the one or more tangible computer-readable storage media, the program instructions executable by a processor, the program instructions comprising:

generating one or more unit-test cases from a monolingual code corpus;

filtering the generated unit-test cases to generate a corpus of unit-test cases which have acceptability scores exceeding one or more predefined thresholds;

translating one or more code samples of the monolingual code corpus from a source language to a target language using a pretrained Large Language Model (LLM);

translating the generated unit-test cases from the source language to the target language;

validating the LLM-translated code samples using the translated unit-test cases;

creating a parallel-data training corpus comprising the LLM-translated code samples that pass the validation;

fine-tuning the pretrained large language model (LLM) using the parallel-data training corpus;

translating a given code segment using the fine-tuned large language model (LLM);

testing the translated given code segment; and

facilitating deployment of the tested given code segment in the target language.

15 . An apparatus comprising:

a memory; and

at least one processor, coupled to said memory, and operative to perform operations comprising:

generating one or more unit-test cases from a monolingual code corpus;

filtering the generated unit-test cases to generate a corpus of unit-test cases which have acceptability scores exceeding one or more predefined thresholds;

translating one or more code samples of the monolingual code corpus from a source language to a target language using a pretrained Large Language Model (LLM);

translating the generated unit-test cases from the source language to the target language using a rules-based translator;

validating the LLM-translated code samples using the translated unit-test cases; and

creating a parallel-data training corpus comprising the LLM-translated code samples that pass the validation;

fine-tuning the pretrained large language model (LLM) using the parallel-data training corpus;

translating a given code segment using the fine-tuned large language model (LLM);

testing the translated given code segment; and

facilitating deployment of the tested given code segment in the target language.

16 . The apparatus of claim 15 , the operations further comprising:

collecting one or more code samples in one or more programming languages; and

extracting the monolingual code corpus from the collection of the one or more of the code samples.

17 . The apparatus of claim 16 , wherein the filtering of the generated unit-test cases further comprises retaining the generated unit-test cases that have toolkit-metrics exceeding the one or more predefined thresholds.

18 . The apparatus of claim 16 , the operations further comprising performing code translation of the given segment of code in the source language to the target language using the large language model.

19 . The apparatus of claim 16 , the operations further comprising retraining the large language model using the parallel-data training corpus.

20 . The apparatus of claim 16 , the operations further comprising verifying that the code samples of the monolingual code corpus and the translated code are functionally equivalent using the translated unit-test cases.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2024
From: SHYPULA, ALEXANDER GREGORY; WANG, BAILIN; KIM, YOON HYUNG
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 066710/0107 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2023
From: TANG, ZILU; AGARWAL, MAYANK; CHEN, JIE
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064439/0615 →
Continuity (1)
Related Publication 20250045185A1 · Feb 6, 2025
References Cited (52)
US 10007594B2 · Andrejko · 2018 [cited by examiner]
US 10185713B1 · Denkowski · 2019 [cited by examiner]
US 10324695B2 · Champagne · 2019 [cited by applicant]
US 10606573B2 · Apte · 2020 [cited by applicant]
US 10664381B2 · Walters · 2020 [cited by applicant]
US 10684943B2 · Fei · 2020 [cited by examiner]
US 10740694B2 · Harvill · 2020 [cited by applicant]
US 10783456B2 · Strope · 2020 [cited by applicant]
US 10996935B2 · Jonnadula · 2021 [cited by applicant]
US 11037028B2 · Bojar · 2021 [cited by examiner]
US 11133001B2 · Andreas · 2021 [cited by applicant]
US 11145291B2 · Rusak · 2021 [cited by applicant]
US 11194550B2 · Davis · 2021 [cited by applicant]
US 11257272B2 · Rowell · 2022 [cited by applicant]
US 11269605B1 · Nandanuru · 2022 [cited by examiner]
US 11574017B2 · Boxwell · 2023 [cited by examiner]
US 11893363B2 · Drain · 2024 [cited by examiner]
US 11984116B2 · Haikin · 2024 [cited by examiner]
US 12001325B2 · Adachi · 2024 [cited by examiner]
US 20070043553A1 · Dolan · 2007 [cited by examiner]
US 20190340466A1 · Berseth · 2019 [cited by applicant]
US 20220066747A1 · Drain · 2022 [cited by examiner]
US 20220084510A1 · Peng · 2022 [cited by applicant]
CN 112905188A · 2021 [cited by applicant]
Roziere, Baptiste, et al. “Leveraging automated unit tests for unsupervised code translation.” arXiv preprint arXiv:2110.06773 (2021). pp. 1-20. (Year: 2021). [cited by examiner]
Tufano, Michele, et al. “Unit test case generation with transformers and focal context.” arXiv preprint arXiv:2009.05617 (2020). pp. 1-15. (Year: 2020). [cited by examiner]
Palomba, Fabio, et al. “Automatic test case generation: What if test code quality matters?.” Proceedings of the 25th International Symposium on Software Testing and Analysis. 2016. pp. 130-141. (Year: 2016). [cited by examiner]
Ahmed, Toufique, and Premkumar Devanbu. “Multilingual training for software engineering.” Proceedings of the 44th International Conference on Software Engineering. 2022.pp. 1443-1455. (Year: 2022). [cited by examiner]
Heering, Jan, and Paul Klint. “Towards monolingual programming environments.” ACM Transactions on Programming Languages and Systems (TOPLAS) 7.2 (1985): pp. 183-213. (Year: 1985). [cited by examiner]
Svyatkovskiy, Alexey, et al. “Intellicode compose: Code generation using transformer.” Proceedings of the 28th ACM joint meeting on European software engineering conference and symposium on the foundations of software e… [cited by examiner]
Khanuja, Simran, et al. “GLUECoS: An evaluation benchmark for code-switched NLP.” arXiv preprint arXiv:2004.12376 (2020). pp. 1-11. (Year: 2020). [cited by examiner]
Cai, Deng, et al. “Neural machine translation with monolingual translation memory.” arXiv preprint arXiv:2105.11269 (2021). pp. 1-12 (Year: 2021). [cited by examiner]
Chen, Fuxiang, et al. “On the transferability of pre-trained language models for low-resource programming languages.” Proceedings of the 30th IEEE/ACM international conference on program comprehension. 2022. pp. 401-412… [cited by examiner]
Cobol Blues downloaded from: http://fingfx.thomsonreuters.com/gfx/rngs/USA-BANKS-COBOL/010040KH18J/index.html Mar. 10, 2023 pp. 3. [cited by applicant]
Weisz et al., Better Together? An Evaluation of AI-Supported Code Translation. Feb. 15, 2022 pp. 1-35. [cited by applicant]
Project CodeNet downloaded from: https://github.com/IBM/Project_CodeNet May 5, 2021 pp. 1-8. [cited by applicant]
CodeXGLUE downloaded from: https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/code-to-code-trans Mar. 10, 2023 pp. 5. [cited by applicant]
Pynguin-Python General Unit Test Generator downloaded from: https://pynguin.readthedocs.io/en/latest/ Mar. 10, 2020 pp. 2. [cited by applicant]
EvoSuite | Automatic Test Suite Generation for Java downloaded from:https://www.evosuite.org/ Sep. 16, 2021 pp. 6. [cited by applicant]
Information Technology Agencies Need to Develop Modernization Plans for Critical Legacy Systems https://www.gao.gov/assets/700/699616.pdf Jun. 2019 pp. 79. [cited by applicant]
Introducing IBM Mono2Micro downloaded from: https://www.ibm.com/cloud/blog/announcements/ibm-mono2micro May 6, 2020 pp. 1-10. [cited by applicant]
Banks scramble to fix old systems as IT ‘cowboys’ ride into sunset. downloaded from https://www.reuters.com/article/us-usa-banks-cobol/banks-scramble-to-fix-old-systems-as-it-cowboys-ride-into-sunset-idUSKBN17C0D8 on Ma… [cited by applicant]
Haluptzok et al., Language Models Can Teach Themselves to Program Better (https://arxiv.org/abs/2207.14502) Apr. 12, 2023 pp. 1-23. [cited by applicant]
Large language model. Downloaded from https://en.wikipedia.org/wiki/Large_language_model on May 10, 2023. pp. 18. [cited by applicant]
Conneau et al. “Cross-Lingual Language Model Pretraining”, Advances in neural information processing systems, Dec. 2019, Article No. 634, pp. 7059-7069. [cited by applicant]
Hartvigsen et al. “ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection”, arXiv:2203.09509v4 [cs.CL], Jul. 14, 2022, 18 pages. [cited by applicant]
Ian King. “An Ancient Computer Language Is Slowing America's Giant Stimulus”, Technology, Economics, Apr. 13, 2020, 2 pages. [cited by applicant]
IBM. “Application Modernization Consulting and Services”, downloaded from https://web.archive.org/web/20230328011150/https://www.ibm.com/consulting/application-modernization, Mar. 28, 2023, 10 pages. [cited by applicant]
IBM. “Application Modernization”, downloaded from https://www.ibm.com/support/pages/application-modernization, May 18, 2022, 2 pages. [cited by applicant]
IBM. “Modernize applications for interoperability and ROI”, downloaded from https://web.archive.org/web/20210425081107/https://www.ibm.com/cloud/application-modernization, Apr. 25, 2021, 09 Pages. [cited by applicant]
Lachaux et al. “Unsupervised Translation of Programming Languages”, arXiv:2006.03511v3 [cs.CL], Sep. 22, 2020, 21 pages. [cited by applicant]
Roziere et al. “Leveraging Automated Unit Tests for Unsupervised Code Translation”, arXiv preprint arXiv:2110.06773, Feb. 16, 2022, 20 pages. [cited by applicant]