IP Library Granted Patent US 12,436,745
Granted Patent B2
US 12,436,745 · App. 18/382,018 · Granted Oct 7, 2025

Developing a programming language model for machine learning tasks

Inventors: Mahinthan Chandramohan (Brisbane, AU); Behnaz Hassanshahi (Brisbane, AU); Padmanabhan Krishnan (Brisbane, AU); Dai Nguyen (Kelvin Grove, AU)
Assignee: Oracle International Corporation
G06F8/35G06F40/284G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,436,745
App. No.
18/382,018
Granted
Oct 7, 2025
Kind
B2
Abstract

A method develops a programming language model for machine learning tasks. The method includes adjusting a token list to include a language token used by a tokenizer for a pretrained language model. The pretrained language model includes a set of layers. The set of layers includes a set of initial layers, an embedding layer, and an output layer. The method further includes performing an output layer modification of the output layer to replace the output vector with the embedding vector. The method further includes freezing the set of initial layers to generate a set of frozen layers of the pretrained language model that do not update during training. The method further includes training the pretrained language model using the language token, the output layer modification, and the set of frozen layers to form a fine-tuned model from the pretrained language model.

Claims (49)

1. A method comprising:

adjusting a token list to include a language token used by a tokenizer for a pretrained language model,

wherein the pretrained language model comprises a set of layers,

wherein the set of layers comprises a set of initial layers, an embedding layer, and an output layer, and

wherein the output layer generates an output vector from an embedding vector generated by the embedding layer;

performing an output layer modification of the output layer to replace the output vector with the embedding vector;

freezing the set of initial layers to generate a set of frozen layers of the pretrained language model that do not update during training; and

training the pretrained language model using the language token, the output layer modification, and the set of frozen layers to form a fine-tuned model from the pretrained language model, wherein training the pretrained language model comprises backpropagating a difference between a training output vector and an expected vector to a set of end layers of the pretrained language model to form the fine-tuned model.

2. The method of claim 1 , further comprising:

processing an input text using the fine-tuned model to generate an output text that is responsive to the input text.

3. The method of claim 1 , wherein the language token corresponds to a keyword of a programming language.

4. The method of claim 1 , wherein the language token corresponds to a name from an application programming interface of a programming language.

5. The method of claim 1 , wherein the language token corresponds to a name from a standard library of a programming language.

6. The method of claim 1 , further comprising:

training the pretrained language model with a set of training input tokens comprising one or more of a set of programming language tokens, a set of natural language tokens, and a set of syntax tree tokens.

7. The method of claim 1 , wherein the set of frozen layers includes the set of initial layers and does not include a set of end layers of the pretrained language model.

8. The method of claim 1 , wherein the set of frozen layers includes the embedding layer.

9. The method of claim 1 , wherein the set of frozen layers does not include the output layer.

10. The method of claim 1 , wherein training the pretrained language model using the language token, the output layer modification, and the set of frozen layers comprises:

processing a training input vector using the pretrained language model to generate the training output vector; and

comparing the training output vector to the expected vector.

11. A system comprising:

at least one processor;

an application executing on the at least one processor to perform:

adjusting a token list to include a language token used by a tokenizer for a pretrained language model,

wherein the pretrained language model comprises a set of layers,

wherein the set of layers comprises a set of initial layers, an embedding layer, and an output layer, and

wherein the output layer generates an output vector from an embedding vector generated by the embedding layer;

performing an output layer modification of the output layer to replace the output vector with the embedding vector;

freezing the set of initial layers to generate a set of frozen layers of the pretrained language model that do not update during training; and

training the pretrained language model using the language token, the output layer modification, and the set of frozen layers to form a fine-tuned model from the pretrained language model, wherein training the pretrained language model comprises backpropagating a difference between a training output vector and an expected vector to a set of end layers of the pretrained language model to form the fine-tuned model.

12. The system of claim 11 , wherein the application is further configured to perform:

processing an input text using the fine-tuned model to generate an output text that is responsive to the input text.

13. The system of claim 11 , wherein the language token corresponds to a keyword of a programming language.

14. The system of claim 11 , wherein the language token corresponds to a name from an application programming interface of a programming language.

15. The system of claim 11 , wherein the language token corresponds to a name from a standard library of a programming language.

16. The system of claim 11 , wherein the application is further configured to perform:

training the pretrained language model with a set of training input tokens comprising one or more of a set of programming language tokens, a set of natural language tokens, and a set of syntax tree tokens.

17. The system of claim 11 , wherein the set of frozen layers includes the set of initial layers and does not include a set of end layers of the pretrained language model.

18. The system of claim 11 , wherein the set of frozen layers includes the embedding layer.

19. The system of claim 11 , wherein the set of frozen layers does not include the output layer.

20. A non-transitory computer readable storage medium storing computer readable program code which, when executed by a processor, performs:

adjusting a token list to include a language token used by a tokenizer for a pretrained language model,

wherein the pretrained language model comprises a set of layers,

wherein the set of layers comprises a set of initial layers, an embedding layer, and an output layer, and

wherein the output layer generates an output vector from an embedding vector generated by the embedding layer;

performing an output layer modification of the output layer to replace the output vector with the embedding vector;

freezing the set of initial layers to generate a set of frozen layers of the pretrained language model that do not update during training; and

training the pretrained language model using the language token, the output layer modification, and the set of frozen layers to form a fine-tuned model from the pretrained language model, wherein training the pretrained language model comprises backpropagating a difference between a training output vector and an expected vector to a set of end layers of the pretrained language model to form the fine-tuned model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2023
From: CHANDRAMOHAN, MAHINTHAN; HASSANSHAHI, BEHNAZ; KRISHNAN, PADMANABHAN; NGUYEN, DAI
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 065312/0826 →
Continuity (1)
Related Publication 20250130780A1 · Apr 24, 2025
References Cited (44)
US 11429352B2 · Pujar · 2022 [cited by examiner]
US 11562142B2 · Nijkamp · 2023 [cited by examiner]
US 11893347B2 · Klein · 2024 [cited by examiner]
US 12249252B1 · Ghosh · 2025 [cited by examiner]
US 12361215B2 · Wei · 2025 [cited by examiner]
US 20200312301A1 · Polovets · 2020 [cited by examiner]
US 20220107828A1 · Mostafa · 2022 [cited by examiner]
US 20220245350A1 · Ormerod · 2022 [cited by examiner]
Liu, Fang, et al. “Multi-task learning based pre-trained language model for code completion.” Proceedings of the 35th IEEE/ACM international conference on automated software engineering. 2020. pp. 473-483. (Year: 2020). [cited by examiner]
Feng, Zhangyin, et al. “Codebert: A pre-trained model for programming and natural languages.” arXiv preprint arXiv:2002.08155 (2020). pp. 1-12. (Year: 2020). [cited by examiner]
Gururangan, Suchin, et al. “Don't stop pretraining: Adapt language models to domains and tasks.” arXiv preprint arXiv:2004. 10964 (2020). pp. 1-19. (Year: 2020). [cited by examiner]
Dehaerne, Enrique, et al. “Code generation using machine learning: A systematic review.” Ieee Access 10 (2022): pp. 82434-82455. (Year: 2022). [cited by examiner]
Perez, Luis, Lizi Ottens, and Sudharshan Viswanathan. “Automatic code generation using pre-trained language models.” arXiv preprint arXiv:2102.10535 (2021). pp. 1-9. (Year: 2021). [cited by examiner]
Rahmani, Kia, et al. “Multi-modal program inference: a marriage of pre-trained language models and component-based synthesis.” Proceedings of the ACM on Programming Languages 5.OOPSLA (2021): pp. 158:1-158:29. (Year: 20… [cited by examiner]
Ahmadi, M. et al., “Finding bugs using your own code: Detecting functionally-similar yet inconsistent code”, In USENIX Security Symposium, pp. 2025-2040, Aug. 11-13, 2021 (17 pages). [cited by applicant]
Al-Omari, F. et al., “SemanticCloneBench: A semantic code clone benchmark using crowd-source knowledge”, In IEEE 14th International Workshop on Software Clones (IWSC), pp. 57-63, Feb. 1, 2020 (7 pages). [cited by applicant]
Devlin, J. et al., “BERT: Pretraining of deep bidirectional transformers for language understanding”, arXiv preprint arXiv:1810.04805, May 24, 2019 (16 pages). [cited by applicant]
Feng, Z. et al., “CodeBERT: A pre-trained model for programming and natural languages”, In Findings of EMNLP, pp. 1536-1547, Nov. 16-20, 2020 (12 pages). [cited by applicant]
Guo, D. et al., “UniXcoder: Unified cross-modal pre-training for code representation” arXiv preprint arXiv:2203.03850, Mar. 8, 2022 (14 pages). [cited by applicant]
Guo, D. et al., Graphcodebert: Pre-training code representations with data flow. In ICLR, Sep. 13, 2021 (18 pages). [cited by applicant]
Hochreiter, S. et al., “Long short-term memory”, Neural Computation, 9:1735-1780, Nov. 15, 1997 (32 pages). [cited by applicant]
Husain, H., et al., “Codesearchnet challenge: Evaluating the state of semantic code search”, arXiv preprint arXiv:1909.09436, Jun. 8, 2020 (6 pages). [cited by applicant]
Kim, Y. “Convolutional neural networks for sentence classification” In EMNLP, pp. 1746-1751, Oct. 25-29, 2014 (6 pages). [cited by applicant]
Kipf, T. N. et al., “Semi-supervised classification with graph convolutional networks” In ICLR, Feb. 22, 2017 14 pages). [cited by applicant]
Kim, S. et al., “Vuddy: A scalable approach for vulnerable code clone discovery”, In IEEE Symposium on Security and Privacy, May 22-26, 2017 (20 pages). [cited by applicant]
Liu, Y. et al., RoBERTa: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, Jul. 26, 2019 (13 pages). [cited by applicant]
Li, Y. et al., “Gated Graph Sequence Neural Networks”, In ICLR, Sep. 22, 2017 (20 pages). [cited by applicant]
Li, Y. et al., Do pre-trained language models indeed understand software engineering tasks? arXiv:2211.10623, Nov. 19, 2022 (16 pages). [cited by applicant]
Li, Z. et al., “VulDeePecker: A deep learning-based system for vulnerability detection”, arXiv preprint arXiv:1801.01681, Jan. 5, 2018 (15 pages). [cited by applicant]
Vandermaaten, L. et al., “Visualizing Data using t-SNE”, Journal of machine learning research, pp. 2579-2605, Nov. 1, 2008 (27 pages). [cited by applicant]
Mou, L. et al., “Convolutional neural networks over tree structures for programming language processing”, In AAAI, Dec. 5, 2015 (8 pages). [cited by applicant]
Neuhaus, S. et al., “Predicting vulnerable software components”, In ACM CCS, pp. 529-540, Oct. 29, 2007 (12 pages). [cited by applicant]
Ren, S. et al., “CodeBLEU: a method for automatic evaluation of code synthesis” arXiv preprint arXiv:2009.10297, Sep. 27, 2020 (8 pages). [cited by applicant]
Russell, R. et al., “Automated vulnerability detection in source code using deep representation learning”, In ICMLA, Dec. 1, 2018 (6 pages). [cited by applicant]
Sennrich, R. et al., “Neural machine translation of rare words with subword units”, In ACL, pp. 1715-1725, Aug. 7-12, 2016 (11 pages). [cited by applicant]
Svajlenko, J. et al., “Towards a big data curated benchmark of inter-project code clones”, In ICSME, pp. 476-480, Jan. 1, 2014 (5 pages). [cited by applicant]
Shin, Y. et al., “Evaluating complexity, code churn, and developer activity metrics as indicators of software vulnerabilities”, IEEE Transactions on Software Engineering, 37, Mar. 18, 2009 (34 pages). [cited by applicant]
Woo, S. et al., “Movery: A precise approach for modified vulnerable code clone discovery from modified open-source software components”, In USENIX Security Symposium, pp. 3037-3053, Aug. 10-12, 2022 (18 pages). [cited by applicant]
Wang, W. et al., “Detecting code clones with graph neural network and flow-augmented abstract syntax tree”, In SANER, pp. 261-271, Feb. 20, 2020 (11 pages). [cited by applicant]
White, M. et al., “Deep learning code fragments for code clone detection”, In ASE, pp. 87-98, Sep. 3-7, 2016 (12 pages). [cited by applicant]
Wan, Y. et al., “What do they capture?—A structural analysis of pre-trained language models for source code”, In CSE, Feb. 14, 2022 (12 pages). [cited by applicant]
Zhao, G. et al., “Deepsim: deep learning code functional similarity”, In ESEC/FSE, pp. 141-151, Oct. 1, 2018 (11 pages). [cited by applicant]
Zhou, Y. et al., “Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks”, In NeurIPS, Sep. 8, 2019 (11 pages). [cited by applicant]
Zhang, Y. et al., “ASTRO: An AST-assisted approach for generalizable neural clone detection”, In Proceedings of Automated Software Engineering Industrial Showcase (ASE), Aug. 17, 2022 (5 pages). [cited by applicant]