IP Library › Granted Patent US 12,498,910
Granted Patent B2
US 12,498,910 · App. 18/202,564 · Granted Dec 16, 2025

Training syntax-aware language models with AST path prediction

Inventors: Pritam Dash (Vancouver, CA); Arno Schneuwly (Effretikon, CH); Saeid Allahdadian (Vancouver, CA); Matteo Casserini (Zurich, CH); Felix Schmidt (Baden-Dattwi, CH)
Assignee: Oracle International Corporation
G06F8/427
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,498,910
App. No.
18/202,564
Granted
Dec 16, 2025
Kind
B2
Abstract

In an embodiment, a computer stores and operates a logic encoder that is an artificial neural network that infers a fixed-size encoded logic from textual or tokenized source logic. Without machine learning, a special parser generates a parse tree that represents the source logic and a fixed-size correctly encoded tree that represents the parse tree. For finetuning the logic encoder, an encoded tree generator is an artificial neural network that accepts the fixed-size encoded logic as input and responsively infers a fixed-size incorrectly encoded tree that represents the parse tree. The neural weights of the logic encoder (and optionally of the encoded tree generator) are adjusted based on backpropagation of error (i.e. loss) as a numerically measured difference between the fixed-size incorrectly encoded tree and the fixed-size correctly encoded tree.

Claims (39)

1 . A method comprising:

increasing accuracy of a logic encoder that is a bidirectional encoder by finetuning the logic encoder, including repeating in a predetermined count of repetitions:

a) inferring, by the logic encoder, a fixed-size encoded logic from a source logic;

b) generating a parse tree that represents the source logic;

c) generating a fixed-size correctly encoded tree that represents the parse tree that represents the source logic;

d) inferring, based on the fixed-size encoded logic, by an encoded tree generator, a fixed-size incorrectly encoded tree that represents the parse tree that represents the source logic; and

e) neural backpropagating, in the logic encoder, a difference between the fixed-size incorrectly encoded tree and the fixed-size correctly encoded tree; and

performing after said increasing the accuracy of the logic encoder:

deploying, into a production environment, the logic encoder without the encoded tree generator;

generating, by the logic encoder, a second fixed-size encoded logic for a second source logic in an integrated development environment (IDE); and

performing, by the IDE based on the second fixed-size encoded logic, source logic completion;

wherein the method is performed by one or more computers.

2 . The method of claim 1 wherein:

said inferring the fixed-size incorrectly encoded tree is performed by the encoded tree generator;

the method further comprises adjusting the encoded tree generator based on said difference between the fixed-size incorrectly encoded tree and the fixed-size correctly encoded tree.

3 . The method of claim 1 further comprising sizing, based on a count of distinct tree paths in parse trees in a training corpus, the fixed-size correctly encoded tree and the fixed-size incorrectly encoded tree.

4 . The method of claim 1 wherein said inferring the fixed-size incorrectly encoded tree comprises inferring a plurality of distinct tree paths in the parse tree that represents the source logic.

5 . The method of claim 4 wherein said inferring the plurality of distinct tree paths comprises inferring a respective count of each tree path in the plurality of distinct tree paths.

6 . The method of claim 1 wherein said generating the fixed-size correctly encoded tree is based on at least one selected from a group consisting of: a predefined minimum path length, a predefined maximum path length, and non-terminal nodes of the parse tree that represents the source logic but not terminal nodes.

7 . The method of claim 1 further comprising based on the second fixed-size encoded logic, performing at least one selected from a group consisting of: source logic translation, source logic documentation generation, and programing reference material recommendation.

8 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

increasing accuracy of a logic encoder that is a bidirectional encoder by finetuning the logic encoder, including repeating in a predetermined count of repetitions:

a) inferring, by the logic encoder, a fixed-size encoded logic from a source logic;

b) generating a parse tree that represents the source logic;

c) generating a fixed-size correctly encoded tree that represents the parse tree that represents the source logic;

d) inferring, based on the fixed-size encoded logic, by an encoded tree generator, a fixed-size incorrectly encoded tree that represents the parse tree that represents the source logic; and

e) neural backpropagating, in the logic encoder, a difference between the fixed-size incorrectly encoded tree and the fixed-size correctly encoded tree; and

performing after said increasing the accuracy of the logic encoder:

deploying, into a production environment, the logic encoder without the encoded tree generator;

generating, by the logic encoder, a second fixed-size encoded logic for a second source logic in an integrated development environment (IDE); and

performing, by the IDE based on the second fixed-size encoded logic, source logic completion.

9 . The one or more non-transitory computer-readable media of claim 8 wherein:

said inferring the fixed-size incorrectly encoded tree is performed by the encoded tree generator;

the instructions further cause adjusting the encoded tree generator based on said difference between the fixed-size incorrectly encoded tree and the fixed-size correctly encoded tree.

10 . The one or more non-transitory computer-readable media of claim 8 wherein the instructions further cause sizing, based on a count of distinct tree paths in parse trees in a training corpus, the fixed-size correctly encoded tree and the fixed-size incorrectly encoded tree.

11 . The one or more non-transitory computer-readable media of claim 8 wherein said inferring the fixed-size incorrectly encoded tree comprises inferring a plurality of distinct tree paths in the parse tree that represents the source logic.

12 . The one or more non-transitory computer-readable media of claim 11 wherein said inferring the plurality of distinct tree paths comprises inferring a respective count of each tree path in the plurality of distinct tree paths.

13 . The one or more non-transitory computer-readable media of claim 8 wherein said generating the fixed-size correctly encoded tree is based on at least one selected from a group consisting of: a predefined minimum path length, a predefined maximum path length, and non-terminal nodes of the parse tree that represents the source logic but not terminal nodes.

14 . The one or more non-transitory computer-readable media of claim 8 wherein the instructions further cause based on the second fixed-size encoded logic, performing at least one selected from a group consisting of: source logic translation, source logic documentation generation, and programing reference material recommendation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 26, 2023
From: DASH, PRITAM; SCHNEUWLY, ARNO; ALLAHDADIAN, SAEID; CASSERINI, MATTEO; SCHMIDT, FELIX
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 063775/0126 →
Continuity (2)
Provisional Application 63459014 · Apr 13, 2023
Related Publication 20240345815A1 · Oct 17, 2024
References Cited (65)
US 11656851B2 · Clement · 2023 [cited by examiner]
US 12073195B2 · Duan · 2024 [cited by applicant]
US 12130809B2 · Miller · 2024 [cited by applicant]
US 20170255611A1 · Kubosawa · 2017 [cited by applicant]
US 20210185066A1 · Shah · 2021 [cited by applicant]
US 20220197917A1 · Schneuwly et al. · 2022 [cited by applicant]
US 20220198294A1 · Schneuwly · 2022 [cited by examiner]
US 20220261228A1 · Schneuwly et al. · 2022 [cited by applicant]
US 20240289606A1 · Wang · 2024 [cited by examiner]
CN 112381280B · 2023 [cited by examiner]
EP 2128798 · 2009 [cited by applicant]
Zeng, Jie, et al., “Fast Code Clone Detection Based on Weighted Recursive Autoencoders”, IEEE Access, vol. 7; pp. 125062-125078, 2019. doi: 10.1109/ACCESS.2019.2938825, published Sep. 2, 2019, 17pgs. [cited by applicant]
White, Martin, et al., “Deep Learning Code Fragments for Code Clone Detection”, 31st IEEE/ACM ICASE 2016, dx.doi.org/10.1145/2970276.2970326, pp. 87-98, publ Aug. 25, 2016, 12pgs. [cited by applicant]
Wang, Xin, et al., “SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation”, https://arxiv.org/abs/2108.04556, Sep. 9, 2021, 9pgs. [cited by applicant]
Jiang, Lingxiao, et al., “Deckard: Scalable and Accurate Tree-based Detection of Code Clones”, 29th ICSE '07, doi: 10.1109/ICSE.2007.30, pp. 96-105, May 24, 2007, 10pgs. [cited by applicant]
Gao, Yi, et al., “TECCD: A Tree Embedding Approach for Code Clone Detection”, 2019 IEEE CSME, doi: 10.1109/ICSME.2019.00025, pp. 145-156, 2019, 12pgs. [cited by applicant]
Fang, Chunrong, et al., “Functional code clone detection with syntax and semantics fusion learning”, Proc of the 29th ACM SIGSOFT ISSTA '20, pp. 516-527, doi: 10.1145/3395363.3397362, Jul. 18, 2020, 12pgs. [cited by applicant]
Johnson, Jeff, et al., “Billion-scale similarity search with GPUs”, https://arxiv.org/pdf/1702.08734.pdf, Feb. 28, 2017, 12pgs. [cited by applicant]
Feng, Zhangyin, et al., “CodeBERT: A Pre-Trained Model for Programming and Natural Languages”, Findings of the Assoctn for Computnl Linguistics: EMNLP 2020, pp. 1536-1547, https://doi.org/10.18653/v1/2020.findings-emnlp… [cited by applicant]
Chen, Mark, et al., “Evaluating Large Language Models Trained on Code”, https://arxiv.org/abs/2107.03374, Jul. 14, 2021, 35pgs. [cited by applicant]
“Stack Overflow—Where Developers Learn, Share, & Build Careers”, Stack Exchange Inc., https://stackoverflow.com/, 2023, 13pgs. [cited by applicant]
“Stack Exchange Data Dump: Stack Exchange, Inc.”, https://archive.org/details/stackexchange, 2023, 8pgs. [cited by applicant]
“ChatGPT: Optimizing Language Models for Dialogue”, GOpenAI, LLC, https://chat.openai.com/, 2023, 1pg. [cited by applicant]
Li, Raymond et al., “Starcoder: may the source be with you!”, CoRR, abs/2305.06161, 2023, available: https://doi.org/10.48550/arXiv.2305.06161. [cited by applicant]
Alon, Uri et al., “code2vec: learning distributed representations of code”, Proc. Acm Program. Lang., 3(POPL):40:1-40:29, 2019. [cited by applicant]
Ben Allal, Loubna et al., “Santacoder: don't reach for the stars!”, CoRR, abs/2301.03988, 2023, available: https://doi.org/10.48550/arXiv.2301.03988. [cited by applicant]
Brown et al. “Language models are few- shot learner”, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020. [cited by applicant]
Chen, Mark et al., “Evaluating large language models trained on code”, CoRR, abs/2107.03374, 2021, available: https://arxiv.org/ abs/2107.03374. [cited by applicant]
Chowdhery, Aakanksha et al., “Palm: Scaling language modeling with pathways”, CORR, abs/2204.02311, 2022, available: https://doi.org/10.48550/arXiv.2204.02311. [cited by applicant]
Devlin, Jacob et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics… [cited by applicant]
Guo, Daya et al., “Graphcodebert: Pre-training code representations with data flow”, 9th International Conference on Learning Representations, ICLR 2021, Austria, May 3-7, 2021, available: https:// openreview.net/forum?… [cited by applicant]
Guo, Daya et al., “Unixcoder: Unified cross-modal pre-training for code representation”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistic, ACL 2022, May 22-27, 2022, pp. 7212-7225. [cited by applicant]
Izadi, Maliheh et al., “Codefill: Multi-token code completion by jointly learning from structure and naming sequences”, Proceedings of the 44th International Conference On Software Engineering, ICSE '22, New York, NY, U… [cited by applicant]
Jiang, Xue et al., “A tree-based pre-trained model for programming language”, Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence, Jul. 27-30, 2021, vol. 161 Of Proceedings of Machine … [cited by applicant]
Alon, Uri et al., “code2seq: Generating sequences from structured representations of code”, 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, OpenReview.net, 2019. [cited by applicant]
Kim, Seohyun et al., “Code prediction by feeding trees to transformers”, 43rd IEEE/ACM International Conference on Software Engineering, ICSE 2021, Madrid, Spain, May 22-30, 2021, pp. 150-162, available: https://doi.org… [cited by applicant]
Zu{umlaut over ( )}gner, Daniel et al., “Language-Agnostic Representation Learning of Source Code from Structure and Context”, 9th International Conference on Learning Representations, ICLR 2021, May 3-7, 2021, availabl… [cited by applicant]
Liu, Yapeng et al., “Improving code completion by sequence features and structural features”, Proceedings of the 4th World Symposium on Software Engineering, WSSE '22, 2022, pp. 51-58. [cited by applicant]
OpenAI, “GPT-4 technical report”, CoRR, abs/2303.08774, 2023, available: https://doi.org/10.48550/arXiv.2303.08774. [cited by applicant]
Peng, Han et al., “Integrating Tree Path in Transformer for Code Representation”, Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 9343-9354. [cited by applicant]
Radford, Alec et al., “Improving language understanding by generative pre-training”, 2018. [cited by applicant]
Radford, Alec et al., “Language models are unsupervised multitask learners”, Openai Blog, 1(8):9, 2019. [cited by applicant]
Tang, Ze, et al., “Ast-trans: Code summarization with efficient tree-structured attention”, 2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE), 2022, pp. 150-162. [cited by applicant]
Touvron, Hugo et al., “Llama: Open and efficient foundation language models”, CoRR, abs/2302.13971, 2023, available: https://doi.org/10.48550/arXiv.2302.13971. [cited by applicant]
Vaswani, Ashish et al., “Attention is all you need”, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, Dec. 4-9, 2017, Long Beach, CA, USA, pp. 5998-6… [cited by applicant]
Wang, Xiao et al., “Heloc: hierarchical contrastive learning of source code representation”, Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension, ICPC 2022, Virtual Event, May 16-17, 2022,… [cited by applicant]
Wang, Xin et al., “SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation”, 2021. [cited by applicant]
Wang, Yanlin et al., “Code completion by modeling flattened abstract syntax trees as graphs”, Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Feb. 2-9, 2021, pp. 14015-14023. [cited by applicant]
Kanade, Aditya et al., “Learning and Evaluating Contextual Embedding of Source Code”, Proceedings of the 37th International Conference on Machine Learning, Jul. 13-18, 2020, vol. 119 of Proceedings of Machine Learning R… [cited by applicant]
He et al., “A Reusable SQL Injection Detection Method for Java Web Applications”, KSII Transactions on Internet and Information Systems vol. 14, No. 6, dated Jun. 2020, 15 pages. [cited by applicant]
Alon et al., “A General Path-Based Representation for Predicting Program Properties” dated Apr. 22, 2018, 16 pages. [cited by applicant]
Alon, et al., “code2vec: Learning Distributed Representations of Code”, Proc. ACM Program. Lang., vol. 3, No. POPL, Article 40. Publication date: Jan. 2019, 29 pages. [cited by applicant]
Apel et al., “Learning SQL for Database Intrusion Detection using Context-sensitive Modelling”, dated 2009, 33 pages. [cited by applicant]
Bockermann et al., “Learning SQL for Database Intrusion Detection Using Context-Sensitive Modelling”, DIMVA dated 2009, LNCS 5587, 10 pages. [cited by applicant]
Cai et al., “An Abstract Syntax Tree Encoding Method for Cross-Project Defect Prediction”, IEEE, dated Nov. 15, 2019, 10 pages. [cited by applicant]
“Tree-sitter”, Introduction, https://tree-sitter.github.io/tree-sitter/, last accessed Jun. 9, 2023, 4pgs. [cited by applicant]
Guo, Daya, et al., “UniXcoder: Unified Cross-Modal Pre-training for Code Representation”, https://arxiv.org/abs/2203.03850, Mar. 8, 2022, 14pgs. [cited by applicant]
Zugner, Daniel, et al., “Language-Agnostic Representation Learning of Source Code from Structure and Context”, ICLR 2021, https://arxiv.org/abs/2103.11318, Mar. 21, 2021, 22pgs. [cited by applicant]
Hinton, Geoffrey, et al., “Distilling the Knowledge in a Neural Network”, https://arxiv.org/abs/1503.02531, Mar. 9, 2015, 9pgs. [cited by applicant]
Jimenez et al., “On the Impact of Tokenizer and Parameters on N-Gram Based Code Analysis”, dated 2018, 12 pages. [cited by applicant]
Mou et al., “Building Program Vector Representations for Deep Learning”, dated Sep. 11, 2014, 11 pages. [cited by applicant]
Parr, Terence, “The Definitive ANTLR 4 Reference”, The Pragmatic Bookshelf, copyright 2012The Pragmatic Programmers, LLC, www.it-ebooks.info, book version Jan. 2013, 322pgs. [cited by applicant]
Yamaguchi et al., “Generalized Vulnerability Extrapolation using Abstract Syntax Trees”, ACSAC '12 Dec. 3-7, 2012, Orlando, Florida USA, 10 pages. [cited by applicant]
Zhang, Jian, et al., “A Novel Neural Source Code Representation Based on Abstract Syntax Tree”, 2019 IEEE/ACM 41st ICSE, doi: 10.1109/ICSE.2019.00086, pp. 783-794, May 25, 2019, 12pgs. [cited by applicant]
Devlin, Jacob, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, https://arxiv.org/abs/1810.04805, May 24, 2019, 16pgs. [cited by applicant]