IP Library Granted Patent US 12,566,596
Granted Patent B2
US 12,566,596 · App. 18/235,461 · Granted Mar 3, 2026

Graph path prediction and masked language modelling joint training algorithm for language models

Inventors: Tomas Feith (Zurich, CH); Arno Schneuwly (Effretikon, CH); Saeid Allahdadian (Vancouver, CA); Matteo Casserini (Zurich, CH); Felix Schmidt (Baden-Dattwil, CH)
Assignee: Oracle International Corporation
G06F8/425G06F8/427G06F16/9024G06F16/9027G06N3/045G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,596
App. No.
18/235,461
Granted
Mar 3, 2026
Kind
B2
Abstract

In an embodiment providing natural language processing (NLP), a computer generates a histogram that correctly represents a graph that represents a lexical text, and generates a token sequence encoder that is trainable and untrained. During training such as pretraining, the token sequence encoder infers an encoded sequence that incorrectly represents the lexical text, and the encoded sequence is dense and saves space. To increase the accuracy of the sequence encoder by learning, the token sequence encoder is adjusted based on, as discussed herein, an indirectly measured numeric difference between the encoded sequence that incorrectly represents the lexical text and the histogram that correctly represents the graph.

Claims (67)

1 . A method comprising:

training a token sequence encoder that is untrained by:

a) generating a histogram that correctly represents a graph that represents a lexical text;

b) inferring from the lexical text, by the token sequence encoder, an encoded sequence that incorrectly represents the lexical text;

c) generating a decoded histogram by decoding the encoded sequence that incorrectly represents the lexical text;

d) measuring a difference between the decoded histogram and the histogram that correctly represents the graph; and

e) adjusting the token sequence encoder based on said difference between said histograms; and

inferring from a second lexical text, by the token sequence encoder, a second encoded sequence that correctly represents the second lexical text;

wherein the method is performed by one or more computers.

2 . The method of claim 1 wherein:

said decoding the encoded sequence that incorrectly represents the lexical text is performed by a linear decoder;

the token sequence encoder is nonlinear.

3 . The method of claim 1 wherein:

the method further comprises multitask learning by the token sequence encoder;

said inferring the encoded sequence that incorrectly represents the lexical text and said adjusting the token sequence encoder occur during said multitask learning by the token sequence encoder.

4 . The method of claim 1 wherein:

the method further comprises:

generating a lossy representation of the lexical text, and

inferring, by the token sequence encoder, an encoded sequence that incorrectly represents the lossy representation of the lexical text;

said adjusting the token sequence encoder is further based on the encoded sequence that incorrectly represents the lossy representation of the lexical text.

5 . The method of claim 4 wherein:

said generating the lossy representation of the lexical text comprises excluding at least one token from a token sequence that represents the lexical text;

the method further comprises inferring, by a machine learning model, the at least one token;

said adjusting based on the encoded sequence that incorrectly represents the lossy representation of the lexical text is based on said inferring the at least one token.

6 . The method of claim 1 wherein said generating the histogram that correctly represents the graph comprises counting occurrences of a particular n-gram in the graph.

7 . The method of claim 1 wherein the graph is at least one selected from a group consisting of a directed graph and a dataflow graph.

8 . The method of claim 1 wherein:

the encoded sequence that incorrectly represents the lexical text has a fixed size;

said inferring the encoded sequence that incorrectly represents the lexical text comprises accepting a variable length sequence of lexical tokens.

9 . The method of claim 8 wherein each lexical token in the variable length sequence of lexical tokens represents multiple vertices.

10 . A method comprising:

detecting that a parse error in a lexical text prevents generation of a graph that represents the lexical text;

generating, by a token sequence encoder, a fixed-size encoded sequence that represents the lexical text; and

inferring, from the fixed-size encoded sequence that represents the lexical text, a code completion of an expression in a logic statement in the lexical text;

wherein the method is performed by an integrated development environment (IDE).

11 . One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause:

training a token sequence encoder that is untrained by:

a) generating a histogram that correctly represents a graph that represents a lexical text;

b) inferring from the lexical text, by the token sequence encoder, an encoded sequence that incorrectly represents the lexical text;

c) generating a decoded histogram by decoding the encoded sequence that incorrectly represents the lexical text;

d) measuring a difference between the decoded histogram and the histogram that correctly represents the graph; and

e) adjusting the token sequence encoder based on said difference between said histograms; and

inferring from a second lexical text, by the token sequence encoder, a second encoded sequence that correctly represents the second lexical text.

12 . The one or more non-transitory computer-readable media of claim 11 wherein:

said decoding the encoded sequence that incorrectly represents the lexical text is performed by a linear decoder;

the token sequence encoder is nonlinear.

13 . The one or more non-transitory computer-readable media of claim 11 wherein:

the instructions further cause multitask learning by the token sequence encoder;

said inferring the encoded sequence that incorrectly represents the lexical text and said adjusting the token sequence encoder occur during said multitask learning by the token sequence encoder.

14 . The one or more non-transitory computer-readable media of claim 11 wherein:

the instructions further cause:

generating a lossy representation of the lexical text, and

inferring, by the token sequence encoder, an encoded sequence that incorrectly represents the lossy representation of the lexical text;

said adjusting the token sequence encoder is further based on the encoded sequence that incorrectly represents the lossy representation of the lexical text.

15 . The one or more non-transitory computer-readable media of claim 14 wherein:

said generating the lossy representation of the lexical text comprises excluding at least one token from a token sequence that represents the lexical text;

the instructions further cause inferring, by a machine learning model, the at least one token;

said adjusting based on the encoded sequence that incorrectly represents the lossy representation of the lexical text is based on said inferring the at least one token.

16 . The one or more non-transitory computer-readable media of claim 11 wherein said generating the histogram that correctly represents the graph comprises counting occurrences of a particular n-gram in the graph.

17 . The one or more non-transitory computer-readable media of claim 11 wherein

the instructions further cause:

after said training the token sequence encoder, receiving a second graph that is a portion of a third graph by incompletely receiving a data stream that consists of the third graph, and

inferring, by the token sequence encoder and without completely receiving the third graph, an encoded sequence that represents the second graph.

18 . One or more non-transitory computer-readable media storing instructions that, when executed by an integrated development environment (IDE), cause:

detecting that a parse error in a lexical text prevents generation of a graph that represents the lexical text;

generating, by a token sequence encoder, a fixed-size encoded sequence that represents the lexical text; and

inferring, from the fixed-size encoded sequence that represents the lexical text, a code completion of an expression in a logic statement in the lexical text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2023
From: FEITH, TOMAS; SCHNEUWLY, ARNO; ALLAHDADIAN, SAEID; CASSERINI, MATTEO; SCHMIDT, FELIX
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 064638/0058 →
Continuity (1)
Related Publication 20250060951A1 · Feb 20, 2025
References Cited (51)
US 20220198294A1 · Schneuwly et al. · 2022 [cited by applicant]
CN 111046630B · 2021 [cited by applicant]
GB 2605652A · 2022 [cited by examiner]
Dasgupta, Arpan et al., “Review of extreme multilabel classification”, 2023, 47 pages. [cited by applicant]
Kementchedjhieva, Yova et al., “An exploration of encoder-decoder approaches to multi-label classification for legal and biomedical text”, CoRR, abs/2305.05627, 2023, available: https://doi.org/10.48550/arXiv.2305.05627… [cited by applicant]
Kanade, Aditya et al., “Learning and Evaluating Contextual Embedding of Source Code”, Proceedings of the 37th International Conference on Machine Learning, Jul. 13-18, 2020, vol. 119 of Proceedings of Machine Learning R… [cited by applicant]
Jiang, Xue et al., “A tree-based pre-trained model for programming language”, Proceedings of the Thirty-Seventh Conference on Uncertainty in Artificial Intelligence, Jul. 27-30, 2021, vol. 161 of Proceedings of Machine … [cited by applicant]
Izadi, Maliheh et al., “Codefill: Multi-token code completion by jointly learning from structure and naming sequences”, Proceedings of the 44th International Conference on Software Engineering, ICSE '22, New York, NY, U… [cited by applicant]
Husain, Hamel et al., “Code—searchnet challenge: Evaluating the state of semantic code search”, CoRR, abs/1909.09436, 2019, available: http://arxiv.org/abs/1909.09436, 6 pages. [cited by applicant]
Guo, Daya et al., “Unixcoder: Unified cross-modal pre-training for code representation”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistic, ACL 2022, May 22-27, 2022, pp. 7212-7225. [cited by applicant]
Abramovich, Felix et al., “Classification with many classes: Challenges and pluses”, J. Multivar, Anal, 174, 2019, available: https://doi.org/10. 1016/j.jmva.2019.104536, 25 pages. [cited by applicant]
Devlin, Jacob et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Proceedings of The 2019 Conference of the North American Chapter of the Association for Computational Linguistics… [cited by applicant]
Liu, Feng et al., “MLRF: multi-label classification through random forest with label-set partition”, Advanced Intelligent Computing Theories and Applications—11th International Conference, Aug. 20-23, 2015, pp. 407-418. [cited by applicant]
Chowdhery, Aakanksha et al., “Palm: Scaling language modeling with pathways”, CoRR, abs/2204.02311, 2022, available: https://doi.org/10.48550/arXiv.2204.02311, 87 pages. [cited by applicant]
Chen, Mark et al., “Evaluating large language models trained on code”, CoRR, abs/2107.03374, 2021, available: https://arxiv.org/ abs/2107.03374, 35 pages. [cited by applicant]
Caruana, Rich, “Multitask Learning”, Machine Learning, 28, Jul. 1997, 35 pages. [cited by applicant]
Brown et al. “Language models are few-shot learner”, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, 75 pages. [cited by applicant]
Ben Allal, Loubna et al., “Santacoder: don't reach for the stars!”, CoRR, abs/2301.03988, 2023, available: https://doi.org/10.48550/arXiv.2301.03988, 22 pages. [cited by applicant]
Alon, Uri et al., “code2vec: learning distributed representations of code”, Proc. ACM Program. Lang., 3(POPL):40:1-40:29, 2019. [cited by applicant]
Alon, Uri et al., “code2seq: Generating sequences from structured representations of code”, 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019, OpenReview.net, 2019,… [cited by applicant]
Guo, Daya et al., “Graphcodebert: Pre-training code representations with data flow”, 9th International Conference on Learning Representations, ICLR 2021, Austria, May 3-7, 2021, available: https:// openreview.net/forum?… [cited by applicant]
Read, Jesse et al., “Deep learning for multi-label classification”, CoRR, abs/1502.05988, 2015, available: http://arxiv.org/abs/1502.05988, 8 pages. [cited by applicant]
Zhang, Min-Ling et al., “A review on multi-label learning algorithms”, IEEE Transactions on Knowledge and Data Engineering, 26(8), 2014, 1819-1837. [cited by applicant]
Weston, Jason et al., “Label partitioning for sublinear ranking”, Proceedings of the 30th International Conference on Machine Learning, vol. 28 of Proceedings of Machine Learning Research, Jun. 17-19, 2013, pp. 181-189. [cited by applicant]
Wang, Yanlin et al., “Code completion by modeling flattened abstract syntax trees as graphs”, Thirty-Fifth AAAI Conference on Artificial Intelligence, AAAI 2021, Feb. 2-9, 2021, pp. 14015-14023. [cited by applicant]
Wang, Xin et al., “SynCoBERT: Syntax-Guided Multi-Modal Contrastive Pre-Training for Code Representation”, 2021, 9 pages. [cited by applicant]
Wang, Xiao et al., “Heloc: hierarchical contrastive learning of source code representation”, Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension, ICPC 2022, Virtual Event, May 16-17, 2022,… [cited by applicant]
Vaswani, Ashish et al., “Attention is all you need”, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, Dec. 4-9, 2017, Long Beach, CA, USA, pp. 5998-6… [cited by applicant]
Touvron, Hugo et al., “Llama: Open and efficient foundation language models”, CoRR, abs/2302.13971, 2023, available: https://doi.org/10.48550/arXiv.2302.13971, 27 pages. [cited by applicant]
Kim, Seohyun et al., “Code prediction by feeding trees to transformers”, 43rd IEEE/ACM International Conference on Software Engineering, ICSE 2021, Madrid, Spain, May 22-30, 2021, pp. 150-162, available: https://doi.org… [cited by applicant]
Sultana, Farhana et al., “Advancements in image classification using convolutional neural network”, CoRR, abs/1905.03288, 2019, available: http://arxiv.org/abs/1905.03288, 9 pages. [cited by applicant]
Li, Raymond et al., “Starcoder: may the source be with you!”, CoRR, abs/2305.06161, 2023, available: https://doi.org/10.48550/arXiv.2305.06161, 55 pages. [cited by applicant]
Radford, Alec et al., “Language models are unsupervised multitask learners”, OPENAI Blog, 1(8):9, 2019, 24 pages. [cited by applicant]
Radford, Alec et al., “Improving language understanding by generative pre-training”, 2018, 12 pages. [cited by applicant]
Prabhu, Yashoteja et al., “FastXML: A fast, accurate and stable tree-classifier for extreme multi-label learning”, Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2014… [cited by applicant]
Peng, Han et al., “Integrating Tree Path in Transformer for Code Representation”, Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 9343-9354. [cited by applicant]
OpenAI, “GPT-4 technical report”, CoRR, abs/2303.08774, 2023, available: https://doi.org/10.48550/arXiv.2303.08774, 100 pages. [cited by applicant]
Liu, Yapeng et al., “Improving code completion by sequence features and structural features”, Proceedings of the 4th World Symposium on Software Engineering, WSSE '22, 2022, pp. 51-58. [cited by applicant]
Zu{umlaut over ( )}gner, Daniel et al., “Language-Agnostic Representation Learning of Source Code from Structure and Context”, 9th International Conference on Learning Representations, ICLR 2021, May 3-7, 2021, availabl… [cited by applicant]
Tang, Ze, et al., “Ast-trans: Code summarization with efficient tree-structured attention”, 2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE), 2022, pp. 150-162. [cited by applicant]
Maduko, “Graph summaries for optimizing graph pattern queries on rdf databases” downloaded Jan. 4, 2024 https://getd.libs.uga.edu/pdfs/maduko_angela_i_200905_phd.pdf 106 pages. [cited by applicant]
Blume, “Semantic Structural Graph Summaries for Evolving and Distributed Graphs” OPARU, 2 (Nov. 18, 2022), 162 pages. [cited by applicant]
Zugner, Daniel, et al., “Language-Agnostic Representation Learning of Source Code from Structure and Context”, ICLR 2021, https://arxiv.org/abs/2103.11318, Mar. 21, 2021, 22pgs. [cited by applicant]
Zhang, Jian, et al., “A Novel Neural Source Code Representation Based on Abstract Syntax Tree”, 2019 IEEE/ACM 41st ICSE, doi: 10.1109/ICSE.2019.00086, pp. 783-794, May 25, 2019, 12pgs. [cited by applicant]
Mou et al., “Building Program Vector Representations for Deep Learning”, dated Sep. 11, 2014, 11 pages. [cited by applicant]
Jimenez et al., “On the Impact of Tokenizer and Parameters on N-Gram Based Code Analysis”, dated 2018, 12 pages. [cited by applicant]
Guo, Daya, et al., “UniXcoder: Unified Cross-Modal Pre-training for Code Representation”, https://arxiv.org/abs/2203.03850, Mar. 8, 2022, 14pgs. [cited by applicant]
Devlin, Jacob, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, https://arxiv.org/abs/1810.04805, May 24, 2019, 16pgs. [cited by applicant]
Cai et al., “An Abstract Syntax Tree Encoding Method for Cross-Project Defect Prediction”, IEEE, dated Nov. 15, 2019, 10 pages. [cited by applicant]
Alon et al., “A General Path-Based Representation for Predicting Program Properties” dated Apr. 22, 2018, 16 pages. [cited by applicant]
“Tree-sitter”, Introduction, https://tree-sitter.github.io/tree-sitter/, last accessed Jun. 9, 2023, 4pgs. [cited by applicant]