IP Library › Granted Patent US 12,608,181
Granted Patent B2
US 12,608,181 · App. 18/165,254 · Granted Apr 21, 2026

Generation of synthetic training data using grammar mapping

Inventors: Konstantin Andreyevich Golobokov (Seattle, WA); Zeqi Lin (Beijing, CN); Haizhen Zhang (Bothell, WA); Yu Hu (Sammamish, WA); Yousef Ahmed Al-Kofahi (Niskayuna, NY); Jonathan Richard Malsan (Redmond, WA); Haiyuan Cao (Issaquah, WA); Daniel Akintola Fatade (Seattle, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F8/37G06F8/42G06F40/211
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,181
App. No.
18/165,254
Filed
Feb 6, 2023
Granted
Apr 21, 2026
Kind
B2
Art Unit
2191
USPC
717/143
Abstract

The automatic generation of synthetic training data that can be used to train a language model to generate code examples following a code language based on a natural language input. Thus, new language models may be created, or existing language models may be fine-tuned, to adapt to automatically generate code without having to manually generate bulk quantities of training data. Rather, a many-to-many grammar mapping is navigated to generate training data. Specifically, the many-to-many grammar mapping maps code grammar to natural grammar. Then, each training data is generated by navigating the many-to-many grammar mapping definition to generate a mapping of a respective code expression to a respective natural language expression.

Claims (47)

1 . A computing system that generates synthetic training data, said computing system comprising:

one or more processors; and

one or more hardware storage devices that store instructions that are executable by the one or more processors to cause the computing system to:

access a many-to-many grammar mapping definition for mapping code grammar and natural grammar, the code grammar being associated with a code language and defining how to form code strings from an alphabet of the code language in a way that the code strings are valid according to a syntax of the code language, the natural grammar being associated with a natural language and defining how to form natural language strings from an alphabet of the natural language in a way that the natural language strings at least approximate a syntax of the natural language;

generate a plurality of mappings between code expressions and natural language expressions by navigating the many-to-many grammar mapping definition to respectively generate a respective mapping of a respective code expression to a respective natural language expression, wherein the navigation includes a random sampling of the many-to-many grammar mappings definition such that each mapping in the plurality of mappings is different than at least most of the other mappings in the plurality of mappings, and wherein, as a result of the random sampling of the many-to-many grammar mappings definition, a different navigation results in a different intermediate form of an expression mapping, which is associated with the plurality of mappings;

select a first pair from among the plurality of mappings, the first pair comprising a first mapping between a first code expression and a first natural language expression;

syntactically validate the first code expression by executing an argument checker for a specific application development tool against the first code expression; and

in response to the argument checker syntactically validating the first code expression, include the first pair in a plurality of training data.

2 . The computing system in accordance with claim 1 , the instructions being further executable to cause the computing system to:

identify a code language that the training data is to follow; and

selecting the many-to-many grammar mapping definition corresponding to the identified code language, wherein there are different applicable many-to-many grammar mapping definition depending at least upon the identity of the code language.

3 . The computing system in accordance with claim 2 , the instructions being further executable to cause the computing system to:

identify a natural language that the training data is to follow; the selecting of the many-to-many grammar mapping definition that also corresponds to the identified natural language, wherein there are different applicable many-to-many grammar mapping definition depending at least upon the identity of the code language and the identity of the natural language.

4 . The computing system in accordance with claim 1 , the instructions being further executable to cause the computing system to:

identify a natural language that the training data is to be have; and

select the many-to-many grammar mapping definition corresponding to the identified natural language, wherein there are different applicable many-to-many grammar mapping definition depending at least upon the identity of the natural language.

5 . The computing system in accordance with claim 1 , wherein there are different applicable many-to-many grammar mapping definition depending at least upon one of the identity of the code language and the identity of the natural language, the different many-to-many grammar mapping definitions following a common grammar definition schema.

6 . The computing system in accordance with claim 1 , the many-to-many grammar mapping definition comprising a tree structure that is navigable downward from root node to a leaf node, the navigating of the many-to-many grammar mapping definition performed by randomly navigating the tree structure downward from the root node to formulate at least an intermediate form of the mapping of a respective code expression to a respective natural language expression.

7 . The computing system in accordance with claim 6 , wherein the intermediate form of the mapping of a respective code expression is further subject to application of context in the form of name-value pairs.

8 . The computing system in accordance with claim 1 , the instructions being further executable to cause the computing system to automatically generate the many-to-many grammar mapping based on a plurality of seed mappings between natural language expressions and code expressions.

9 . The computing system in accordance with claim 1 , wherein generating the plurality of training data is further performed by perturbing a representation of each of at least some of the mappings of respective code expressions and respective natural language expressions.

10 . The computing system in accordance with claim 1 , wherein generating the plurality of training data is also performed by using validation rules to filter a representation of each of at least some of the mappings of respective code expressions and respective natural language expressions.

11 . The computing system in accordance with claim 1 , wherein generating the plurality of training data is also performed by using validation rules to alter a representation of each of at least some of the mappings of respective code expressions and respective natural language expressions.

12 . A method for generating synthetic training data to train a language model to generate code examples following a code language based on a natural language input, the method comprising:

accessing a many-to-many grammar mapping definition for mapping code grammar and natural grammar, the code grammar being associated with a code language and defining how to form code strings from an alphabet of the code language in a way that the code strings are valid according to a syntax of the code language, the natural grammar being associated with a natural language and defining how to form natural language strings from an alphabet of the natural language in a way that the natural language strings at least approximate a syntax of the natural language;

generating a plurality of mappings between code expressions and natural language expressions by navigating the many-to-many grammar mapping definition to respectively generate a respective mapping of a respective code expression to a respective natural language expression, wherein the navigation includes a random sampling of the many-to-many grammar mappings definition such that each mapping in the plurality of mappings is different than at least most of the other mappings in the plurality of mappings, and wherein, as a result of the random sampling of the many-to-many grammar mappings definition, a different navigation results in a different intermediate form of an expression mapping, which is associated with the plurality of mappings;

selecting a first pair from among the plurality of mappings, the first pair comprising a first mapping between a first code expression and a first natural language expression;

syntactically validating the first code expression by executing an argument checker for a specific application development tool against the first code expression; and

in response to the argument checker syntactically validating the first code expression, including the first pair in a plurality of training data.

13 . The method in accordance with claim 12 , further comprising

identifying a code language that the training data is to have; and

selecting the many-to-many grammar mapping definition corresponding to the identified code language, wherein there are different applicable many-to-many grammar mapping definition depending at least upon the identity of the code language.

14 . The method in accordance with claim 13 , further comprising

identifying a natural language that the training data is to have; and

selecting the many-to-many grammar mapping definition corresponding to the identified natural language, wherein there are different applicable many-to-many grammar mapping definition depending at least upon the identity of the natural language.

15 . The method in accordance with claim 12 , the many-to-many grammar mapping definition comprising a tree structure that is navigable downward from root node to a leaf node, the navigating of the many-to-many grammar mapping definition performed by randomly navigating the tree structure downward from the root node to formulate at least an intermediate form of the mapping of a respective code expression to a respective natural language expression.

16 . The method in accordance with claim 15 , further comprising:

further subjecting the intermediate form of the mapping of a respective code expression to application of context in the form of name-value pairs.

17 . The method in accordance with claim 12 , wherein automatically generating the many-to-many grammar mapping is based on a plurality of seed mappings between natural language expressions and code expressions.

18 . The method in accordance with claim 12 , the generation of the plurality of training data is also performed by including perturbing a representation of each of at least some of the mappings of respective code expressions and respective natural language expressions.

19 . The method in accordance with claim 12 , wherein the generation of the plurality of training data is also performed by including using validation rules to filtering a representation of each of at least some of the mappings of respective code expressions and respective natural language expressions.

20 . One or more hardware storage devices that store instructions that are executable by one or more processors to cause the one or more processors to:

access a many-to-many grammar mapping definition for mapping code grammar and natural grammar, the code grammar being associated with a code language and defining how to form code strings from an alphabet of the code language in a way that the code strings are valid according to a syntax of the code language, the natural grammar being associated with a natural language and defining how to form natural language strings from an alphabet of the natural language in a way that the natural language strings at least approximate a syntax of the natural language; and

generate a plurality of mappings between code expressions and natural language expressions by navigating the many-to-many grammar mapping definition to respectively generate a respective mapping of a respective code expression to a respective natural language expression, wherein the navigation includes a random sampling of the many-to-many grammar mappings definition such that each mapping in the plurality of mappings is different than at least most of the other mappings in the plurality of mappings, and wherein, as a result of the random sampling of the many-to-many grammar mappings definition, a different navigation results in a different intermediate form of an expression mapping, which is associated with the plurality of mappings;

select a first pair from among the plurality of mappings, the first pair comprising a first mapping between a first code expression and a first natural language expression;

syntactically validate the first code expression by executing an argument checker for a specific application development tool against the first code expression; and

in response to the argument checker syntactically validating the first code expression, include the first pair in a plurality of training data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2023
From: GOLOBOKOV, KONSTANTIN ANDREYEVICH; LIN, ZEQI; ZHANG, HAIZHEN; HU, YU; AL-KOFAHI, YOUSEF AHMED; MALSAN, JONATHAN RICHARD; CAO, HAIYUAN; FATADE, DANIEL AKINTOLA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063622/0590 →
Continuity (1)
Related Publication 20240264809A1 · Aug 8, 2024
References Cited (70)
US 10217059B2 · Yang · 2019 [cited by applicant]
US 11327722B1 · Bahrami · 2022 [cited by examiner]
US 11720346B2 · Wu · 2023 [cited by examiner]
US 20170220327A1 · Allen · 2017 [cited by examiner]
US 20200097261A1 · Smith · 2020 [cited by examiner]
US 20200110803A1 · Djalali · 2020 [cited by examiner]
US 20210142005A1 · Shankar · 2021 [cited by applicant]
US 20210208855A1 · Zhang · 2021 [cited by examiner]
US 20220138240A1 · Bahrami · 2022 [cited by examiner]
US 20220236964A1 · Bahrami · 2022 [cited by examiner]
US 20220365776A1 · Ramsl · 2022 [cited by examiner]
US 20230046961A1 · Kurabayashi · 2023 [cited by examiner]
US 20230100208A1 · Bahrami · 2023 [cited by examiner]
US 20230119613A1 · Lin · 2023 [cited by examiner]
US 20240184555A1 · De Toni · 2024 [cited by examiner]
US 20240330170A1 · Mesde · 2024 [cited by examiner]
Ansari et al; NLI-GSC: A Natural Language Interface for Generating SourceCode, 12 pages (Year: 2022). [cited by examiner]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/013903 (MS#412363- 1 PCT01) May 29, 2024, 15 pages. [cited by applicant]
“About GitHub Copilot”, Retrieved from: https://docs.github.com/en/copilot/overview-of-github-copilot/about-github-copilot, Retrieved on: Sep. 15, 2022, 3 Pages. [cited by applicant]
“Azure OpenAI Service PREVIEW”, Retrieved from: https://azure.microsoft.com/en-us/products/cognitive-services/openai-service/#overview, Retrieved on: Sep. 15, 2022, 14 Pages. [cited by applicant]
“huggingface/CodeBERTa-small-v1”, Retrieved from: https://huggingface.co/huggingface/CodeBERTa-small-v1, Retrieved on: Sep. 15, 2022, 6 Pages. [cited by applicant]
“Programmatically Build Training Data”, Retrieved from: https://www.snorkel.org/, Retrieved on: Sep. 15, 2022, 5 Pages. [cited by applicant]
“Spider 1.0 Yale Semantic Parsing and Text-to-SQL Challenge”, Retrieved from: https://yale-lily.github.io/spider, Retrieved on: Sep. 15, 2022, 15 Pages. [cited by applicant]
Andrew, Ng, “MLOps: From Model-Centric to Data-Centric AI”, Retrieved from: https://docplayer.net/213213717-Mlops-from-model-centric-to-data-centric-ai-andrew-ng.html, Jun. 2021, 29 Pages. [cited by applicant]
Belinkov, et al., “Synthetic and Natural Noise Both Break Neural Machine Translation”, In Proceedings of International Conference on Learning Representations, Feb. 16, 2018, 13 Pages. [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners”, In Proceedings of 34th Conference on Neural Information Processing Systems, Dec. 6, 2020, 25 Pages. [cited by applicant]
Chen, et al., “CodeT: Code Generation with Generated Tests”, In Repository of arXiv:2207.10397v1, Jul. 21, 2022, 12 Pages. [cited by applicant]
Chen, et al., “Evaluating Large Language Models Trained on Code”, In Repository of arXiv:2107.03374v1, Jul. 7, 2021, 35 Pages. [cited by applicant]
Cheng, et al., “AdvAug: Robust Adversarial Augmentation for Neural Machine Translation”, In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistic, Jul. 5, 2020, pp. 5961-5970. [cited by applicant]
Dehouck, et al., “Data Augmentation via Subtree Swapping for Dependency Parsing of Low-Resource Languages”, In Proceedings of the 28th International Conference on Computational Linguistics, Dec. 8, 2013, pp. 3818-3830. [cited by applicant]
Deng, et al., “Structure-Grounded Pretraining for Text-to-SQL”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun. 6, 2021,… [cited by applicant]
Goswell, et al., “Common Data Model”, Retrieved from: https://learn.microsoft.com/en-us/common-data-model/, Apr. 8, 2022, 4 Pages. [cited by applicant]
Guo, et al., “Revisiting Iterative Back-Translation from the Perspective of Compositional Generalization”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, Issue 9, May 18, 2021, pp. 7601-7609. [cited by applicant]
Herbert, et al., “List of CDM Data Types”, Retrieved from: https://learn.microsoft.com/en-us/common-data-model/sdk/list-of-datatypes, Jul. 1, 2022, 65 Pages. [cited by applicant]
Herzig, et al., “Don't Paraphrase, Detect! Rapid and Effective Data Collection for Semantic Parsing”, In Proceedings of Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conf… [cited by applicant]
Hu, et al., “LoRA: Low-Rank Adaptation of Large Language Models”, In Proceedings of International Conference on Learning Representations, Apr. 2022, 13 Pages. [cited by applicant]
Jiao, et al., “TinyBERT: Distilling BERT for Natural Language Understanding”, In Proceedings of Findings of the Association for Computational Linguistics: Empirical Methods in Natural Language Processing, Nov. 16, 2020,… [cited by applicant]
Karimi, et al., “AEDA: An Easier Data Augmentation Technique for Text Classification”, In Proceedings of Findings of the Association for Computational Linguistics: Empirical Methods in Natural Language Processing, Nov. … [cited by applicant]
Keysers, et al., “Measuring Compositional Generalization: A Comprehensive Method on Realistic Data”, In Proceedings of International Conference on Learning Representations, Sep. 26, 2019, 38 Pages. [cited by applicant]
Kim, et al., “COGS: A Compositional Generalization Challenge Based on Semantic Interpretation”, In Proceedings of Conference on Empirical Methods in Natural Language Processing, Nov. 16, 2020, pp. 9087-9105. [cited by applicant]
Kummerfeld, et al., “Improving Text-to-SQL Evaluation Methodology”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, vol. 1 Long Papers, Jul. 15, 2018, pp. 351-360. [cited by applicant]
Lake, et al., “Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks”, In Proceedings of the 35th International Conference on Machine Learning, vol. 80, Jul. 10, 20… [cited by applicant]
Li, et al., “Data Augmentation Approaches in Natural Language Processing: A Survey”, In Repository of arXiv:2110.01852v1, Oct. 5, 2021, 42 Pages. [cited by applicant]
Liang, Percy, “Semantic Parsing for Natural Language Interfaces”, Retrieved from: https://crossminds.ai/video/semantic-parsing-for-natural-language-interfaces-606fec85f43a7f2f827c1135/, Dec. 6, 2020, 3 Pages. [cited by applicant]
Lindhorst, Greg, “What is Microsoft Power Fx?”, Retrieved from: https://powerapps.microsoft.com/en-us/blog/what-is-microsoft-power-fx/, Mar. 2, 2021, 17 pages. [cited by applicant]
Pelletier, Francisj. , “The Principle of Semantic Compositionality”, In Journal of Topol, vol. 13, Issue 1, Mar. 1994, pp. 11-24. [cited by applicant]
Pl, et al., “Towards Robustness of Text-to-SQL Models Against Natural and Realistic Adversarial Table Perturbation”, In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, vol. 1: Lo… [cited by applicant]
Radford, et al., “Language Models are Unsupervised Multitask Learners”, In Journal of OpenAI Blog, vol. 1, Issue 8, Feb. 24, 2019, 24 Pages. [cited by applicant]
Raffel, et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer”, In Journal of Machine Learning Research, vol. 21, Issue 140, Jun. 2020, 67 Pages. [cited by applicant]
Ratner, et al., “Snorkel: Rapid Training Data Creation with Weak Supervision”, In Proceedings of the VLDB Endowment, vol. 11, Issue 3, Nov. 2017, pp. 269-282. [cited by applicant]
Shen, et al., “Mixture Models for Diverse Machine Translation: Tricks of the Trade”, In Proceedings of the 36th International Conference on Machine Learning, May 24, 2019, 10 Pages. [cited by applicant]
Shin, et al., “Constrained Language Models Yield Few-Shot Semantic Parsers”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Nov. 7, 2021, pp. 7699-7715. [cited by applicant]
Shin, et al., “Few-Shot Semantic Parsing with Language Models Trained on Code”, In Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie… [cited by applicant]
Shu, et al., “Logic-Consistency Text Generation from Semantic Parses”, In Proceedings of Findings of the Association for Computational Linguistics: ACL-IJCNLP, Aug. 1, 2021, pp. 4414-4426. [cited by applicant]
Sun, et al., “Mixup-Transformer: Dynamic Data Augmentation for NLP Tasks”, In Proceedings of the 28th International Conference on Computational Linguistics, Dec. 8, 2020, pp. 3436-3440. [cited by applicant]
Thakur, et al., “Augmented SBERT: Data Augmentation Method for Improving Bi-Encoders for Pairwise Sentence Scoring Tasks”, In Proceedings of the Conference of the North American Chapter of the Association for Computatio… [cited by applicant]
Wang, et al., “Building a Semantic Parser Overnight”, In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing … [cited by applicant]
Wang, et al., “CodeT5: Identifier-Aware Unified Pre-Trained Encoder-Decoder Models for Code Understanding and Generation”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Nov. 7, 20… [cited by applicant]
Wei, et al., “EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing and the 9th International… [cited by applicant]
Wong, et al., “Machine Translation using Constraint-Based Synchronous Grammar”, In Journal of Tsinghua Science and Technology, vol. 11, Issue 3, Jun. 2006, pp. 295-306. [cited by applicant]
Xie, et al., “Unsupervised Data Augmentation for Consistency Training”, In Journal of Advances in Neural Information Processing Systems, vol. 33, Dec. 6, 2020, 13 Pages. [cited by applicant]
Xu, et al., “AutoQA: From Databases To QA Semantic Parsers With Only Synthetic Training Data”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Nov. 16, 2020, pp. 422-434. [cited by applicant]
Xu, et al., “Diversity-Promoting GAN: A Cross-Entropy Based Generative Adversarial Network for Diversified Text Generation”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Oct. 31,… [cited by applicant]
Yang, et al., “Neural Retrieval for Question Answering with Cross-Attention Supervised Data Augmentation”, In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internat… [cited by applicant]
Yu, et al., “GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing”, In Repository of arXiv:2009.13845v2, May 29, 2021, 16 Pages. [cited by applicant]
Yu, et al., “Score: Pre-Training for Context Representation in Conversational Semantic Parsing”, In Proceedings of International Conference on Learning Representations, Sep. 28, 2020, 16 Pages. [cited by applicant]
Yu, et al., “Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task”, In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Oct. 31… [cited by applicant]
Zhang, et al., “Character-level Convolutional Networks for Text Classification”, In Journal of Advances in Neural Information Processing Systems, vol. 28, Dec. 7, 2015, 9 Pages. [cited by applicant]
Zhu, et al., “Texygen: A Benchmarking Platform for Text Generation Models”, In Proceedings of the 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, Jul. 8, 2018, pp. 1097-1100. [cited by applicant]
International Preliminary Report on Patentability received for PCT Application No. PCT/US2024/013903 (MS#412363-PCT01) mailed on Aug. 21, 2025, 10 pages. [cited by applicant]