IP Library › Granted Patent US 12,572,755
Granted Patent B2
US 12,572,755 · App. 18/531,223 · Granted Mar 10, 2026

System and method for augmenting training data for natural language to meaning representation language systems

Inventors: Philip Arthur (Sydney, AU); Gioacchino Tangari (Sydney, AU); Nitika Mathur (Melbourne, AU); Aashna Devang Kanuga (Foster City, CA); Cong Duy Vu Hoang (Melbourne, AU); Poorya Zaremoodi (Melbourne, AU); Thanh Long Duong (Melbourne, AU); Mark Edward Johnson (Sydney, AU)
Assignee: Oracle International Corporation
G06F40/40G06F16/245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,755
App. No.
18/531,223
Granted
Mar 10, 2026
Kind
B2
Abstract

Techniques for augmenting training data include accessing training data comprising a plurality of training examples comprising a first training example comprising a first natural language utterance and a first logical form for the first natural language utterance. A second natural language utterance is generated by adding or replacing one or more values in the first natural language utterance. A logical form for the second natural language utterance is generated. A second training example is generated, comprising the second natural language utterance and the logical form for the second natural language utterance. The training data is augmented by adding the second training example to the plurality of training examples to generate an augmented training data set. A machine learning model is trained to generate logical forms for utterances using the augmented training data set.

Claims (49)

1 . A computer-implemented method comprising:

(a) accessing training data comprising a plurality of training examples comprising a first training example, the first training example comprising a first natural language utterance, a first logical form for the first natural language utterance, and first metadata associated with the first natural language utterance, the first metadata including information about a database schema for a database to be queried using a logical form;

(b) generating a second natural language utterance by adding or replacing one or more values in the first natural language utterance;

(c) generating the logical form for the second natural language utterance;

(d) producing updated metadata based on the first metadata and the second natural language utterance;

(e) generating a second training example comprising the second natural language utterance, the logical form for the second natural language utterance, and the updated metadata;

(f) augmenting the training data by adding the second training example to the plurality of training examples to generate an augmented training data set; and

(g) training a machine learning model to generate logical forms for utterances using the augmented training data set.

2 . The computer-implemented method of claim 1 , further comprising:

repeating steps (b)-(f) to generate and add a configured number of additional training examples to the augmented training data set, wherein a type of augmentation and a set of replacement values are further configured.

3 . The computer-implemented method of claim 2 , wherein the replacement values are selected, based on the configuration, by randomly generating data points.

4 . The computer-implemented method of claim 1 , wherein producing the updated metadata comprises:

adjusting an offset value for schema linking to reflect the one or more replacement values.

5 . The computer-implemented method of claim 1 , wherein the logical forms correspond to database query representations.

6 . The computer-implemented method of claim 1 , further comprising:

deploying the machine learning model to generate an output logical form for an input natural language utterance.

7 . A system comprising:

one or more processors; and

one or more computer-readable media storing instructions which, when executed by the one or more processors, cause the system to perform operations comprising:

(a) accessing training data comprising a plurality of training examples comprising a first training example, the first training example comprising a first natural language utterance, a first logical form for the first natural language utterance, and first metadata associated with the first natural language utterance, the first metadata including information about a database schema for a database to be queried using a logical form;

(b) generating a second natural language utterance by adding or replacing one or more values in the first natural language utterance;

(c) generating the logical form for the second natural language utterance;

(d) producing updated metadata based on the first metadata and the second natural language utterance;

(e) generating a second training example comprising the second natural language utterance, the logical form for the second natural language utterance;

(f) augmenting the training data by adding the second training example to the plurality of training examples to generate an augmented training data set, and the updated metadata; and

(g) training a machine learning model to generate logical forms for utterances using the augmented training data set.

8 . The system of claim 7 , the operations further comprising:

repeating steps (b)-(f) to generate and add a configured number of additional training examples to the augmented training data set, wherein a type of augmentation and a set of replacement values are further configured.

9 . The system of claim 8 , wherein the replacement values are selected, based on the configuration, by randomly generating data points.

10 . The system of claim 7 , wherein producing the updated metadata comprises:

adjusting an offset value for schema linking to reflect the one or more replacement values.

11 . The system of claim 7 , wherein the logical forms correspond to database query representations.

12 . The system of claim 7 , further comprising:

deploying the machine learning model to generate an output logical form for an input natural language utterance.

13 . One or more non-transitory computer-readable media storing instructions which, when executed by one or more processors, cause a system to perform operations comprising:

(a) accessing training data comprising a plurality of training examples comprising a first training example, the first training example comprising a first natural language utterance, and a first logical form for the first natural language utterance, and first metadata associated with the first natural language utterance, the first metadata including information about a database schema for a database to be queried using a logical form;

(b) generating a second natural language utterance by adding or replacing one or more values in the first natural language utterance;

(c) generating the logical form for the second natural language utterance;

(d) producing updated metadata based on the first metadata and the second natural language utterance;

(e) generating a second training example comprising the second natural language utterance, the logical form for the second natural language utterance, and the updated metadata;

(f) augmenting the training data by adding the second training example to the plurality of training examples to generate an augmented training data set; and

(g) training a machine learning model to generate logical forms for utterances using the augmented training data set.

14 . The one or more non-transitory computer-readable media of claim 13 , the operations further comprising:

repeating steps (b)-(f) to generate and add a configured number of additional training examples to the augmented training data set, wherein a type of augmentation and a set of replacement values are further configured.

15 . The one or more non-transitory computer-readable media of claim 13 , wherein producing the updated metadata comprises:

adjusting an offset value for schema linking to reflect the one or more replacement values.

16 . The one or more non-transitory computer-readable media of claim 13 , wherein the logical forms correspond to database query representations.

17 . The one or more non-transitory computer-readable media of claim 13 , the operations further comprising:

deploying the machine learning model to generate an output logical form for an input natural language utterance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 7, 2023
From: ARTHUR, PHILIP; TANGARI, GIOACCHINO; MATHUR, NITIKA; KANUGA, AASHNA DEVANG; HOANG, CONG DUY VU; ZAREMOODI, POORYA; DUONG, THANH LONG; JOHNSON, MARK EDWARD
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 065802/0835 →
Continuity (1)
Related Publication 20250190710A1 · Jun 12, 2025
References Cited (36)
US 6901399B1 · Corston · 2005 [cited by examiner]
US 9244976B1 · Zhang · 2016 [cited by examiner]
US 9448995B2 · Kurz · 2016 [cited by applicant]
US 10796104B1 · Lee et al. · 2020 [cited by applicant]
US 11429607B2 · Anand et al. · 2022 [cited by applicant]
US 11508360B2 · Peng et al. · 2022 [cited by applicant]
US 20180075015A1 · Bennett · 2018 [cited by examiner]
US 20180375951A1 · Dhere · 2018 [cited by examiner]
US 20200134032A1 · Lin et al. · 2020 [cited by applicant]
US 20200210525A1 · Yang et al. · 2020 [cited by applicant]
US 20200334233A1 · Lee et al. · 2020 [cited by applicant]
US 20200334252A1 · Lee · 2020 [cited by applicant]
US 20210357409A1 · Rodriguez et al. · 2021 [cited by applicant]
US 20220067277A1 · Pentyala et al. · 2022 [cited by applicant]
US 20220245134A1 · Aalipour Hafshejani et al. · 2022 [cited by applicant]
US 20220318311A1 · Wang · 2022 [cited by examiner]
US 20230186025A1 · John et al. · 2023 [cited by applicant]
US 20230315856A1 · Lee et al. · 2023 [cited by applicant]
US 20230334309A1 · Streltsov et al. · 2023 [cited by applicant]
U.S. Appl. No. 18/218,385, “Ex Parte Quayle Action”, mailed Jun. 2, 2025, 10 pages. [cited by applicant]
“Answering Business Questions with Amazonquicksight Q”, Available online at: https://docs.aws.amazon.com/quicksight/latest/user/working-with-quicksight-q.html, 2023, 2 pages. [cited by applicant]
“Querying Your Data with Simply Ask (NLQ)”, Available Online at: https://docs.sisense.com/main/SisenseLinux/simply-ask-query-in-natural-language.htm#:˜:text=Simply%20Ask%20is%20Sisense's%20Natural,provide%20you%20with%2… [cited by applicant]
“What is Amazon QuickSight?”, Available online at: https://docs.aws.amazon.com/quicksight/latest/user/welcome.html, 2023, 3 pages. [cited by applicant]
Cheng et al., “Learning Structured Natural Language Representations for Semantic Parsing”, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguisti… [cited by applicant]
Dayananda , “Joining Across Data Sources on Amazon Quicksight”, Amazon QuickSight, Analytics, AWS Big Data, Nov. 1, 2019, 11 pages. [cited by applicant]
Desai et al., “Calibration of Pre-Trained Transformers”, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Nov. 16-20, 2020, pp. 295-302. [cited by applicant]
Guo et al., “On Calibration of Modern Neural Networks”, Proceedings of the 34th International Conference on Machine Learning, vol. 70, Aug. 3, 2017, 14 pages. [cited by applicant]
He et al., “DebertaV3: Improving Deberta using Electra-style Pre-training with Gradient-disentangled Embedding Sharing”, Conference paper at ICLR 2023, Mar. 24, 2023, 16 pages. [cited by applicant]
Kingma et al., “ADAM: A Method for Stochastic Optimization”, International Conference on Learning Representations, Conference paper at ICLR 2015, Jan. 30, 2017, 15 pages. [cited by applicant]
Montgomery et al., “Towards a Natural Language Query Processing System”, 1st International Conference on Big Data Analytics and Practices (IBDAP), arXiv:2009.12414, Sep. 2020, 6 pages. [cited by applicant]
Nassiri et al., “An Intermediate Representation-based Approach for Query Translation using a Syntax-Directed Method”, (IJACSA) International Journal of Advanced Computer Science and Applications, vol. 11, No. 8, 2020, 7… [cited by applicant]
Nie et al., “GraphQ IR: Unifying the Semantic Parsing of Graph Query Languages with One Intermediate Representation”, Computer Science, ArXiv, Nov. 7, 2022, 18 pages. [cited by applicant]
Rubin et al., “SmBoP: Semi-autoregressive Bottom-up Semantic Parsing”, Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun.… [cited by applicant]
Saha et al., “ATHENA: An Ontology-Driven System for Natural Language Querying over Relational Data Stores”, Proceedings of the VLDB Endowment, vol. 9, No. 12, Aug. 1, 2016, pp. 1209-1220. [cited by applicant]
Sun et al., “A Survey of Pretrained Language Models”, Lecture Notes in Computer Science book series (LNAI) vol. 13369, Jul. 19, 2022. [cited by applicant]
Wang et al., “RAT-SQL: Relation-Aware Schema Encoding and Linking for Text-to-SQL Parsers”, Available Online at: https://arxiv.org/pdf/1911.04942.pdf, Aug. 24, 2021, 12 pages. [cited by applicant]