IP Library › Granted Patent US 12,530,344
Granted Patent B2
US 12,530,344 · App. 18/725,248 · Granted Jan 20, 2026

Method for enhancing learning data set in natural language processing system

Inventors: Wook Shin Han (Pohang-si, KR); Hyuk Kyu Kang (Pohang-si, KR); Hyeon Ji Kim (Ulsan, KR)
Assignee: POSTECH RESEARCH AND BUSINESS DEVELOPMENT FOUNDATION
G06F16/242G06F16/211
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,344
App. No.
18/725,248
Granted
Jan 20, 2026
Kind
B2
Abstract

A learning data set enhancing method comprises the steps of: determining a first initial learning data set comprising a first learning-purpose natural language query, database information regarding the database, and a first learning-purpose SQL query corresponding to the first learning-purpose natural language query; generating first novel database information regarding a first novel database having a different schema from the database by applying a first schema modification operation to the database; generating a first novel learning-purpose SQL query regarding the first novel database by applying a first SQL query synchronization operation corresponding to the first schema modification operation to the first learning-purpose SQL query; and determining a first novel learning data set comprising the first learning-purpose natural language query, the first novel database information, and the first novel learning-purpose SQL query.

Claims (105)

1 . A method of augmenting training datasets used for training a predetermined neural network in a natural language processing system, the natural language processing system including a translator for translating a natural language query into a structured query language (SQL) query based on the predetermined neural network, a data augmentation engine augmenting the training datasets for the translator, and a training engine training the translator using the augmented training datasets, the method comprising:

determining, with the data augmentation engine, a first initial training dataset including a first training natural language query, database information of a database, and a first training SQL query corresponding to the first training natural language query;

determining, with the data augmentation engine, a first schema modification operation from among predetermined schema modification operations;

generating, with the data augmentation engine, first new database information of a first new database having a different schema than the database by applying the first schema modification operation to the database;

determining, with the data augmentation engine, whether the first schema modification operation transforms the first training SQL query;

when the first schema modification operation is determined to transform the first training SQL query, generating, with the data augmentation engine, a first new training SQL query for the first new database by applying a first SQL query synchronization operation to the first training SQL query, the first SQL query synchronization operation being determined based on predetermined SQL query synchronization operation definition information corresponding to the first schema modification operation;

when the first schema modification operation is determined not to transform the first training SQL query, determining, with the data augmentation engine, the first training SQL query as the first new training SQL query for the first new database;

determining, with the data augmentation engine, a first new training dataset including the first training natural language query, the first new database information, and the generated or determined first new training SQL query; and

providing, with the data augmentation engine, the first new training dataset as input to the training engine.

2 . The method of claim 1 , further comprising:

generating, with the data augmentation engine, second new database information of a second new database having a different schema than the database by applying a second schema modification operation, determined from among the predetermined schema modification operations, to the database;

generating, with the data augmentation engine, a second new training SQL query for the second new database by applying a second SQL query synchronization operation corresponding to the second schema modification operation to the first training SQL query; and

determining, with the data augmentation engine, a second new training dataset including the first training natural language query, the second new database information, and the second new training SQL query.

3 . The method of claim 2 , further comprising:

determining, with the data augmentation engine, a second initial training dataset including a second training natural language query, the database information, and a second training SQL query corresponding to the second training natural language query;

generating, with the data augmentation engine, third new database information of a third new database having a different schema than the database by applying the first schema modification operation to the database;

generating, with the data augmentation engine, a third new training SQL query for the third new database by applying the first SQL query synchronization operation to the second training SQL query; and

determining, with the data augmentation engine, a third new training dataset including the second training natural language query, the third new database information, and the third new training SQL query.

4 . The method of claim 3 , further comprising:

generating, with the data augmentation engine, fourth new database information of a fourth new database having a different schema than the first new database by applying the first schema modification operation to the first new database;

generating, with the data augmentation engine, a fourth new training SQL query for the fourth new database by applying the first SQL query synchronization operation to the first new training SQL query; and

determining, with the data augmentation engine, a fourth new training dataset including the first training natural language query, the fourth new database information, and the fourth new training SQL query.

5 . The method of claim 1 , further comprising:

determining, with the data augmentation engine, a second initial training dataset including a second training natural language query, the database information, and a second training SQL query corresponding to the second training natural language query;

generating, with the data augmentation engine, second new database information of a second new database having a different schema than the database by applying the first schema modification operation to the database;

generating, with the data augmentation engine, a second new training SQL query for the second new database by applying the first SQL query synchronization operation corresponding to the first schema modification operation to the second training SQL query; and

determining, with the data augmentation engine, a second new training dataset including the second training natural language query, the second new database information, and the second new training SQL query.

6 . The method of claim 1 , wherein the first schema modification operation includes at least one of a first modification operation of changing a schema structure of the database or a second modification operation of changing a name of a schema element of the database.

7 . A method of training a predetermined neural network in a natural language processing system, the natural language processing system including a translator for translating a natural language query into a structured query language (SQL) query based on the predetermined neural network, a data augmentation engine augmenting training datasets for the translator, and a training engine training the translator using the augmented training datasets, the method comprising:

determining, with the data augmentation engine, a first initial training dataset including a first training natural language query, database information of a database, and a first training SQL query corresponding to the first training natural language query;

determining, with the data augmentation engine, a first schema modification operation from among predetermined schema modification operations;

generating, with the data augmentation engine, first new database information of a first new database having a different schema than the database by applying the first schema modification operation to the database;

determining, with the data augmentation engine, whether the first schema modification operation transforms the first training SQL query;

when the first schema modification operation is determined to transform the first training SQL query, generating, with the data augmentation engine, a first new training SQL query for the first new database by applying a first SQL query synchronization operation to the first training SQL query, the first SQL query synchronization operation being determined based on predetermined SQL query synchronization operation definition information corresponding to the first schema modification operation;

when the first schema modification operation is determined not to transform the first training SQL query, determining, with the data augmentation engine, the first training SQL query as the first new training SQL query for the first new database;

determining, with the data augmentation engine, a first new training dataset including the first training natural language query, the first new database information, and the generated or determined first new training SQL query;

training, with the training engine, the neural network using the first new training dataset received from the data augmentation engine to optimize parameters of the translator; and

providing, with the training engine, the parameters output by the trained neural network and translation reference data generated during training of the neural network for reference in translation to the translator.

8 . The method of claim 7 , further comprising:

generating, with the data augmentation engine, second new database information of a second new database having a different schema than the database by applying a second schema modification operation, determined from among the predetermined schema modification operations, to the database;

generating, with the data augmentation engine, a second new training SQL query for the second new database by applying a second SQL query synchronization operation corresponding to the second schema modification operation to the first training SQL query; and

determining, with the data augmentation engine, a second new training dataset including the first training natural language query, the second new database information, and the second new training SQL query,

wherein the training of the neural network is performed using the first and second new training datasets.

9 . The method of claim 8 , further comprising:

determining, with the data augmentation engine, a second initial training dataset including a second training natural language query, the database information, and a second training SQL query corresponding to the second training natural language query;

generating, with the data augmentation engine, third new database information of a third new database having a different schema than the database by applying the first schema modification operation to the database;

generating, with the data augmentation engine, a third new training SQL query for the third new database by applying the first SQL query synchronization operation to the second training SQL query; and

determining, with the data augmentation engine, a third new training dataset including the second training natural language query, the third new database information, and the third new training SQL query,

wherein the training of the neural network is performed using the first to third new training datasets.

10 . The method of claim 9 , further comprising:

generating, with the data augmentation engine, fourth new database information of a fourth new database having a different schema than the first new database by applying the first schema modification operation to the first new database;

generating, with the data augmentation engine, a fourth new training SQL query for the fourth new database by applying the first SQL query synchronization operation to the first new training SQL query; and

determining, with the data augmentation engine, a fourth new training dataset including the first training natural language query, the fourth new database information, and the fourth new training SQL query,

wherein the training of the neural network is performed using the first to fourth new training datasets.

11 . The method of claim 7 , further comprising:

determining, with the data augmentation engine, a second initial training dataset including a second training natural language query, the database information, and a second training SQL query corresponding to the second training natural language query;

generating, with the data augmentation engine, second new database information of a second new database having a different schema than the database by applying the first schema modification operation to the database;

generating, with the data augmentation engine, a second new training SQL query for the second new database by applying the first SQL query synchronization operation corresponding to the first schema modification operation to the second training SQL query; and

determining, with the data augmentation engine, a second new training dataset including the second training natural language query, the second new database information, and the second new training SQL query,

wherein the training of the neural network is performed using the first and second new training datasets.

12 . The method of claim 7 , wherein the first schema modification operation includes at least one of a first modification operation of changing a schema structure of the database or a second modification operation of changing a name of a schema element of the database.

13 . An information search device for providing a search result corresponding to a natural language query, the information search device comprising:

a memory configured to store program commands; and

a processor connected to the memory and configured to execute the program commands stored in the memory,

wherein, when the program commands are executed by the processor, the program commands cause the processor to perform the operations of:

training, with a training engine, a predetermined neural network;

receiving the natural language query from a user;

translating, with a translator, the natural language query into a structured query language (SQL) query based on the predetermined neural network;

acquiring, with a relational database management system, the search result corresponding to the natural language query using the SQL query; and

providing the search result to the user,

wherein the program commands causing the processor to perform the operation of training the predetermined neural network cause the processor to perform the operations of:

determining, with the data augmentation engine, a first initial training dataset including a first training natural language query, database information of a database, and a first training SQL query corresponding to the first training natural language query;

determining, with the data augmentation engine, a first schema modification operation from among predetermined schema modification operations;

generating, with the data augmentation engine, first new database information of a first new database having a different schema than the database by applying the first schema modification operation to the database;

determining, with the data augmentation engine, whether the first schema modification operation transforms the first training SQL query;

when the first schema modification operation is determined to transform the first training SQL query, generating, with the data augmentation engine, a first new training SQL query for the first new database by applying a first SQL query synchronization operation to the first training SQL query, the first SQL query synchronization operation being determined based on predetermined SQL query synchronization operation definition information corresponding to the first schema modification operation;

when the first schema modification operation is determined not to transform the first training SQL query, determining, with the data augmentation engine, the first training SQL query as the first new training SQL query for the first new database;

determining, with the data augmentation engine, a first new training dataset including the first training natural language query, the first new database information, and the generated or determined first new training SQL query;

training, with the training engine, the neural network using the first new training dataset received from the data augmentation engine to optimize parameters of the translator;

providing, with the training engine, the parameters output by the trained neural network and translation reference data generated during training of the neural network for reference in translation to the translator; and

wherein the program commands, when executed by the processor, further cause the processor, when translating, with the translator, the natural language query into the structured query language (SQL) query based on the predetermined neural network, to translate the natural language query into the SQL query using the parameters and the translation reference data.

14 . The information search device of claim 13 , wherein the program commands causing the processor to perform the operation of training the neural network cause the processor to further perform the operations of:

generating, with the data augmentation engine, second new database information of a second new database having a different schema than the database by applying a second schema modification operation, determined from among the predetermined schema modification operations, to the database;

generating, with the data augmentation engine, a second new training SQL query for the second new database by applying a second SQL query synchronization operation corresponding to the second schema modification operation to the first training SQL query; and

determining, with the data augmentation engine, a second new training dataset including the first training natural language query, the second new database information, and the second new training SQL query,

wherein the training of the neural network is performed using the first and second new training datasets.

15 . The information search device of claim 14 , wherein the program commands causing the processor to perform the operation of training the neural network cause the processor to further perform the operations of:

determining, with the data augmentation engine, a second initial training dataset including a second training natural language query, the database information, and a second training SQL query corresponding to the second training natural language query;

generating, with the data augmentation engine, third new database information of a third new database having a different schema than the database by applying the first schema modification operation to the database;

generating, with the data augmentation engine, a third new training SQL query for the third new database by applying the first SQL query synchronization operation to the second training SQL query; and

determining, with the data augmentation engine, a third new training dataset including the second training natural language query, the third new database information, and the third new training SQL query,

wherein the training of the neural network is performed using the first to third new training datasets.

16 . The information search device of claim 14 , wherein the program commands causing the processor to perform the operation of training the neural network cause the processor to further perform the operations of:

generating, with the data augmentation engine, fourth new database information of a fourth new database having a different schema than the first new database by applying the first schema modification operation to the first new database;

generating, with the data augmentation engine, a fourth new training SQL query for the fourth new database by applying the first SQL query synchronization operation to the first new training SQL query; and

determining, with the data augmentation engine, a fourth new training dataset including the first training natural language query, the fourth new database information, and the fourth new training SQL query,

wherein the training of the neural network is performed using the first to fourth new training datasets.

17 . The information search device of claim 13 , wherein the program commands causing the processor to perform the operation of training the neural network cause the processor to further perform the operations of:

determining, with the data augmentation engine, a second initial training dataset including a second training natural language query, the database information, and a second training SQL query corresponding to the second training natural language query;

generating, with the data augmentation engine, second new database information of a second new database having a different schema than the database by applying the first schema modification operation to the database;

generating, with the data augmentation engine, a second new training SQL query for the second new database by applying the first SQL query synchronization operation corresponding to the first schema modification operation to the second training SQL query; and

determining, with the data augmentation engine, a second new training dataset including the second training natural language query, the second new database information, and the second new training SQL query,

wherein the training of the neural network is performed using the first and second new training datasets.

18 . The information search device of claim 13 , wherein the first schema modification operation includes at least one of a first modification operation of changing a schema structure of the database or a second modification operation of changing a name of a schema element of the database, and

when the first schema modification operation only includes the first modification operation, the first SQL query synchronization operation is a null operation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2024
From: HAN, WOOK SHIN; KANG, HYUK KYU; KIM, HYEON JI
To: POSTECH RESEARCH AND BUSINESS DEVELOPMENT FOUNDATION
Reel/Frame 067870/0026 →
Priority Claims (1)
KR 10-2021-0192460 · Dec 30, 2021 · national
Continuity (1)
Related Publication 20250061111A1 · Feb 20, 2025
References Cited (14)
US 10747761B2 · Zhoug et al. · 2020 [cited by applicant]
US 11205099B2 · Shlens et al. · 2021 [cited by applicant]
US 20170185921A1 · Zhang · 2017 [cited by applicant]
US 20210357409A1 · Rodriguez et al. · 2021 [cited by applicant]
US 20230185799A1 · Hoang · 2023 [cited by examiner]
US 20230186161A1 · Arthur · 2023 [cited by examiner]
CN 112783921A · 2021 [cited by applicant]
KR 102261199 · 2021 [cited by applicant]
KR 1020210156964 · 2021 [cited by applicant]
International Search Report issued for corresponding International Patent Application No. PCT/KR2021/020301 on Sep. 21, 2022, along with an English Translation (5 pages). [cited by applicant]
Written Opinion issued for corresponding International Patent Application No. PCT/KR2021/020301 on Sep. 21, 2022 (3 pages). [cited by applicant]
Florin Brad et al., “Dataset for a Neural Natural Language Interface for Databases (NNLIDB)”, arXiv:1707.03172v1, Jul. 11, 2017, <https://arxiv.org/abs/1707.03172v1>, 14 pages, Cited in NPL Nos. 1 and 2. [cited by applicant]
Hyeonji Kim et al, “Natural language to SQL: Where are we today?”, Jun. 1, 2020, DOI: https://doi.org/10.14778/3401960.3401970, pp. 1737-1750. [cited by applicant]
Loredana Caruccio et al., “Synchronization of Queries and Views Upon Schema Evolutions: A Survey”, ACM Transactions on Database Systems, vol. 41, No. 2, Article 9, May 2016, 41 pages. [cited by applicant]