IP Library › Granted Patent US 12,730,809
Granted Patent B2
US 12,730,809 · App. 18/934,077 · Granted Sep 8, 2026

Automated prompt augmentation and engineering using ML automation in SQL query engine

Inventors: Anatoly Yakovlev (Hayward, CA); Sandeep R. Agrawal (San Jose, CA); Sanjay Jinturkar (Basking Ridge, NJ); Nipun Agarwal (Saratoga, CA)
Assignee: Oracle International Corporation
G06F16/24542G06F16/243G06F40/40G06N3/0475
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,730,809
App. No.
18/934,077
Granted
Sep 8, 2026
Kind
B2
Abstract

A database system generates a prompt for an LLM or other machine learning (ML) model to narrow the search space to highly relevant information about a database. A distinct instance of a classifier, a clustering algorithm, or a topic modeling model can be trained based on information from ML automation within the database system, respectively for each column or table in the database. Model instances can then be used during generative LLM inferencing to identify relevant sources of data to answer the user's question. Thus, the prompt generation combines ML automation and other ML models or an LLM for topic modeling and schema description.

Claims (77)

1 . A method, comprising:

receiving a natural language query from a user;

retrieving metadata from a machine learning (ML) automation component of a database system, the metadata comprising at least one of:

workload statistics of one or more users of the database system,

schematic details of a data source, or

dynamic content statistics of the data source;

generating a linguistic prompt, based on the metadata and the natural language query received from the user, for a generative artificial intelligence (AI) model; and

generating, by the generative AI model, a response based on the linguistic prompt and a search of the data source, wherein the linguistic prompt limits a scope of the search of the data source based on the metadata,

wherein the method is performed by one or more computing devices.

2 . The method of claim 1 , wherein the metadata is generated by one or more machine learning models for predicting resource usage and query performance in the database system.

3 . The method of claim 1 , wherein:

the data source is a relational database comprising one or more database tables,

the generative AI model comprises a large language model (LLM), and

the natural language query is a query about the relational database.

4 . The method of claim 1 , wherein:

the data source is a relational database comprising one or more database tables,

the generative AI model comprises a natural language to structured query language (NL2SQL) generative model, and

the NL2SQL generative model is configured to generate one or more SQL queries for searching the relational database based on the linguistic prompt.

5 . The method of claim 4 , further comprising:

providing the linguistic prompt as input to the NL2SQL generative model to generate the one or more SQL queries.

6 . The method of claim 5 , further comprising:

causing the one or more SQL queries to be displayed to the user.

7 . The method of claim 5 , further comprising:

executing the one or more SQL queries against the one or more database tables to generate a search result.

8 . The method of claim 7 , further comprising:

providing the search result as input to a large language model to generate a natural language description of the search result.

9 . The method of claim 1 , wherein:

the data source is a relational database comprising one or more database tables, and

generating the linguistic prompt comprises adding a set of one or more schema descriptions by selecting a subset of database tables from the one or more database tables based at least in part on the metadata and adding schema descriptions for the subset of database tables to the linguistic prompt.

10 . The method of claim 9 , wherein adding the set of one or more schema descriptions further comprises:

generating one or more query embeddings based on the natural language query; and

selecting the subset of database tables based at least in part on similarity of the one or more query embeddings and per-table topic modeling embeddings of the one or more database tables.

11 . The method of claim 1 , wherein:

the data source comprises an object store comprising one or more vector stores representing a plurality of documents using semantic encodings, and

generating the linguistic prompt comprises filtering the one or more vector stores based on the metadata.

12 . The method of claim 1 , wherein the ML automation component performs at least one of:

auto provisioning,

auto parallel loading,

auto data placement,

auto encoding,

auto query plan improvement,

auto query time estimation,

auto change propagation,

auto scheduling, or

auto error recovery.

13 . One or more non-transitory computer-readable media storing instructions which, when executed by one or more processors, cause:

receiving a natural language query from a user;

retrieving metadata from a machine learning (ML) automation component of a database system, the metadata comprising at least one of:

workload statistics of one or more users of the database system,

schematic details of a data source, or

dynamic content statistics of the data source;

generating a linguistic prompt, based on the metadata and the natural language query received from the user, for a generative artificial intelligence (AI) model; and

generating, by the generative AI model, a response based on the linguistic prompt and a search of the data source, wherein the linguistic prompt limits a scope of the search of the data source based on the metadata.

14 . The one or more non-transitory computer-readable media of claim 13 , wherein the metadata is generated by one or more machine learning models for predicting resource usage and query performance in the database system.

15 . The one or more non-transitory computer-readable media of claim 13 ,

wherein:

the data source is a relational database comprising one or more database tables,

the generative AI model comprises a large language model (LLM), and

the natural language query is a query about the relational database.

16 . The one or more non-transitory computer-readable media of claim 13 ,

wherein:

the data source is a relational database comprising one or more database tables,

the generative AI model comprises a natural language to structured query language (NL2SQL) generative model, and

the NL2SQL generative model is configured to generate one or more SQL queries for searching the relational database based on the linguistic prompt.

17 . The one or more non-transitory computer-readable media of claim 13 ,

wherein:

the data source is a relational database comprising one or more database tables, and

generating the linguistic prompt comprises adding a set of one or more schema descriptions by selecting a subset of database tables from the one or more database tables based at least in part on the metadata and adding schema descriptions for the subset of database tables to the linguistic prompt.

18 . The one or more non-transitory computer-readable media of claim 17 , wherein

adding the set of one or more schema descriptions further comprises:

generating one or more query embeddings based on the natural language query; and

selecting the subset of database tables based at least in part on similarity of the one or more query embeddings and per-table topic modeling embeddings of the one or more database tables.

19 . The one or more non-transitory computer-readable media of claim 13 ,

wherein:

the data source comprises an object store comprising one or more vector stores representing a plurality of documents using semantic encodings, and

generating the linguistic prompt comprises filtering the one or more vector stores based on the metadata.

20 . The method of claim 1 , wherein generating the linguistic prompt comprises generating a context portion of the linguistic prompt based on the metadata, wherein the context portion of the linguistic prompt limits the scope of the search of the data source.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2024
From: AGARWAL, NIPUN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 069284/0868 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2024
From: YAKOVLEV, ANATOLY; AGRAWAL, SANDEEP R.; JINTURKAR, SANJAY
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 069264/0987 →
Continuity (2)
Provisional Application 63563211 · Mar 8, 2024
Related Publication 20250284688A1 · Sep 11, 2025
References Cited (24)
US 5991733A · Aleia · 1999 [cited by examiner]
US 12135890B1 · Lim · 2024 [cited by examiner]
US 12488014B1 · Sharma · 2025 [cited by examiner]
US 20160140177A1 · Chamberlin · 2016 [cited by examiner]
US 20180336198A1 · Zhong · 2018 [cited by examiner]
US 20200104733A1 · Bart · 2020 [cited by applicant]
US 20200302018A1 · Turkkan et al. · 2020 [cited by applicant]
US 20210406717A1 · Tauheed · 2021 [cited by examiner]
US 20230410801A1 · Mishra · 2023 [cited by examiner]
US 20250104132A1 · Dasher · 2025 [cited by examiner]
US 20250139138A1 · Hays · 2025 [cited by examiner]
US 20250165473A1 · Lin · 2025 [cited by examiner]
US 20250284688A1 · Yakovlev · 2025 [cited by examiner]
CM 117435713A · 2024 [cited by applicant]
CN 116383349A · 2023 [cited by applicant]
CN 117609516A · 2024 [cited by applicant]
Ziegler, Albert et al., “A developer's guide to prompt engineering and LLMs”, available: https://github.blog/ai-and-ml/generative-ai/prompt-engineering-guide-generative-ai-llms/. [cited by applicant]
Yang, Hui et al., “Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions”, Jun. 4, 2023, 14 pages. [cited by applicant]
Wei, Jason et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models”, Jan. 10, 2023, 43 pages. [cited by applicant]
Wang, Xuezhi et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models” Mar. 7, 2023, 24 pages. [cited by applicant]
Shin, et al. “AUTOPROMPT: Eliciting Knowledge from Language Models with Automatically Generated Prompts”, Nov. 7, 2020, 15 pages. [cited by applicant]
Liu, Nelson F. et al., “Lost in the Middle: How Language Models Use Long Contexts”, Nov. 20, 2023, 18 pages. [cited by applicant]
Gao et al., “Retrieval-Augmented Generation for Large Language Models: A Survey”, Jan. 5, 2024; Available: https://arxiv.org/pdf/2312.10997v4, 26 pages. [cited by applicant]
Anonymous, “HeatWave User Guide”, Retrieved on Jan. 5, 2026; Available: URL:https://web.archive.org/web/20211019155144/https://downloads.mysql.com/docs/heatwave-en.a4.pdf, 98 pages. [cited by applicant]