IP Library › Granted Patent US 12,505,095
Granted Patent B1
US 12,505,095 · App. 19/255,270 · Granted Dec 23, 2025

Method and system for constructing vector databases used for converting free text queries to cyber language queries

Inventors: Eli Rozen (Tel Aviv, IL); Gil Knafo (Tel Aviv, IL); Yarden Sasson (Ramat Gan, IL)
Assignee: Vega Cyber Solutions LTD
G06F16/2433G06F16/213G06F16/24522H04L63/1425
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,095
App. No.
19/255,270
Granted
Dec 23, 2025
Kind
B1
Abstract

A system and method for querying data sources for cybersecurity analysis is presented. The system and method include: receiving security logs from at least one data source, wherein the security logs lack pre-defined schema; generating a schema of the security logs based on at least a type of data of the security logs, wherein the generated schema includes fields of the security logs and values of the fields; embedding field vectors, wherein each field vector is a vector representation of a value of each respective field; embedding value vectors, wherein each value vector is a vector representation of a natural language description of each value in each respective field; and generating a query in a cyber language query, using an AI system, for execution on at least one target data source based, in part, on the generated schema, the embedded field vectors, and the embedded value vectors.

Claims (76)

1 . A method for querying data sources for cybersecurity analysis, comprising:

receiving, by a device, security logs from at least one data source, wherein the security logs lack pre-defined schema;

generating, using an artificial intelligence (AI) embedding system, a schema of the security logs based on at least a type of data of the security logs, wherein the generated schema includes fields of the security logs and values of the fields;

embedding, using the AI embedding system, field vectors, wherein each field vector is a vector representation of a value of each respective field;

embedding, using the AI embedding system, value vectors, wherein each value vector is a vector representation of a natural language description of each value in each respective field;

generating a query in a cyber language query, using an AI system; and

executing the generated query on at least one target data source based, in part, on the generated schema, the embedded field vectors, and the embedded value vectors.

2 . The method of claim 1 , further comprising:

generating pairs of NLQs and cyber language queries, wherein the pairs are potentially valid matches between NLQs and the cyber language queries;

validating that the generated cyber language queries are properly matched to the generated NLQs; and

embedding vectors of each of the pairs.

3 . The method of claim 1 , wherein generating the cyber language query, using an AI system further comprises:

identifying fields that are relevant to a natural language query (NLQ) based on a field vector similarity search using the embedded field vectors;

computing values for the identified fields based on actual values that are relevant to the identified fields and the NLQ based on a value vector similarity search using the embedded value vectors;

refining the generated schema based on the identified fields and the computed values;

generating, based on at least the refined schema, a prompt for a Large Language Model (LLM), wherein the LLM is executed by an AI-based query generation system; and

feeding the prompt to the LLM, wherein the LLM executes the prompt using the AI-based query generation system, to convert the NLQ into the cyber language query.

4 . The method of claim 3 , wherein the field vector similarity search further comprises:

comparing, using a semantic distance metric, a feature vector representing the NLQ and each embedded field vector;

identifying at least one field for which the embedded field vector is below a semantic distance threshold from the feature vector representing the NLQ; and

validating the identified at least one field against ground-truth pairs of natural language field descriptions and fields searchable in the cyber language for at least one target data source.

5 . The method of claim 3 , wherein the value vector similarity search further comprises:

comparing, using a semantic distance metric, a feature vector representing the NLQ and each embedded value vector;

identifying at least one value for which the embedded value vector is below a semantic distance threshold from the value vector representing the NLQ; and

validating the at least one value against ground-truth pairs of natural language value descriptions and actual values in the at least one target data source.

6 . The method of claim 3 , wherein refining the schema of the logs based on the identified fields and computed values further comprises:

narrowing the schema to the most pertinent elements of the schema based on the identified fields and computed values.

7 . The method of claim 1 , further comprising:

generating, using the AI embedding system, the natural language descriptions of the received fields, wherein the AI embedding system is configured to gather relevant information on the fields from official documentation and other available data sources to generate the natural language descriptions.

8 . The method of claim 1 , further comprising:

validating syntactical correctness of the generated cyber language query.

9 . The method of claim 1 , wherein the cyber language query is a read-only request to process data stored hierarchically in databases, tables, and columns.

10 . The method of claim 1 , wherein the cyber language is at least Kusto Query Language.

11 . A non-transitory computer-readable medium storing a set of instructions for querying data sources for cybersecurity analysis, the set of instructions comprising:

one or more instructions that, when executed by one or more processors of a device, cause the device to:

receive, by a device, security logs from at least one data source, wherein the security logs lack pre-defined schema;

generate, using an artificial intelligence (AI) embedding system, a schema of the security logs based on at least a type of data of the security logs, wherein the generated schema includes fields of the security logs and values of the fields;

embed, using the AI embedding system, field vectors, wherein each field vector is a vector representation of a value of each respective field;

embed, using the AI embedding system, value vectors, wherein each value vector is a vector representation of a natural language description of each value in each respective field;

generate a query in a cyber language query, using an AI system and

execute the query on at least one target data source based, in part, on the generated schema, the embedded field vectors, and the embedded value vectors.

12 . A system for querying data sources for cybersecurity analysis comprising:

one or more processors; and

a memory including instructions configured to:

receive, by a device, security logs from at least one data source, wherein the security logs lack pre-defined schema;

generate, using an artificial intelligence (AI) embedding system, a schema of the security logs based on at least a type of data of the security logs, wherein the generated schema includes fields of the security logs and values of the fields;

embed, using the AI embedding system, field vectors, wherein each field vector is a vector representation of a value of each respective field;

embed, using the AI embedding system, value vectors, wherein each value vector is a vector representation of a natural language description of each value in each respective field; and

generate a query in a cyber language query, using an AI system; and

execute the query on at least one target data source based, in part, on the generated schema, the embedded field vectors, and the embedded value vectors.

13 . The system of claim 12 , further comprising:

generating pairs of NLQs and cyber language queries, wherein the pairs are potentially valid matches between NLQs and the cyber language queries

validating that the generated cyber language queries are properly matched to the generated NLQs; and

embedding vectors of each of the pairs.

14 . The system of claim 12 , wherein generating the cyber language query, using an AI system further comprises:

identifying fields that are relevant to a natural language query (NLQ) based on a field vector similarity search using the embedded field vectors

computing values for the identified fields based on actual values that are relevant to the identified fields and the NLQ based on a value vector similarity search using the embedded value vectors

refining the generated schema based on the identified fields and the computed values

generating, based on at least the refined schema, a prompt for a Large Language Model (LLM), wherein the LLM is executed by an AI-based query generation system; and

feeding the prompt to the LLM, wherein the LLM executes the prompt using the AI-based query generation system, to convert the NLQ into the cyber language query.

15 . The system of claim 14 , wherein the field vector similarity search further comprises:

comparing, using a semantic distance metric, a feature vector representing the NLQ and each embedded field vector

identifying at least one field for which the embedded field vector is below a semantic distance threshold from the feature vector representing the NLQ; and

validating the identified at least one field against ground-truth pairs of natural language field descriptions and fields searchable in the cyber language for at least one target data source.

16 . The system of claim 14 , wherein the value vector similarity search further comprises:

comparing, using a semantic distance metric, a feature vector representing the NLQ and each embedded value vector

identifying at least one value for which the embedded value vector is below a semantic distance threshold from the value vector representing the NLQ; and

validating the at least one value against ground-truth pairs of natural language value descriptions and actual values in the at least one target data source.

17 . The system of claim 14 , wherein refining the schema of the logs based on the identified fields and computed values further comprises:

narrowing the schema to the most pertinent elements of the schema based on the identified fields and computed values.

18 . The system of claim 12 , further comprising:

generating, using the AI embedding system, the natural language descriptions of the received fields, wherein the AI embedding system is configured to gather relevant information on the fields from official documentation and other available data sources to generate the natural language descriptions.

19 . The system of claim 12 , further comprising:

validating syntactical correctness of the generated cyber language query.

20 . The system of claim 12 , wherein the cyber language query is a read-only request to process data stored hierarchically in databases, tables, and columns.

21 . The system of claim 12 , wherein the cyber language is at least Kusto Query Language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2025
From: ROZEN, ELI; KNAFO, GIL; SASSON, YARDEN
To: VEGA CYBER SOLUTIONS LTD
Reel/Frame 071579/0331 →
Continuity (1)
Provisional Application 63770086 · Mar 11, 2025
References Cited (21)
US 9680779B2 · Marovets · 2017 [cited by examiner]
US 11736526B2 · Jeong · 2023 [cited by examiner]
US 11954102B1 · Palaniappan et al. · 2024 [cited by applicant]
US 20200120102A1 · Cybulski · 2020 [cited by examiner]
US 20220035775A1 · Sriharsha · 2022 [cited by examiner]
US 20230291743A1 · Shua · 2023 [cited by examiner]
US 20240070270A1 · Mace et al. · 2024 [cited by applicant]
US 20240259435A1 · Giralte · 2024 [cited by examiner]
US 20240265913A1 · Laptev · 2024 [cited by examiner]
US 20240364712A1 · Chandana · 2024 [cited by examiner]
US 20240419803A1 · Blum et al. · 2024 [cited by applicant]
US 20250028746A1 · Cantu et al. · 2025 [cited by applicant]
US 20250086308A1 · Crume et al. · 2025 [cited by applicant]
US 20250156384A1 · Rajagopalan · 2025 [cited by examiner]
US 20250217346A1 · Black · 2025 [cited by examiner]
US 20250245446A1 · Ayed et al. · 2025 [cited by applicant]
US 20250247400A1 · Verma · 2025 [cited by examiner]
2025 3rd International Conference on Self Sustainable Artificial Intelligence Systems (ICSSAS) Year: 2025 | Conference Paper | Publisher: IEEENarang et al., “Project Pulse: A Scalable, Secure, and Smart Project Manageme… [cited by examiner]
Ji et al., “Leveraging Large Language Model for Intelligent Log Processing and Autonomous Debugging in Cloud AI Platforms,” 2025 8th International Conference on Advanced Electronic Materials, Computers and Software Engi… [cited by examiner]
Tang, X., Abdi, A. H., Eichelbaum, J., Das, M., Klein, A., Pakis, N. I., Blum, W., Mace, D. L., Raja, T., Padmanabhan, N., & Xing, Y. (2024). NL2KQL: From natural language to kusto query. In arXiv [cs.DB]. http://arxiv.… [cited by applicant]
J. Yi, G. Chen and Z. Shen, “RH-SQL: Refined Schema and Hardness Prompt for Text-to-SQL,”, 2024, 2024 6th International Conference on Electronic Engineering and Informatics (EEi), Chongqing, China, pp. 1340-1343 (Year: … [cited by applicant]
Cited By (2)
US 12,675,446 US 12,694,128