Systems and methods for multidimensional data structuring and visualization
Systems and methods provide a multidimensional data structure to tailor data storage and corresponding user output. Tailored graph data structures reduce the level of hallucination in composing an answer to a user using a large language model. In response to a user query of documents, the user query is translated to a vector. From the vector, the most relevant pieces of documents that are as similar as possible in semantics to the question of a person can be determined and communicated to the user. In one aspect, using the vector base (vector query applied to the documents), a graph can be generated in the form of nodes and connections between data in the documents.
1 . A method for multidimensional data indexing, the method comprising:
receiving a plurality of documents from a user;
translating the plurality of documents into a first vector query;
querying a large language model (LLM) with the first vector query to generate first keywords for the plurality of documents;
generating an initial graph with the first keywords for the plurality of documents;
receiving, from the user, a knowledge area comprising a specific domain field, a research goal comprising at least one task, and a desired result comprising a final vision associated with the research goal;
generating a system message template for a keywords query;
embedding the knowledge area, the research goal, and the desired result to the system message template to create a system message;
translating the system message into a second vector query;
querying the LLM with the second vector query to generate second keywords;
determining a difference in keywords between the first keywords and the second keywords, wherein the difference includes at least one keyword that is in the first keywords but not in the second keywords or at least one keyword that is in the second keywords but not in the first keywords;
rebuilding the graph including adding at least one new connection between nodes of the graph based on the difference in keywords.
2 . The method of claim 1 , wherein rebuilding the graph further comprises: performing keywords reconciliation including:
finding a correspondence for each of the difference in keywords in organized data including by conducting a fuzzy search for a given keyword in the organized data, and
leaving only reconciled keywords having correspondence in organized data in the difference in keywords by when the fuzzy search returns 0, filtering out the given keyword and when the fuzzy search returns one or more results, using a first fuzzy search result.
3 . The method of claim 2 , wherein rebuilding the graph further comprises:
performing keywords relevance including:
generating a description of each keyword with the LLM,
comparing the description of each keyword from the LLM with a knowledge area defined in the organized data, and
removing keywords that are not relevant based on the comparing.
4 . The method of claim 2 , wherein rebuilding the graph further comprises:
performing keywords relevance including:
generating a first similarity vector for the keyword and the first fuzzy search result,
generating a second similarity vector for the knowledge area,
performing a vector comparison between the first vector and the second vector, and
removing keywords that are not relevant based on the vector comparison.
5 . The method of claim 1 , wherein rebuilding the graph further comprises:
for each relevant keyword, obtaining a taxonomy from organized data, wherein the taxonomy includes all parent items and all child items the relevant keyword; and
integrating each relevant keyword and the taxonomy in the graph through common connections and between new and existing keywords in the graph.
6 . The method of claim 5 , wherein rebuilding the graph further comprises:
forming a new node for each relevant keyword and each new taxonomy node;
forming a new edge for a first relationship in the plurality of documents; and
forming a new edge for a second relationship according to the taxonomy.
7 . The method of claim 1 , wherein generating the initial graph comprises:
creating a first plurality of nodes from an inner context of the plurality of documents;
clustering the first plurality of nodes according to a category of knowledge; and
forming the initial graph having at least three levels based on the first keywords and the clustering of the first plurality of nodes.
8 . The method of claim 7 , further comprising:
creating an updated plurality of nodes based on the inner context of the plurality of documents and the difference in keywords;
provided additional context about clustering the updated plurality of nodes according to a category of knowledge; and
forming a new graph based on the additional context.
9 . The method of claim 8 , wherein determining the difference between the first keywords and the second keywords further comprises comparing the first keywords and the second keywords to find at least one difference, and wherein providing the difference to the user further comprises modifying only the at least one difference in the graphical user interface.
10 . The method of claim 1 , further comprising providing the graph in a navigable graphical user interface.
11 . The method of claim 10 , wherein providing the graph in a navigable graphical user interface further comprises:
defining a plurality of graph depths;
extracting all possible connections in the plurality documents with the keyword and keywords with concept;
processing the possible connections by filtering and sorting most frequently used first keywords and second keywords;
visualizing the graph in the graphical user interface as a tree mindmap using the plurality of graph depths and the possible connections.
12 . The method of claim 1 , wherein the knowledge area, the research goal, and the desired result are captured from a user interface chat with the user.
13 . The method of claim 1 , further comprising:
indexing the plurality of documents as an indexed plurality of documents before generating the initial graph, wherein
generating the initial graph further comprises uploading the indexed plurality of documents to the LLM.
14 . A system for multidimensional data indexing, the system comprising:
a graph database configured to store a plurality of graphs, each graph comprising nodes, edges, node attributes, and edge attributes; and
a computing device including:
at least one processor and memory operably coupled to the at least one processor, instructions that, when executed, cause the at least one processor to implement:
a large language model (LLM) configured to receive a plurality of documents from a user,
a graph building engine configured to generate an initial graph including first keywords for the plurality of documents,
wherein the LLM is further configured to:
receive, from the user, a knowledge area comprising a specific domain field, a research goal comprising at least one task, and a desired result comprising a final vision associated with the research goal,
generate a system message template for a keywords query, and
embed the knowledge area, the research goal, and the desired result to the system message template to create a system message,
an agent configured to translate a query of the plurality of documents into a first vector query of the LLM to generate first keywords and translate the system message into a second vector query of the LLM to generate second keywords, and
wherein the graph building engine is further configured to:
determine a difference in keywords between the first keywords and the second keywords wherein the difference includes at least one keyword that is in the first keywords but not in the second keywords or at least one keyword that is in the second keywords but not in the first keywords, and
rebuild the graph including adding at least one new connection between nodes of the graph based on the difference in keywords.
15 . The system of claim 14 , wherein the graph building engine is further configured to rebuild the graph including:
performing keywords reconciliation including:
finding a correspondence for each of the difference in keywords in organized data including by conducting a fuzzy search for a given keyword in the organized data, and leaving only reconciled keywords having correspondence in organized data in the difference in keywords by when the fuzzy search returns 0, filtering out the given keyword and when the fuzzy search returns one or more results, using a first fuzzy search result.
16 . The system of claim 15 , wherein the graph building engine is further configured to rebuild the graph including:
performing keywords relevance including:
generating a description of each keyword with the LLM,
comparing the description of each keyword from the LLM with the knowledge area, and removing keywords that are not relevant based on the comparing.
17 . The system of claim 15 , wherein the graph building engine is further configured to rebuild the graph including:
performing keywords relevance including:
generating a first similarity vector for the keyword and the first fuzzy search result;
generating a second similarity vector for the knowledge area;
performing a vector comparison between the first vector and the second vector;
removing keywords that are not relevant based on the vector comparison.
18 . The system of claim 14 , wherein the graph building engine is further configured to rebuild the graph including:
for each relevant keyword, obtaining a taxonomy from organized data, wherein the taxonomy includes all parent items and all child items the relevant keyword; and
integrating each relevant keyword and the taxonomy in the graph through common connections and between new and existing keywords in the graph.
19 . The system of claim 18 , wherein the graph building engine is further configured to rebuild the graph including:
forming a new node for each relevant keyword and each new taxonomy node;
forming a new edge for a first relationship in the plurality of documents; and
forming a new edge for a second relationship according to the taxonomy.
20 . The system of claim 14 , the graph building engine is further configured to generate the initial graph including:
creating a first plurality of nodes from an inner context of the plurality of documents;
clustering the first plurality of nodes according to a category of knowledge; and
forming the initial graph having at least three levels based on the first keywords and the clustering of the first plurality of nodes.