IP Library Granted Patent US 12,423,265
Granted Patent B1
US 12,423,265 · App. 18/893,703 · Granted Sep 23, 2025

Prompting a large language model for vector embeddings and metadata to generate an indexed computing file

Inventors: Eric B. Hensley (San Francisco, CA); Mark C. Wolochuk (Portland, OR)
Assignee: Aravo Solutions, Inc.
G06F16/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,423,265
App. No.
18/893,703
Granted
Sep 23, 2025
Kind
B1
Abstract

Disclosed are methods and systems for prompting a large language model (LLM) for vector embeddings and metadata to generate an indexed computing file. An exemplary method includes: receiving a file from a source; extracting data from the file; transmitting the data to an LLM; determining metadata associated with the file using the LLM; chunking the file, thereby generating a chunked file; transmitting the chunked file to the LLM; determining at least one vector embedding associated with the chunked file using the LLM; configuring a vector database; and inserting the metadata and the at least one vector embedding into the vector database, thereby generating an indexed computing file.

Claims (59)

1. A method for generating an indexed computing file based on inserting vector embeddings and metadata, received from a large language model (LLM), into a configured vector database, the method comprising:

receiving, using one or more computing device processors, a file from a first file source, wherein the file comprises unstructured data;

extracting, using the one or more computing device processors, first text from the file;

transmitting, using the one or more computing device processors, the first text from the file to an LLM;

receiving, using the one or more computing device processors, at a first time, metadata associated with the file from the LLM, wherein the metadata associated with the file comprises or is based on file quality data, wherein the file quality data comprises or is based on at least three of: a nature of the file, a credibility of the file, a freshness of the file, and a file quality indicator of the file, wherein the credibility of the file comprises a first indicator associated with a source of the file, wherein the metadata associated with the file further comprises a citation, wherein the citation comprises: at least some second text from the file, a file name corresponding with the file, and a page number associated with the at least some second text from the file;

executing, using the one or more computing device processors, at a second time or at the first time, a chunking computing operation using the file, thereby resulting in a chunked file;

transmitting, using the one or more computing device processors, third text associated with the chunked file to the LLM;

receiving, using the one or more computing device processors, at least one vector embedding for the third text associated with the chunked file from the LLM, wherein the at least one vector embedding comprises or is based on a semantic structure of at least some of the third text associated with the chunked file, wherein the semantic structure is based on or associated with a conceptual meaning of the at least some of the text associated with the chunked file, wherein the at least some of the third text associated with the chunked file comprises or is associated with a first computing prompt;

configuring, using the one or more computing device processors, a vector database to store vector embeddings and metadata, thereby resulting in a configured vector database;

first inserting, using the one or more computing device processors, at a third time following the first time and the second time, the at least one vector embedding for the third text associated with the chunked file, into the configured vector database;

second inserting, using the one or more computing device processors, at the third time following the first time and the second time, the metadata associated with the file into the configured vector database; and

generating, based on the first inserting the at least one vector embedding for the third text associated with the chunked file into the configured vector database, and the second inserting the metadata associated with the file into the configured vector database, an indexed computing file.

2. The method of claim 1 , wherein the file quality data comprising or being based on the at least three of: the nature of the file, the credibility of the file, the freshness of the file, and the file quality indicator of the file, further comprises or is based on: the nature of the file, the credibility of the file, the freshness of the file, and the file quality indicator of the file.

3. The method of claim 1 , wherein the nature of the file comprises a second indicator associated with a classification of the file.

4. The method of claim 1 , wherein the freshness of the file comprises a second indicator associated with a creation time of the file.

5. A system for generating an indexed computing file based on inserting vector embeddings and metadata, received from a large language model (LLM), into a configured vector database, the system comprising:

one or more computing system processors; and

memory storing instructions that, when executed by the one or more computing system processors, cause the system to:

receive a file from a first file source, wherein the file comprises unstructured data;

extract first text from the file;

transmit the first text from the file to a LLM;

receive, at a first time, metadata associated with the file from the LLM, wherein the metadata associated with the file comprises or is based on file quality data, wherein the file quality data comprises or is based on at least three of: a nature of the file, a credibility of the file, a freshness of the file, and a file quality indicator of the file, wherein the metadata associated with the file comprises on a citation, wherein the citation comprises: at least some second text from the file, a file name corresponding with the file, and a page number associated with the at least some second text from the file:

execute, at a second time or at the first time, a chunking computing operation using the file, thereby resulting in a chunked file;

transmit third text associated with the chunked file to the LLM;

receive at least one vector embedding for the third text associated with the chunked file from the LLM, wherein the at least one vector embedding comprises or is based on a semantic structure of at least some of the third text associated with the chunked file, wherein the semantic structure is based on or associated with a conceptual meaning of the at least some of the third text associated with the chunked file, wherein the at least some of the third text associated with the chunked file comprises or is associated with a first computing prompt;

configure a vector database to store vector embeddings and metadata, thereby resulting in a configured vector database;

first insert, at a third time following the first time and the second time, the at least one vector embedding for the third text associated with the chunked file, into the configured vector database;

second insert, at the third time following the first time and the second time, the metadata associated with the file into the configured vector database; and

generate, based on the first inserting the at least one vector embedding for the third text associated with the chunked file into the configured vector database, and the second inserting the metadata associated with the file into the configured vector database, an indexed computing file.

6. The system of claim 5 , wherein the metadata associated with the file further comprises third party source data.

7. The system of claim 5 , wherein the system comprises or is comprised in one or more computing systems associated with one or more locations.

8. The system of claim 5 , wherein the LLM is hosted on a third-party server.

9. The system of claim 5 , wherein the LLM is hosted on a local server.

10. The system of claim 5 , wherein the semantic structure of the at least some of the third text associated with the chunked file comprises the conceptual meaning of the at least some of the third text associated with the chunked file.

11. A method for generating an indexed computing file based on inserting vector embeddings and metadata, received from a large language model (LLM), into a configured vector database, the method comprising:

receiving, using one or more computing device processors, a file from a first file source, wherein the file comprises unstructured data;

extracting, using the one or more computing device processors, first data from the file;

transmitting, using the one or more computing device processors, the first data from the file to a LLM;

receiving, using the one or more computing device processors, at a first time, metadata associated with the file from the LLM, wherein the metadata associated with the file comprises or is based on file quality data, wherein the file quality data comprises or is based on at least three of: a nature of the file, a credibility of the file, a freshness of the file, and a file quality indicator of the file, wherein the metadata associated with the file comprises a citation, wherein the citation comprises: at least some second data from the file, a file name corresponding with the file, and a page number associated with the at least some second data from the file;

executing, using the one or more computing device processors, at a second time or at the first time, a chunking computing operation using the file, thereby resulting in a chunked file;

transmitting, using the one or more computing device processors, third data associated with the chunked file to the LLM;

receiving, using the one or more computing device processors, at least one vector embedding for the third data associated with the chunked file from the LLM, wherein the at least one vector embedding comprises or is based on a semantic structure of at least some of the third data associated with the chunked file, wherein the semantic structure is based on or associated with a conceptual meaning of the at least some of the third data associated with the chunked file, wherein the at least some of the third data associated with the chunked file comprises or is associated with a first computing prompt;

configuring, using the one or more computing device processors, a vector database to store vector embeddings and metadata, thereby resulting in a configured vector database;

first inserting, using the one or more computing device processors, at a third time following the first time and the second time, the at least one vector embedding for the third data associated with the chunked file, into the configured vector database;

second inserting, using the one or more computing device processors, at the third time following the first time and the second time, the metadata associated with the file into the configured vector database,

wherein the first inserting the at least one vector embedding for the third data associated with the chunked file into the configured vector database, and the second inserting the metadata associated with the file into the configured vector database, result in an indexed computing file; and

storing, using the one or more computing device processors, the indexed computing file in a file repository.

12. The method of claim 11 , wherein the at least some of the third data associated with the chunked file comprises or is based on at least one of: a word from the chunked file, a phrase from the chunked file, a sentence from the chunked file, a paragraph from the chunked file, or the chunked file.

13. The method of claim 11 , wherein the file quality indicator of the file is based on the nature of the file, the credibility of the file, and the freshness of the file.

14. The method of claim 11 , wherein the one or more computing device processors are comprised in one or more computing systems, wherein the one or more computing systems are located in one or more locations.

15. The method of claim 11 , wherein the first data from the file comprises at least one of: text, an image, a figure, a table, or a diagram.

16. The method of claim 11 , wherein the file from the first file source comprises at least one of: an audit document, a Service Organization Control (SOC) 2 report, a policy document, a 10K financial report, a technical description document, a SOC 1 report, a data security document, a corporate charter, an information technology procedure document, a financial report, a questionnaire, a 10Q report, a human resources document, or a screenshot of an internal system.

17. The method of claim 1 , wherein:

the nature of the file comprises a second indicator associated with a classification of the file;

the freshness of the file comprises a third indicator associated with a creation time of the file; and

the file quality indicator of the file is based on the nature of the file, the credibility of the file, and the freshness of the file.

18. The method of claim 1 , further comprising executing, using the one or more computing device processors, based on filter data, a filtering operation on the configured vector database, wherein the filter data is based on at least one of: the nature of the file, the credibility of the file, the freshness of the file, or the file quality indicator of the file.

19. The method of claim 18 , wherein the filter data based on the at least one of: the nature of the file, the credibility of the file, the freshness of the file, or the file quality indicator of the file is based on at least two of: the nature of the file, the credibility of the file, the freshness of the file, or the file quality indicator of the file.

20. The method of claim 18 , wherein the filter data based on the at least one of: the nature of the file, the credibility of the file, the freshness of the file, or the file quality indicator of the file is based on at least three of: the nature of the file, the credibility of the file, the freshness of the file, or the file quality indicator of the file.

Assignments (2)
SECURITY INTEREST Recorded Jan 27, 2026
From: ARAVO SOLUTIONS, INC.
To: COMERICA BANK
Reel/Frame 073599/0351 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 26, 2024
From: HENSLEY, ERIC B.; WOLOCHUK, MARK C.
To: ARAVO SOLUTIONS, INC.
Reel/Frame 068708/0381 →
References Cited (41)
US 8316292B1 · Verstak · 2012 [cited by examiner]
US 11960514B1 · Taylert · 2024 [cited by examiner]
US 11971914B1 · Watson · 2024 [cited by examiner]
US 12056003B1 · Ramos · 2024 [cited by examiner]
US 12105729B1 · Haq · 2024 [cited by examiner]
US 12111858B1 · Radhakrishnan · 2024 [cited by examiner]
US 12141539B1 · Nichol · 2024 [cited by examiner]
US 12155742B1 · Zafar · 2024 [cited by examiner]
US 20110282888A1 · Koperski · 2011 [cited by examiner]
US 20140075282A1 · Shah · 2014 [cited by examiner]
US 20220222289A1 · Srinivasan · 2022 [cited by examiner]
US 20230259705A1 · Tunstall-Pedoe et al. · 2023 [cited by applicant]
US 20230274086A1 · Tunstall-Pedoe et al. · 2023 [cited by applicant]
US 20230274089A1 · Tunstall-Pedoe et al. · 2023 [cited by applicant]
US 20230274094A1 · Tunstall-Pedoe et al. · 2023 [cited by applicant]
US 20230315955A1 · Karadimitriou et al. · 2023 [cited by applicant]
US 20230316006A1 · Tunstall-Pedoe et al. · 2023 [cited by applicant]
US 20240104305A1 · Glesinger · 2024 [cited by examiner]
US 20240256678A1 · Thompson · 2024 [cited by applicant]
US 20240273227A1 · Thompson · 2024 [cited by applicant]
US 20240281472A1 · LaRhette et al. · 2024 [cited by applicant]
US 20240289365A1 · Beauchamp · 2024 [cited by examiner]
US 20240289863A1 · Smith Lewis · 2024 [cited by examiner]
US 20240311407A1 · Barron · 2024 [cited by examiner]
US 20240330589A1 · Kotaru · 2024 [cited by examiner]
US 20240362286A1 · He · 2024 [cited by examiner]
US 20240370479A1 · Hudetz · 2024 [cited by examiner]
US 20240370517A1 · DeVos · 2024 [cited by examiner]
US 20240370570A1 · Betthauser · 2024 [cited by examiner]
US 20240378390A1 · Korganyan · 2024 [cited by examiner]
US 20240394291A1 · Nelson · 2024 [cited by examiner]
US 20240396920A1 · Bonney et al. · 2024 [cited by applicant]
Non-Final Office Action dated Dec. 18, 2024 in connection with U.S. Appl. No. 18/893,699, 25 pages. [cited by applicant]
Notice of Allowance dated Dec. 18, 2024 in connection with U.S. Appl. No. 18/893,706, 9 pages. [cited by applicant]
Daqqah, Bilal H. “Leveraging Large Language Models (LLMs) for Automated Extraction and Processing of Complex Ordering Forms.” PhD diss., Massachusetts Institute of Technology, 2024. (Year: 2024). [cited by applicant]
Liu, Xiaoxia, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, and Dongxia Wang. “Prompting frameworks for large language models: A survey.” arXiv preprint arXiv:2311.12785 (2023). (Year: 2023). [cited by applicant]
Wedholm, William. “Exploring the Influence of Data Formats on the Consistency of Large Language Models Outputs.” (2024). (Year: 2024). [cited by applicant]
Schilling-Wilhelmi, Mara, Martino Rfos-Garcfa, Sherjeel Shabih, Marfa Victoria Gil, Santiago Miret, Christoph T. Koch, Jose A Marquez, and Kevin Maik Jablonka. “From text to insight: large language models for materials … [cited by applicant]
De Bellis, A, Structuring the unstructured: an LLM-guided transition, Doctoral Consortium at ISWC 2023 co-located with 22st International Semantic Web Conference (ISWC 2023), pp. 1-8. (Year: 2023). [cited by applicant]
Aishwarya, V, A Prompt Engineering Approach for Structured Data Extraction from Unstructured Text Using Conversational LLMs, ACAi 2023: 2023 6th International Conference on Algorithms, Computing and Artificial Intellige… [cited by applicant]
Peng et al., Embedding-based Retrieval with LLM for Effective Agriculture Information Extracting from Unstructured Data, R. Peng, K. Liu, P. Yang, Z. Yuan, and S. Li, “Embedding-basedretrieval with LLM for effective agr… [cited by applicant]