IP Library Patent Application 18972115
Patent Application
App. No. 18/972,115

Techniques for Generating Structured Data from Unstructured Data

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/972,115
Abstract

This disclosure describes techniques for generating structured data, such as for client system intelligence, based on unstructured data in an efficient, valuable, automated, and intelligent manner. Content may be extracted from unstructured data and then processed and stored in a manner to facilitate correlating the content with a query. For example, content may be embedded into a vector space. When a query is received, the query may similarly be embedded into the vector space in order to identify content that is relevant to the query. A prompt for a machine learning (ML) model (e.g., a large language model (LLM)) may then be automatically generated based on the query and the relevant content. The output of the ML model may then be validated and integrated into various downstream systems and subsystems, such as to recognize client development opportunities.

Claims (63)

1 . A method for generation of structured data for query execution, the method comprising:

identifying, by a server computer system, a query requesting information regarding a client system;

generating, with a first machine learning (ML) model executed by the server computer system, an embedding in vector space based on the query;

identifying, by the server computer system, a similar embedding in the vector space and located in a vector store, wherein the vector store includes a plurality of embeddings generated based on data scraped from a website associated with the client system;

retrieving, from a database, content associated with the similar embedding;

generating, by the server computer system, a prompt based on the query and the content associated with the similar embedding;

providing, by the server computer system, the prompt to a second ML model;

identifying, by the server computer system, response data generated by the second ML model based on the prompt; and

transforming, by the server computer system, the response data into structured data corresponding to the information requested regarding the client system.

2 . The method of claim 1 , further comprising:

generating a plurality of content snapshots based on data obtained from scraping a website associated with the client system; and

generating the plurality of embeddings in the vector store based on the plurality of content snapshots and the first ML model.

3 . The method of claim 2 , wherein generating the plurality of embeddings based on the plurality of content snapshots and the first ML model comprises:

removing a script from at least one of the plurality of content snapshots to produce cleaned content;

extracting a text chunk from the cleaned content; and

providing the text chunk to the first ML model.

4 . The method of claim 2 , wherein each of the plurality of content snapshots comprises an HTML snapshot.

5 . The method of claim 1 , wherein providing the response to the second ML model comprises:

generating an application program interface (API) request comprising the prompt; and

communicating the API request to the second ML model.

6 . The method of claim 5 , wherein identifying the response data generated by the second ML model based on the prompt comprises receiving an API response corresponding to the API request.

7 . The method of claim 1 , wherein the first ML model comprises a bi-encoder or a bi-direction encoder.

8 . The method of claim 1 , further comprising updating a profile corresponding to the client system based on the structured data corresponding to the information requested regarding the client system.

9 . The method of claim 1 , further comprising configuring a service of the server computer system based on the structured data corresponding to the information requested regarding the client system.

10 . A server computer system, comprising:

a memory; and

a processor coupled to the memory configured to:

identify a query requesting information regarding a client system;

generate, with a first machine learning (ML) model, an embedding in vector space based on the query;

identify a similar embedding in the vector space and located in a vector store, wherein the vector store includes a plurality of embeddings generated based on data scraped from a website associated with the client system;

retrieve, from a database, content associated with the similar embedding;

generating a prompt based on the query and the content associated with the similar embedding;

provide the prompt to a second ML model;

identify response data generated by the second ML model based on the prompt; and

transform the response data into structured data corresponding to the information requested regarding the client system.

11 . The server computer system of claim 10 , wherein the processor coupled to the memory is further configured to:

generate a plurality of content snapshots based on data obtained from scraping a website associated with the client system; and

generate the plurality of embeddings in the vector store based on the plurality of content snapshots and the first ML model.

12 . The server computer system of claim 11 , wherein to generate the plurality of embeddings based on the plurality of content snapshots and the first ML model the processor coupled to the memory is further configured to:

remove a script from at least one of the plurality of content snapshots to produce cleaned content;

extract a text chunk from the cleaned content; and

provide the text chunk to the first ML model.

13 . The server computer system of claim 11 , wherein each of the plurality of content snapshots comprises an HTML snapshot.

14 . The server computer system of claim 10 , wherein the processor coupled to the memory is further configured to update a profile corresponding to the client system based on the structured data corresponding to the information requested regarding the client system.

15 . The server computer system of claim 10 , wherein the processor coupled to the memory is further configured to configure a service system of the server computer system based on the structured data corresponding to the information requested regarding the client system.

16 . A non-transitory computer readable storage medium including instructions that, when executed by a processor, cause the processor to perform operations, the operations comprising:

identifying, by a server computer system, a query requesting information regarding a client system;

generating, with a first machine learning (ML) model executed by the server computer system, an embedding in vector space based on the query;

identifying, by the server computer system, a similar embedding in the vector space and located in a vector store, wherein the vector store includes a plurality of embeddings generated based on data scraped from a website associated with the client system;

retrieving, from a database, content associated with the similar embedding;

generating, by the server computer system, a prompt based on the query and the content associated with the similar embedding;

providing, by the server computer system, the prompt to a second ML model;

identifying, by the server computer system, response data generated by the second ML model based on the prompt; and

transforming, by the server computer system, the response data into structured data corresponding to the information requested regarding the client system.

17 . The non-transitory computer readable storage medium of claim 16 , the operations further comprising:

generating a plurality of content snapshots based on data obtained from scraping a website associated with the client system; and

generating the plurality of embeddings in the vector store based on the plurality of content snapshots and the first ML model.

18 . The non-transitory computer readable storage medium of claim 17 , the operations to generate the plurality of embeddings based on the plurality of content snapshots and the first ML model further comprising:

removing a script from at least one of the plurality of content snapshots to produce cleaned content;

extracting a text chunk from the cleaned content; and

providing the text chunk to the first ML model.

19 . The non-transitory computer readable storage medium of claim 16 , the operations further comprising updating a profile corresponding to the client system based on the structured data corresponding to the information requested regarding the client system.

20 . The non-transitory computer readable storage medium of claim 16 , the operations further comprising configuring a service system of the server computer system based on the structured data corresponding to the information requested regarding the client system.

Assignments (2)
CHANGE OF NAME Recorded Mar 13, 2026
From: STRIPE, INC.
To: STRIPE, LLC
Reel/Frame 075093/0754 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2024
From: KEDIA, GAUTAM; KWAK, DONG (DAVID); SINGH, AMANPREET; HWANG, MICHELLE
To: STRIPE, INC.
Reel/Frame 069513/0835 →