IP Library Granted Patent US 12664220
Granted Patent B2
US 12664220 · App. 18/539,521 · Granted Jun 23, 2026

Systems and methods for AI-based content extraction and generation

Inventor: Aaron Flores (Milpitas, CA)
Assignee: YAHOO ASSETS LLC
G06F16/951G06F40/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664220
App. No.
18/539,521
Granted
Jun 23, 2026
Kind
B2
Abstract

Disclosed are systems and methods that provide a decision-intelligence (DI)-based, computerized framework for deterministically identifying and extracting content from network resources, and generating focused content based therefrom for delivery to electronic users. The framework enables real-time customization of extraction tasks, which can cause tailored, accurate results that are contextually relevant and tied into the purpose of the extraction task. This data can then be compiled and/or leveraged to generate content campaigns that can target specific sets of users, geographies, time periods, trends and the like. The framework can leverage a large language model (LLM) to seamlessly extract relevant information from network resources, which enables the generation and execution of extraction requests that can be dynamically executed and updated, which can enable the framework to “drill-down” on contextual and/or topical aspects of categories of data.

Claims (63)

1 . A method comprising:

receiving, by a device, a request related to a network resource, the request comprising information related to a topic;

determining, by the device, a taxonomy based on the request, the taxonomy comprising information related to the topic and the network resource, the taxonomy indicating an organization of topical content associated with the network resource;

providing, by the device, an input to a large language model (LLM) comprising text corresponding to the taxonomy;

executing, by the device, the LLM based on the taxonomy, the execution comprising the LLM retrieving information from the network resource according to the indicated organization;

generating, by the device, via the execution of the LLM, an LLM output, the LLM output comprising a dataset extracted from the network resource via analysis of the network resource by the LLM based on the input;

analyzing, by the device, the LLM output to determine whether a fine-tuned LLM input is required, the determination corresponding to whether the extracted dataset corresponds, at least to a threshold similarity value, to the taxonomy;

generating, by the device, in response to a determination that the extracted dataset does not correspond to the taxonomy by at least the threshold similarity value, another input, the other input comprising text corresponding to a modified taxonomy, the modified taxonomy indicating an updated organization of topical content associated with the network resource;

executing, by the device, the LLM based on the other input to generate another LLM output, the other LLM output comprising another dataset extracted from the network resource via analysis of the network resource by the LLM based on the other input; and

generating, by the device, a set of content items for the network resource based on the other LLM output.

2 . The method of claim 1 , further comprising analyzing the other LLM output to determine whether the other extracted dataset corresponds, at least to the threshold similarity value, to the modified taxonomy, wherein the generation of the set of content items is performed when the other extracted dataset satisfies the threshold similarity value.

3 . The method of claim 1 , further comprising:

providing an LLM prompt for the other input; and

generating the other input based on the LLM prompt.

4 . The method of claim 3 , wherein the LLM is provided recursive inputs until a satisfactory LLM output is generated.

5 . The method of claim 1 , further comprising:

analyzing the other LLM output; and

determining a set of content parameters, the set of content parameters comprising indicators as to a type and form of digital content.

6 . The method of claim 5 , wherein the generation of the set of content items is based on the set of content parameters.

7 . The method of claim 1 , further comprising:

analyzing the request via an executed artificial intelligence/machine learning model; and

determining information related to the request, the topic and a requestor, the determined information providing an intent of the request, wherein the determination of the taxonomy is based on the determined information related to the request.

8 . A device comprising:

a processor configured to:

receive a request related to a network resource, the request comprising information related to a topic;

determine a taxonomy based on the request, the taxonomy comprising information related to the topic and the network resource, the taxonomy indicating an organization of topical content associated with the network resource;

provide an input to a large language model (LLM) comprising text corresponding to the taxonomy;

execute the LLM based on the taxonomy, the execution comprising the LLM retrieving information from the network resource according to the indicated organization;

generate, via the execution of the LLM, an LLM output, the LLM output comprising a dataset extracted from the network resource via analysis of the network resource by the LLM based on the input;

analyze the LLM output to determine whether a fine-tuned LLM input is required, the determination corresponding to whether the extracted dataset corresponds, at least to a threshold similarity value, to the taxonomy;

generate, in response to a determination that the extracted dataset does not correspond to the taxonomy by at least the threshold similarity value, another input, the other input comprising text corresponding to a modified taxonomy, the modified taxonomy indicating an updated organization of topical content associated with the network resource;

execute the LLM based on the other input to generate another LLM output, the other LLM output comprising another dataset extracted from the network resource via analysis of the network resource by the LLM based on the other input; and

generate a set of content items for the network resource based on the other LLM output.

9 . The device of claim 8 , wherein the processor is further configured to analyze the other LLM output to determine whether the other extracted dataset corresponds, at least to the threshold similarity value, to the modified taxonomy, wherein the generation of the set of content items is performed when the other extracted dataset satisfies the threshold similarity value.

10 . The device of claim 8 , wherein the processor is further configured to:

provide an LLM prompt for the other input; and

generate the other input based on the LLM prompt.

11 . The device of claim 8 , wherein the processor is further configured to:

analyze the other LLM output; and

determine a set of content parameters, the set of content parameters comprising indicators as to a type and form of digital content, wherein the generation of the set of content items is based on the set of content parameters.

12 . The device of claim 8 , wherein the processor is further configured to:

analyze the request via an executed artificial intelligence/machine learning model; and

determine information related to the request, the topic and a requestor, the determined information providing an intent of the request, wherein the determination of the taxonomy is based on the determined information related to the request.

13 . A non-transitory computer-readable storage medium tangibly encoded with computer-executable instructions that when executed by a device, perform a method comprising:

receiving, by the device, a request related to a network resource, the request comprising information related to a topic;

determining, by the device, a taxonomy based on the request, the taxonomy comprising information related to the topic and the network resource, the taxonomy indicating an organization of topical content associated with the network resource;

providing, by the device, an input to a large language model (LLM) comprising text corresponding to the taxonomy;

executing, by the device, the LLM based on the taxonomy, the execution comprising the LLM retrieving information from the network resource according to the indicated organization;

generating, by the device, via the execution of the LLM, an LLM output, the LLM output comprising a dataset extracted from the network resource via analysis of the network resource by the LLM based on the input;

analyzing, by the device, the LLM output to determine whether a fine-tuned LLM input is required, the determination corresponding to whether the extracted dataset corresponds, at least to a threshold similarity value, to the taxonomy;

generating, by the device, in response to a determination that the extracted dataset does not correspond to the taxonomy by at least the threshold similarity value, another input, the other input comprising text corresponding to a modified taxonomy, the modified taxonomy indicating an updated organization of topical content associated with the network resource;

executing, by the device, the LLM based on the other input to generate another LLM output, the other LLM output comprising another dataset extracted from the network resource via analysis of the network resource by the LLM based on the other input; and

generating, by the device, a set of content items for the network resource based on the other LLM output.

14 . The non-transitory computer-readable storage medium of claim 13 , further comprising analyzing the other LLM output to determine whether the other extracted dataset corresponds, at least to the threshold similarity value, to the modified taxonomy, wherein the generation of the set of content items is performed when the other extracted dataset satisfies the threshold similarity value.

15 . The non-transitory computer-readable storage medium of claim 13 , further comprising:

providing an LLM prompt for the other input; and

generating the other input based on the LLM prompt.

16 . The non-transitory computer-readable storage medium of claim 13 , further comprising:

analyzing the other LLM output; and

determining a set of content parameters, the set of content parameters comprising indicators as to a type and form of digital content, wherein the generation of the set of content items is based on the set of content parameters.

17 . The non-transitory computer-readable storage medium of claim 13 , further comprising:

analyzing the request via an executed artificial intelligence/machine learning model; and

determining information related to the request, the topic and a requestor, the determined information providing an intent of the request, wherein the determination of the taxonomy is based on the determined information related to the request.