IP Library Granted Patent US 12,572,555
Granted Patent B2
US 12,572,555 · App. 17/981,746 · Granted Mar 10, 2026

Method and system for data mining

Inventors: Sainath Vellal (Sunnyvale, CA); Kostas Tsioutsiouliklis (Sunnyvale, CA)
Assignee: YAHOO ASSETS LLC
G06F16/2465G06F16/24578G06F16/9024
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,555
App. No.
17/981,746
Granted
Mar 10, 2026
Kind
B2
Abstract

The present teaching relates to method and system for generating a stream of content items. A plurality of entities associated with a time-window are obtained, wherein each entity of the plurality of entities is associated with at least one content item. For each entity, a first parameter with respect to the time-window, a second parameter with respect to previous time-windows, and a trendiness score based on a function of the first parameter and the second parameter are respectively calculated. A graph based on one or more entity-pairs is generated, wherein each entity-pair of the one or more entity-pairs satisfies a first criterion. A stream of content items is generated based on the graph, wherein each content item in the stream of content items corresponds to at least one of the one or more entity pairs.

Claims (49)

1 . A method, implemented on a machine having at least one processor, storage, and a communication platform capable of connecting to a network for generating a stream of web-based content items, the method comprising:

crawling, by an engine, in accordance with a configuration model controlling a number of hyperlinks crawled from a webpage, web-based published content items;

extracting, by the engine, in accordance with a machine learning model, a plurality of entities from each of the web-based published content items, wherein each entity is within a corresponding time window;

determining, for each of the published content items, a corresponding category from a plurality of categories;

appending, via a web-application embedded in a webpage, each of the published content items with metadata in the web-application, wherein the metadata includes one or more of: a corresponding publisher, the corresponding entities, the corresponding category, and a count of the corresponding entities;

clustering the appended content items based on the categories to generate one or more clusters;

storing the time windows, the extracted entities, and the clustered content items in a repository;

obtaining, from the repository, entities stored in their corresponding time-windows to calculate, for each of the extracted entities, a corresponding trendiness score as a function including a numerical difference between a number of occurrences of the entity in the time window and an average number of occurrences of the entity in previous time windows;

determining one or more entity-pairs with respect to the extracted entities, wherein a co-occurrence count of entities included in each of the one or more entity-pairs exceeds a threshold;

ranking stored content items in each of the one or more clusters based on an average trendiness score of entities included in each of the one or more entity-pairs in the cluster;

generating, based on a query for content items that are trending, a stream of web-based content items by selecting a highest ranked content item from each of the one or more clusters to be included in the stream; and

scaling up the stream of web-based content items across the plurality of categories.

2 . The method of claim 1 , wherein the clustering the appended content items is in accordance with a clustering model.

3 . The method of claim 1 , wherein the metadata further comprises a publish time of the corresponding content item.

4 . The method of claim 3 , wherein the generating the stream of content items is further based on the publish time.

5 . The method of claim 1 , wherein the plurality of categories comprise at least one of finance, politics, sports, and entertainment.

6 . A non-transitory machine-readable medium having information recorded thereon for generating a stream of web-based content items, wherein the information, when read by a machine, causes the machine to perform operations comprising:

crawling, by an engine, in accordance with a configuration model controlling a number of hyperlinks crawled from a webpage, web-based published content items;

extracting, by the engine, in accordance with a machine learning model, a plurality of entities from each of the web-based published content items, wherein each entity is within a corresponding time window;

determining for each of the published content items, a corresponding category from a plurality of categories;

appending, via a web-application embedded in a webpage, each of the published content items with metadata in the web-application, wherein the metadata includes one or more of: a corresponding publisher, the corresponding entities, the corresponding category, and a count of the corresponding entities;

clustering the appended content items based on the categories to generate one or more clusters;

storing the time windows, the extracted entities, and the clustered content items in a repository;

obtaining, from the repository, entities stored in their corresponding time-windows to calculate, for each of the extracted entities, a corresponding trendiness score as a function including a numerical difference between a number of occurrences of the entity in the time window and an average number of occurrences of the entity in previous time windows;

determining one or more entity-pairs with respect to the extracted entities, wherein a co-occurrence count of entities included in each of the one or more entity-pairs exceeds a threshold;

ranking stored content items in each of the one or more clusters based on an average trendiness score of entities included in each of the one or more entity-pairs in the cluster;

generating, based on a query for content items that are trending, a stream of web-based content items by selecting a highest ranked content item from each of the one or more clusters to be included in the stream; and

scaling up the stream of web-based content items across the plurality of categories.

7 . The medium of claim 6 , wherein the clustering the appended content items is in accordance with a clustering model.

8 . The medium of claim 6 , wherein the metadata further comprises a publish time of the corresponding content item.

9 . The medium of claim 8 , wherein the generating the stream of content items is further based on the publish time.

10 . The medium of claim 6 , wherein the plurality of categories comprise at least one of finance, politics, sports, and entertainment.

11 . A system for generating a stream of web-based content items, the system comprising:

memory storing computer program instructions; and

one or more processors that, in response to executing the computer program instructions, effectuate operations comprising:

crawling, by an engine, in accordance with a configuration model controlling a number of hyperlinks crawled from a webpage, web-based published content items;

extracting, by the engine, in accordance with a machine learning model, a plurality of entities from each of the web-based published content items, wherein each entity is within a corresponding time window;

determining, for each of the published content items, a corresponding category from a plurality of categories;

appending, via a web-application embedded in a webpage, each of the published content items with metadata in the web-application, wherein the metadata includes one or more of: a corresponding publisher, the corresponding entities, the corresponding category, and a count of the corresponding entities;

clustering the appended content items based on the categories to generate one or more clusters;

storing the time windows, the extracted entities, and the clustered content items in a repository;

obtaining, from the repository, entities stored in their corresponding time-windows to calculate, for each of the extracted entities, a corresponding trendiness score as a function including a numerical difference between a number of occurrences of the entity in the time window and an average number of occurrences of the entity in previous time windows;

determining one or more entity-pairs with respect to the extracted entities, wherein a co-occurrence count of entities included in each of the one or more entity-pairs exceeds a threshold;

ranking stored content items in each of the one or more clusters based on an average trendiness score of entities included in each of the one or more entity-pairs in the cluster;

generating, based on a query for content items that are trending, a stream of web-based content items by selecting a highest ranked content item from each of the one or more clusters to be included in the stream; and

scaling up the stream of web-based content items across the plurality of categories.

12 . The system of claim 11 , wherein the clustering the appended content items is in accordance with a clustering model.

13 . The system of claim 11 , wherein the metadata further comprises a publish time of the corresponding content item.

14 . The system of claim 13 , wherein the generating the stream of content items is further based on the publish time.

Assignments (4)
SUPPLEMENTAL PATENT SECURITY AGREEMENT Recorded Sep 17, 2025
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 072915/0540 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2022
From: VELLAL, SAINATH; TSIOUTSIOULIKLIS, KOSTAS
To: OATH INC.
Reel/Frame 061674/0504 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2022
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 061887/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2022
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 061887/0350 →
Continuity (3)
Continuation 16597173 · Oct 9, 2019
Provisional Application 62890407 · Aug 22, 2019
Related Publication 20230066149A1 · Mar 2, 2023
References Cited (20)
US 7827170B1 · Horling · 2010 [cited by examiner]
US 20070100875A1 · Chi · 2007 [cited by examiner]
US 20080033587A1 · Kurita · 2008 [cited by examiner]
US 20080140471A1 · Ramsey · 2008 [cited by examiner]
US 20100100537A1 · Druzgalski et al. · 2010 [cited by applicant]
US 20110179017A1 · Meyers · 2011 [cited by examiner]
US 20110282874A1 · Xu et al. · 2011 [cited by applicant]
US 20110314007A1 · Dassa et al. · 2011 [cited by applicant]
US 20110320715A1 · Ickman et al. · 2011 [cited by applicant]
US 20140040387A1 · Spivack · 2014 [cited by examiner]
US 20140283048A1 · Howes et al. · 2014 [cited by applicant]
US 20140337451A1 · Choudhary et al. · 2014 [cited by applicant]
US 20150154278A1 · Allen et al. · 2015 [cited by applicant]
US 20150220946A1 · Horesh et al. · 2015 [cited by applicant]
US 20160359993A1 · Hendrickson et al. · 2016 [cited by applicant]
US 20180349352A1 · Mabbu · 2018 [cited by applicant]
US 20200159787A1 · Chernenkov · 2020 [cited by examiner]
US 20200267112A1 · Jain et al. · 2020 [cited by applicant]
“Definition: web application (web app).” Published Jan. 2023 by TechTarget. Accessed Oct. 17, 2024 from https://www.techtarget.com/searchsoftwarequality/definition/Web-application-Web-app (Year: 2023). [cited by examiner]
“Ontology”, Computer Desktop Encyclopedia, the Computer Language Company Inc. 2021. Retrieved on Sep. 15, 2021 from https ://www.computerlanguage.com/results .php?definition=ontology. [cited by applicant]