IP Library Granted Patent US 8,560,548
Granted Patent B2
US 8,560,548 · App. 12/544,090 · Granted Oct 15, 2013

System, method, and apparatus for multidimensional exploration of content items in a content store

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,560,548
App. No.
12/544,090
Granted
Oct 15, 2013
Kind
B2
Abstract

A computer-implemented method for accessing content items in a content store are described. In one embodiment, the computer-implemented method includes maintaining a text index of content items in a content store to enable a keyword search on the content items, receiving a query having a keyword and generating a hit list from the text index using the keyword, and extracting frequent phrases from text within content items of the hit list. The computer-implemented method also includes assigning a relative relevance to the frequent phrases and grouping content items into topics based on presence of relevant phrases within the content items of the hit list. The hit list includes one or more content items of the content store. The frequent phrases having a relatively high relevance are relevant phrases.

Claims (44)

1. A computer-implemented method for accessing content items in a content store comprising:

maintaining a text index of content items in a content store to enable a keyword search on the content items;

receiving a query having a keyword and generating a hit list from the text index using the keyword, the hit list comprising two or more content items of the content store;

extracting frequent phrases from text within content items of the hit list by estimating an intersection size, wherein the estimating comprises executing an algorithm to intersect a first posting list generated from globally frequent phrases with a second posting list generated from the hit list, wherein the algorithm terminates the executing in response to the earlier of identifying a predetermined M maximum number of comparisons or a predetermined I maximum number of common points;

assigning a relative relevance to the frequent phrases wherein frequent phrases having a relatively high relevance are relevant phrases; and

grouping content items into topics based on presence of relevant phrases within the content items of the hit list.

2. The computer-implemented method of claim 1 , further comprising ranking content items of the hit list by importance.

3. The computer-implemented method of claim 2 wherein importance is based on associations with other content items in the content store.

4. The computer-implemented method of claim 3 , wherein the associations are citations to one or more different content items.

5. The computer-implemented method of claim 2 wherein grouping content items into topics is further based on importance.

6. The computer-implemented method of claim 1 , further comprising:

maintaining a content model for metadata of the content items, wherein the metadata includes links between two or more content items;

using metadata of the content items in the hit list to populate static dimensions of a multidimensional schema;

analyzing content items of the hit list to determine dynamic dimensions; and

providing, based on multidimensional schema, OLAP-style analysis of the content items of the hit list, including at least one analysis selected from the group consisting of aggregation, navigation and reporting.

7. The computer-implemented method of claim 1 , further comprising:

preprocessing the content items to enrich the text index, wherein the preprocessing comprises generating a list of globally frequent phrases, wherein the globally frequent phrases are phrases that consist of at least two words that appear in more than a first threshold number of different content items in the content store; and

adding the globally frequent phrases to the text index.

8. The computer-implemented method of claim 1 , further comprising:

maintaining a priority queue of the top-k most frequent phrases while intersecting the hit list with globally frequent phrases in descending order based upon length of the globally frequent phrases;

terminating the intersecting in response to identifying a globally frequent phrase with a length less than a current minimum intersection size in the priority queue; and

identifying phrases in the queue as the top-k dynamically frequent phrases.

9. The computer-implemented method of claim 1 , further comprising:

maintaining a priority queue of the top-k most frequent phrases while intersecting the hit list with the globally frequent phrases, wherein the estimating comprises:

randomizing the globally frequent phrases in the text index to produce the first randomized posting list; and

randomizing the hit list to produce the second randomized posting list, wherein the algorithm is a modified zipper algorithm.

10. The computer-implemented method of claim 9 , wherein the estimating further comprises multiplying a Jaccard Distance estimator with a union estimator to produce an estimate for the intersection of the first randomized posting list and the second randomized posting list.

11. The computer-implemented method of claim 1 , further comprising extracting labels for topics based on relevant phrases within the content items of the hit list.

12. The computer-implemented method of claim 1 , further comprising:

receiving an input selecting a topic;

extracting frequent phrases from text within content items of the topic;

assigning a relative relevance to the frequent phrases wherein frequent phrases having a relatively high relevance are relevant phrases;

grouping content items into subtopics based on presence of relevant phrases within the content items of the topic.

13. A system comprising:

a content management system (CMS) comprising:

a plurality of content items;

a content store to store content items;

a text index of text within the content items; and

an exploration server coupled to the content store and the text index and configured to:

generate a hit list from the text index using the keyword, the hit list comprising two or more content items of the content store;

extract frequent phrases from text within content items of the hit list by estimating an intersection size, wherein the estimating comprises executing an algorithm to intersect a first posting list generated from globally frequent phrases with a second posting list generated from the hit list, wherein the algorithm terminates the executing in response to the earlier of identifying a predetermined M maximum number of comparisons or a predetermined I maximum number of common points;

assign a relative relevance to the frequent phrases wherein frequent phrases having a relatively high relevance are relevant phrases; and

group content items into topics based on presence of relevant phrases within the content items of the hit list; and

a multidimensional schema manager to manage a multidimensional schema comprising a schema of a fact table and schemata of static dimensions, wherein dynamic dimensions of the multidimensional schema are populated in response to the exploration server returning the hit list, and wherein dynamic dimensions of the multidimensional schema are identified in response to the exploration server returning the hit list based upon content dynamically extracted from the subset of content items identified in the hit list.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2024
From: BEIJING PIANRUOJINGHONG TECHNOLOGY CO., LTD.
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 066565/0952 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2023
From: AWEMANE LTD.
To: BEIJING PIANRUOJINGHONG TECHNOLOGY CO., LTD.
Reel/Frame 064501/0498 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2021
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: AWEMANE LTD.
Reel/Frame 057991/0960 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2009
From: BAID, AKANKSHA; REINWALD, BERTHOLD; SIMITSIS, ALKIS; SISMANIS, JOHN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 023118/0920 →