IP Library Granted Patent US 8,396,864
Granted Patent B1
US 8,396,864 · App. 11/478,843 · Granted Mar 12, 2013

Categorizing documents

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,396,864
App. No.
11/478,843
Granted
Mar 12, 2013
Kind
B1
Abstract

Categorizing documents is disclosed. A hierarchy of topics is received. A seed for each topic is determined. One or more documents is received. The seed is used to evaluate the relevance of each document to one or more of the received topics. One or more topics is associated with each document.

Claims (83)

1. A method comprising:

selecting a topic from a hierarchy of topics;

receiving, from a user, a seed set including one or more seed pages for the topic;

receiving a document that is not associated with the topic;

using the seed to determine, with a processor, a topic destination score and a topic source score for the document relative to the document, the topic destination score indicating an amount of content in the document relating to the topic and the topic source score indicating reachability of content relating to the topic through the document;

receiving a query; and

returning the document as a result for the query based at least in part on the topic source score and the topic destination score for the document;

wherein returning the document as a result for the query based at least in part on the topic source score and the topic destination score for the document further comprises

calculating a topic score according to a weighted average of the topic source score and topic destination score; and

returning the document as a result for the query based at least in part on the topic score.

2. The method of claim 1 wherein the document is received as a result of a search.

3. The method of claim 1 wherein the seed is one or more seed sites.

4. The method of claim 1 wherein the seed is one or more seed terms.

5. The method of claim 1 wherein the seed represents a flavor.

6. The method of claim 1 wherein the seed is dynamic.

7. The method of claim 1 wherein the received document includes a portion of content accessible via the World Wide Web.

8. The method of claim 1 wherein the received document is included in an arbitrary document collection.

9. The method of claim 1 wherein the received document is included among a plurality of documents located on an intranet.

10. The method of claim 1 further comprising calculating a topic score according to the topic source score and topic destination score and determining whether to display an ad on the document based at least in part on the topic score.

11. The method of claim 1 further comprising calculating a topic score according to the topic source score and topic destination score and ordering the document among a set of search results based at least in part on the topic score.

12. The method of claim 1 further comprising calculating a topic score according to the topic source score and topic destination score and populating a directory based at least in part on the topic score.

13. The method of claim 1 further comprising calculating a topic score according to the topic source score and topic destination score and partitioning a collection of documents based at least in part on the topic score.

14. The method of claim 1 further comprising calculating a topic score according to the topic source score and topic destination score and categorizing a blog based at least in part on the topic score.

15. The method of claim 1 further comprising calculating a topic score according to the topic source score and topic destination score and identifying a spam site based at least in part on the topic score.

16. The method of claim 1 further comprising calculating a topic score according to the topic source score and topic destination score and identifying a pornographic site based at least in part on the topic score.

17. A system comprising:

a processor; and

a memory coupled to the processor, wherein the memory is configured to provide the processor with executable instructions to:

receive a topic from a hierarchy of topics;

receiving, from a user, a seed set including one or more seed pages for the topic;

receive a document that is not associated with the topic;

use the seed set to determine a source score and a destination score, the destination score indicating an amount content relating to the topic in the document and the source score indicating reachability of content relating to the topic through the document;

receive a query; and

return the document as a result for the query based at least in part on the source score and destination score;

wherein the memory is further configured to provide the processor with executable instructions to return the document as a result for the query based at least in part on the topic source score and the topic destination score for the document by

calculating a topic score according to a weighted average of the topic source score and topic destination score; and

returning the document as a result for the query based at least in part on the topic score.

18. The system of claim 17 wherein the seed is dynamic.

19. A computer program product for categorizing documents, the computer program product tangibly embodied in a non-transitory computer readable storage medium and comprising computer executable instructions for:

selecting a topic from a hierarchy of topics;

receiving, from a user, a seed set including one or more seed pages for a topic included in the hierarchy of topics;

receiving a document that is not associated with the topic;

using the seed to determine a source score and a destination score, the destination score indicating an amount of content relating to the topic in the document and the source score indicating reachability of content relating to the topic through the document;

receiving a query; and

returning the document as a result for the query based at least in part on the source score and destination score;

wherein returning the document as a result for the query based at least in part on the topic source score and the topic destination score for the document further comprises

calculating a topic score according to a weighted average of the topic source score and topic destination score; and

returning the document as a result for the query based at least in part on the topic score.

20. A method comprising:

selecting a topic from a hierarchy of topics;

receiving, from a user, a seed set including one or more seed pages for the topic;

receiving a document that is not associated with the topic;

using the seed to determine, with a processor, a topic destination score and a topic source score for the document relative to the document, the topic destination score indicating an amount of content in the document relating to the topic and the topic source score indicating reachability of content relating to the topic through the document;

receiving a query; and

returning the document as a result for the query based at least in part on the topic source score and the topic destination score for the document;

wherein the document is part of a document set; and

wherein using the seed to determine the topic source score and topic destination score for the document further comprises;

determining a destination score according to a probability of arrival of a random surfer at the document during a random walk biased according to the seed set using a random surfer model; and

determining a source score for the document according to a contribution of the document to destination scores of other documents in the document set.

21. A system comprising:

a processor; and

a memory coupled to the processor, wherein the memory is configured to provide the processor with executable instructions to:

receive a topic from a hierarchy of topics;

receiving, from a user, a seed set including one or more seed pages for the topic;

receive a document that is not associated with the topic;

use the seed set to determine a source score and a destination score, the destination score indicating an amount of content relating to the topic in the document and the source score indicating the likelihood of reachability of content relating to the topic through the document;

receive a query; and

return the document as a result for the query based at least in part on the source score and destination score

wherein the document is part of a document set; and

wherein the memory is further configured to provide the processor with executable instructions to use the seed to determine a source score and a destination score by:

determining a destination score according to a probability of arrival of a random surfer at the document during a random walk biased according to the seed set using a random surfer model; and

determining a source score for the document according to a contribution of the document to destination scores of other documents in the document set.

22. A computer program product for categorizing documents, the computer program product tangibly embodied in a non-transitory computer readable storage medium and comprising computer executable instructions for:

selecting a topic from a hierarchy of topics;

receiving, from a user, a seed set including one or more seed pages for a topic included in the hierarchy of topics;

receiving a document that is not associated with the topic;

using the seed to determine a source score and a destination score, the destination score indicating an amount of content relating to the topic in the document and the source score indicating reachability of content relating to the topic through the document;

receiving a query; and

returning the document as a result for the query based at least in part on the source score and destination score

wherein the document is part of a document set; and

wherein the computer program product further comprises computer executable instructions for:

determining a destination score according to a probability of arrival of a random surfer at the document during a random walk biased according to the seed set using a random surfer model; and

determining a source score for the document according to a contribution of the document to destination scores of other documents in the document set.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2018
From: WAL-MART STORES, INC.
To: WALMART APOLLO, LLC
Reel/Frame 045817/0115 →
MERGER Recorded Apr 19, 2012
From: KOSMIX CORPORATION
To: WAL-MART STORES, INC.
Reel/Frame 028074/0001 →
MERGER Recorded Aug 14, 2008
From: COSMIX CORPORATION
To: KOSMIX CORPORATION
Reel/Frame 021391/0753 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2006
From: HARINARAYAN, VENKY; RAJARAMAN, ANAND
To: COSMIX CORPORATION
Reel/Frame 018309/0415 →