IP Library Granted Patent US 8,285,701
Granted Patent B2
US 8,285,701 · App. 09/920,615 · Granted Oct 9, 2012

Video and digital multimedia aggregator remote content crawler

Assignee: Comcast IP Holdings I, LLC
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,285,701
App. No.
09/920,615
Granted
Oct 9, 2012
Kind
B2
Abstract

A remote content crawler continually crawls a digital communication network looking for content to provide to a content aggregator. The content provided to the aggregator may be stored in a form of an entire content file. The content may include an entire movie, television program or electronic book. Alternatively, the content provided to the aggregator may be a reference to a content file that is stored at, or that will be available at one of the remote locations. The content may be a reference to a future, scheduled live sports event that will be made available to system users. The sports event may be provided for a one time fee, as part of a sports package, for which a fee is collected, or as a free event.

Claims (88)

1. An apparatus, comprising:

a remote content crawler processor that controls the apparatus;

a network resource processor that acquires data related to resources coupled to one or more communications networks;

a crawling criteria processor that acquires crawling criteria, said crawling criteria comprising a plurality of conditions, said plurality of conditions comprising content subject matter, content type, and a delivery method, said acquiring of crawling criteria comprising analyzing hypertext associated with desired hyperlinks and analyzing text proximate to the desired hyperlinks;

a crawler content provider processor that receives, processes and stores content provider listings; and

a network crawler comprising at least one server, wherein the network crawler crawls content providers to acquire data related to available content in accordance with the crawling criteria,

wherein said remote content crawler processor determines that the acquired data related to the resources coupled to one or more communications networks meets said plurality of conditions of said crawling criteria.

2. The apparatus of claim 1 , further comprising:

a content crawler results processor;

a metadata acquisition processor; and

one or more databases, the one or more databases storing information and data generated in and received by the remote content crawler.

3. The apparatus of claim 2 , wherein the one or more databases, comprise:

a content provider listing database;

a crawling criteria database; and

a network resources database.

4. A method comprising:

acquiring network resource data, wherein the network resource data comprises address data for content servers;

acquiring crawling criteria, wherein the crawling criteria comprises a plurality of conditions and are used during a crawling operation to search for content, said plurality of conditions comprising content subject matter, content type, and a delivery method;

crawling network resources via at least one processor in accordance with the crawling criteria;

determining if a content provider data meets said plurality of conditions of said crawling criteria;

providing said content provider data to a user terminal; and

receiving content provider rankings from users of a network,

wherein acquiring the crawling criteria comprises automatically acquiring the crawling criteria, and

wherein automatically acquiring the crawling criteria, comprises:

analyzing and importing metadata schemes for standardized and proprietary content formats;

parsing metadata field names and descriptive terms;

analyzing hypertext associated with desired hyperlinks; and

analyzing text proximate to the desired hyperlinks, wherein analyzing hypertext identifies terms that relate to a data type or content category.

5. The method of claim 4 , further comprising storing the network resource data, the crawling criteria, and the content provider data in one or more databases.

6. The method of claim 4 , wherein acquiring network resource data comprises indexing the address data according to one or more address types.

7. The method of claim 6 , wherein the address types include top-level domain and subdomain names, Universal Resource Identifiers, Universal Resource Locators (URLs), and Internet Protocol (IP) address numbers.

8. The method of claim 4 , further comprising updating the address data.

9. The method of claim 8 , wherein updating the address data, comprises:

receiving hyperlinked domain names for the network resources;

downloading domain name records from public and private domain name registration sources;

synchronizing local Domain Name Service (DNS) databases with one or more DNS databases over one or more communications networks;

performing reverse domain name resolution, comprising locating URLs associated with allowable IP address numbers;

verifying DNS aliases and duplicate URLs against IP addresses; and

eliminating any duplicate URLs identified by the verifying step.

10. The method of claim 4 , wherein the network resource data comprises:

a URL owner identity;

a URL owner contact information;

available content types;

an expiration time of a domain name; and

subdomain names to be excluded during crawling.

11. The method of claim 4 , wherein the crawling criteria, comprises:

terms, phrases and keywords;

data type descriptions;

metadata field names; and

metadata type descriptors, wherein the metadata type descriptors are associated with eligible content as one or more of hypertext descriptions and embedded file and data stream attributes and metadata.

12. The method of claim 4 , wherein acquiring the crawling criteria comprises acquiring the crawling criteria through manual input.

13. The method of claim 4 , wherein the ranking of the content providers is based on one or more of quantity of available content, provider professional association membership, and content provider ratings.

14. The method of claim 13 , further comprising determining a frequency of crawling a content provider based on the ranking of the content provider.

15. The method of claim 4 , wherein crawling the network resources comprises crawling with one or more crawling servers.

16. The method of claim 15 , further comprising:

subdividing the network resources;

assigning the subdivided network resources to the one or more crawling servers; and

at a crawler server:

reading data from the assigned network resources,

communicating with the assigned network resources,

downloading data from the assigned network resources.

17. The method of claim 16 , further comprising:

following links from a first network resource to subsequent network resources, wherein following the links comprises:

analyzing hypertext structure of the first network resource to determine if the links have been crawled,

determining if a network resource has been downloaded or updated since a previous crawl of the network resource, and

analyzing the hypertext structure to determine if the link points to a network resource comprising a web page or other hypertext file.

18. The method of claim 17 , further comprising:

caching hypertext files containing the data related to the content;

caching the links from the first network resource to subsequent network resources; and

indexing web pages or other hypertext files of interest.

19. The method of claim 17 , wherein comparing the content to the crawling criteria comprises using a comparison algorithm that compares elements in a hypertext file to the crawling criteria.

20. The method of claim 4 , further comprising:

acquiring and processing metadata related to a network resource included in the network resources; and

processing content results from the network resource.

21. A method comprising:

acquiring network resource data, wherein the network resource data comprises address data for content servers;

acquiring crawling criteria, wherein the crawling criteria comprises a plurality of conditions and are used during a crawling operation to search for content, said plurality of conditions comprising content subject matter, content type, and a delivery method;

crawling network resources via at least one processor in accordance with the crawling criteria;

determining if a content provider data meets said plurality of conditions of said crawling criteria;

providing said content provider data to a user terminal;

receiving content provider rankings from users of a network; and

updating the address data, wherein updating the address data, comprises:

receiving hyperlinked domain names for the network resources;

downloading domain name records from public and private domain name registration sources;

synchronizing local Domain Name Service (DNS) databases with one or more DNS databases over one or more communications networks;

performing reverse domain name resolution, comprising locating URLs associated with allowable IP address numbers;

verifying DNS aliases and duplicate URLs against IP addresses; and

eliminating any duplicate URLs identified by the verifying step.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2008
From: SEDNA PATENT SERVICES, LLC (F/K/A TVGATEWAY, LLC)
To: COMCAST IP HOLDINGS I, LLC
Reel/Frame 021570/0353 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2004
From: DISCOVERY COMMUNICATIONS, INC.
To: SEDNA PATENT SERVICES, LLC
Reel/Frame 015239/0350 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2001
From: SWART, WILLIAM D.; ASMUSSEN, MICHAEL L.; MCCOSKEY, JOHN S.
To: DISCOVERY COMMUNICATIONS, INC.
Reel/Frame 012353/0042 →
Continuity (1)
Related Publication 20030028896A1 · Feb 6, 2003