IP Library › Granted Patent US 7,496,581
Granted Patent B2
US 7,496,581 · App. 10/621,474 · Granted Feb 24, 2009

Information search system, information search method, HTML document structure analyzing method, and program product

Assignee: International Business Machines Corporation
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,496,581
App. No.
10/621,474
Granted
Feb 24, 2009
Kind
B2
Abstract

In an information search using a computer, a flexible information search based on a variety of strategies that depend on a purpose of use of information is effectively realized. An information search system comprises a document structure analyzing section for analyzing a structure of an HTML document taking into account a meaning in a prescribed web page, a significance calculating section for calculating the degrees of significance of other web sites linking from the web page, based on a result of the analysis and according to predetermined strategies, and a crawling executing section for crawling the web sites depending on the degrees of significance calculated by the significance calculating section.

Claims (36)

1. A method comprising:

reading an HTML document of a web page as an analyzing object;

conducting a temporary block analysis based on a description of HTML tags of the HTML document;

using the HTML tags to temporarily divide the HTML document into blocks;

identifying unnecessary information elements in the HTML document, wherein the unnecessary information elements include:

plural information elements that include an OBJECT_IMAGE having a same Uniform Resource Locator (URL), wherein the OBJECT_IMAGE describes a type of media used to display the HTML document,

a block of text in the HTML document that is shorter than a maximum predetermined length, and wherein the block of text appears in the HTML document more than a predetermined frequency,

multiple anchors having a same title,

image tags that only perform a role of punctuation for text in the HTML document, and

multiple text blocks having a same description;

defining any block in the HTML document that is deemed to be meaningless as an OBJECT_DELIMITER, wherein a block is deemed to be meaningless if that block contains only said unnecessary information elements and at least one anchor; and

crawling only anchors found in blocks that have not been defined as OBJECT_DELIMITERs.

2. The method of claim 1 , wherein the maximum predetermined length is 12 bytes.

3. The method of claim 2 , wherein the predetermined frequency is ten times.

4. A computer-readable medium encoded with a computer program, wherein the computer program, when executed, performs the steps of:

reading an HTML document of a web page as an analyzing object;

conducting a temporary block analysis based on a description of HTML tags of the HTML document;

using the HTML tags to temporarily divide the HTML document into blocks;

identifying unnecessary information elements in the HTML document, wherein the unnecessary information elements include:

plural information elements that include an OBJECT_IMAGE having a same Uniform Resource Locator (URL), wherein the OBJECT_IMAGE describes a type of media used to display the HTML document,

a block of text in the HTML document that is shorter than a maximum predetermined length, and wherein the block of text appears in the HTML document more than a predetermined frequency,

multiple anchors having a same title,

image tags that perform a role of punctuation for text in the HTML document, and multiple text blocks having a same description;

defining any block in the HTML document that is deemed to be meaningless as an OBJECT_DELIMITER, wherein a block is deemed to be meaningless if that block contains only said unnecessary information elements; and

crawling only anchors found in blocks that have not been defined as OBJECT_DELIMITERs.

5. The computer-readable medium of claim 4 , wherein, the maximum predetermined length is 12 bytes.

6. The computer-readable medium of claim 4 , wherein the predetermined frequency is ten times.

7. A method comprising:

dividing an HTML document into blocks;

identifying unnecessary information elements in the HTML document, wherein the unnecessary information elements include:

a block of text in the HTML document that is shorter than a maximum predetermined length, and wherein the block of text appears in the HTML document more than a predetermined frequency,

multiple anchors having a same title,

image tags that only perform a role of punctuation for text in the HTML document, and

multiple text blocks having a same description;

defining any block in the HTML document that is deemed to be meaningless, wherein a block is deemed to be meaningless if that block contains only the unnecessary information elements and at least one anchor; and

crawling only anchors found in blocks that have not been deemed meaningless due to containing only the unnecessary information elements.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 27, 2003
From: NOMIYAMA, HIROSHI; IWAO, TOSHITAKA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 014642/0595 →
Priority Claims (1)
JP 2002-211634 · Jul 19, 2002 · national
Continuity (1)
Related Publication 20040054654A1 · Mar 18, 2004