IP Library › Granted Patent US 7,440,938
Granted Patent B2
US 7,440,938 · App. 10/838,231 · Granted Oct 21, 2008

Method and apparatus for calculating similarity among documents

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,440,938
App. No.
10/838,231
Granted
Oct 21, 2008
Kind
B2
Abstract

Information that individual elements (characteristic character strings) indicative of characteristics of a registered document appear in the registered document is stored in advance. When calculating similarity of the registered document, a query designated by a searcher is analyzed. The query is represented by a characteristic vector having the individual elements which take the relation between a plurality of words into consideration. Pieces of appearance information of the individual words contained in the query are counted. The counted appearance information is compared with a searching index to calculate similarity between documents.

Claims (59)

1. A similarity calculation method for calculating similarity among documents in a document search system adapted to search documents registered in advance, comprising:

storing, in a storage as a searching index, each of keywords contained in a previously registered document;

extracting each of keywords contained in a query inputted by a searcher;

deciding an attribute of each of said keywords by referring to an element-type dictionary which provides correspondence relationships between keywords and attributes;

classifying said extracted keywords on an attribute basis in accordance with a result of decision and counting a number of keywords belonging to each of the attributes thus classified to determine a count number of each of the attributes;

comparing said count numbers of attributes thus determined with count numbers of corresponding attributes of said previously registered document;

calculating similarity between said query and said previously registered document on a basis of comparison of said count numbers; and

wherein the searching index stores character strings with respect to a plurality of previously registered documents;

wherein at least one keyword of the keywords contained in the query is decided as a first group of the keywords, and at least one different keyword of the keywords is decided as a second group of the keywords;

wherein the deciding, classifying, counting, comparing and calculating operations are performed for the first group of the keywords to determine a first subset of the previously registered documents having similarity to the at least one keyword, and are separately performed for the second group of the keywords to determine a second subset of the previously registered documents having similarity to the at least one different keyword;

the similarity calculation method further comprising:

determining any previously registered documents which are common to both the first and second subsets to narrow-down to documents having similarity to said query.

2. A computer-readable medium having computer-readable code therein, which, when extended by a computer, causes the computer to implement a similarity calculation method for calculating similarity among documents in a document search system adapted to search documents registered in advance, the similarity calculation method comprising:

storing, in a storage as a searching index, each of keywords contained in a previously registered document;

extracting each of keywords contained in a query inputted by a searcher;

deciding an attribute of each of said keywords by referring to an element-type dictionary which provides correspondence relationships between keywords and attributes;

classifying said extracted keywords on an attribute basis in accordance with a result of decision and counting a number of keywords belonging to each of the attributes thus classified to determine a count number of each of the attributes;

comparing said count numbers of attributes thus determined with count numbers of corresponding attributes of said previously registered document;

calculating similarity between said query and said previously registered document on a basis of comparison of said count numbers; and

wherein the searching index stores character strings with respect to a plurality of previously registered documents;

wherein at least one keyword of the keywords contained in the query is decided as a first group of the keywords, and at least one different keyword of the keywords is decided as a second group of the keywords;

wherein the deciding, classifying, counting, comparing and calculating operations are performed for the first group of the keywords to determine a first subset of the previously registered documents having similarity to the at least one keyword, and are separately performed for the second group of the keywords to determine a second subset of the previously registered documents having similarity to the at least one different keyword;

the similarity calculation method further comprising:

determining any previously registered documents which are common to both the first and second subsets to narrow-down to documents having similarity to said query.

3. A document search system adapted to search documents registered in advance, comprising:

a processor adapted with software and supportive hardware, for calculating similarity among documents, the processor adapted to effect operations comprising:

storing, in a storage as a searching index, each of keywords contained in a previously registered document;

extracting each of keywords contained in a query inputted by a searcher;

deciding an attribute of each of said keywords by referring to an element-type dictionary which provides correspondence relationships between keywords and attributes;

classifying said extracted keywords on an attribute basis in accordance with a result of decision and counting a number of keywords belonging to each of the attributes thus classified to determine a count number of each of the attributes;

comparing said count numbers of attributes thus determined with count numbers of corresponding attributes of said previously registered document;

calculating similarity between said query and said previously registered document on a basis of comparison of said count numbers; and

wherein the searching index stores character strings with respect to a plurality of previously registered documents;

wherein at least one keyword of the keywords contained in the query is decided as a first group of the keywords, and at least one different keyword of the keywords is decided as a second group of the keywords;

wherein the deciding, classifying, counting, comparing and calculating operations are performed for the first group of the keywords to determine a first subset of the previously registered documents having similarity to the at least one keyword, and are separately performed for the second group of the keywords to determine a second subset of the previously registered documents having similarity to the at least one different keyword;

the processor adapted to effect further operations comprising:

determining any previously registered documents which are common to both the first and second subsets to narrow-down to documents having similarity to said query.

4. A similarity calculation method for calculating similarity among documents in a document search system adapted to search documents registered in advance, comprising:

storing, in a storage as a searching index, each of keywords contained in a previously registered document;

extracting each of keywords contained in a query inputted by a searcher;

deciding an attribute of each of said keywords by referring to an element-type dictionary which provides correspondence relationships between keywords and attributes;

classifying said extracted keywords on an attribute basis in accordance with a result of decision and counting a number of keywords belonging to each of the attributes thus classified to determine a count number of each of the attributes;

comparing said count numbers of attributes thus determined with count numbers of corresponding attributes of said previously registered document; and

calculating similarity between said query and said previously registered document on a basis of comparison of said count numbers.

5. A computer-readable medium having computer-readable code therein, which, when executed by a computer, causes the computer to implement a similarity calculation method for calculating similarity among documents in a document search system adapted to search documents registered in advance, the similarity calculation method comprising:

storing, in a storage as a searching index, each of keywords contained in a previously registered document;

extracting each of keywords contained in a query inputted by a searcher;

deciding an attribute of each of said keywords by referring to an element-type dictionary which provides correspondence relationships between keywords and attributes;

classifying said extracted keywords on an attribute basis in accordance with a result of decision and counting a number of keywords belonging to each of the attributes thus classified to determine a count number of each of the attributes;

comparing said count numbers of attributes thus determined with count numbers of corresponding attributes of said previously registered document; and

calculating similarity between said query and said previously registered document on a basis of comparison of said count numbers.

6. A document search system adapted to search documents registered in advance, comprising:

a processor adapted with software and supportive hardware, for calculating similarity among documents, the processor adapted to effect operations comprising:

storing, in a storage as a searching index, each of keywords contained in a previously registered document;

extracting each of keywords contained in a query inputted by a searcher;

deciding an attribute of each of said keywords by referring to an element-type dictionary which provides correspondence relationships between keywords and attributes;

classifying said extracted keywords on an attribute basis in accordance with a result of decision and counting a number of keywords belonging to each of the attributes thus classified to determine a count number of each of the attributes;

comparing said count numbers of attributes thus determined with count numbers of corresponding attributes of said previously registered document; and

calculating similarity between said query and said previously registered document on a basis of comparison of said count numbers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2004
From: MATSUBAYASHI, TADATAKA; SUGAYA, NATSUKO; IIJIMA, MICHIO; OGAWA, YUICHI; WATANABE, YUUKI; YAMAMOTO, SHINYA; SUDOU, TSUYOSHI
To: HITACHI, LTD.
Reel/Frame 015662/0431 →
Priority Claims (1)
JP 2003-200193 · Jul 23, 2003 · national
Continuity (1)
Related Publication 20050021508A1 · Jan 27, 2005