IP Library Granted Patent US 8,041,718
Granted Patent B2
US 8,041,718 · App. 11/673,252 · Granted Oct 18, 2011

Processing apparatus and associated methodology for keyword extraction and matching

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,041,718
App. No.
11/673,252
Granted
Oct 18, 2011
Kind
B2
Abstract

An information processing apparatus includes an acquisition unit acquiring keywords extracted from text data representing a first content to be a base of a search and scores of the respective keywords, and keywords extracted from text data representing a second content for calculating a degree of matching with the first content, and scores of the respective keywords, a matching-degree calculation unit calculating the degree of matching between the first content and the second content based on scores of keywords commonly included in the acquired keywords relating to the first content and the acquired keywords relating to the second content, and an output unit outputting, as a search result, information on a predetermined number of the second content which has a high degree of matching with the first content based on a result of calculation performed by the matching-degree calculation unit.

Claims (45)

1. An information processing apparatus comprising:

a keyword extraction unit configured to extract keywords from text data representing a first content to be a base of a search, to set scores of the respective keywords extracted from the text data representing the first content, to extract keywords from text data representing a second content for calculating a degree of matching with the first content, and to set scores of the respective keywords extracted from the text data representing the second content, the scores of the respective keywords extracted from the text data representing the first and second content set in accordance with occurrence positions of the keywords within the first and second content;

a keyword expansion unit configured to expand the keywords extracted from the text data representing the first content data and to set scores of the expanded keywords based upon respective user-defined scoring coefficients of the expanded keywords;

a matching-degree calculation unit configured to calculate a degree of matching between the first content and the second content based on scores of keywords included in both the keywords extracted from the text data representing the first content and the keywords extracted from the text data representing the second content; and

an output unit configure to output, as a search result, information on a predetermined number of the second content which has a high degree of matching with the first content, based on a result of calculation performed by the matching-degree calculation unit, wherein

the matching-degree calculation unit is further configured to multiply the scores of the keywords included in both the keywords extracted from the text data representing the first content and the keywords extracted from the text data representing the second content, and to calculate, as the degree of matching between the first content and the second content, a value obtained by adding results of multiplications of the scores of the keywords included in both the keywords extracted from the text data representing the first content and the keywords extracted from the text data representing the second content.

2. The information processing apparatus according to claim 1 , wherein

the keyword extraction unit is further configured to calculate the score of each keyword based on a frequency of occurrence of the keyword in text data.

3. The information processing apparatus according to claim 1 , wherein

with a second content having a calculated degree of matching with the first content being a third content to be a base of a search,

the keyword extraction unit is further configured to extract keywords from text data representing the third content, to set scores of the respective keywords extracted from the text data representing the third content, to extract keywords from text data representing the second content, and to set scores of the respective keywords extracted from the text data representing the second content,

the matching-degree calculation unit is further configured to calculate a degree of matching between the third content and the second content based on scores of keywords included in both the keywords extracted from the text data representing the third content and the keywords extracted from the text data representing the second content, and

the output unit is further configured to output, as a search result, information on a predetermined number of the second content which has a high degree of matching with the third content, based on a result of calculation of the degree of matching between the third content and the second content performed by the matching-degree calculation unit.

4. The information processing apparatus according to claim 1 , wherein

the output unit is further configured to output, as the search result, a list of information on the predetermined number of the second content which has the high degree of matching with the first content based on the result of calculation performed by the matching-degree calculation unit.

5. The information processing apparatus according to claim 1 , wherein

based on the result of calculation performed by the matching-degree calculation unit, the output unit is further configured to output, as a search result, information on the predetermined number of the second content which has the high degree of matching with the first content in a descending order of a degree of matching.

6. The information processing apparatus according to claim 1 , wherein

each of the occurrence positions are prioritized such that keywords which occur within an occurrence position of higher priority are assigned a higher score than keywords which occur within an occurrence position of lower priority.

7. The information processing apparatus according to claim 1 , wherein

the keyword extraction unit is further configured to calculate the score of each keyword based on an attribute of the keyword such that keywords including a proper noun or a name are assigned a higher score than keywords including a general noun or verb.

8. The information processing apparatus according to claim 1 , wherein

the keyword expansion unit is further configured to expand the keywords extracted from the text data representing the first content data to determine at least one of synonym keywords, broader keywords, narrower keywords, and related keywords, and to set the scores of the expanded keywords based upon respective user-defined scoring coefficients of the synonym keywords, broader keywords, narrower keywords, and related keywords.

9. The information processing apparatus according to claim 1 , wherein

the matching-degree calculation unit is further configured to calculate the degree of matching between the first content and the second content based on the scores of the expanded keywords.

10. An information processing method comprising:

extracting keywords from text data representing a first content to be a base of a search;

setting scores of the respective keywords extracted from the text data representing the first content;

extracting keywords from text data representing a second content for calculating a degree of matching with the first content;

setting scores of the respective keywords extracted from the text data representing the second content, the scores of the respective keywords extracted from the text data representing the first and second content set in accordance with occurrence positions of the keywords within the first and second content;

expanding the keywords extracted from the text data representing the first content;

setting scores of the expanded keywords based upon respective user-defined scoring coefficients of the expanded keywords;

calculating a degree of matching between the first content and the second content based on scores of keywords included in both the keywords extracted from the text data representing the first content and the keywords extracted from the text data representing the second content; and

outputting, as a search result, information on a predetermined number of the second content which has a high degree of matching with the first content, based on a result of the calculating, wherein

the calculating includes multiplying the scores of the keywords included in both the keywords extracted from the text data representing the first content and the keywords extracted from the text data representing the second content, and calculating, as the degree of matching between the first content and the second content, a value obtained by adding results of multiplications of the keywords included in both the keywords extracted from the text data representing the first content and the keywords extracted from the text data representing the second content.

11. A computer readable storage medium storing computer readable instructions thereon, that, when executed by a processor, cause the processor to execute a process comprising:

extracting keywords from text data representing a first content to be a base of a search;

setting scores of the respective keywords extracted from the text data representing the first content;

extracting keywords from text data representing a second content for calculating a degree of matching with the first content;

setting scores of the respective keywords extracted from the text data representing the second content, the scores of the respective keywords extracted from the text data representing the first and second content set in accordance with occurrence positions of the keywords within the first and second content;

expanding the keywords extracted from the text data representing the first content;

setting scores of the expanded keywords based upon respective user-defined scoring coefficients of the expanded keywords;

calculating a degree of matching between the first content and the second content based on scores of keywords included in both the keywords extracted from the text data representing the first content and the keywords extracted from the text data representing the second content; and

outputting, as a search result, information on a predetermined number of the second content which has a high degree of matching with the first content, based on a result of the calculating, wherein

the calculating includes multiplying the scores of the keywords included in both the keywords extracted from the text data representing the first content and the keywords extracted from the text data representing the second content, and calculating, as the degree of matching between the first content and the second content, a value obtained by adding results of multiplications of the scores of the keywords included in both the keywords extracted from the text data representing the first content and the keywords extracted from the text data representing the second content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2017
From: SONY CORPORATION
To: SATURN LICENSING LLC
Reel/Frame 043177/0794 →