IP Library Granted Patent US 9,249,287
Granted Patent B2
US 9,249,287 · App. 14/002,692 · Granted Feb 2, 2016

Document evaluation apparatus, document evaluation method, and computer-readable recording medium using missing patterns

Inventors: Yusuke Muraoka (Tokyo, JP); Dai Kusui (Tokyo, JP); Yukitaka Kusumura (Tokyo, JP); Hironori Mizuguchi (Tokyo, JP)
Assignee: NEC Corporation
C08L23/06B29C43/24B29C45/0013B29C47/0004B29C47/0038C08J3/226G06N3/08B29C47/0009B29C47/0019B29C47/0021B29C47/0023
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,249,287
App. No.
14/002,692
Granted
Feb 2, 2016
Kind
B2
Abstract

In order to accurately learn a function for evaluating documents, even in the case where sample documents having missing feature values are included as training data, a document evaluation apparatus is provided with a data classification unit ( 3 ) that classifies a set of sample documents based on missing patterns of a first feature vector, a first learning unit ( 4 ) that uses feature values that are not missing in the first feature vector and evaluation values to learn a first function for calculating a first score which is a weighted evaluation value for each classification, a feature vector generation unit ( 5 ) that computes a feature value corresponding to each classification using the first score, and generates a second feature vector having the computed feature values, and a second learning unit ( 6 ) that uses the second feature vector and the evaluation values to learn a second function for calculating a second score for evaluating documents targeted for evaluation.

Claims (25)

1. A document evaluation apparatus, realized by a computer, for evaluating a document using a set of sample documents having a first feature vector consisting of a plurality of feature values and evaluation values of the documents, comprising:

a processor, wherein the processor is configured to

classify the set of sample documents, based on a missing pattern indicating a set of indices whose feature values are missing in the first feature vector;

use feature values that are not missing in the first feature vector and the evaluation values to learn, for each classification, a first function for calculating a first score which is a weighted evaluation value for each classification;

compute a feature value corresponding to each classification using the first score, and generate a second feature vector having the computed feature values;

use the second feature vector and the evaluation values to learn a second function for calculating a second score for evaluating a document targeted for evaluation; and

measure an appearance frequency of the sample documents for each missing pattern, based on the set of sample documents, and match a missing pattern whose appearance frequency is less than or equal to a set threshold to a missing pattern that is most similar to said missing pattern and whose appearance frequency is greater than the threshold.

2. The document evaluation apparatus according to claim 1 , wherein the processor is further configured to receive input of a set of documents targeted for evaluation, compute the second score for each document based on the second function, and rank the documents based on the second scores.

3. The document evaluation apparatus according to claim 1 , wherein the processor is configured to generate the feature values constituting the second feature vector by normalizing the first scores so as to fall within a set range.

4. A document evaluation method for evaluating a document using a set of sample documents having a first feature vector consisting of a plurality of feature values and evaluation values of the documents, comprising the steps of:

(a) classifying the set of sample documents, based on a missing pattern indicating a set of indices whose feature values are missing in the first feature vector;

(b) using feature values that are not missing in the first feature vector and the evaluation values to learn, for each classification, a first function for calculating a first score which is a weighted evaluation value for each classification;

(c) computing a feature value corresponding to each classification using the first score, and generating a second feature vector having the computed feature values;

(d) using the second feature vector and the evaluation values to learn a second function for calculating a second score for evaluating a document targeted for evaluation; and

(e) measuring an appearance frequency of the sample documents for each missing pattern, based on the set of sample documents, and matching a missing pattern whose appearance frequency is less than or equal to a set threshold to a missing pattern that is most similar to said missing pattern and whose appearance frequency is greater than the threshold.

5. A non-transitory computer-readable recording medium storing a program for evaluating by computer a document using a set of sample documents having a first feature vector consisting of a plurality of feature values and evaluation values of the documents, the program including commands for causing the computer to execute the steps of:

(a) classifying the set of sample documents, based on a missing pattern indicating a set of indices whose feature values are missing in the first feature vector;

(b) using feature values that are not missing in the first feature vector and the evaluation values to learn, for each classification, a first function for calculating a first score which is a weighted evaluation value for each classification;

(c) computing a feature value corresponding to each classification using the first score, and generating a second feature vector having the computed feature values;

(d) using the second feature vector and the evaluation values to learn a second function for calculating a second score for evaluating a document targeted for evaluation; and

(e) measuring an appearance frequency of the sample documents for each missing pattern, based on the set of sample documents, and matching a missing pattern whose appearance frequency is less than or equal to a set threshold to a missing pattern that is most similar to said missing pattern and whose appearance frequency is greater than the threshold.

6. The document evaluation method according to claim 4 further includes the step of (f) receiving input of a set of documents targeted for evaluation, computing the second score for each document based on the second function, and ranking the documents based on the second scores.

7. The document evaluation method according to claim 4 in which, in the step (c), the feature values constituting the second feature vector are generated by normalizing the first scores so as to fall within a set range.

8. The computer-readable recording medium according to claim 5 further includes the step of (f) receiving input of a set of documents targeted for evaluation, computing the second score for each document based on the second function, and ranking the documents based on the second scores.

9. The computer-readable recording medium according to claim 5 in which, in the step (c), the feature values constituting the second feature vector are generated by normalizing the first scores so as to fall within a set range.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2014
From: MURAOKA, YUSUKE; KUSUI, DAI; KUSUMURA, YUKITAKA; MIZUGUCHI, HIRONORI
To: NEC CORPORATION
Reel/Frame 033977/0346 →
Priority Claims (1)
JP 2012-038286 · Feb 24, 2012 · national
Continuity (1)
Related Publication 20130332401A1 · Dec 12, 2013