IP Library › Granted Patent US 9,311,390
Granted Patent B2
US 9,311,390 · App. 12/362,380 · Granted Apr 12, 2016

System and method for handling the confounding effect of document length on vector-based similarity scores

Inventor: Derrick C. Higgins (Philadelphia, PA)
Assignee: Educational Testing Service
G06F17/3069
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,311,390
App. No.
12/362,380
Granted
Apr 12, 2016
Kind
B2
Abstract

A computer-implemented method, system, and computer program product for generating vector-based similarity scores in text document comparisons considering confounding effects of document length. Vector-based methods for comparing the semantic similarity between texts (such as Content Vector Analysis and Random Indexing) have a characteristic which may reduce their usefulness for some applications: the similarity estimates they produce are strongly correlated with the lengths of the texts compared. The statistical basis for this confound is described, and suggests the application of a pivoted normalization method from information retrieval to correct for the effect of document length. In two text categorization experiments, Random Indexing similarity scores using pivoted normalization are shown to perform significantly better than standard vector-based similarity estimation methods.

Claims (32)

1. A computer-implemented method of generating vector-based similarity scores in text document comparisons considering document length, comprising:

computing a mean of a number of word types of two text documents to be compared;

determining a similarity score with a vector-based similarity model, wherein the vector-based similarity model is a Random Indexing model and wherein a normalization slope parameter has a value of 10;

performing pivoted document length normalization on the similarity score using the mean of the number of word types of the two text documents as a normalization affected by both text documents and using the normalization slope parameter; and

outputting a normalized similarity score.

2. A computer-implemented method of generating vector-based similarity scores in text document comparisons considering document length, comprising:

computing a mean of a number of word types of two text documents to be compared;

determining a similarity score with a vector-based similarity model, wherein the vector-based similarity model is a Content-Vector Analysis model and wherein a normalization slope parameter has a value of 5;

performing pivoted document length normalization on the similarity score using the mean of the number of word types of the two text documents as a normalization affected by both text documents and using the normalization slope parameter; and

outputting a normalized similarity score.

3. A computer system for generating vector-based similarity scores in text document comparisons considering document length, comprising:

a computer programmed with instructions that, when executed, cause the computer to execute steps comprising:

computing a mean of a number of word types of two text documents;

determining a similarity score with a vector-based similarity model, wherein the vector-based similarity model is a Random Indexing model and wherein a normalization slope parameter has a value of 10;

performing pivoted document length normalization on the similarity score using the mean of the number of word types of the two text documents as a normalization affected by both text documents and using the normalization slope parameter; and

outputting a normalized similarity score.

4. A computer system for generating vector-based similarity scores in text document comparisons considering document length, comprising:

a computer programmed with instructions that, when executed, cause the computer to execute steps comprising:

computing a mean of a number of word types of two text documents to be compared;

determining a similarity score with a vector-based similarity model, wherein the vector-based similarity model is a Content-Vector Analysis model and wherein a normalization slope parameter has a value of 5;

performing pivoted document length normalization on the similarity score using the mean of the number of word types of the two text documents as a normalization affected by both text documents and using the normalization slope parameter; and

outputting a normalized similarity score.

5. An article of manufacture comprising a non-transitory computer-readable storage medium for causing a computer to generate vector-based similarity scores in text document comparisons considering document length, said computer readable medium including programming instructions that, when executed, cause the computer to execute steps comprising:

computing a mean of a number of word types of two text documents;

determining a similarity score with a vector-based similarity model, wherein the vector-based similarity model is a Random Indexing model and wherein a normalization slope parameter has a value of 10;

performing pivoted document length normalization on the similarity score using the mean of the number of word types of the two text documents as a normalization affected by both text documents and using the normalization slope parameter; and

outputting a normalized similarity score.

6. An article of manufacture comprising a non-transitory computer-readable storage medium for causing a computer to generate vector-based similarity scores in text document comparisons considering document length, said computer readable medium including programming instructions that, when executed, cause the computer to execute steps comprising:

computing a mean of a number of word types of two text documents to be compared;

determining a similarity score with a vector-based similarity model, wherein the vector-based similarity model is a Content-Vector Analysis model and wherein a normalization slope parameter has a value of 5;

performing pivoted document length normalization on the similarity score using the mean of the number of word types of the two text documents as a normalization affected by both text documents and using the normalization slope parameter; and

outputting a normalized similarity score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2009
From: HIGGINS, DERRICK C.
To: EDUCATIONAL TESTING SERVICE
Reel/Frame 022176/0890 →
Continuity (2)
Provisional Application 61024303 · Jan 29, 2008
Related Publication 20090190839A1 · Jul 30, 2009