IP Library Granted Patent US 8,533,148
Granted Patent B1
US 8,533,148 · App. 13/632,943 · Granted Sep 10, 2013

Document relevancy analysis within machine learning systems including determining closest cosine distances of training examples

Inventors: Christian Feuersänger (Rheinbach, DE); Dietrich Wettschereck (Bonn, DE); Jan Puzicha (Rheinbach, DE)
Assignee: Recommind, Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,533,148
App. No.
13/632,943
Granted
Sep 10, 2013
Kind
B1
Abstract

Systems and methods that quantify document relevance for a document relative to a training corpus and select a best match or best matches are provided herein. Methods may include generating an example-based explanation for relevancy of a document to a training corpus by executing a support vector machine classifier, the support vector machine classifier performing a centroid classification of a relevant document in a term frequency-inverse document frequency features space relative to training examples in a training corpus, and generating an example-based explanation by selecting a best match for the relevant document from the training examples based upon the centroid classification. Determining the training example having the closest cosine distance to the relevant document includes ranking the training examples by stretching the internal best match scores for the training examples linearly to cover a complete unit interval.

Claims (39)

1. A method for quantifying relevancy of a document to a training corpus, the method comprising:

calculating, by a processor, an internal best match score for a relevant document relative to each training example in the training corpus by:

determining cosine distances between the relevant document and training examples in the training corpus relative to term frequency-inverse document frequency weights associated with the training examples;

determining the training example having a closest cosine distance to the relevant document; and

outputting the training example having the closest cosine distance to the relevant document;

wherein determining the training example having the closest cosine distance to the relevant document further comprises ranking the training examples by stretching the internal best match scores for the training examples linearly to cover a complete unit interval.

2. The method according to claim 1 , further comprising converting each of the training examples into a high-dimensional feature space using term frequencies.

3. The method according to claim 1 , further comprising calculating an internal best match score for each of the training examples by multiplying a square root of term frequencies by an inverse document frequency.

4. The method according to claim 1 , further comprising:

training a classifier on a training corpus that comprises training examples;

classifying a set of documents; and

determining relevant documents in the set.

5. The method according to claim 4 , wherein the classifier comprises a support vector machine.

6. The method according to claim 4 , wherein determining relevant documents in the set comprises determining distances between each document within the set of documents relative to a support vector machine model.

7. The method according to claim 1 , further comprising outputting a list of the training examples based upon ranked cosine distances between the relevant document and the training examples.

8. The method according to claim 7 , further comprising applying a relevancy threshold to affect an amount of training examples that are included in the list.

9. A machine learning system that quantifies relevancy of a document to a training corpus, the system comprising:

at least one server comprising a processor configured to execute instructions that reside in memory, the instructions comprising:

a classifier module that:

calculates an internal best match score for a relevant document relative to each training example in the training corpus by:

determining cosine distances between the document and training examples in the training corpus relative to term frequency-inverse document frequency weights associated with the training examples; and

determines the training example having a closest cosine distance to the relevant document; and

a user interface module that outputs the training example having the closest cosine distance to the relevant document;

wherein the classifier module determines the training example having the closest cosine distance to the relevant document by ranking the training examples by stretching the internal best match scores for the training examples linearly to cover a complete unit interval.

10. The machine learning system according to claim 9 , wherein each of the training examples has been converted into a high-dimensional feature space using term frequencies.

11. The machine learning system according to claim 9 , wherein the classifier module calculates the internal best match score for each of the training examples by multiplying a square root of term frequencies by an inverse document frequency.

12. The machine learning system according to claim 9 , wherein the classifier module further:

classifies a set of documents; and

determines relevant documents in the set, the classifier module being trained on a training corpus that comprises training examples.

13. The machine learning system according to claim 12 , wherein the classifier module comprises a support vector machine.

14. The machine learning system according to claim 12 , wherein the classifier module determines relevant documents in the set by determining cosine distances between each document within the set of documents using a support vector machine model.

15. The machine learning system according to claim 9 , wherein the user interface module further outputs a list of the training examples based upon ranked cosine distances between the relevant document and the training examples.

16. The machine learning system according to claim 15 , wherein the classifier module further applies a relevancy threshold to affect an amount of training examples that are included in the list.

17. A method for generating an example-based explanation for relevancy of a document to a training corpus, the method comprising:

executing, by a processor, a support vector machine classifier that:

generates a classification model using the training corpus; and

classifies subject documents using the classification model;

performing a centroid classification of a selected relevant document in a term frequency-inverse document frequency feature space; and

generating an example-based explanation by selecting a best match for the selected relevant document from the training corpus.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2026
From: OPEN TEXT HOLDINGS, INC.
To: OPEN TEXT INC.
Reel/Frame 074362/0745 →
MERGER Recorded Jul 20, 2018
From: RECOMMIND, INC.
To: OPEN TEXT HOLDINGS, INC.
Reel/Frame 046414/0847 →
CHANGE OF ADDRESS Recorded Jul 13, 2017
From: RECOMMIND, INC.
To: RECOMMIND, INC.
Reel/Frame 043187/0308 →
CHANGE OF ADDRESS Recorded Jun 13, 2017
From: RECOMMIND, INC.
To: RECOMMIND, INC.
Reel/Frame 042801/0885 →
CHANGE OF ADDRESS Recorded Feb 13, 2017
From: RECOMMIND, INC.
To: RECOMMIND, INC.
Reel/Frame 041693/0805 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2012
From: FEUERSANGER, CHRISTIAN; WETTSCHERECK, DIETRICH; PUZICHA, JAN
To: RECOMMIND, INC.
Reel/Frame 029123/0883 →