IP Library Granted Patent US 9,256,649
Granted Patent B2
US 9,256,649 · App. 13/920,803 · Granted Feb 9, 2016

Method and system of filtering and recommending documents

Inventors: Robert M. Patton (Knoxville, TN); Thomas E. Potok (Oak Ridge, TN)
Assignee: UT-Battelle LLC
G06F17/3053G06F17/3097
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,256,649
App. No.
13/920,803
Granted
Feb 9, 2016
Kind
B2
Abstract

Disclosed is a method and system for discovering documents using a computer and providing a small set of the most relevant documents to the attention of a human observer. Using the method, the computer obtains a seed document from the user and generates a seed document vector using term frequency-inverse corpus frequency weighting. A keyword index for a plurality of source documents can be compared with the weighted terms of the seed document vector. The comparison is then filtered to reduce the number of documents, which define an initial subset of the source documents. Initial subset vectors are generated and compared to the seed document vector to obtain a similarity value for each comparison. Based on the similarity value, the method then recommends one or more of the source documents.

Claims (24)

1. A computer programmed with a series of instructions that, when executed by a processor, cause the computer to perform the method steps comprising:

obtaining a keyword index for a plurality of source documents;

obtaining a seed document;

generating a seed document vector with a plurality of weighted terms using term frequency-inverse corpus frequency weighting;

determining a plurality of significant search terms from the plurality of weighted terms;

filtering the significant search terms found in the keyword index to retain only the significant search terms that occur in less than a predetermined measure of the total number of source documents to define an initial subset of source documents;

generating initial subset vectors and comparing each of the initial subset vectors to the seed document vector to obtain a similarity value for each comparison;

recommending one or more of the source documents from the initial subset based on the similarity value;

wherein recommending one or more of the source documents includes generating secondary subset vectors using term frequency-inverse corpus frequency weighting;

comparing each secondary subset vector to the seed document vector, and recording a vector similarity value and a number of common terms for each comparison, thereby creating a recommendation set;

generating recommendation set vectors using term frequency-inverse corpus frequency weighting;

comparing each of the recommendation set vectors to the seed document vector to obtain a similarity value for each comparison;

sorting the similarity values;

selecting source documents based on a predetermined measure of similarity values;

rejecting source documents having less than a predetermined measure of common terms relative to the seed document; and

displaying one or more of the source documents.

2. The computer of claim 1 , wherein determining the plurality of significant search terms includes selecting a predetermined measure of the highest weighted terms from the seed document.

3. The computer of claim 1 , wherein filtering the significant search terms includes:

determining for each of the significant search terms the number of source documents that contain the significant search term; and

retaining only the significant search terms for which the number is less than a predetermined measure of the total number of source documents.

4. The computer of claim 3 , wherein the predetermined measure is a percentage of the total number of source documents.

5. The computer of claim 3 , wherein the initial subset includes only the source documents containing the retained significant search terms.

6. The computer of claim 1 , further including creating a secondary subset from the initial subset by comparing each of the initial subset vectors to the seed document vector to obtain a comparison value, and selecting source documents based on a predetermined measure of similarity values.

7. The computer of claim 1 , wherein the initial subset vector is generated using term frequency-inverse corpus frequency weighting.

Assignments (3)
CONFIRMATORY LICENSE Recorded Sep 30, 2014
From: UT-BATTELLE, LLC
To: U.S. DEPARTMENT OF ENERGY
Reel/Frame 033847/0691 →
CONFIRMATORY LICENSE Recorded Sep 29, 2014
From: UT-BATTELLE, LLC
To: U.S. DEPARTMENT OF ENERGY
Reel/Frame 033845/0490 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2013
From: POTOK, THOMAS E.; PATTON, ROBERT M.
To: UT-BATTELLE, LLC
Reel/Frame 030867/0988 →
Continuity (2)
Provisional Application 61661038 · Jun 18, 2012
Related Publication 20130339373A1 · Dec 19, 2013