IP Library Granted Patent US 7,249,312
Granted Patent B2
US 7,249,312 · App. 10/241,982 · Granted Jul 24, 2007

Attribute scoring for unstructured content

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,249,312
App. No.
10/241,982
Granted
Jul 24, 2007
Kind
B2
Abstract

The present invention provides for a system and method of assigning attribute scores, including those for abstract and diverse concepts such as sentiment (e.g., happiness, anger, sadness) to documents containing unstructured content (natural language content). Such content may include, but is not limited to, Web pages, e-mails, word processing documents, computer logs, chat logs, audio files, graphical images, text files, books, magazines, articles, etc. The attributes that are scored are specialized views of the unstructured content. Accordingly, attribute scoring denotes the processes of assigning one or more new measures to a document containing unstructured content. The attribute score or measure indicates the degree of the attributes found within the document.

Claims (72)

1. A computer implemented method for scoring of an attribute exhibited in a corpus of unstructured documents, the method comprising:

storing a plurality of scoring models, each of the scoring models associated with a different method of combining subdocument scores for a document, wherein the plurality of scoring models comprises:

a first scoring model for determining an average degree to which the attribute is present in the document;

a second scoring model for determining a maximum degree to which the attribute is present in the document;

a third scoring model for determining a total degree to which the attribute is present in the document; and

a fourth scoring model for determining an amount of change in the attribute in the subdocuments of the document over a period of time;

decomposing each unstructured document into subdocuments;

determining an attribute score for each subdocument in the corpus corresponding to the level said subdocument exhibits a feature associated with the attribute; and

combining subdocument attribute scores for each document to produce a document attribute score, further comprising:

selecting a scoring model for scoring the documents; and

for each document, determining a document attribute score for the document by applying the selected scoring model to the subdocument scores of the subdocuments of the document.

2. The method of claim 1 , further comprising combining the document attribute scores for each document to produce a corpus attribute score.

3. The method of claim 1 , further comprising determining which features of interest in said corpus of unstructured documents affect the document attribute score.

4. The method of claim 3 , wherein determining which features of interest affect the document attribute score comprises:

identifying candidate features in a corpus of training documents; and

storing each candidate feature in association with the training document in which it occurs.

5. The method of claim 1 , wherein the attribute is a sentiment expressed within the document.

6. The method of claim 1 , wherein at least one unstructured document comprises a plurality of subdocuments, each subdocument created independently and separately of the other subdocuments over a period of time.

7. The method of claim 1 , wherein determining each subdocument's attribute score comprises determining if any weighted features are present in each subdocument.

8. The method of claim 7 , wherein each subdocument's attribute score corresponds to a combination of values of weighted features.

9. A computer implemented method for scoring of an attribute exhibited in a corpus of unstructured documents, the method comprising:

storing a plurality of scoring models, each of the scoring models associated with a different method of combining subdocument scores for a document;

for each of a plurality of information needs, associating the information need with one of the scoring models for determining a document attribute score from the subdocument attribute scores of the subdocuments of the document, wherein the information needs comprise:

determining an average degree to which the attribute is present in the document;

determining a maximum degree to which the attribute is present in the document;

determining a total degree to which the attribute is present in the document; and

determining an amount of change in the attribute in the subdocuments of the document over a period of time;

decomposing each unstructured document into subdocuments;

determining an attribute score for each subdocument in the corpus corresponding to the level the subdocument exhibits a feature associated with the attribute; and

combining subdocument attribute scores for each document to produce a document attribute score, wherein combining subdocument attribute scores for each document to produce a document attribute score, further comprising:

determining a current information need with respect to the attribute; and

for each document, applying the scoring model associated with the current information need to the subdocument attribute scores of the document to determine the document attribute score.

10. A computer readable medium containing computer readable instructions executable by a processor for scoring of an attribute exhibited in a corpus of unstructured documents, the computer readable instructions comprising:

a training component for determining weights of features are associated with the attribute;

an attribute scoring component for determining attribute scores by:

decomposing each unstructured document into subdocuments;

determining an attribute score for each subdocument in said corpus in response to the weights of features associated with the attribute and present in the subdocument; and

a document scoring component for combining the subdocument attribute scores for each document to produce a document attribute score, the document scoring component including for each of a plurality of information needs, a scoring model for determining the document attribute score from the subdocument attribute scores of the subdocuments of the document, wherein the information needs comprise:

determining an average degree to which the attribute is present in the document;

determining a maximum degree to which the attribute is present in the document;

determining a total degree to which the attribute is present in the document; and

determining an amount of change in the attribute in the subdocuments of the document over a period of time.

11. The computer readable medium of claim 10 , further comprising:

a corpus scoring component for combining the document attribute scores for each document to produce a corpus attribute score.

12. The computer readable medium of claim 10 , wherein the training component is further adapted for determining which features of interest in said corpus of unstructured documents affect the document attribute score.

13. The computer readable medium of claim 12 , wherein the training component is adapted to determine which features of interest affect the document attribute score by:

identifying training features in a corpus of training documents; and

storing each training feature in association with the training document in which it occurs.

14. The computer readable medium of claim 10 , wherein the attribute scoring component is adapted to determine each subdocument's attribute score by determining if any weighted features are exhibited by each subdocument.

15. The computer readable medium of claim 14 , wherein each subdocument's attribute score corresponds to a combination of values of weighted features.

16. A computer implemented method for scoring a document with respect to an attribute, the method comprising:

identifying a plurality of subdocuments which comprise the document;

for each subdocument, determining a subdocument score as a function of features of the subdocument that match features associated with the attribute, to produce a plurality of subdocument scores for the document;

providing a plurality of scoring models, wherein each scoring model is associated with a different method of combining the subdocument scores and wherein each scoring model is associated with an information need, the scoring models comprising:

a first scoring model for determining an average degree to which the attribute is present in the document;

a second scoring model for determining a maximum degree to which the attribute is present in the document;

a third scoring model for determining a total degree to which the attribute is present in the document; and

a fourth scoring model for determining an amount of change in the attribute in the subdocuments of the document over a period of time;

selecting one of the scoring models by determining a current information need, and selecting the scoring model associated with the current information need; and

determining a document score for the document using by applying the selected scoring model to the subdocument scores of the subdocuments.

17. A computer implemented method for scoring a document with respect to an attribute, the method comprising:

identifying a plurality of subdocuments which comprise the document;

for each subdocument, determining a subdocument score as a function of features of the subdocument that match features associated with the attribute, to produce a plurality of subdocument scores for the document;

providing a plurality of scoring models, wherein each scoring model is associated with a different method of combining the subdocument scores and wherein each scoring model is associated with an information need, the scoring models comprising:

a first scoring model for determining an average degree to which the attribute is present in the document;

a second scoring model for determining a maximum degree to which the attribute is present in the document;

a third scoring model for determining a total degree to which the attribute is present in the document; and

a fourth scoring model for determining an amount of change in the attribute in the subdocuments of the document over a period of time;

selecting one of the scoring models; and

determining a document score for the document using by applying the selected scoring model to the subdocument scores of the subdocuments.

18. The computer implemented method of claim 17 , wherein the attribute is a sentiment expressed within the subdocuments.

19. The computer implemented method of claim 17 , wherein the subdocuments that comprise the document are each created independently and separately of the other subdocuments over a period of time.

Assignments (6)
MERGER Recorded Jan 10, 2020
From: FIRST DATA SOLUTIONS INC.
To: FIRST DATA RESOURCES, LLC
Reel/Frame 051480/0269 →
MERGER Recorded Jan 8, 2020
From: INTELLIGENT RESULTS (A/K/A INTELLIGENT RESULTS, INC.)
To: FIRST DATA SOLUTIONS L.L.C. (A/K/A FIRST DATA SOLUTIONS INC.)
Reel/Frame 051456/0126 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Aug 19, 2019
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: FIRST DATA CORPORATION; DW HOLDINGS, INC.; FIRST DATA RESOURCES, INC. (K/N/A FIRST DATA RESOURCES, LLC); FUNDSXPRESS FINANCIAL NETWORKS, INC.; INTELLIGENT RESULTS, INC. (K/N/A FIRST DATA SOLUTIONS, INC.); LINKPOINT INTERNATIONAL, INC.; MONEY NETWORK FINANCIAL, LLC; SIZE TECHNOLOGIES, INC.; TASQ TECHNOLOGY, INC.; TELECHECK INTERNATIONAL, INC.
Reel/Frame 050090/0060 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Aug 19, 2019
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: FIRST DATA CORPORATION; DW HOLDINGS, INC.; FIRST DATA RESOURCES, LLC; FUNDSXPRESS FINANCIAL NETWORK, INC.; FIRST DATA SOLUTIONS, INC.; LINKPOINT INTERNATIONAL, INC.; MONEY NETWORK FINANCIAL, LLC; SIZE TECHNOLOGIES, INC.; TASQ TECHNOLOGY, INC.; TELECHECK INTERNATIONAL, INC.
Reel/Frame 050091/0474 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENT RIGHTS Recorded Aug 19, 2019
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: FIRST DATA CORPORATION
Reel/Frame 050094/0455 →
RELEASE OF SECURITY INTEREST Recorded Jul 30, 2019
From: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: CARDSERVICE INTERNATIONAL, INC.; DW HOLDINGS INC.; FIRST DATA CORPORATION; FIRST DATA RESOURCES, LLC; FUNDSXPRESS, INC.; INTELLIGENT RESULTS, INC.; LINKPOINT INTERNATIONAL, INC.; SIZE TECHNOLOGIES, INC.; TASQ TECHNOLOGY, INC.; TELECHECK INTERNATIONAL, INC.; TELECHECK SERVICES, INC.
Reel/Frame 049902/0919 →