IP Library Granted Patent US 10,929,453
Granted Patent B2
US 10,929,453 · App. 16/522,727 · Granted Feb 23, 2021

Verifying textual claims with a document corpus

Inventor: Christopher Malon (Fort Lee, NJ)
G06F16/355G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,929,453
App. No.
16/522,727
Granted
Feb 23, 2021
Kind
B2
Abstract

A system verifies textual claims using a document corpus. The system includes a memory for storing program code and a processor device for running the code to retrieve documents from the corpus based on Term Frequency Inverse Document Frequency (TFIDF) similarity to a set of textual claims. The processor extracts named entities and capitalized phrases from the textual claims. The processor retrieves documents from the corpus with titles matching any of the extracted named entities and capitalized phrases. The processor extracts premise sentences from the retrieved documents. The processor classifies the premise sentences together with sources of the premises sentences against the textual claims to obtain classifications from among possible classifications including a supported, an unverified, or a contradicted classification. The processor aggregates the classifications over the premise sentences to selectively output, for each textual claim, an overall decision of the supported classification, the unverified classification, or the contradicted classification.

Claims (39)

1. A system for verifying textual claims using a document corpus, comprising:

a memory for storing program code; and

a processor device for running the program code to

retrieve documents from the document corpus based on Term Frequency Inverse Document Frequency (TFIDF) similarity to a set of textual claims;

extract named entities and capitalized phrases from the textual claims;

retrieve documents from the document corpus with titles matching any of the extracted named entities and capitalized phrases;

extract premise sentences from the retrieved documents;

classify the premise sentences together with sources of the premises sentences against the textual claims to obtain classifications from among possible classifications including a supported classification, an unverified classification, or a contradicted classification; and

aggregate the classifications over the premise sentences to selectively output, for each of the textual claims, an overall decision of the supported classification, the unverified classification, or the contradicted classification.

2. The system of claim 1 , wherein the processor device further outputs supporting statements for the overall decision from the document corpus.

3. The system of claim 1 , wherein the processor adds a title of a source document to each of the premise sentences.

4. The system of claim 1 , wherein the title of the source document is added before a corresponding one of the premise statements.

5. The system of claim 1 , wherein the system is comprised in a news summarization system.

6. The system of claim 1 , wherein the processor retrieves the documents from the document corpus using a threshold on an overall number of the documents retrieved from the document corpus to limit a number of processed document from the document corpus for each of multiple retrievals.

7. The system of claim 1 , wherein each of the premise sentences is individually classified together with a corresponding one of the sources against the textual claims to obtain the classifications for each of the textual claims.

8. The system of claim 1 , wherein concatenated sets of the premise statements are classified together with corresponding ones of the sources against the textual claims to obtain the classifications for each of the textual claims.

9. The system of claim 1 , wherein the processor device biases the overall decision by resolving conflicts between supporting and refuting information in favor of the supporting information.

10. A computer-implemented method for verifying textual claims using a document corpus, comprising:

retrieving, by a processor device, documents from the document corpus based on Term Frequency Inverse Document Frequency (TFIDF) similarity to a set of textual claims;

extracting, by the processor device, named entities and capitalized phrases from the textual claims;

retrieving, by the processor device, documents from the document corpus with titles matching any of the extracted named entities and capitalized phrases;

extracting, by the processor device, premise sentences from the retrieved documents;

classifying, by the processor device, the premise sentences together with sources of the premises sentences against the textual claims to obtain classifications from among possible classifications including a supported classification, an unverified classification, or a contradicted classification; and

aggregating, by the processor device, the classifications over the premise sentences to selectively output, for each of the textual claims, an overall decision of the supported classification, the unverified classification, or the contradicted classification.

11. The computer-implemented method of claim 10 , wherein the processor device further outputs supporting statements for the overall decision from the document corpus.

12. The computer-implemented method of claim 10 , wherein the processor adds a title of a source document to each of the premise sentences.

13. The computer-implemented method of claim 10 , wherein the title of the source document is added before a corresponding one of the premise statements.

14. The computer-implemented method of claim 10 , wherein the system is comprised in a news summarization system.

15. The computer-implemented method of claim 10 , wherein each of the retrieving steps involve a respective threshold on an overall number of the documents retrieved from the document corpus to limit a number of processed document from the document corpus.

16. The computer-implemented method of claim 10 , wherein each of the premise sentences is individually classified together with a corresponding one of the sources against the textual claims to obtain the classifications for each of the textual claims.

17. The computer-implemented method of claim 10 , wherein concatenated sets of the premise statements are classified together with corresponding ones of the sources against the textual claims to obtain the classifications for each of the textual claims.

18. The computer-implemented method of claim 10 , further comprising biasing the overall decision by resolving conflicts between supporting and refuting information in favor of the supporting information.

19. A computer program product for verifying textual claims using a document corpus, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

retrieving, by a processor device, documents from the document corpus based on Term Frequency Inverse Document Frequency (TFIDF) similarity to a set of textual claims;

extracting, by the processor device, named entities and capitalized phrases from the textual claims;

retrieving, by the processor device, documents from the document corpus with titles matching any of the extracted named entities and capitalized phrases;

extracting, by the processor device, premise sentences from the retrieved documents;

classifying, by the processor device, the premise sentences together with sources of the premises sentences against the textual claims to obtain classifications from among possible classifications including a supported classification, an unverified classification, or a contradicted classification; and

aggregating, by the processor device, the classifications over the premise sentences to selectively output, for each of the textual claims, an overall decision of the supported classification, the unverified classification, or the contradicted classification.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2021
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 054954/0596 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2019
From: MALON, CHRISTOPHER
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 049866/0823 →
Continuity (2)
Provisional Application 62716664 · Aug 9, 2018
Related Publication 20200050621A1 · Feb 13, 2020
Cited By (1)
US 12,327,078