IP Library Granted Patent US 11,144,579
Granted Patent B2
US 11,144,579 · App. 16/272,239 · Granted Oct 12, 2021

Use of machine learning to characterize reference relationship applied over a citation graph

Inventors: Brendan Bull (Durham, NC); Andrew Hicks (Durham, NC); Scott Robert Carrier (Apex, NC); Dwi Sianto Mansjur (Cary, NC)
Assignee: International Business Machines Corporation
G06F16/31G06F16/353G06F16/93G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,144,579
App. No.
16/272,239
Granted
Oct 12, 2021
Kind
B2
Abstract

Techniques for document analysis using machine learning are provided. A selection of an index is received document, and a plurality of documents that refer to the index document is identified. For each respective document in the plurality of documents, a respective portion of the respective document is extracted, where the respective portion refers to the index document, and a respective vector representation is generated for the respective portion. A plurality of groupings is generated for the plurality of documents based on how each of the plurality of documents relate to the index document, by processing the vector representations using a trained classifier. Finally, at least an indication of the plurality of groupings is provided, along with the index document.

Claims (82)

1. A method comprising:

receiving a selection of an index document;

identifying a plurality of documents that refer to the index document;

for each respective document in the plurality of documents:

extracting a respective portion of the respective document, wherein the respective portion refers to the index document; and

generating, by operation of one or more processors, a respective vector representation for the respective portion;

generating a plurality of groupings for the plurality of documents based on how each of the plurality of documents relate to the index document, comprising, for each respective document in the plurality of documents:

processing the respective vector representation using a trained classifier to assign the respective document to a respective category of a plurality of categories, wherein the plurality of categories includes (i) a category for documents that support the index document, and (ii) a category for documents that do not support the index document; and

providing at least an indication of the plurality of groupings, along with the index document.

2. The method of claim 1 , wherein the trained classifier is trained using a curated training set of documents, wherein each respective document in the curated training set is labeled based on how it relates to a respective reference document.

3. The method of claim 1 , wherein identifying the plurality of documents that refer to the index document comprises:

identifying documents that cite the index document; and

identifying documents that refer to the index document without explicitly citing the index document.

4. The method of claim 1 , wherein extracting the respective portion of each respective document comprises:

identifying a location in the respective document where the index document is referenced; and

extracting a predefined number of sentences before and after the identified location.

5. The method of claim 1 , wherein each of the plurality of groupings correspond to a respective category of a plurality of categories, wherein the plurality of categories further includes at least one of:

(i) a category for documents expanding on the index document;

(ii) a category for documents criticizing the index document;

(iii) a category for documents showing similar findings as the index document;

(iv) a category for documents that failed to show similar findings as the index document; or

(v) a category for documents that rely on the index document for support.

6. The method of claim 1 , the method further comprising:

providing a number of documents that are included in each of the plurality of groupings;

providing a link to each of the plurality of documents; and

providing the respective portion of each respective document in the plurality of documents.

7. The method of claim 1 , wherein the documents included in a first category of the plurality of groupings are sorted based in part on a respective importance measure of each respective document.

8. The method of claim 1 , wherein identifying the plurality of documents that refer to the index document comprises accessing a citation graph that includes the index document.

9. A computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code executable by one or more computer processors to perform an operation comprising:

receiving a selection of an index document;

identifying a plurality of documents that refer to the index document;

for each respective document in the plurality of documents:

extracting a respective portion of the respective document, wherein the respective portion refers to the index document; and

generating a respective vector representation for the respective portion;

generating a plurality of groupings for the plurality of documents based on how each of the plurality of documents relate to the index document, comprising, for each respective document in the plurality of documents:

processing the respective vector representation using a trained classifier to assign the respective document a respective category of a plurality of categories, wherein the plurality of categories includes (i) a category for documents that support the index document, and (ii) a category for documents that do not support the index document; and

providing at least an indication of the plurality of groupings, along with the index document.

10. The computer-readable storage medium of claim 9 , wherein the trained classifier is trained using a curated training set of documents, wherein each respective document in the curated training set is labeled based on how it relates to a respective reference document.

11. The computer-readable storage medium of claim 9 , wherein identifying the plurality of documents that refer to the index document comprises:

identifying documents that cite the index document; and

identifying documents that refer to the index document without explicitly citing the index document.

12. The computer-readable storage medium of claim 9 , wherein extracting the respective portion of each respective document comprises:

identifying a location in the respective document where the index document is referenced; and

extracting a predefined number of sentences before and after the identified location.

13. The computer-readable storage medium of claim 9 , wherein each of the plurality of groupings correspond to a respective category of a plurality of categories, wherein the plurality of categories further includes at least one of:

(i) a category for documents expanding on the index document;

(ii) a category for documents criticizing the index document;

(iii) a category for documents showing similar findings as the index document;

(iv) a category for documents that failed to show similar findings as the index document; or

(v) a category for documents that rely on the index document for support.

14. The computer-readable storage medium of claim 9 , the operation further comprising:

providing a number of documents that are included in each of the plurality of groupings;

providing a link to each of the plurality of documents; and

providing the respective portion of each respective document in the plurality of documents.

15. A system comprising:

one or more computer processors; and

a memory containing a program which when executed by the one or more computer processors performs an operation, the operation comprising:

receiving a selection of an index document;

identifying a plurality of documents that refer to the index document;

for each respective document in the plurality of documents:

extracting a respective portion of the respective document, wherein the respective portion refers to the index document; and

generating a respective vector representation for the respective portion;

generating a plurality of groupings for the plurality of documents based on how each of the plurality of documents relate to the index document, comprising, for each respective document in the plurality of documents:

processing the respective vector representation using a trained classifier to assign the respective document a respective category of a plurality of categories, wherein the plurality of categories includes (i) a category for documents that support the index document, and (ii) a category for documents that do not support the index document; and

providing at least an indication of the plurality of groupings, along with the index document.

16. The system of claim 15 , wherein the trained classifier is trained using a curated training set of documents, wherein each respective document in the curated training set is labeled based on how it relates to a respective reference document.

17. The system of claim 15 , wherein identifying the plurality of documents that refer to the index document comprises:

identifying documents that cite the index document; and

identifying documents that refer to the index document without explicitly citing the index document.

18. The system of claim 15 , wherein extracting the respective portion of each respective document comprises:

identifying a location in the respective document where the index document is referenced; and

extracting a predefined number of sentences before and after the identified location.

19. The system of claim 15 , wherein each of the plurality of groupings correspond to a respective category of a plurality of categories, wherein the plurality of categories further includes at least one of:

(i) a category for documents expanding on the index document;

(ii) a category for documents criticizing the index document;

(iii) a category for documents showing similar findings as the index document;

(iv) a category for documents that failed to show similar findings as the index document; or

(v) a category for documents that rely on the index document for support.

20. The system of claim 15 , the operation further comprising:

providing a number of documents that are included in each of the plurality of groupings;

providing a link to each of the plurality of documents; and

providing the respective portion of each respective document in the plurality of documents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2019
From: BULL, BRENDAN; HICKS, ANDREW; CARRIER, SCOTT ROBERT; MANSJUR, DWI SIANTO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048293/0972 →
Continuity (1)
Related Publication 20200257709A1 · Aug 13, 2020
Cited By (2)
US 12,340,565 US 12,482,540