IP Library Granted Patent US 10,909,320
Granted Patent B2
US 10,909,320 · App. 16/270,431 · Granted Feb 2, 2021

Ontology-based document analysis and annotation generation

Inventors: Brendan Bull (Durham, NC); Paul Lewis Felt (Utah County, UT); Andrew Hicks (Durham, NC)
Assignee: International Business Machines Corporation
G06F40/289G06F16/93G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,909,320
App. No.
16/270,431
Granted
Feb 2, 2021
Kind
B2
Abstract

Techniques for cognitive annotation are provided. An electronic document including textual data is received. A plurality of importance scores are generated for a plurality of words included in the electronic document by processing the electronic document using a trained passage encoder. Important words are identified based on the plurality of importance scores. One or more clusters of words are generated, where each of the one or more clusters of words includes at least one of the plurality of important words. A representative word is selected for a first cluster, and the representative word is mapped to one or more concepts from a predefined list of concepts. The one or more concepts are disambiguated to identify a set of relevant concepts for the electronic document. An annotated version of the electronic document is generated based at least in part on the set of relevant concepts.

Claims (58)

1. A computer-implemented method comprising:

receiving an electronic document including textual data;

generating a plurality of importance scores for a plurality of words included in the electronic document by processing the electronic document using a trained passage encoder;

identifying a plurality of important words, from the plurality of words, based on the plurality of importance scores;

generating one or more clusters of words, from the plurality of important words, wherein each of the one or more clusters of words includes at least one of the plurality of important words;

selecting, for a first cluster of the one or more clusters of words, a representative word;

mapping the representative word for the first cluster to one or more concepts from a predefined list of concepts;

disambiguating, by operation of one or more computer processors, the one or more concepts to identify a set of relevant concepts for the electronic document;

generating an annotated version of the electronic document based at least in part on the set of relevant concepts; and

generating a plurality of terms that summarize the electronic document based on mapping the set of relevant concepts to a predefined set of search terms, wherein the predefined set of search terms comprises medical subject heading (MeSH) terms.

2. The method of claim 1 , wherein identifying the plurality of important words comprises:

generating an importance score for each word in the electronic document; and

determining an expected importance score for the electronic document; and

selecting words from the electronic document with an importance score exceeding the expected importance score.

3. The method of claim 1 , wherein generating the one or more clusters of words comprises:

generating, for each respective word in the plurality of important words, a respective vector; and

clustering vectors that exceed a predefined threshold of similarity.

4. The method of claim 1 , wherein the passage encoder is a recurrent neural network (RNN).

5. The method of claim 1 , wherein the predefined list of concepts comprises unified medical language system (UMLS) concept unique identifiers (CUIs).

6. A computer-readable storage medium having computer-readable program code embodied therewith, the computer-readable program code executable by one or more computer processors to perform an operation comprising:

receiving an electronic document including textual data;

generating a plurality of importance scores for a plurality of words included in the electronic document by processing the electronic document using a trained passage encoder;

identifying a plurality of important words, from the plurality of words, based on the plurality of importance scores;

generating one or more clusters of words, from the plurality of important words, wherein each of the one or more clusters of words includes at least one of the plurality of important words;

selecting, for a first cluster of the one or more clusters of words, a representative word;

mapping the representative word for the first cluster to one or more concepts from a predefined list of concepts;

disambiguating the one or more concepts to identify a set of relevant concepts for the electronic document;

generating an annotated version of the electronic document based at least in part on the set of relevant concepts; and

generating a plurality of terms that summarize the electronic document based on mapping the set of relevant concepts to a predefined set of search terms, wherein the predefined set of search terms comprises medical subject heading (MeSH) terms.

7. The computer-readable storage medium of claim 6 , wherein identifying the plurality of important words comprises:

generating an importance score for each word in the electronic document; and

determining an expected importance score for the electronic document; and

selecting words from the electronic document with an importance score exceeding the expected importance score.

8. The computer-readable storage medium of claim 6 , wherein generating the one or more clusters of words comprises:

generating, for each respective word in the plurality of important words, a respective vector; and

clustering vectors that exceed a predefined threshold of similarity.

9. The computer-readable storage medium of claim 6 , wherein the passage encoder is a recurrent neural network (RNN).

10. The computer-readable storage medium of claim 6 , wherein the predefined list of concepts comprises unified medical language system (UMLS) concept unique identifiers (CUIs).

11. A system comprising:

one or more computer processors; and

a memory containing a program which when executed by the one or more computer processors performs an operation, the operation comprising:

receiving an electronic document including textual data;

generating a plurality of importance scores for a plurality of words included in the electronic document by processing the electronic document using a trained passage encoder;

identifying a plurality of important words, from the plurality of words, based on the plurality of importance scores;

generating one or more clusters of words, from the plurality of important words, wherein each of the one or more clusters of words includes at least one of the plurality of important words;

selecting, for a first cluster of the one or more clusters of words, a representative word;

mapping the representative word for the first cluster to one or more concepts from a predefined list of concepts;

disambiguating the one or more concepts to identify a set of relevant concepts for the electronic document;

generating an annotated version of the electronic document based at least in part on the set of relevant concepts; and

generating a plurality of terms that summarize the electronic document based on mapping the set of relevant concepts to a predefined set of search terms, wherein the predefined set of search terms comprises medical subject heading (MeSH) terms.

12. The system of claim 11 , wherein identifying the plurality of important words comprises:

generating an importance score for each word in the electronic document; and

determining an expected importance score for the electronic document; and

selecting words from the electronic document with an importance score exceeding the expected importance score.

13. The system of claim 11 , wherein generating the one or more clusters of words comprises:

generating, for each respective word in the plurality of important words, a respective vector; and

clustering vectors that exceed a predefined threshold of similarity.

14. The system of claim 11 , wherein the passage encoder is a recurrent neural network (RNN).

Assignments (3)
SECURITY INTEREST Recorded Oct 1, 2025
From: MERATIVE US L.P.; MERGE HEALTHCARE INCORPORATED
To: TCG SENIOR FUNDING L.L.C., AS COLLATERAL AGENT
Reel/Frame 072808/0442 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2022
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: MERATIVE US L.P.
Reel/Frame 061496/0752 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2019
From: BULL, BRENDAN; FELT, PAUL LEWIS; HICKS, ANDREW
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 048271/0326 →