IP Library › Granted Patent US 11,366,966
Granted Patent B1
US 11,366,966 · App. 16/512,652 · Granted Jun 21, 2022

Named entity recognition and disambiguation engine

Inventors: Michael Ramsey (Arlington, VA); Gabriel Parlato-Altay (Washington, DC); Aron Szanto (New York, NY); Anthony Liu (Cambridge, MA); Sireesh Gururaja (Cambridge, MA); Benjamin Hsu (Alexandria, VA); Domenic Puzio (Arlington, VA)
Assignee: Kensho Technologies, LLC
G06F40/295G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,366,966
App. No.
16/512,652
Filed
Jul 16, 2019
Granted
Jun 21, 2022
Kind
B1
Art Unit
2652
USPC
704/9
Abstract

Methods, and systems for named entity recognition and disambiguation. One of the methods includes: receiving a document; extracting a plurality of mentions within the document; clustering the plurality of mentions to produce at least one cluster; identifying a plurality of candidate entities for the cluster; pairing each candidate entity and the cluster to create a plurality of candidate entity-cluster pairings; generating features for each candidate entity-cluster pairing; and selecting an entity, from the plurality of candidate entities, for the cluster based at least in part on features of a candidate entity pairing.

Claims (53)

1. A method comprising:

receiving a document;

extracting a plurality of mentions within the document;

clustering the plurality of mentions to produce a cluster;

identifying a plurality of candidate entities for the cluster;

pairing each candidate entity and the cluster to create a plurality of candidate entity-cluster pairings;

generating features for the plurality of candidate entity-cluster pairings; and

selecting an entity, from the plurality of candidate entities, for the cluster based at least in part on the generated features of the plurality of candidate entity-cluster pairings, wherein selecting the entity comprises:

providing the generating features as an input to a machine learning model that has been trained to generate output data identifying an entity based on processing of a document having only a subset of mentions labeled of a plurality of mentions;

processing, by the machine learning model, the provided features through each layer of the machine learning model to generate output data identifying an entity; and

obtaining the output data generated, by the machine learning model, wherein the obtained output data identifies the selected entity.

2. The method of claim 1 , wherein extracting a plurality of mentions within the document comprises

extracting a number of mentions within a document; and

breaking the document up into at least one sub-document, the sub-document comprising less text than the document, when the number of mentions within the document exceeds a threshold.

3. The method of claim 1 , wherein the cluster has a cluster representative.

4. The method of claim 3 , wherein identifying a plurality of candidate entities for the cluster further comprises filtering the plurality of candidates to just those with aliases having a similarity score for similarity to a cluster representative above a threshold.

5. The method of claim 1 , wherein clustering the plurality of mentions to produce at least one cluster further comprises clustering using a coreferencer.

6. The method of claim 5 , wherein the coreferencer clusters the plurality of mentions using at least one of a glossary or a string similarity score.

7. The method of claim 1 , wherein the method further comprises performing the extracting, clustering, identifying and pairing steps for a plurality of clusters and wherein generating features for each candidate entity-cluster pairing comprises generating a coherence score, wherein the coherence score indicates how coherent the candidate entity is given other entities that have been paired with clusters.

8. The method of claim 7 , wherein the coherence score comprises entity similarities.

9. The method of claim 1 , wherein each candidate entity-cluster pairing has a candidate entity and wherein generating features for a candidate entity-cluster pairing comprises obtaining a probability that an arbitrary text span of a mention in the cluster is referring to the candidate entity in a candidate entity-cluster pairing.

10. The method of claim 1 , wherein each candidate entity-cluster pairing has a candidate entity and wherein generating features for a candidate entity-cluster pairing comprises obtaining a conditional probability that a text string of a mention of the cluster is used to refer to the candidate entity.

11. The method of claim 1 , wherein each candidate entity-cluster pairing has a candidate entity and wherein generating features for each candidate entity-cluster pairing comprises a conditional probability that the candidate entity is referenced by a text string of a mention of the cluster.

12. The method of claim 1 , wherein each candidate entity-cluster pairing has a candidate entity and wherein generating features for each candidate entity-cluster pairing comprises obtaining a topic model vector similarity score between the document and text content for the candidate entity.

13. The method of claim 1 , wherein each candidate entity-cluster pairing has a candidate entity and wherein generating features for each candidate entity-cluster pairing comprises generating a similarity score between a predicted type of a mention in the cluster and a type of the candidate entity for a candidate entity-cluster pairing.

14. The method of claim 1 , wherein each candidate entity-cluster pairing has a candidate entity and wherein generating features for each candidate entity-cluster pairing comprises generating a similarity score between embeddings of the document of a mention of the cluster and textual content of a candidate entity.

15. The method of claim 1 , wherein the method further comprises using individual classes to handle access to documents, mentions, clusters and candidates.

16. The method of claim 1 , wherein the method further comprises tagging the document with the selected entity.

17. The method of claim 1 , wherein at least one of the plurality of candidate entity-cluster pairings comprises a candidate-entity-mention pairing.

18. The method of claim 1 , wherein extracting a plurality of mentions within the document comprises extracting a plurality of mentions within the document using at least one of named entity recognition, a glossary or a gazetteer.

19. A system comprising:

one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving a document;

extracting a plurality of mentions within the document;

clustering the plurality of mentions to produce a cluster;

identifying a plurality of candidate entities for the cluster;

pairing each candidate entity and the cluster to create a plurality of candidate entity-cluster pairings;

generating features for the plurality of candidate entity-cluster pairings; and

selecting an entity, from the plurality of candidate entities, for the cluster based at least in part on the generated features of the plurality of candidate-entity cluster pairings, wherein selecting the entity comprises:

providing the generating features as an input to a machine learning model that has been trained to generate output data identifying an entity based on processing of a document having only a subset of mentions labeled of a plurality of mentions;

processing, by the machine learning model, the provided features through each layer of the machine learning model to generate output data identifying an entity; and

obtaining the output data generated, by the machine learning model, wherein the obtained output data identifies the selected entity.

20. A method comprising:

obtaining a document;

extracting a plurality of mentions within the document;

clustering the plurality of mentions to produce a cluster;

identifying a plurality of candidate entities for the cluster;

pairing each candidate entity with the cluster to create a plurality of candidate entity-cluster pairings, wherein each candidate entity-cluster pairing has a candidate entity;

generating features for the plurality of candidate entity-cluster pairings, wherein the features include a coherence score and the coherence score indicates how coherent the candidate entity is given a set of mentions in the document for the cluster of the candidate entity-cluster pairing; and

selecting an entity, from the plurality of candidate entities, for each cluster based at least in part on the generated features of the plurality of candidate entity pairings, wherein selecting the entity comprises:

providing the generating features as an input to a machine learning model that has been trained to generate output data identifying an entity based on processing of a document having only a subset of mentions labeled of a plurality of mentions;

processing, by the machine learning model, the provided features through each layer of the machine learning model to generate output data identifying an entity; and

obtaining the output data generated, by the machine learning model, wherein the obtained output data identifies the selected entity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 24, 2019
From: RAMSEY, MICHAEL; PARLATO-ALTAY, GABRIEL; SZANTO, ARON; LIU, ANTHONY; GURURAJA, SIREESH; HSU, BENJAMIN; PUZIO, DOMENIC
To: KENSHO TECHNOLOGIES, LLC
Reel/Frame 049849/0013 →
Cited By (12)
US 12,236,195 US 12,323,262 US 12,468,765 US 12,499,163 US 12,518,106 US 12,547,631 US 12,596,735 US 12,619,825 US 12,645,670 US 12,670,505 US 12,724,968 US 12,730,966