IP Library Granted Patent US 11,562,008
Granted Patent B2
US 11,562,008 · App. 15/334,009 · Granted Jan 24, 2023

Detection of entities in unstructured data

Inventor: Samuel Roy Carter (Pleasanton, CA)
Assignee: MICRO FOCUS LLC
G06F16/313
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,562,008
App. No.
15/334,009
Filed
Oct 25, 2016
Granted
Jan 24, 2023
Kind
B2
Art Unit
2152
USPC
707/751
Abstract

Examples herein involve detection of entities in unstructured data. Terms are extracted from unstructured data. Entities scores for the terms are calculated using information from a name probability source, a known entity database, and historical context information. The entity scores indicate a probability that the respective terms refer to entities. The presence of detected entities are indicated based on the entity scores.

Claims (68)

1. A method comprising:

extracting terms from a corpus of unstructured data, the terms including a first term;

calculating a first entity score that scores the first term based on a frequency of occurrence of the first term in a name probability source and to indicate a probability that the first term refers to an entity of interest, the first term being extracted from the corpus of unstructured data, including:

applying a first weighting factor to data received from the name probability source; and

applying a second weighting factor to the probability that the first term refers to an entity of interest;

determining whether the first term is found in a known entity database that includes a plurality of known entity names;

adjusting the first entity score based on whether the first term is found in the known entity database, wherein the probability that the first term refers to the entity of interest is increased if the first term is found in the known entity database;

determining that the entity of interest has been detected in the corpus of unstructured data based on a comparison of the first entity score to a threshold entity score; and

indicating the entity of interest has been detected in the corpus of unstructured data.

2. The method as defined in claim 1 , wherein the name probability source indicates a first frequency that the first term refers to a name in a first language or first locality and a second frequency that the first term refers to the name in a second language or second locality, and wherein determining the frequency of occurrence of the first term in the name probability source comprises determining that the first frequency should be used.

3. The method as defined in claim 1 , wherein the known entity database comprises a directory of known entities or an organization structure of known entities.

4. The method as defined in claim 1 , further comprising:

from the corpus of unstructured data or outside the corpus of unstructured data;

determining that the first term is an alternate term for the entity of interest based on historical context information; and

adjusting the first entity score based on determining that the first term is the alternate term.

5. The method as defined in claim 1 , wherein determining whether the first term is found in the known entity database comprises:

determining whether a portion of the first term matches an identifier of an entity in the known entity database.

6. The method as defined in claim 1 , further comprising:

obtaining historical context information from the corpus of unstructured data or outside the corpus of unstructured data;

determining that the first term is associated with the entity of interest based on content of the historical context information; and

adjusting the first entity score based on the determination that the first term is associated with the entity of interest.

7. The method as defined in claim 6 , wherein the first term does not match a term for an entity in the known entity database.

8. A non-transitory machine readable storage medium comprising instructions that, when executed, cause a machine to:

extract a term from a corpus of unstructured data;

calculate an entity score that scores the term based on a frequency of occurrence of the term in a name probability source and to indicate a probability that the term is a name or non-name, the term being extracted from the corpus of unstructured data, wherein the instructions when executed further cause the machine to:

apply a first weighting factor to data received from the name probability source; and

apply a second weighting factor to the probability that the term is a name or non-name;

determine whether the term is found in a known entity database of known entities associated with a matter;

adjust the entity score based on whether the term is found in the known entity database;

obtain historical context information associated with the matter;

adjust the entity score based on information from the historical context information; and

indicate a presence of an entity in the corpus of unstructured data based on the entity score.

9. The non-transitory machine readable medium of claim 8 , wherein the instructions when executed, further cause the machine to:

indicate the entity in the corpus of unstructured data is an entity of interest based on the entity score, the entity of interest comprising an entity associated with the matter.

10. The non-transitory machine readable medium of claim 8 , wherein the entity in the corpus of unstructured data is not referred to in the known entity database.

11. An apparatus comprising:

a processor to:

extract terms from unstructured data, the terms including a first term;

calculate a first entity score that scores the first term based on a frequency of occurrence of the first term in a name probability source, the first term being extracted from the unstructured data, and wherein the processor is further to:

apply a first weighting factor to data received from the name probability source; and

apply a second weighting factor to the probability that the first term refers to an entity of interest;

determine whether a portion of the first term matches content in a known entity database that includes a plurality of known entity names;

adjust the first entity score based on whether a portion of the first term matches content in the known entity database;

identify an entity in the unstructured data based on the first entity score; and

indicate a presence of the identified entity in the unstructured data.

12. The apparatus of claim 11 , wherein the processor is to use a threshold entity score to determine whether the first term refers to the identified entity, wherein the identified entity is not in the known entity database.

13. The apparatus of claim 11 , wherein the processor is to use a threshold entity score to determine whether the first term refers to the identified entity, the identified entity being an entity of interest associated with a matter.

14. The apparatus of claim 13 , wherein the entity of interest is not in the known entity database.

15. The apparatus of claim 11 , wherein the processor is to indicate relationship information between the identified entity and another entity detected in the unstructured data.

16. The apparatus of claim 11 , wherein the processor is to:

obtain historical context information associated with a matter; and

adjust the first entity score based on the historical context information, and wherein the name probability source comprises a probability that the terms are a name or non-name and the known entity database comprises known entities associated with the matter.

17. The apparatus of claim 11 , wherein the processor is to:

obtain historical context information;

determine that the first term is an alternate term for the identified entity based on the historical context information; and

adjust the first entity score based on the determination that the first term is the alternate term for the identified entity.

18. The apparatus of claim 17 , wherein to determine that the first term is an alternate term for the identified entity, the processor is to:

determine that the first term is a nickname for the identified entity based on text from the historical context information.

19. The apparatus of claim 11 , wherein the processor is to:

obtain historical context information;

identify, from the historical context information, one or more communications between the identified entity and another entity;

generate a social graph that associates the identified entity and the another entity based on the one or more communications; and

adjust the first entity score based on the social graph.

20. The apparatus of claim 11 , wherein the processor is to:

obtain historical context information;

identify, from the historical context information, a contextual term;

determine that the contextual term relates to a first entity of interest; and

determine that the contextual term refers to a second entity of interest based on the determination that the contextual term relates to the first entity of interest.

Assignments (8)
RELEASE OF SECURITY INTEREST REEL/FRAME 044183/0577 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC (F/K/A ENTIT SOFTWARE LLC)
Reel/Frame 063560/0001 →
RELEASE OF SECURITY INTEREST REEL/FRAME 044183/0718 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC (F/K/A ENTIT SOFTWARE LLC); BORLAND SOFTWARE CORPORATION; MICRO FOCUS (US), INC.; SERENA SOFTWARE, INC; ATTACHMATE CORPORATION; MICRO FOCUS SOFTWARE INC. (F/K/A NOVELL, INC.); NETIQ CORPORATION
Reel/Frame 062746/0399 →
CHANGE OF NAME Recorded Aug 8, 2019
From: ENTIT SOFTWARE LLC
To: MICRO FOCUS LLC
Reel/Frame 050004/0001 →
SECURITY INTEREST Recorded Oct 11, 2017
From: ATTACHMATE CORPORATION; BORLAND SOFTWARE CORPORATION; NETIQ CORPORATION; MICRO FOCUS (US), INC.; MICRO FOCUS SOFTWARE, INC.; ENTIT SOFTWARE LLC; ARCSIGHT, LLC; SERENA SOFTWARE, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 044183/0718 →
SECURITY INTEREST Recorded Oct 11, 2017
From: ENTIT SOFTWARE LLC; ARCSIGHT, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 044183/0577 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2017
From: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
To: ENTIT SOFTWARE LLC
Reel/Frame 042746/0130 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 5, 2016
From: CARTER, SAMUEL ROY
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 040520/0541 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2016
From: CARTER, SAMUEL ROY
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 040123/0195 →
Continuity (1)
Related Publication 20180113931A1 · Apr 26, 2018
Cited By (1)
US 12,619,957