IP Library Granted Patent US 11,138,377
Granted Patent B2
US 11,138,377 · App. 16/730,122 · Granted Oct 5, 2021

Automated document analysis comprising company name recognition

Inventors: David A. Cook (Barrington, IL); Andrzej H. Jachowicz (Tower Lakes, IL); Phillip Karl Jones (Bartlett, IL)
Assignee: Freedin Solutions Group, LLC
G06F40/295G06F3/0481G06F40/205G06F40/232G06F40/247G06F40/284G06F40/289G06F40/30G06F40/40G06F40/106G06F40/253
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,138,377
App. No.
16/730,122
Granted
Oct 5, 2021
Kind
B2
Abstract

At least two processing device-implemented company name recognition components, operating upon a body of text in a document, identify at least one company name occurrence in the body of text based at least in part on a company identifier list. The company name recognition techniques implemented by each of the at least two company name recognition components are different from each other. The at least one company name occurrence is used to update the company identifier list. The updated company identifier list is then used by the at least two company name recognition components to identify at least one additional name occurrence in the same body of text. This process of repeatedly identifying occurrences of company names in the body of text and updating the company identifier list is performed until such time that no further company name occurrences are identified in the body of text.

Claims (81)

1. A method for performing, by at least one processing device, automated document analysis of a document comprising a body of text, the method comprising:

accessing a token among a sequence of tokens constituting the body of the text in the document;

comparing the token with company names included in a company identifier list;

determining, based on the comparison, a potential match between the token and at least one company name in the company identifier list;

analyzing, based on a determination of no match, whether the token constitutes a possessive form of the at least one company name;

determining, based on a determination of no match, whether the token is a synonym or a substitute of the at least one company name;

further determining, based on a determination of no match, whether the token includes a punctuation included in the at least one company name; and

establishing a match with the at least one company name when the token constitutes at least one of a possessive form, a synonym, a substitute, or a punctuation.

2. The method of claim 1 , further comprising:

determining that the token constitutes a possessive form of the at least one company name; and

indicating, based on an identification, a determination of a match.

3. The method of claim 1 , further comprising:

determining that the token is a synonym or a substitute of the at least one company name; and

indicating, based on an identification, a determination of a match.

4. The method of claim 1 , further comprising

determining that the token includes a punctuation included in the at least one company name; and

indicating, based on an identification, a determination of a match.

5. The method of claim 1 , further comprising:

categorizing, based on the match, the token according to the at least one company name.

6. The method of claim 1 , further comprising:

determining if the match is part of a potential larger match or indicative of a presence of an additional company name occurrence nearby.

7. The method of claim 1 , further comprising:

determining whether the at least one company name indicates an additional company name occurrence.

8. A system comprising:

at least one processing device; and

memory operatively connected to the at least one processing device, the memory comprising executable instructions that when executed by the at least one processing device cause the at least one processing device to:

access a token among a sequence of tokens constituting a body of a text in a document;

compare the token with company names included in a company identifier list;

determine, based on the comparison, a potential match between the token and at least one company name in the company identifier list;

analyze, based on a determination of no match, whether the token constitutes a possessive form of the at least one company name;

determine, based on a determination of no match, whether the token is a synonym or a substitute of the at least one company name;

further determine, based on a determination of no match, whether the token includes a punctuation included in the at least one company name; and

establish a match with the at least one company name when the token constitutes at least one of a possessive form, a synonym, a substitute, or a punctuation.

9. The system of claim 8 , further comprising executable instructions that, when executed by the at least one processing device, cause the at least one processing device to:

determine that the token constitutes a possessive form of the at least one company name; and indicate, based on an identification, a determination of a match.

10. The system of claim 8 , further comprising executable instructions that, when executed by the at least one processing device, cause the at least one processing device to:

determine that the token is a synonym or a substitute of the at least one company name; and indicate, based on an identification, a determination of a match.

11. The system of claim 8 , further comprising executable instructions that, when executed by the at least one processing device, cause the at least one processing device to:

determine that the token includes a punctuation included in the at least one company name; and indicate, based on an identification, a determination of a match.

12. The system of claim 8 , further comprising executable instructions that, when executed by the at least one processing device, cause the at least one processing device to:

categorize, based on the match, the token according to the at least one company name.

13. The system of claim 8 , further comprising executable instructions that, when executed by the at least one processing device, cause the at least one processing device to:

determine that the match is part of a potential larger match or indicative of a presence of an additional company name occurrence nearby.

14. The system of claim 8 , further comprising executable instructions that, when executed by the at least one processing device, cause the at least one processing device to:

determine whether the at least one company name indicates an additional company name occurrence.

15. A non-transitory computer readable medium comprising executable instructions that when executed by at least one processing device cause the at least one processing device to perform automated document analysis of a document comprising a body of text in which the at least one processing device is caused to:

access a token among a sequence of tokens constituting the body of the text in the document;

compare the token with company names included in a company identifier list;

determine, based on the comparison, a potential match between the token and at least one company name in the company identifier list;

analyze, based on a determination of no match, whether the token constitutes a possessive form of the at least one company name;

determine, based on a determination of no match, whether the token is a synonym or a substitute of the at least one company name;

further determine, based on a determination of no match, whether the token includes a punctuation included in the at least one company name; and

establish a match with the at least one company name when the token constitutes at least one of a possessive form, a synonym, a substitute, or a punctuation.

16. The non-transitory computer readable medium of claim 15 , wherein those executable instructions are further operative to:

determine that the token constitutes a possessive form of the at least one company name; and indicate, based on an identification, a determination of a match.

17. The non-transitory computer readable medium of claim 15 , wherein those executable instructions are further operative:

to determine that the token is a synonym or a substitute of the at least one company name; and indicate, based on an identification, a determination of a match.

18. The non-transitory computer readable medium of claim 15 , wherein those executable instructions are further operative to:

determine that the token includes a punctuation included in the at least one company name; and indicate, based on an identification, a determination of a match.

19. A method for performing, by at least one processing device, automated document analysis of a document comprising a body of text, the method comprising:

accessing a token among a sequence of tokens constituting the body of the text in the document;

comparing the token with company names included in a company identifier list;

determining, based on the comparison, a potential match between the token and at least one company name in the company identifier list;

analyzing, based on a determination of no match, whether the token constitutes a possessive form of the at least one company name;

determining, based on a determination of no match, whether the token is a synonym or a substitute of the at least one company name;

further determining, based on a determination of no match, whether the token includes a punctuation included in the at least one company name;

establishing a match with the at least one company name when the token constitutes at least one of a possessive form, a synonym or a substitute, and a punctuation;

determining if the match is part of a potential larger match or indicative of a presence of an additional company name occurrence nearby; and

prioritizing, when the match is determined as part of a potential larger match, the potential larger match over the match.

20. A system comprising:

at least one processing device; and

memory operatively connected to the at least one processing device, the memory comprising executable instructions that when executed by the at least one processing device cause the at least one processing device to:

access a token among a sequence of tokens constituting a body of a text in a document;

compare the token with company names included in a company identifier list;

determine, based on the comparison, a potential match between the token and at least one company name in the company identifier list;

analyze, based on a determination of no match, whether the token constitutes a possessive form of the at least one company name;

determine, based on a determination of no match, whether the token is a synonym or a substitute of the at least one company name;

further determine, based on a determination of no match, whether the token includes a punctuation included in the at least one company name;

establish a match with the at least one company name when the token constitutes at least one of a possessive form, a synonym or a substitute, and a punctuation;

determine that the match is part of a potential larger match or indicative of a presence of an additional company name occurrence nearby; and

prioritize, responsive to a determination that the match is determined as part of a potential larger match, the potential larger match over the match.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2020
From: COOK, DAVID A.; JACHOWICZ, ANDRZEJ H.; JONES, PHILLIP K.
To: FREEDOM SOLUTIONS GROUP, LLC D/B/A MICROSYSTEMS
Reel/Frame 052802/0643 →
Continuity (4)
Continuation 16375845 · Apr 4, 2019
Continuation 15249374 · Aug 27, 2016
Provisional Application 62211097 · Aug 28, 2015
Related Publication 20200134261A1 · Apr 30, 2020