IP Library › Granted Patent US 11,914,583
Granted Patent B2
US 11,914,583 · App. 17/022,594 · Granted Feb 27, 2024

Utilizing regular expression embeddings for named entity recognition systems

Inventors: Jeremy Edward Goodsitt (Champaign, IL); Austin Grant Walters (Savoy, IL); Reza Farivar (Champaign, IL); Mark Louis Watson (Sedona, AZ); Anh Truong (Champaign, IL); Galen Rafferty (Mahomet, IL); Vincent Pham (Champaign, IL)
Assignee: Capital One Services, LLC
G06F16/243G06F16/3329G06F16/35G06N20/00G06F16/374
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,914,583
App. No.
17/022,594
Granted
Feb 27, 2024
Kind
B2
Abstract

Various embodiments are directed to a system that utilizes regular expression (regex) to recognize at least portions of characters, words, text, numbers, etc. in a structured or unstructured dataset, any patterns associated therewith, and/or similarities between the determined patterns. In examples, a regex-based pattern recognition platform may receive a dataset and determine whether at least a first regex pattern and a second regex pattern can be identified. The occurrences of the first and second regex patterns and the frequency of those occurrences may reveal something about the dataset itself or any patterns contained therein.

Claims (24)

1. A method comprising:

determining, via one or more processors, whether a dataset comprises unstructured text;

determining, via the one or more processors, that at least a portion of the unstructured text corresponds to a regex pattern of a regex list, wherein the regex pattern comprises at least one metacharacter, the at least one metacharacter associated with a non-literal meaning;

replacing, via the one or more processors, the portion of the unstructured text with an encoding that represents the regex pattern to generate a modified dataset; and

providing at least the modified dataset to at least one entity recognition system.

2. The method of claim 1 , wherein the portion of the unstructured text is a first character, the first character being an alphanumeric character or special character.

3. The method of claim 2 , wherein the encoding is a second character, the second character being an alphanumeric character or a special character.

4. The method of claim 1 , wherein the portion of the unstructured text is a string of two or more consecutive characters, the two or more consecutive characters of the strings being alphanumeric characters or special characters.

5. The method of claim 4 , wherein the encoding is a word of at least one character, the at least one character of the word being an alphanumeric character or a special character.

6. The method of claim 1 , further comprising:

determining, via the one or more processors, any false matches between the regex pattern of the regex list and the dataset;

refining, based on any determined false matches, the regex list via a machine learning model or a classification model.

7. At least one non-transitory computer-readable storage medium storing computer-readable program code executable by at least one processor to:

determine whether a dataset comprises unstructured text;

determine that at least a portion of the unstructured text corresponds to a regex pattern of a regex list, wherein the regex pattern comprises at least one metacharacter, the at least one metacharacter associated with a non-literal meaning;

replace the portion of the unstructured text with an encoding that represents the regex pattern to generate a modified dataset; and

provide at least the modified dataset to at least one entity recognition system.

8. The at least one non-transitory computer-readable storage medium of claim 7 , wherein the portion of the unstructured text is a first character, the first character being an alphanumeric character or special character.

9. The at least one non-transitory computer-readable storage medium of claim 8 , wherein the encoding is a second character, the second character being an alphanumeric character or a special character.

10. The at least one non-transitory computer-readable storage medium of claim 7 , wherein the portion of the unstructured text is a string of two or more consecutive characters, the two or more consecutive characters of the strings being alphanumeric characters or special characters.

11. The at least one non-transitory computer-readable storage medium of claim 10 , wherein the encoding is a word of at least one character, the at least one character of the word being an alphanumeric character or a special character.

12. The at least one non-transitory computer-readable storage medium of claim 7 , wherein the computer-readable program code further causes the at least one processor to:

determine any false matches between the regex pattern of the regex list and the dataset; and

refine, based on any determined false matches, the regex list via a machine learning model or a classification model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2020
From: GOODSITT, JEREMY EDWARD; WALTERS, AUSTIN GRANT; FARIVAR, REZA; WATSON, MARK LOUIS; TRUONG, ANH; RAFFERTY, GALEN; PHAM, VINCENT
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 053790/0143 →
Continuity (2)
Continuation 16549786 · Aug 23, 2019
Related Publication 20210056099A1 · Feb 25, 2021