IP Library › Granted Patent US 11,941,018
Granted Patent B2
US 11,941,018 · App. 16/904,298 · Granted Mar 26, 2024

Regular expression generation for negative example using context

Inventors: Michael Malak (Denver, CO); Luis E. Rivas (Denver, CO); Mark L. Kreider (Arvada, CO)
Assignee: Oracle International Corporation
G06F16/258G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,941,018
App. No.
16/904,298
Granted
Mar 26, 2024
Kind
B2
Abstract

Techniques for generated regular expressions are disclosed. In some embodiments, a regular expression generator may receive input data comprising one or more character sequences. The regular expression generator may convert character sequences into a sets of regular expression codes and/or span data structures. The regular expression generator may identify a longest common subsequence shared by the sets of regular expression codes and/or spans, and may generate a regular expression based upon the longest common subsequence. A negative example may be used to generate the regular expression. Context from the negative example may be determined in order to generate the regular expression.

Claims (62)

1. A method of generating regular expressions, comprising:

receiving, by a regular expression generator comprising one or more processors, a first selection comprising one or more positive character sequences in a first field of a data structure, each of the one or more positive character sequences corresponding to a positive example that is to be matched by a regular expression generated by the regular expression generator;

after receiving the first selection, generating, by the regular expression generator, a first regular expression, wherein the first regular expression matches the positive example;

in response to the generation of the first regular expression, displaying, by the regular expression generator, the first regular expression that is generated based on matching the positive example;

receiving, by the regular expression generator, a second selection comprising one or more negative character sequences in a second field of the data structure, each of the one or more negative character sequences corresponding to a negative example that is not to be matched by the regular expression generated by the regular expression generator;

in response to receiving the second selection, determining a context of the selected one or more negative character sequences in the second field of the data structure corresponding to the negative example;

updating, in real-time, the first regular expression that was generated to match the positive example to include the context that was determined based on the one or more negative character sequences; and

displaying, by the regular expression generator, the updated first regular expression.

2. The method according to claim 1 , wherein the receiving the first selection comprises receiving, via a user interface, a selection of the one or more positive character sequences in a first data cell of a data set.

3. The method according to claim 2 , further comprising automatically selecting, by the regular expression generator, character sequences in a plurality of data cells in the data set corresponding to the first selection comprising one or more positive character sequences.

4. The method according to claim 3 , wherein the receiving the second selection comprises receiving, via the user interface, the selection of the one or more negative character sequences in a second data cell of the data set.

5. The method according to claim 4 , further comprising automatically selecting, by the regular expression generator, character sequences in the plurality of data cells in the data set corresponding to the second selection comprising one or more negative character sequences.

6. The method according to claim 3 , wherein the first selection is highlighted in a first highlight format and the second selection is highlighted in a second highlight format that is different from the first highlight format.

7. The method according to claim 6 , wherein the determining the context of the one or more negative character sequences corresponding to the negative example comprises:

identifying an embedded highlighting location of the second selection;

determining context from data to a left of the embedded highlighting location of the second selection; and

determining context from data to a right of the embedded highlighting location of the highlighted second selected.

8. The method according to claim 7 , wherein the determining the context of the one or more negative character sequences corresponding to the negative example further comprises:

filtering the character sequences in the plurality of data cells in the data set corresponding to the first selection comprising the one or more negative character sequences that were automatically selected based on the determined context from data to the left of the embedded highlighting location and based on the determined context from data to the right of the embedded highlighting location; and

removing the filtered character sequences from the selected character sequences in the plurality of data cells in the data set corresponding to the selected one or more negative character sequences.

9. The method according to claim 8 , wherein the determining the context from data to the left of an embedded highlighting location comprises identifying a first span to the left of the embedded highlighting location; and

wherein filtering the character sequences in the plurality of data cells in the data set corresponding to the selected one or more negative character sequences further comprises identifying spans in the character sequences in the plurality of data cells corresponding to the selected one or more negative character sequences that do not match the first span to the left of the embedded highlighting location.

10. The method according to claim 9 , wherein the determining the context from data to the left of an embedded highlighting location further comprises identifying a second span to the left of the embedded highlighting; and

wherein filtering the character sequences in the plurality of data cells in the data set corresponding to the selected one or more negative character sequences further comprises identifying spans in the character sequences in the plurality of data cells corresponding to the selected one or more negative character sequences that do not match the second span to the left of the embedded highlighting location.

11. The method according to claim 7 , wherein the determining the context from data to the right of an embedded highlighting location comprises identifying a first span to the right of the embedded highlighting location; and

wherein filtering the character sequences in the plurality of data cells in the data set corresponding to the second selection comprising one or more negative character sequences further comprises identifying spans in the character sequences in the plurality of data cells corresponding to the second selection comprising one or more negative character sequences that do not match the first span to the right of the embedded highlighting location.

12. A regular expression generator server computer comprising:

a processor;

a memory;

a computer readable medium coupled to the processor, the computer readable medium storing instructions executable by the processor for implementing a method comprising:

receiving, by a regular expression generator comprising one or more processors, a first selection comprising one or more positive character sequences in a first field of a data structure, each of the one or more positive character sequences corresponding to a positive example that is to be matched by a regular expression generated by the regular expression generator;

after receiving the first selection, generating, by the regular expression generator, a first regular expression, wherein the first regular expression matches the positive example;

in response to the generation of the first regular expression, displaying, by the regular expression generator, the first regular expression that is generated based on matching the positive example;

receiving, by the regular expression generator, a second selection comprising one or more negative character sequences in a second field of the data structure, each of the one or more negative character sequences corresponding to a negative example that is not to be matched by the regular expression generated by the regular expression generator;

in response to receiving the second selection, determining a context of the selected one or more negative character sequences in the second field of the data structure corresponding to the negative example;

updating, in real-time, the first regular expression that was generated to match the positive example to include the context that was determined based on the one or more negative character sequences; and

displaying, by the regular expression generator, the updated first regular expression.

13. The server computer according to claim 12 , wherein the receiving the first selection comprises receiving, via a user interface, a selection of the one or more positive character sequences in a first data cell of a data set.

14. The server computer according to claim 13 , wherein the first selection is highlighted in a first highlight format and the second selection is highlighted in a second highlight format that is different from the first highlight format.

15. The server computer according to claim 13 , wherein the determining the context of the one or more negative character sequences corresponding to the negative example comprises:

identifying an embedded highlighting location of the second selection;

determining context from data to a left of the embedded highlighting location of the second selection; and

determining context from data to a right of the embedded highlighting location of the highlighted second selected.

16. The server computer according to claim 15 , wherein the determining the context of the one or more negative character sequences corresponding to the negative example further comprises:

filtering the one or more negative character sequences in a plurality of data cells in the data set corresponding to the second selection comprising the one or more negative character sequences that were automatically selected based on the determined context from data to the left of the embedded highlighting location and based on the determined context from data to the right of the embedded highlighting location; and

removing the filtered character sequences from the selected character sequences in the plurality of data cells in the data set corresponding to the second selection comprising the one or more negative character sequences.

17. A non-transitory computer readable medium including instructions configured to cause one or more processors to perform operations comprising:

receiving, by a regular expression generator comprising one or more processors, a first selection comprising one or more positive character sequences in a first field of a data structure, each of the one or more positive character sequences corresponding to a positive example that is to be matched by a regular expression generated by the regular expression generator;

after receiving the first selection, generating, by the regular expression generator, a first regular expression, wherein the first regular expression matches the positive example;

in response to the generation of the first regular expression, displaying, by the regular expression generator, the first regular expression that is generated based on matching the positive example;

receiving, by the regular expression generator, a second selection comprising one or more negative character sequences in a second field of the data structure, each of the one or more negative character sequences corresponding to a negative example that is not to be matched by the regular expression generated by the regular expression generator;

in response to receiving the second selection, determining a context of the selected one or more negative character sequences in the second field of the data structure corresponding to the negative example;

updating, in real-time, the first regular expression that was generated to match the positive example to include the context that was determined based on the one or more negative character sequences; and

displaying, by the regular expression generator, the first regular expression.

18. The computer readable medium according to claim 17 , wherein the receiving the first selection comprises receiving, via a user interface, a selection of the one or more positive character sequences in a first data cell of a data set.

19. The computer readable medium according to claim 18 , wherein the determining the context of the one or more negative character sequences corresponding to the negative example comprises:

identifying an embedded highlighting location of the second selection;

determining context from data to a left of the embedded highlighting location of the second selection; and

determining context from data to a right of the embedded highlighting location of the highlighted second selected.

20. The computer readable medium according to claim 19 , wherein the determining the context of the one or more negative character sequences corresponding to the negative example further comprises:

filtering the one or more negative character sequences in a plurality of data cells in the data set corresponding to the second selection comprising the one or more negative character sequences that were automatically selected based on the determined context from data to a left of the embedded highlighting location and based on the determined context from data to a right of the embedded highlighting location; and

removing the filtered character sequences from the selected character sequences in the plurality of data cells in the data set corresponding to the second selection comprising the one or more negative character sequences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 17, 2020
From: MALAK, MICHAEL; RIVAS, LUIS E.; KREIDER, MARK L.
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 052967/0956 →
Continuity (5)
Continuation In Part 16438325 · Jun 11, 2019
Provisional Application 62865797 · Jun 24, 2019
Provisional Application 62749001 · Oct 22, 2018
Provisional Application 62684498 · Jun 13, 2018
Related Publication 20200320092A1 · Oct 8, 2020