IP Library › Granted Patent US 12,361,278
Granted Patent B2
US 12,361,278 · App. 17/214,294 · Granted Jul 15, 2025

Automated generation and integration of an optimized regular expression

Inventors: Vinu Varghese (Bangalore, IN); Nirav Jagdish Sampat (Mumbai, IN); Balaji Janarthanam (Chennai, IN); Anil Kumar (Bangalore, IN); Shikhar Srivastava (Bangalore, IN); Kunal Jaiwant Kharsadia (Pune, IN); Saran Prasad (Nagar, IN)
Assignee: ACCENTURE GLOBAL SOLUTIONS LIMITED
G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,278
App. No.
17/214,294
Filed
Mar 26, 2021
Granted
Jul 15, 2025
Kind
B2
Art Unit
2121
USPC
706/25
Abstract

A system for obtaining optimized regular expression may convert a received input data into a plurality of embeddings. The system may receive a generated regular expression corresponding to the plurality of embeddings, wherein the generated regular expression is at least one of an existing regular expression from a database and a newly generated regular expression. The system may parse the generated regular expression into a plurality of sub-blocks. The system may classify the plurality of sub-blocks to obtain a plurality of classified sub-blocks. The system may evaluate a quantifier class for each classified sub-block to identify a corresponding computationally expensive class. The system may perform an iterative analysis to obtain a plurality of optimized sub-blocks associated with a minimum computation time. The system may combine the plurality of optimized sub-blocks to obtain the optimized regular expression.

Claims (50)

1. A system comprising:

a decoder is configured to:

receive an input data corresponding to a programming language;

obtain a plurality of embeddings using Embeddings from Language Models (ELMO) which facilitates computing word vectors on top of a two-layered bidirectional language model (biLM) which includes forward pass and backward pass;

convert the input data into the plurality of embeddings, wherein each embedding corresponds to a vector based representation of an attribute corresponding to the input data;

a processor including an artificial intelligence (AI) engine to generate an optimized regular expression (Regex) based on at least one of the type of the programming language and a corresponding library in the database, wherein the AI engine comprises a learning model associated with the AI engine for reinforced learning, wherein the learning model of the AI engine is automatically updated based on an information of the corresponding library in the database, the processor comprising:

a Regex optimizer to:

receive a generated regular expression corresponding to the plurality of embeddings, wherein the generated regular expression is at least one of an existing regular expression from a database and a newly generated regular expression from a Regex generator associated with the processor;

parse the generated regular expression into a plurality of sub-blocks;

classify, using a neural network based classifier associated with the AI engine, the plurality of sub-blocks to obtain a plurality of classified sub-blocks;

evaluate, using the neural network based classifier, a quantifier class for each classified sub-block to identify a corresponding computationally expensive class, wherein the neural network based classifier includes an input layer, multiple hidden layers and an output layer;

perform, using at least two reinforcement learning based deep neural network associated with the AI engine, an iterative analysis based on actor-critic algorithm to obtain a plurality of optimized sub-blocks associated with a minimum computation time, wherein one of the at least two reinforcement learning pertains to an actor and the another reinforcement learning pertains to a critic for executing the actor-critic algorithm, wherein the iterative analysis comprises replacing the quantifier class with a substitute class and performing iterative computation cycles based on an updated learning model for evaluating the corresponding computation time with respect to the substitute class by using a replace function corresponding to the programming language, wherein evaluating the corresponding computation time comprises:

verifying if computation loops of the iterative computation cycles corresponding to the substitute class exceed a pre-defined threshold; and

discarding the computation loops and identifying a new substitute class that requires minimum time or has minimum computation loops if the computation time for the substitute class exceeds the pre-defined threshold; and

combine the plurality of optimized sub-blocks to obtain the optimized regular expression.

2. The system as claimed in claim 1 , wherein the system comprises a simulator coupled to the processor, the simulator to:

perform simulation to generate, based on the optimized regular expression, an output in the form of filtered results comprising an endorsed result corresponding to at least value of corresponding processing time, wherein the system facilitates automated integration of the endorsed result within an application module pertaining to the programming language.

3. The system as claimed in claim 2 , wherein the simulation comprises generation of a plurality of test cases and execution of the test cases based on the optimized regular expression received from the Regex optimizer wherein the simulation includes key performance indicators (KPI) of the optimized regular expression including calculation of at least one of a performance and number of matches based on a file type of the input data.

4. The system as claimed in claim 3 , wherein the simulator generates the plurality of test cases based on at least one of a file type, a file size, a type of the programming language, and an optimization type corresponding to the input data, and wherein the plurality of test cases are executed, using the AI engine, based on the optimized regular expression, to obtain the output, wherein the system facilitates evaluating the processing time for each execution based on which the endorsed result is selected.

5. The system as claimed in claim 1 , wherein the system comprises the (Regex) generator to generate the newly generated regular expression using the AI engine.

6. The system as claimed in claim 1 , wherein the Regex generator comprises a transformer architecture including a trained model utilizing a softmax function to generate the generated regular expression in the programming language.

7. The system as claimed in claim 1 , wherein the system comprises a pre-execution recommender coupled to the decoder and the processor, the pre-execution recommender to:

generate, using the plurality of embeddings, an automated recommendation pertaining to a requirement for generating the optimized regular expression, wherein the automated recommendation pertains to a recommendation of the existing regex in the database, wherein based on the recommended existing regex, the AI engine provides at least one of a negative recommendation to perform the automated generation of the optimized regular expression and a positive recommendation to utilize the existing regular expression.

8. The system as claimed in claim 7 , wherein the automated recommendation is based on verification of one or more parameters of the plurality of embeddings with pre-stored parameters in the database.

9. The system as claimed in claim 1 , wherein the input data corresponds to at least one of a plain language text, a file type, a file size, a type of the programming language, an existing regular expression, and an application content, wherein the application content includes a uniform resource locator, and wherein the plain language text is English language text.

10. The system as claimed in claim 1 , wherein the decoder executes a pre-trained artificial intelligence model to obtain the plurality of embeddings.

11. The system as claimed in claim 1 , wherein the Regex optimizer parses the generated regex by using a split function corresponding to the programming language, and wherein the Regex optimizer replaces the quantifier class with the substitute class by using a replace function corresponding to the programming language.

12. The system as claimed in claim 1 , wherein the optimization is performed by an optimizer algorithm associated with the Regex optimizer, wherein the iterative analysis is performed using reinforcement learning embedded within the reinforcement learning based deep neural network, and wherein the reinforced learning comprises the learning model associated with the Al engine.

13. The system as claimed in claim 12 , wherein information pertaining to the optimized regular expression is stored in the database such that based on the stored information, the learning model of the AI engine is automatically updated, and wherein based on the updated learning model, the Regex optimizer automatically prioritizes and selects a suitable library for each sub-block during the iterative computation cycles, such that number of required iterative computation cycles reduces by using the updated learning model.

14. The system as claimed in claim 1 , wherein the corresponding computation time in the iterative analysis is evaluated to verify if computation loops of the iterative computation cycles corresponding to the substitute class exceeds the pre-defined threshold, wherein the computation loops are discarded upon exceeding the pre-defined threshold and the new substitute class is evaluated to identify the substitute class that requires minimum time for the corresponding computation loops to obtain the optimized sub-blocks including minimum computation loops, the computation loops are discarded based on penalization of the computation loops by a reinforcement logic associated with the reinforcement learning based deep neural network.

15. The system as claimed in claim 1 , wherein the reinforcement learning based deep neural network is trained with multiple libraries corresponding to the programming language, and wherein the reinforcement learning based deep neural network comprises plurality of hidden layers including optimizer and loss function corresponding to a sparse categorical cross entropy function.

16. The system as claimed in claim 15 , wherein the database is a knowledge database comprising the multiple libraries corresponding to the programming language, and wherein the database comprises optimization-based data and execution-based data corresponding to pre-stored regular expression.

17. A method for obtaining optimized regular expression, the method comprising:

receiving, by a processor, an input data corresponding to a programming language;

obtaining, by the processor, a plurality of embeddings using Embeddings from Language Models (ELMO) which facilitates computing word vectors on top of a two-layered bidirectional language model (biLM) which includes forward pass and backward pass;

converting, by the processor, the input data into the plurality of embeddings, wherein each embedding corresponds to a vector based representation of an attribute corresponding to the input data;

generating, by an Artificial Intelligence (AI) engine associated with the processor, an optimized regular expression (Regex) based on at least one of the type of the programming language and a corresponding library in the database, wherein the AI engine comprises a learning model associated with the AI engine for reinforced learning, wherein the learning model of the AI engine is automatically updated based on an information of the corresponding library in the database, wherein generating the optimized regular expression comprises:

receiving, by the processor, the generated regular expression, from the Artificial Intelligence (AI) engine, corresponding to the plurality of embeddings, wherein the generated regular expression is at least one of an existing regular expression from a database and a newly generated regular expression;

parsing, by the processor, the generated regular expression into a plurality of sub-blocks;

classifying, by the processor associated with a neural network-based classifier, the plurality of sub-blocks to obtain a plurality of classified sub-blocks;

evaluating, by the processor, a quantifier class for each classified sub-block to identify a corresponding computationally expensive class;

performing, by the processor executing using at least a two reinforcement learning based deep neural network associated with the AI engine, an iterative analysis based on an actor-critic algorithm to obtain a plurality of optimized sub-blocks associated with a minimum computation time, wherein one of the at least two reinforcement learning pertains to an actor and the another reinforcement learning pertains to a critic for executing the actor-critic algorithm, wherein the iterative analysis comprises replacing the quantifier class with a substitute class and performing iterative computation cycles based on an updated learning model for evaluating the corresponding computation time with respect to the substitute class by using a replace function corresponding to the programming language, wherein evaluating the corresponding computation time comprises:

verifying if computation loops of the iterative computation cycles corresponding to the substitute class exceed a pre-defined threshold;

discarding the computation loops and identifying a new substitute class that requires minimum time or has minimum computation loops if the computation time for the substitute class exceeds the pre-defined threshold; and

combining, by the processor, the plurality of optimized sub-blocks to obtain the optimized regular expression.

18. The method as claimed in claim 17 , wherein the method comprises: generating, using the plurality of embeddings, an automated recommendation pertaining to a requirement for generating the optimized regular expression, wherein the automated recommendation pertains to a recommendation of the existing regex in the database,

wherein based on the recommended existing regex, the AI engine provides at least one of a negative recommendation to perform the automated generation of the optimized regular expression and a positive recommendation to utilize the existing regular expression.

19. The method as claimed in claim 18 , wherein the method comprises:

performing simulation, by the processor, based on the optimized regular expression, to generate an output in the form of filtered results comprising an endorsed result corresponding to a least value of corresponding processing time; and

integrating automatically, by the processor, the endorsed result within an application module pertaining to the programming language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2021
From: VARGHESE, VINU; SAMPAT, NIRAV JAGDISH; JANARTHANAM, BALAJI; KUMAR, ANIL; SRIVASTAVA, SHIKHAR; KHARSADIA, KUNAL JAIWANT; PRASAD, SARAN
To: ACCENTURE GLOBAL SOLUTIONS LIMITED
Reel/Frame 055882/0244 →
Continuity (1)
Related Publication 20220309335A1 · Sep 29, 2022
References Cited (18)
US 7779049B1 · Phillips · 2010 [cited by examiner]
US 9898467B1 · Pitzel · 2018 [cited by examiner]
US 20100192225A1 · Ma · 2010 [cited by examiner]
US 20120192163A1 · Glendenning · 2012 [cited by applicant]
US 20140310291A1 · Huang · 2014 [cited by examiner]
US 20190384783A1 · Malak · 2019 [cited by examiner]
US 20210365802A1 · Sofer · 2021 [cited by examiner]
NPL: Zhong, Z., Guo, J., Yang, W., Peng, J., Xie, T., Lou, J. G., . . . & Zhang, D. (Oct. 2018). Semregex: A semantics-based approach for generating regular expressions from natural language specifications. (Year: 2018). [cited by examiner]
NPL: Zhong, Zexuan, et al. “Generating regular expressions from natural language specifications: Are we there yet?” (2018). (Year: 2018). [cited by examiner]
NPL: Backurs, Arturs, et al. “Which Regular Expression Patterns are Hard to Match? ”(2016). (Year: 2016). [cited by examiner]
NPL: Park, J. U., Ko, S. K., Cognetta, M., & Han, Y. S. (Nov. 2019). Softregex: Generating regex from natural language descriptions using softened regex equivalence. (Year: 2019). [cited by examiner]
Tu Chaofan et al., “Learning Regular Expression for Interpretable Medical Text Classification Using a Pool-based Simulated Annealing and Word-vector Models”, 2020 IEEE Congress on Evolutionary Computation (CEC), IEEE Ju… [cited by applicant]
Matthew E. Peters et al., “Deep contextualized word representations”, Mar. 22, 2018, 15 pages. <https://arxiv.org/pdf/1802.05365.pdf>. [cited by applicant]
Josh Taylor, “ELMo: Contextual language embedding”, towards data science, Jan. 6, 2019, 10 pages. <https://towardsdatascience.com/elmo-contextual-language-embedding-335de2268604>. [cited by applicant]
Ashish Vaswani, “Attention Is All You Need”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 6, 2017, 15 pages. <https://arxiv.org/pdf/1706.03762.pdf>. [cited by applicant]
Jay Alammar, “The Illustrated Transformer”, retrieved from the Internet on Feb. 23, 2021, 23 pages. <http://jalammar.github.io/illustrated-transformer/>. [cited by applicant]
European Search Report Application No. 22151574.5 dated Jun. 15, 2022, 6 Pages. [cited by applicant]
First Examination Report Indian Application No. 202214006730 dated Nov. 2, 2022, 8 Pages. [cited by applicant]