IP Library › Granted Patent US 12,518,229
Granted Patent B2
US 12,518,229 · App. 17/677,257 · Granted Jan 6, 2026

Workflow extraction and execution based on encoded flow graphs

Inventors: Monika Gupta (Gurugram, IN); Sampath Dechu (Bangalore, IN); Hussain Jagirdar (Ujjain, IN); Akansha Khanna (Bangalore, IN); Naveen Eravimangalath Purushothaman (Thrissur, IN)
Assignee: International Business Machines Corporation
G06Q10/0633G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,229
App. No.
17/677,257
Filed
Feb 22, 2022
Granted
Jan 6, 2026
Kind
B2
Art Unit
3625
USPC
705/7.11
Abstract

Methods, systems, and computer program products for generating automation recommendations for ad hoc processes are provided herein. A computer-implemented method includes obtaining workflow data comprising descriptions associated with one or more dynamic processes; creating event logs based at least in part on the descriptions; applying a graph extraction process to derive process flow graphs from the created event logs; generating embeddings of the process flow graphs, wherein the embeddings encode at least one of: one or more structural features and one or more attribute features of the process flow graphs; and identifying at least one of the process flow graphs to be automated based on the generated embeddings; and outputting the identified at least one process flow graph to at least one of: a user and a robotic process automation tool.

Claims (74)

1 . A computer-implemented method for improving automated process discovery by generating an executable process variant from unstructured textual data, the method comprising:

obtaining, by at least one computing device, workflow data associated with one or more dynamic processes, wherein the workflow data comprise unstructured textual data corresponding to at least one of textual descriptions of the one or more dynamic processes and textual descriptions of one or more ad hoc tasks corresponding to the one or more dynamic processes;

creating, by the at least one computing device, event logs based at least in part on the textual descriptions;

applying, by the at least one computing device, a graph extraction process to derive process flow graphs from the created event logs, wherein the process flow graphs comprise nodes corresponding to the one or more ad hoc tasks and edges indicating at least one of an order and a duration of the one or more ad hoc tasks;

executing, by the at least one computing device, a first machine learning embedding process, wherein the first machine learning embedding process generates a first set of vector embeddings by processing the process flow graphs to encode a plurality of structural features of the process flow graphs, the plurality of structural features comprising at least a number of nodes and a number of edges;

executing, by the at least one computing device, a second machine learning embedding process, wherein the second machine learning embedding process generates a second set of vector embeddings by processing the textual descriptions, wherein the second set of vector embeddings encode one or more semantic features of the textual descriptions;

generating, by the at least one computing device, a unified set of attributed process graph vector embeddings by combining, for each process flow graph, the corresponding vector embedding encoding the plurality of structural features with the corresponding vector embedding encoding the one or more semantic features;

applying, by the at least one computing device, at least one unsupervised machine learning process to the unified set of attributed process graph vector embeddings to cluster the process flow graphs into one or more clusters;

computing, by the at least one computing device, automation scores for the process flow graphs based at least in part on the unified set of attributed process graph vector embeddings, wherein the automation scores indicate a complexity of automating the process flow graphs;

identifying, by the at least one computing device, at least one of the process flow graphs to be automated based at least in part on the computed automation scores;

generating, by the at least one computing device, at least one executable process variant of the identified at least one process flow graph, the executable process variant comprising a new process flow graph comprising a modified sequence of computer-executable tasks, wherein generating the at least one executable process variant comprises: (i) identifying a portion of the identified at least one process flow graph corresponding to an activity based on a comparison of an average time to perform the activity relative to other activities within a given one of the one or more clusters comprising the identified at least one process flow graph and a frequency of occurrence of the activity across multiple process flow graphs within the cluster, and (ii) modifying the identified portion based on the plurality of structural features and the one or more semantic features encoded in the attributed process graph vector embeddings corresponding to at least one other process flow graph within the cluster;

outputting, by the at least one computing device, the generated at least one executable process variant to a robotic process automation tool; and

in response to the outputting, using, by the at least one computing device, the robotic process automation tool to automatically execute the modified sequence of computer-executable tasks of the at least one executable process variant.

2 . The computer-implemented method of claim 1 , further comprising:

generating, by the at least one computing device, a respective keyphrase to represent each of the one or more clusters based on the textual descriptions associated with the process flow graphs within each cluster.

3 . The computer-implemented method of claim 1 , wherein the first set of vector embeddings of the process flow graphs generated by the first machine learning embedding process further encode one or more of: at least one temporal feature, and at least one frequency distribution feature.

4 . The computer-implemented method of claim 1 , wherein the plurality of structural features further comprises one or more of:

a maximum node fan-out count;

a maximum node fan-in count;

a number of variants; and

a number of loops.

5 . The computer-implemented method of claim 1 , wherein the identifying comprises:

ranking the process flow graphs based on the computed automation scores.

6 . The computer-implemented method of claim 1 , wherein generating the at least one executable process variant comprises at least one of: adding, deleting, or changing one or more repetitive tasks associated with the identified process flow graph.

7 . A computer program product for improving automated process discovery by generating an executable process variant from unstructured textual data, the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing device to cause the computing device to:

obtain workflow data associated with one or more dynamic processes, wherein the workflow data comprise unstructured data corresponding to at least one of textual descriptions of the one or more dynamic processes and textual descriptions of one or more ad hoc tasks corresponding to the one or more dynamic processes;

create event logs based at least in part on the textual descriptions;

apply a graph extraction process to derive process flow graphs from the created event logs, wherein the process flow graphs comprise nodes corresponding to the one or more ad hoc tasks and edges indicating at least one of an order and a duration of the one or more ad hoc tasks;

execute a first machine learning embedding process, wherein the first machine learning embedding process generates a first set of vector embeddings by processing the process flow graphs to encode a plurality of structural features of the process flow graphs, the plurality of structural features comprising at least a number of nodes and a number of edges;

execute a second machine learning embedding process, wherein the second machine learning embedding process generates a second set of vector embeddings by processing the textual descriptions, wherein the second set of vector embeddings encode one or more semantic features of the textual descriptions;

generate a unified set of attributed process graph vector embeddings by combining, for each process flow graph, the corresponding vector embedding encoding the plurality of structural features with the corresponding vector embedding encoding the one or more semantic features;

apply at least one unsupervised machine learning process to the unified set of attributed process graph vector embeddings to cluster the process flow graphs into one or more clusters;

compute automation scores for the process flow graphs based at least in part on the unified set of attributed process graph vector embeddings, wherein the automation scores indicate a complexity of automating the process flow graphs;

identify at least one of the process flow graphs to be automated based at least in part on the computed automation scores;

generate at least one executable process variant of the identified at least one process flow graph, the executable process variant comprising a new process flow graph comprising a modified sequence of computer-executable tasks, wherein generating the at least one executable process variant comprises: (i) identifying a portion of the identified at least one process flow graph corresponding to an activity based on a comparison of an average time to perform the activity relative to other activities within a given one of the one or more clusters comprising the identified at least one process flow graph and a frequency of occurrence of the activity across multiple process flow graphs within the cluster, and (ii) modifying the identified portion based on the plurality of structural features and the one or more semantic features encoded in the attributed process graph vector embeddings corresponding to at least one other process flow graph within the cluster;

output the generated at least one executable process variant to a robotic process automation tool; and

in response to the outputting, use the robotic process automation tool to automatically execute the modified sequence of computer-executable tasks of the at least one executable process variant.

8 . The computer-implemented method of claim 1 , wherein executing the second machine learning embedding process comprises:

updating pre-trained sentence embeddings using domain-specific training data derived from the textual descriptions of the one or more ad hoc tasks.

9 . The computer-implemented method of claim 8 , wherein updating the pre-trained sentence embeddings comprises tuning the pre-trained sentence embeddings using a task sequence of a process instance associated with the textual descriptions as a context.

10 . The computer program product of claim 7 , wherein the computing device is further caused to:

generate a respective keyphrase to represent each of the one or more clusters based on the textual descriptions associated with the process flow graphs within each cluster.

11 . The computer program product of claim 7 , wherein the first set of vector embeddings of the process flow graphs generated by the first machine learning embedding process further encode one or more of: at least one temporal feature, and at least one frequency distribution feature.

12 . The computer program product of claim 7 , wherein the identifying comprises:

ranking the process flow graphs based on the computed automation scores.

13 . The computer program product of claim 7 , wherein executing the second machine learning embedding process comprises:

updating pre-trained sentence embeddings using domain-specific training data derived from the textual descriptions of the one or more ad hoc tasks, wherein updating the pre-trained sentence embeddings comprises tuning the pre-trained sentence embeddings using a task sequence of a process instance associated with the textual descriptions as a context.

14 . A system for improving automated process discovery by generating an executable process variant from unstructured textual data, the system comprising:

a memory configured to store program instructions;

a processor operatively coupled to the memory to execute the program instructions to:

obtain workflow data associated with one or more dynamic processes, wherein the workflow data comprise unstructured data corresponding to at least one of textual descriptions of the one or more dynamic processes and textual descriptions of one or more ad hoc tasks corresponding to the one or more dynamic processes;

create event logs based at least in part on the textual descriptions;

apply a graph extraction process to derive process flow graphs from the created event logs, wherein the process flow graphs comprise nodes corresponding to the one or more ad hoc tasks and edges indicating at least one of an order and a duration of the one or more ad hoc tasks; execute a first machine learning embedding process, wherein the first machine learning embedding process generates a first set of vector embeddings by processing the process flow graphs to encode a plurality of structural features of the process flow graphs, the plurality of structural features comprising at least a number of nodes and a number of edges;

execute a second machine learning embedding process, wherein the second machine learning embedding process generates a second set of vector embeddings by processing the textual descriptions, wherein the second set of vector embeddings encode one or more semantic features of the textual descriptions;

generate a unified set of attributed process graph vector embeddings by combining, for each process flow graph, the corresponding vector embedding encoding the plurality of structural features with the corresponding vector embedding encoding the one or more semantic features;

apply at least one unsupervised machine learning process to the unified set of attributed process graph vector embeddings to cluster the process flow graphs into one or more clusters;

compute automation scores for the process flow graphs based at least in part on the unified set of attributed process graph vector embeddings, wherein the automation scores indicate a complexity of automating the process flow graphs;

identify at least one of the process flow graphs to be automated based at least in part on the computed automation scores;

generate at least one executable process variant of the identified at least one process flow graph, the executable process variant comprising a new process flow graph comprising a modified sequence of computer-executable tasks, wherein generating the at least one executable process variant comprises: (i) identifying a portion of the identified at least one process flow graph corresponding to an activity based on a comparison of an average time to perform the activity relative to other activities within a given one of the one or more clusters comprising the identified at least one process flow graph and a frequency of occurrence of the activity across multiple process flow graphs within the cluster, and (ii) modifying the identified portion based on the plurality of structural features and the one or more semantic features encoded in the attributed process graph vector embeddings corresponding to at least one other process flow graph within the cluster;

output the generated at least one executable process variant to a robotic process automation tool; and

in response to the outputting, use the robotic process automation tool to automatically execute the modified sequence of computer-executable tasks of the at least one executable process variant.

15 . The system of claim 14 , wherein the processor further executes the program instructions to:

generate a respective keyphrase to represent each of the one or more clusters based on the textual descriptions associated with the process flow graphs within each cluster.

16 . The system of claim 14 , wherein the first set of vector embeddings of the process flow graphs generated by the first machine learning embedding process further encode one or more of: at least one temporal feature, and at least one frequency distribution feature.

17 . The system of claim 14 , wherein the plurality of structural features further comprises one or more of:

a maximum node fan-out count;

a maximum node fan-in count;

a number of variants; and

a number of loops.

18 . The system of claim 14 , wherein the identifying comprises:

ranking the process flow graphs based on the computed automation scores.

19 . The system of claim 14 , wherein generating the at least one executable process variant comprises at least one of: adding, deleting, or changing one or more repetitive tasks associated with the identified process flow graph.

20 . The system of claim 14 , wherein executing the second machine learning embedding process comprises:

updating pre-trained sentence embeddings using domain-specific training data derived from the textual descriptions of the one or more ad hoc tasks, wherein updating the pre-trained sentence embeddings comprises tuning the pre-trained sentence embeddings using a task sequence of a process instance associated with the textual descriptions as a context.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2022
From: GUPTA, MONIKA; DECHU, SAMPATH; JAGIRDAR, HUSSAIN; KHANNA, AKANSHA; PURUSHOTHAMAN, NAVEEN ERAVIMANGALATH
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059065/0136 →
Continuity (1)
Related Publication 20230267396A1 · Aug 24, 2023
References Cited (53)
US 7580911B2 · Sun · 2009 [cited by examiner]
US 8265979B2 · Golani · 2012 [cited by examiner]
US 8321251B2 · Opalach · 2012 [cited by examiner]
US 8688499B1 · Bose · 2014 [cited by examiner]
US 9087236B2 · Dhoolia · 2015 [cited by examiner]
US 9799326B2 · Dhoolia · 2017 [cited by examiner]
US 10523520B2 · Aggarwal · 2019 [cited by examiner]
US 10656979B2 · Ishakian · 2020 [cited by examiner]
US 11200539B2 · Iyer · 2021 [cited by examiner]
US 11270241B2 · Smutko · 2022 [cited by examiner]
US 11500756B2 · Scheepens · 2022 [cited by examiner]
US 11710098B2 · Sridhara · 2023 [cited by examiner]
US 11748682B2 · Smutko · 2023 [cited by examiner]
US 11816596B2 · Pingali · 2023 [cited by examiner]
US 20040260590A1 · Golani · 2004 [cited by examiner]
US 20100042745A1 · Maeda · 2010 [cited by examiner]
US 20120059683A1 · Opalach · 2012 [cited by examiner]
US 20120062574A1 · Dhoolia · 2012 [cited by examiner]
US 20140304027A1 · Wu · 2014 [cited by examiner]
US 20140350983A1 · Lakshmanan · 2014 [cited by examiner]
US 20170111245A1 · Ishakian · 2017 [cited by examiner]
US 20170213544A1 · Dhoolia · 2017 [cited by examiner]
US 20170286190A1 · Ishakian · 2017 [cited by examiner]
US 20180114126A1 · Das · 2018 [cited by examiner]
US 20200065151A1 · Ghosh · 2020 [cited by examiner]
US 20210004711A1 · Gupta · 2021 [cited by examiner]
US 20210216925A1 · Dixit · 2021 [cited by examiner]
US 20210264332A1 · Pingali · 2021 [cited by examiner]
US 20210304139A1 · Sridhara · 2021 [cited by examiner]
US 20210406228A1 · Reminnyi · 2021 [cited by examiner]
US 20220075705A1 · Scheepens · 2022 [cited by examiner]
US 20220075706A1 · Scheepens · 2022 [cited by examiner]
US 20220076147A1 · Scheepens · 2022 [cited by examiner]
US 20220188143A1 · Scheepens · 2022 [cited by examiner]
IN 201841032619A · 2020 [cited by applicant]
Chambers, Alexander J. et al., Automated Business Process Discovery from Unstructured Natural Language Documents BPM Workshops, 2020 (Year: 2020). [cited by examiner]
Khameme, Fatemeh Nikraftar, Recommender System Based on Process Mining University di Padova, Italy, 2022 (Year: 2022). [cited by examiner]
Hugo et al., Business Process Models Clustering Based on Multimodal Search, K-means clustering and Cumulative and Non-Continuous N-Grams, Polibits, vol. 54, 2016 (Year: 2016). [cited by examiner]
Cao, Bin et al, Graph-Based Workflow Recommendation: On Improving Business Process Modeling CIKM'12, ACM, Oct.-Nov. 2012 (Year: 2012). [cited by examiner]
Wang, Huaqing et al., RLRecommender: A Representation-Learning-Based Recommendation Method for Business Process Modeling, ICSOC 2-18, Spring Nature, AG 2018 (Year: 2018). [cited by examiner]
Yu, Xiaoming et al., Workflow Recommendation Based on Graph Embedding 2020 IEEE World Congress on Services, 2020 (Year: 2020). [cited by examiner]
Mell, Peter, et al., The NIST Definition of Cloud Computing, National Institute of Standards and Technology, U.S. Department of Commerce, NIST Special Publication 800-145, Sep. 2011. [cited by applicant]
Gupta, Monika, et al. “Analyzing comments in ticket resolution to capture underlying process interactions.” International Conference on Business Process Management, Springer, Cham, 2020, pp. 219-231. [cited by applicant]
Dustdar, Schahram, et al. “Mining of ad-hoc business processes with TeamLog.” Data & Knowledge Engineering 55.2, Nov. 1, 2005, pp. 129-158. [cited by applicant]
Dorn, Christoph, et al. “Self-adjusting recommendations for people-driven ad-hoc processes.” International conference on business process management. Springer, Berlin, Heidelberg, Sep. 13, 2010, pp. 327-342. [cited by applicant]
Augusto, Adriano, et al. “Split miner: automated discovery of accurate and simple business process models from event logs.” Knowledge and Information Systems 59.2, May 2019, pp. 251-284. [cited by applicant]
Augusto, Adriano, et al. “Automated discovery of process models from event logs: review and benchmark.” IEEE transactions on knowledge and data engineering 31.4, May 2018, pp. 686-705. [cited by applicant]
Bolt, Alfredo, et al., “Finding process variants in event logs.“OTM Confederated International Conferences” On the Move to Meaningful Internet Systems”, Springer, Cham, Oct. 2017, pp. 45-52. [cited by applicant]
Nguyen, Phuc, et al. “TabEAno: Table to Knowledge Graph Entity Annotation.” arXiv preprint arXiv:2010.01829, Oct. 2020. [cited by applicant]
Motahari-Nezhad, et al. “Next best step and expert recommendation for collaborative processes in it service management.” International Conference on Business Process Management. Springer, Berlin, Heidelberg, Aug. 2011, … [cited by applicant]
Albalawi, Rania, Tet Hin Yeap, and Morad Benyoucef. “Using topic modeling methods for short-text data: A comparative analysis.” Frontiers in Artificial Intelligence, vol. 3, Article 42, Jul. 2020. [cited by applicant]
Cirne, Renato, et al. “Data Mining for Process Modeling: A Clustered Process Discovery Approach.” 2020 15th Conference on Computer Science and Information Systems (FedCSIS), IEEE, Sep. 2020, pp. 587-590. [cited by applicant]
Process discovery by using Discovery Bot, Automation Anywhere, Inc., available at https://docs.automationanywhere.com/bundle/enterprise-v2019/page/discovery-bot/topics/discovery-bot-intro.html#, last visited Feb. 22-Feb… [cited by applicant]