IP Library › Granted Patent US 11,507,772
Granted Patent B2
US 11,507,772 · App. 16/591,161 · Granted Nov 22, 2022

Sequence extraction using screenshot images

Inventors: Christian Berg (Seattle, WA); Shirin Feiz Disfani (Port Jeff Station, NY)
Assignee: UIPATH, INC.
G06K9/6218B25J9/163B25J9/1656G06V30/153
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,507,772
App. No.
16/591,161
Granted
Nov 22, 2022
Kind
B2
Abstract

A system and method for sequence extraction using screenshot images to generate a robotic process automation workflow is disclosed. The system and method include capturing a plurality of screenshots of steps performed by a user on an application using a processor, storing the screenshots in memory, determining action clusters from the captured screenshots by randomly clustering actions into an arbitrary predefined number of clusters, wherein screenshots of different variations of a same action is labeled in the clusters, extracting a sequence from the clusters, and discarding consequent events on the screen from the clusters, and generating an automated workflow based on the extracted sequences.

Claims (33)

1. A method for sequence extraction using screenshot images to generate a robotic process automation workflow, the method comprising:

capturing a plurality of screenshots of steps performed by a user on an application using a processor, wherein the capturing includes templating to find a plurality of words and a corresponding location for each of the plurality of words to cluster in forming a template;

storing the screenshots in memory;

determining action clusters from the captured screenshots by randomly clustering actions into an arbitrary predefined number of clusters based on a template applied to the captured screenshots, wherein screenshots of different variations of a same action is labeled in the clusters;

extracting a sequence from the clusters, and discarding consequent events on the screen from the clusters; and

generating an automated workflow based on the extracted sequences.

2. The method of claim 1 , wherein the templating utilizes a threshold in indicating the plurality of words.

3. The method of claim 2 , wherein the threshold comprises approximately 70%.

4. The method of claim 1 , wherein the capturing includes adaptive parameter tuning to iterate a template and tune the capturing for subsequent iterations.

5. The method of claim 1 , wherein the capturing includes random sampling utilizing particle swarm optimization.

6. The method of claim 1 , wherein the capturing includes clustering details incorporating a binary feature vector indicating a presence of template items.

7. The method of claim 1 , wherein the capturing includes novelty by learning sparse representation of screens and tuning cluster granularity.

8. The method of claim 1 , wherein the extracting includes forward link estimation utilizing a forward link prediction module to consider each event and link future events with each event.

9. The method of claim 1 , wherein the extracting includes graphical representation with each graph node corresponding to a screen type discovered in the clustering.

10. The method of claim 9 , wherein the edges of the graph represent each event and the events linked events.

11. The method of claim 1 , wherein the clustering leverages optical character recognition (OCR) data to extract word and location pairs.

12. A system for sequence extraction using screenshot images to generate a robotic process automation workflow, the system comprising:

a processor configured to capture a plurality of screenshots of steps performed by a user on an application, wherein the capturing includes templating to find a plurality of words and a corresponding location for each of the plurality of words to cluster in forming a template; and

a memory module operatively coupled to the processor and configured to store the screenshots; the processor further configured to:

determine action clusters from the captured screenshots by randomly clustering actions into an arbitrary predefined number of clusters based on a template applied to the captured screenshots, wherein screenshots of different variations of a same action is labeled in the clusters;

extract a sequence from the clusters, and discarding consequent events on the screen from the clusters; and

generate an automated workflow based on the extracted sequences.

13. The system of claim 12 , wherein the capturing includes adaptive parameter tuning to iterate a template and tune the capturing for subsequent iterations.

14. The system of claim 12 , wherein the capturing includes clustering details incorporating a binary feature vector indicating a presence of template items.

15. The system of claim 12 , wherein the capturing includes novelty by learning sparse representation of screens and tuning cluster granularity.

16. The system of claim 12 , wherein the extracting includes forward link estimation utilizing a forward link prediction module to consider each event and link future events with each event.

17. The system of claim 12 , wherein the extracting includes graphical representation with each graph node corresponding to a screen type discovered in the clustering.

18. A non-transitory computer-readable medium comprising a computer program product recorded thereon and capable of being run by a processor, including program code instructions for sequence extraction using screenshot images to generate a robotic process automation workflow by implementing the steps comprising:

capturing a plurality of screenshots of steps performed by a user on an application using a processor, wherein the capturing includes templating to find a plurality of words and a corresponding location for each of the plurality of words to cluster in forming a template;

storing the screenshots in memory;

determining action clusters from the captured screenshots by randomly clustering actions into an arbitrary predefined number of clusters based on a template applied to the captured screenshots, wherein screenshots of different variations of a same action is labeled in the clusters;

extracting a sequence from the clusters, and discarding consequent events on the screen from the clusters; and

generating an automated workflow based on the extracted sequences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2020
From: BERG, CHRISTIAN; FEIZ DISFANI, SHIRIN
To: UIPATH, INC.
Reel/Frame 054352/0445 →
Continuity (1)
Related Publication 20210103767A1 · Apr 8, 2021