IP Library › Granted Patent US 11,810,382
Granted Patent B2
US 11,810,382 · App. 17/500,184 · Granted Nov 7, 2023

Training optical character detection and recognition models for robotic process automation

Inventors: Dorin Andrei Laza (Bucharest, RO); Trong Canh Nguyen (Paris, FR)
Assignee: UiPath, Inc.
G06V30/413G06F9/45512G06V10/82G06V30/18057G06V30/262G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,810,382
App. No.
17/500,184
Granted
Nov 7, 2023
Kind
B2
Abstract

Techniques for training an optical character recognition (OCR) model to detect and recognize text in images for robotic process automation (RPA) are disclosed. A text detection model and a text recognition model may be trained separately and then combined to produce the OCR model. Synthetic data and a smaller amount of real, human-labeled data may be used for training to increase the speed and accuracy with which the OCR text detection model and the text recognition model can be trained. After the OCR model has been trained, a workflow may be generated that includes an activity calling the OCR model, and a robot implementing the workflow may be generated and deployed.

Claims (58)

1. A computer-implemented method, comprising:

generating a set of synthetic data, by a computing system;

training a text recognition model for robotic process automation (RPA) on augmented human-labeled data and the set of synthetic data over a first plurality of epochs, by the computing system;

at each epoch of the first plurality of epochs, evaluating performance of the text recognition model against the evaluation dataset until a level of accuracy of the performance begins to decline, by the computing system;

training a text detection model for RPA on the set of synthetic data over a second plurality of epochs, by the computer system; and

at each epoch of the second plurality of epochs, evaluating performance of the text detection model against an evaluation dataset until a level of accuracy of the performance begins to decline, by the computing system, wherein

the text detection model is trained solely using the set of synthetic data.

2. The computer-implemented method of claim 1 , further comprising:

combining the text detection model and the text recognition model into an optical character recognition (OCR) model, by the computing system, wherein

the text detection model and the text recognition model are trained separately and then combined for runtime.

3. The computer-implemented method of claim 2 , further comprising:

training a smaller, faster OCR model using the OCR model, by the computing system.

4. The computer-implemented method of claim 3 , further comprising:

generating a workflow comprising an activity calling the OCR model, by the computing system;

generating an RPA robot implementing the workflow, by the computing system; and

deploying the RPA robot, by the computing system.

5. The computer-implemented method of claim 1 , wherein the text recognition model comprises a residual network and a single long short term (LSTM) memory layer with connectionist temporal classification (TC) for text decoding.

6. The computer-implemented method of claim 1 , wherein the generation of the set of synthetic data comprises placing words in images, adding random noise, configuring the images in blocks, or any combination thereof.

7. The computer-implemented method of claim 1 , wherein the generation of the set of synthetic data comprises adding icons with no text, drawing random polygons in one or more of the images of various types, shapes, and/or sizes, or any combination thereof.

8. The computer-implemented method of claim 1 , wherein the generation of the set of synthetic data comprises building a list of fonts and randomly choosing, mixing, and inserting text in these fonts into the synthetic data.

9. The computer-implemented method of claim 1 , wherein the generation of the set of synthetic data comprises generating negative examples.

10. The computer-implemented method of claim 1 , wherein the text recognition model is configured to determine floating point numbers or dates as single words.

11. The computer-implemented method of claim 1 , wherein a number of character types recognized by the text recognition model is less than or equal to 100.

12. The computer-implemented method of claim 1 , wherein the text recognition model is trained using one or more additional sets of training data over respective epochs.

13. A non-transitory computer-readable medium storing a computer program, the computer program configured to cause at least one processor to:

generate a set of synthetic data;

train a text detection model for robotic process automation (RPA) solely on the set of synthetic data over a first plurality of epochs;

at each epoch of the first plurality of epochs, evaluate performance of the text detection model against an evaluation dataset until a level of accuracy of the performance begins to decline;

train a text recognition model for RPA on augmented human-labeled data and the set of synthetic data over a second plurality of epochs; and

at each epoch of second the plurality of epochs, evaluate performance of the text recognition model against the evaluation dataset until a level of accuracy of the performance begins to decline.

14. The non-transitory computer-readable medium of claim 13 , wherein the computer program is further configured to cause the at least one processor to:

combine the text detection model and the text recognition model into an optical character recognition (OCR) model, wherein

the text detection model and the text recognition model are trained separately and then combined for runtime.

15. The non-transitory computer-readable medium of claim 14 , further comprising:

training a smaller, faster OCR model using the OCR model, by the computing system.

16. The non-transitory computer-readable medium of claim 13 , wherein the generation of the set of synthetic data comprises placing words in images, adding random noise, configuring the images in blocks, or any combination thereof.

17. The non-transitory computer-readable medium of claim 13 , wherein the generation of the set of synthetic data comprises adding icons with no text, drawing random polygons in one or more of the images of various types, shapes, and/or sizes, or any combination thereof.

18. The non-transitory computer-readable medium of claim 13 , wherein the generation of the set of synthetic data comprises building a list of fonts and randomly choosing, mixing, and inserting text in these fonts into the synthetic data.

19. The non-transitory computer-readable medium of claim 13 , wherein the generation of the set of synthetic data comprises generating negative examples.

20. The non-transitory computer-readable medium of claim 13 , wherein the computer program is further configured to cause the at least one processor to train the text detection model, the text recognition model, or both, using one or more additional sets of training data over respective epochs.

21. The non-transitory computer-readable medium of claim 13 , wherein a number of character types recognized by the text recognition model is less than or equal to 100.

22. A computer-implemented method for training an optical character recognition (OCR) model, comprising:

generating a set of synthetic data, by a computing system;

training a text recognition model for robotic process automation (RPA) on augmented human-labeled data and the set of synthetic data over a plurality of epochs, by the computing system;

at each epoch of the plurality of epochs, evaluating performance of the text recognition model against the evaluation dataset until a level of accuracy of the performance begins to decline, by the computing system;

training a text detection model for RPA on the set of synthetic data over a second plurality of epochs, by the computer system; and

at each epoch of the second plurality of epochs, evaluating performance of the text detection model against an evaluation dataset until a level of accuracy of the performance begins to decline, by the computing system;

combining the text detection model and the text recognition model into an optical character recognition (OCR) model, by the computing system; and

training a smaller, faster OCR model using the OCR model, by the computing system.

23. The non-transitory computer-readable medium of claim 22 , wherein the generation of the set of synthetic data comprises:

placing words in images, adding random noise, configuring the images in blocks, or any combination thereof,

adding icons with no text, drawing random polygons in one or more of the images of various types, shapes, and/or sizes, or any combination thereof,

building a list of fonts and randomly choosing, mixing, and inserting text in these fonts into the synthetic data,

generating negative examples, or

any combination thereof.

24. The computer-implemented method of claim 22 , further comprising:

training the text recognition model using one or more additional sets of training data over respective epochs, by the computing system.

25. The computer-implemented method of claim 22 , wherein a number of character types recognized by the text recognition model is less than or equal to 100.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2021
From: LAZA, DORIN ANDREI; NGUYEN, TRONG CANH
To: UIPATH, INC.
Reel/Frame 057779/0312 →
Continuity (2)
Continuation 16700494 · Dec 2, 2019
Related Publication 20220067462A1 · Mar 3, 2022
Cited By (4)
US 12,236,700 US 12,530,666 US 12,572,936 US 12,579,832