IP Library Granted Patent US 11,568,140
Granted Patent B2
US 11,568,140 · App. 17/105,023 · Granted Jan 31, 2023

Optical character recognition using a combination of neural network models

Inventors: Konstantin Anisimovich (Moscow, RU); Aleksei Zhuravlev (Yaroslavl, RU)
Assignee: ABBYY DEVELOPMENT INC.
G06F40/279G06K9/6267G06V10/88G06V30/194
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,568,140
App. No.
17/105,023
Granted
Jan 31, 2023
Kind
B2
Abstract

Embodiments of the present disclosure describe a system and method for optical character recognition. In one embodiment, a system receives an image depicting text. The system extracts features from the image using a feature extractor. The system applies a first decoder to the features to generate a first intermediary output. The system applies a second decoder to the features to generate a second intermediary output, wherein the feature extractor is common to the first decoder and the second decoder. The system determines a first quality metric value for the first intermediary output and a second quality metric value for the second intermediary output based on a language model. Responsive to determining that the first quality metric value is greater than the second quality metric value, the system selects the first intermediary output to represent the text.

Claims (40)

1. A method, comprising:

receiving, by a computer system, an image with text;

extracting, by a feature extractor, a plurality of features from the image;

applying a first decoder to the plurality of features to generate a first intermediary output;

applying a second decoder to the plurality of features to generate a second intermediary output;

determining, based on a language model, a first quality metric value for the first intermediary output and a second quality metric value for the second intermediary output; and

responsive to determining that the first quality metric value is greater than the second quality metric value, selecting the first intermediary output to represent the text.

2. The method of claim 1 , wherein the first decoder includes a connectionist temporal classification (CTC) decoder and the second decoder includes a character decoder with attention.

3. The method of claim 2 , wherein the first decoder includes a first long short-term memory component and the character decoder with attention includes a second long short-term memory component.

4. The method of claim 3 , wherein the first long short-term memory component is bidirectional.

5. The method of claim 2 , wherein the first decoder and the second decoder include a long short-term memory component common to the first decoder and the second decoder.

6. The method of claim 1 , wherein the language model includes at least one of: a dictionary, a morphological model of inflection, a syntactic model, or a model for statistical compatibility of letters or words.

7. The method of claim 1 , wherein the language model includes a recurrent neural network model.

8. A system, comprising:

a memory;

a processor, coupled to the memory, the processor configured to:

receive an image with text;

extract, using a feature extractor, a plurality of features from the image;

apply a first decoder to the plurality of features to generate a first intermediary output;

apply a second decoder to the plurality of features to generate a second intermediary output, wherein the feature extractor is common to the first decoder and the second decoder;

determine, based on a language model, a first quality metric value for the first intermediary output and a second quality metric value for the second intermediary output; and

responsive to determining that the first quality metric value is greater than the second quality metric value, select the first intermediary output to represent the text.

9. The system of claim 8 , wherein the first decoder includes a connectionist temporal classification (CTC) decoder and the second decoder includes a character decoder with attention.

10. The system of claim 9 , wherein the first decoder includes a first long short-term memory component and the character decoder with attention includes a second long short-term memory component.

11. The system of claim 10 , wherein the first long short-term memory component is bidirectional.

12. The system of claim 9 , wherein the first decoder and the second decoder include a long short-term memory component common to the first decoder and the second decoder.

13. The system of claim 8 , wherein the language model includes a dictionary, a morphological model of inflection, a syntactic model, or a model for statistical compatibility of letters or words.

14. The system of claim 8 , wherein the language model includes a recurrent neural network model.

15. A computer-readable non-transitory storage medium comprising executable instructions that, when executed by a computer system, cause the computer system to:

receive an image with text;

extract, using a feature extractor, a plurality of features from the image;

apply a first decoder to the plurality of features to generate a first intermediary output;

apply a second decoder to the plurality of features to generate a second intermediary output, wherein the feature extractor is common to the first decoder and the second decoder;

determine, based on a language model, a first quality metric value for the first intermediary output and a second quality metric value for the second intermediary output; and

responsive to determining that the first quality metric value is greater than the second quality metric value, select the first intermediary output to represent the text.

16. The computer-readable non-transitory storage medium of claim 15 , wherein the first decoder includes a connectionist temporal classification (CTC) decoder and the second decoder includes a character decoder with attention.

17. The computer-readable non-transitory storage medium of claim 16 , wherein the first decoder includes a first long short-term memory component and the character decoder with attention includes a second long short-term memory component.

18. The computer-readable non-transitory storage medium of claim 17 , wherein the first long short-term memory component is bidirectional.

19. The computer-readable non-transitory storage medium of claim 16 , wherein the first decoder and the second decoder include a long short-term memory component common to the first decoder and the second decoder.

20. The computer-readable non-transitory storage medium of claim 15 , wherein the language model includes a dictionary, a morphological model of inflection, a syntactic model, or a model for statistical compatibility of letters or words.

Assignments (3)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2020
From: ANISIMOVICH, KONSTANTIN; ZHURAVLEV, ALEKSEI
To: ABBYY PRODUCTION LLC
Reel/Frame 054686/0335 →
Continuity (1)
Related Publication 20220164533A1 · May 26, 2022
Cited By (1)
US 12,340,360