IP Library Granted Patent US 12,639,966
Granted Patent B2
US 12,639,966 · App. 18/434,892 · Granted May 26, 2026

Method and system of performing domain-based error correction after optical character recognition

Inventors: Mridul Balaraman (Bangalore, IN); Madhusudan Singh (Bangalore, IN); Ashwin Kanth (Bangalore, IN); Pinak Pani Gogoi (Bangalore, IN); Sakshi Singh (Benares, IN)
Assignee: L&T TECHNOLOGY SERVICES LIMITED
G06V30/12G06F40/279G06V10/82G06V30/245
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,966
App. No.
18/434,892
Granted
May 26, 2026
Kind
B2
Abstract

A method and system for performing optical character recognition (OCR) error correction is disclosed. The method includes receiving at least one document image. One or more text entities are determined in the data using an OCR technique. A character embedding, a layout embedding, and a style embedding is determined of each of the one or more text entities in the at least one document image. A concatenated embedding is determined of each of the one or more text entities based on the corresponding character embedding, style embedding, and layout embedding of each of the one or more text entities. A corrected character embedding of each character recognized in the corresponding text entity based on the corresponding concatenated embedding using an encoder-decoder model.

Claims (53)

1 . A method of performing domain-based optical character recognition (OCR) error correction, the method comprising:

receiving, by a processor, at least one document image,

wherein the at least one document image comprises data corresponding to at least one domain;

determining, by the processor, one or more text entities in the data using an Optical Character Recognition (OCR) technique,

wherein the one or more text entities are determined based on recognition of each character of each of the one or more text entities;

for each of the one or more text entities:

determining, by the processor, a character embedding indicative of each character recognized in a corresponding text entity;

determining, by the processor, a style embedding indicative of one or more text-characteristic information of each character recognized in the corresponding text entity;

determining, by the processor, a layout embedding of each of the one or more text entities indicative of a layout hierarchy information of each of the one or more text entities in the at least one document image;

determining, by the processor, a concatenated embedding of each of the one or more text entities based on the corresponding character embedding, the style embedding and the document layout embedding of each of the one or more text entities; and

for each of the one or more text entities:

determining, by the processor, a corrected character embedding of each character recognized in the corresponding text entity based on the corresponding concatenated embedding using an encoder-decoder model,

wherein the encoder-decoder model is trained to recognize data corresponding to the at least one domain.

2 . The method of claim 1 , wherein the text-characteristic information for each of the text entities comprises font information of each character of each of the text entities,

wherein the font information comprises a font type, a font size, case information, and typography information.

3 . The method of claim 1 , wherein the layout hierarchy information comprises a plurality of hierarchy classifications comprising a table, a header, a sub-header, a paragraph, a graphical representation, etc.

4 . The method of claim 1 , wherein the concatenated embedding of each of the one or more text entities is determined by concatenating the corresponding character embedding, the style embedding, and the document layout embedding, respectively.

5 . A system of performing domain-based optical character recognition (OCR) error correction, the system comprising:

a processor; and

a memory communicably coupled to the processor, wherein the memory stores processor-executable instructions, which, on execution by the processor, cause the processor to:

receive at least one document image,

wherein the at least one document image comprises data corresponding to at least one domain;

determine one or more text entities in the data using an Optical Character Recognition (OCR) technique,

wherein the one or more text entities are determined based on recognition of each character of each of the one or more text entities;

for each of the one or more text entities:

determine a character embedding indicative of each character recognized in a corresponding text entity;

determine a style embedding indicative of one or more text-characteristic information of each character recognized in the corresponding text entity;

determine a layout embedding of each of the one or more text entities indicative of a layout hierarchy information of each of the one or more text entities in the at least one document image;

determine a concatenated embedding of each of the one or more text entities based on the corresponding character embedding, style embedding and document layout embedding of each of the one or more text entities; and

for each of the one or more text entities:

determine a corrected character embedding of each character recognized in the corresponding text entity based on the corresponding concatenated embedding using an encoder-decoder model,

wherein the encoder-decoder model is trained to recognize data corresponding to the at least one domain.

6 . The system of claim 5 , wherein the text-characteristic information for each of the text entities comprises font information of each character of each of the text entities,

wherein the font information comprises a font type, a font size, case information, and typography information.

7 . The system of claim 5 , wherein the layout hierarchy information comprises a plurality of hierarchy classifications comprising a table, a header, a sub-header, a paragraph, a graphical representation, etc.

8 . The system of claim 5 , wherein the concatenated embedding of each of the one or more text entities is determined by concatenating the corresponding character embedding, the style embedding, and the document layout embedding, respectively.

9 . A non-transitory computer-readable medium storing computer-executable instructions for performing domain-based optical character recognition (OCR) error correction, the computer-executable instructions configured for:

receiving at least one document image,

wherein the at least one document image comprises data corresponding to at least one domain;

determining one or more text entities in the data using an Optical Character Recognition (OCR) technique,

wherein the one or more text entities are determined based on recognition of each character of each of the one or more text entities;

for each of the one or more text entities:

determining a character embedding indicative of each character recognized in a corresponding text entity;

determining a style embedding indicative of one or more text-characteristic information of each character recognized in the corresponding text entity;

determining a layout embedding of each of the one or more text entities indicative of a layout hierarchy information of each of the one or more text entities in the at least one document image;

determining a concatenated embedding of each of the one or more text entities based on the corresponding character embedding, the style embedding and the document layout embedding of each of the one or more text entities; and

for each of the one or more text entities:

determining a corrected character embedding of each character recognized in the corresponding text entity based on the corresponding concatenated embedding using an encoder-decoder model,

wherein the encoder-decoder model is trained to recognize data corresponding to the at least one domain.

10 . The non-transitory computer-readable medium of claim 9 , wherein the text-characteristic information for each of the text entities comprises font information of each character of each of the text entities, and

wherein the font information comprises a font type, a font size, case information, and typography information.

11 . The non-transitory computer-readable medium of claim 9 , wherein the layout hierarchy information comprises a plurality of hierarchy classifications comprising a table, a header, a sub-header, a paragraph, a graphical representation, etc.

12 . The non-transitory computer-readable medium of claim 9 , wherein the concatenated embedding of each of the one or more text entities is determined by concatenating the corresponding character embedding, the style embedding, and the document layout embedding, respectively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 6, 2024
From: BALARAMAN, MRIDUL; SINGH, MADHUSUDAN; KANTH, ASHWIN; GOGOI, PINAK PANI; SINGH, SAKSHI
To: L&T TECHNOLOGY SERVICES LIMITED
Reel/Frame 066399/0029 →
Priority Claims (1)
IN 202341041053 · Jun 16, 2023 · national
Continuity (1)
Related Publication 20240420493A1 · Dec 19, 2024
References Cited (15)
US 5418864A · Murdock · 1995 [cited by examiner]
US 7664343B2 · Withum · 2010 [cited by examiner]
US 8452099B2 · Reddy · 2013 [cited by examiner]
US 11481605B2 · Nguyen · 2022 [cited by examiner]
US 11842524B2 · Desai · 2023 [cited by examiner]
US 11954139B2 · Paruchuri · 2024 [cited by examiner]
US 12008331B2 · Luan · 2024 [cited by examiner]
US 12505688B2 · Jha · 2025 [cited by examiner]
Sara Salimzadeh; Improving OCR Quality by Post-Correction; Universiteit van Amsterdam; Jul. 12, 2019. [cited by applicant]
Kris Cao, Marek Rei; A Joint Model for Word Embedding and Word Morphology; Proceedings of the 1st Workshop on Representation Learning for NLP, pp. 18-26, University of Cambridge, Berlin, Germany; Aug. 11, 2016. [cited by applicant]
Christian Bartz; Reducing the Annotation Burden: Deep Learning for Optical Character Recognition using less Manual Annotations;HPI; Jul. 28, 2012. [cited by applicant]
Nikita Srivatsan, Jonathan T. Barron, Dan Klein, Taylor Berg-Kirkpatrick; A Deep Factorization of Style and Structure in Fonts; Nov. 2019. [cited by applicant]
Ahmed Hamdi, Elvys Linhares Pontes, Nicolas Sidère, Mickael Coustaty, Antoine Doucet; In-Depth Analysis of the Impact of OCR Errors on Named Entity Recognition and Linking; Natural Language Engineering, 2022, 29 (2), pp… [cited by applicant]
Rui Dong, David Smith; Multi-Input Attention for Unsupervised OCR Correction; College of Computer Information and Science, Northeastern University, Jul. 2018. [cited by applicant]
Kengtao Zheng , Nankai Lin Shengyi Jiang; Unsupervised Character Embedding Correction and Candidate Word Denoising; IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, 2022. [cited by applicant]