IP Library › Granted Patent US 12,573,225
Granted Patent B2
US 12,573,225 · App. 18/506,681 · Granted Mar 10, 2026

Methods and systems of field detection in a document

Inventors: Stanislav Semenov (Moscow, RU); Mikhail Lanin (Moscow, RU)
Assignee: ABBYY Development Inc.
G06V30/412G06F18/214G06F40/284G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,225
App. No.
18/506,681
Granted
Mar 10, 2026
Kind
B2
Abstract

Systems and methods are disclosed to receive a training data set comprising a plurality of document images, wherein each document image of the plurality of document images is associated with respective metadata identifying a document field containing a variable text; generate, by processing the plurality of document images, a first heat map represented by a data structure comprising a plurality of heat map elements corresponding to a plurality of document image pixels, wherein each heat map element stores a counter of a number of document images in which the document field contains a document image pixel associated with the heat map element; receive an input document image; and identify, within the input document image, a candidate region comprising the document field, wherein the candidate region comprises a plurality of input document image pixels corresponding to heat map elements satisfying a threshold condition.

Claims (48)

1 . A method comprising:

receiving a plurality of document images, wherein each document image of the plurality of document images comprises a corresponding document field containing a variable text;

generating, by processing the plurality of document images, a first heat map represented by a data structure comprising a plurality of heat map elements corresponding to a plurality of document image pixels, wherein each heat map element stores a respective value specifying whether the corresponding document field contains a document image pixel associated with the heat map element;

receiving an input document image; and

identifying, using the first heat map, a candidate region of the input document image, the candidate region comprising the document field, wherein the candidate region comprises a plurality of input document image pixels corresponding to heat map elements satisfying a threshold condition.

2 . The method of claim 1 , wherein the candidate region is identified using a plurality of heat maps, the plurality of heat maps comprising the first heat map and one or more additional heat maps, wherein the each of the plurality of heat maps identify a potential document field location corresponding to the document field.

3 . The method of claim 2 , wherein each of the plurality of heat maps identify the potential document field location relative to a respective reference element in each of the plurality of heat maps.

4 . The method of claim 3 , wherein the respective reference element comprises one or more of a predefined word, or a predefined graphical element.

5 . The method of claim 2 , further comprising:

extracting a content of each document image of the plurality of document images, wherein the content is included in the potential document field location; and

analyzing the content of each document image using Byte Pair Encoding (BPE) tokens.

6 . The method of claim 5 , wherein analyzing the content comprises:

representing the content of each document image using BPE tokens to derive tokenized content for each document image;

generating vector representation of the tokenized content for each document image;

calculating a distance between a pair of vector representations of the tokenized content from two document images of the plurality of document images; and

in response to determining that the distance is less than a predefined value, indicating that the potential document field location is likely to be correct.

7 . The method of claim 1 , wherein each document image of the plurality of document images is associated with respective metadata defining a marked up document field location corresponding to the document field for each document image of the plurality of document images.

8 . A system comprising:

a memory device storing instructions;

a processing device coupled to the memory device, the processing device to execute the instructions to:

receive a plurality of document images, wherein each document image of the plurality of document images comprises a corresponding document field containing a variable text;

generate, by processing the plurality of document images, a first heat map represented by a data structure comprising a plurality of heat map elements corresponding to a plurality of document image pixels, wherein each heat map element stores a respective value specifying whether the corresponding document field contains a document image pixel associated with the heat map element;

receive an input document image; and

identify, using the first heat map, a candidate region of the input document image, the candidate region comprising the document field, wherein the candidate region comprises a plurality of input document image pixels corresponding to heat map elements satisfying a threshold condition.

9 . The system of claim 8 , wherein the candidate region is identified using a plurality of heat maps, the plurality of heat maps comprising the first heat map and one or more additional heat maps, wherein the each of the plurality of heat maps identify a potential document field location corresponding to the document field.

10 . The system of claim 9 , wherein each of the plurality of heat maps identify the potential document field location relative to a respective reference element in each of the plurality of heat maps.

11 . The system of claim 10 , wherein the respective reference element comprises one or more of a predefined word, or a predefined graphical element.

12 . The system of claim 9 , wherein the processing device is further to:

extract a content of each document image of the plurality of document images, wherein the content is included in the potential document field location; and

analyze the content of each document image using Byte Pair Encoding (BPE) tokens.

13 . The system of claim 12 , wherein to analyze the content, the processing device is to:

represent the content of each document image using BPE tokens to derive tokenized content for each document image;

generate vector representation of the tokenized content for each document image;

calculate a distance between a pair of vector representations of the tokenized content from two document images of the plurality of document images; and

in response to a determination that the distance is less than a predefined value, indicate that the potential document field location is likely to be correct.

14 . The system of claim 8 , wherein each document image of the plurality of document images is associated with respective metadata defining a marked up document field location corresponding to the document field for each document image of the plurality of document images.

15 . A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to:

receive a plurality of document images, wherein each document image of the plurality of document images comprises a corresponding document field containing a variable text;

generate, by processing the plurality of document images, a first heat map represented by a data structure comprising a plurality of heat map elements corresponding to a plurality of document image pixels, wherein each heat map element stores a respective value specifying whether the corresponding document field contains a document image pixel associated with the heat map element;

receive an input document image; and

identify, using the first heat map, a candidate region of the input document image, the candidate region comprising the document field, wherein the candidate region comprises a plurality of input document image pixels corresponding to heat map elements satisfying a threshold condition.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein the candidate region is identified using a plurality of heat maps, the plurality of heat maps comprising the first heat map and one or more additional heat maps, wherein the each of the plurality of heat maps identify a potential document field location corresponding to the document field.

17 . The non-transitory computer-readable storage medium of claim 16 , wherein each of the plurality of heat maps identify the potential document field location relative to a respective reference element in each of the plurality of heat maps.

18 . The non-transitory computer-readable storage medium of claim 17 , wherein the respective reference element comprises one or more of a predefined word, or a predefined graphical element.

19 . The non-transitory computer-readable storage medium of claim 16 , wherein the processing device is further to:

extract a content of each document image of the plurality of document images, wherein the content is included in the potential document field location; and

analyze the content of each document image using Byte Pair Encoding (BPE) tokens.

20 . The non-transitory computer-readable storage medium of claim 15 , wherein each document image of the plurality of document images is associated with respective metadata defining a marked up document field location corresponding to the document field for each document image of the plurality of document images.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2023
From: SEMENOV, STANISLAV; LANIN, MIKHAIL
To: ABBYY DEVELOPMENT INC.
Reel/Frame 065528/0705 →
Priority Claims (1)
RU RU2020141790 · Dec 17, 2020 · national
Continuity (2)
Continuation 17129906 · Dec 21, 2020
Related Publication 20240078826A1 · Mar 7, 2024
References Cited (84)
US 5638491A · Moed · 1997 [cited by applicant]
US 6886136B1 · Zlotnick et al. · 2005 [cited by applicant]
US 7370034B2 · Franciosa et al. · 2008 [cited by applicant]
US 8265925B2 · Aarskog · 2012 [cited by applicant]
US 8595235B1 · Sampson et al. · 2013 [cited by applicant]
US 8726148B1 · Battilana · 2014 [cited by applicant]
US 8923618B2 · Kutsumi · 2014 [cited by applicant]
US 9613299B2 · Krivosheev et al. · 2017 [cited by applicant]
US 10013643B2 · Yellapragada et al. · 2018 [cited by applicant]
US 10360507B2 · Aravamudan et al. · 2019 [cited by applicant]
US 10467464B2 · Chen et al. · 2019 [cited by applicant]
US 10558712B2 · Zholudev et al. · 2020 [cited by applicant]
US 10679085B2 · Li et al. · 2020 [cited by applicant]
US 10872236B1 · Elor et al. · 2020 [cited by applicant]
US 10936863B2 · Simantov · 2021 [cited by examiner]
US 11074442B2 · Semenov · 2021 [cited by applicant]
US 11861925B2 · Semenov · 2024 [cited by examiner]
US 20060242610A1 · Aggarwal · 2006 [cited by applicant]
US 20070244915A1 · Cha et al. · 2007 [cited by applicant]
US 20080077572A1 · Boyle et al. · 2008 [cited by applicant]
US 20090210406A1 · Freire et al. · 2009 [cited by applicant]
US 20110093464A1 · Cvet et al. · 2011 [cited by applicant]
US 20130262465A1 · Galle et al. · 2013 [cited by applicant]
US 20150112874A1 · Serio et al. · 2015 [cited by applicant]
US 20160004667A1 · Chakerian et al. · 2016 [cited by applicant]
US 20160148074A1 · Jean et al. · 2016 [cited by applicant]
US 20160171627A1 · Lyubarskiy · 2016 [cited by applicant]
US 20170061250A1 · Gao et al. · 2017 [cited by applicant]
US 20170351781A1 · Alexander et al. · 2017 [cited by applicant]
US 20180181808A1 · Sridharan · 2018 [cited by applicant]
US 20180285448A1 · Chia et al. · 2018 [cited by applicant]
US 20180349743A1 · Iurii · 2018 [cited by applicant]
US 20190019503A1 · Henry · 2019 [cited by applicant]
US 20190180094A1 · Zagaynov et al. · 2019 [cited by applicant]
US 20190180154A1 · Orlov et al. · 2019 [cited by applicant]
US 20190205451A1 · Alipov et al. · 2019 [cited by applicant]
US 20190266394A1 · Yu et al. · 2019 [cited by applicant]
US 20190294874A1 · Orlov et al. · 2019 [cited by applicant]
US 20190294921A1 · Kalenkov · 2019 [cited by applicant]
US 20190311194A1 · Zhuravlev · 2019 [cited by applicant]
US 20190361972A1 · Lin · 2019 [cited by applicant]
US 20190385001A1 · Stark · 2019 [cited by applicant]
US 20200327351A1 · Abedini et al. · 2020 [cited by applicant]
US 20200327360A1 · Samala · 2020 [cited by applicant]
US 20200364451A1 · Ammar et al. · 2020 [cited by applicant]
US 20210012102A1 · Cristescu et al. · 2021 [cited by applicant]
US 20210019512A1 · Uppal et al. · 2021 [cited by applicant]
US 20210034853A1 · Matsumoto et al. · 2021 [cited by applicant]
US 20210064861A1 · Semenov · 2021 [cited by applicant]
US 20210064908A1 · Semenov · 2021 [cited by applicant]
US 20210124919A1 · Balakrishnan et al. · 2021 [cited by applicant]
US 20210149993A1 · Torres · 2021 [cited by examiner]
US 20210150338A1 · Semenov · 2021 [cited by applicant]
US 20210165964A1 · Jones · 2021 [cited by applicant]
US 20210182328A1 · Rollings et al. · 2021 [cited by applicant]
US 20210201013A1 · Makhija et al. · 2021 [cited by applicant]
US 20210271872A1 · Gupta · 2021 [cited by examiner]
US 20210295103A1 · Tanniru et al. · 2021 [cited by applicant]
CN 106649853A · 2017 [cited by applicant]
CN 106654853A · 2017 [cited by applicant]
CN 107168955A · 2017 [cited by applicant]
CN 107168955B · 2019 [cited by applicant]
EP 3437019B1 · 2020 [cited by applicant]
RU 2556425C1 · 2015 [cited by applicant]
RU 2661750C1 · 2018 [cited by applicant]
RU 2668717C1 · 2018 [cited by applicant]
RU 2679209C2 · 2019 [cited by applicant]
RU 2691214C1 · 2019 [cited by applicant]
RU 2693332C1 · 2019 [cited by applicant]
RU 2693916C1 · 2019 [cited by applicant]
RU 2737720C1 · 2020 [cited by applicant]
WO 2013135474A1 · 2013 [cited by applicant]
WO 2018126325A1 · 2018 [cited by applicant]
ProvisionalSpecificationofU.S. Appl. No. 62/983,302, filed Feb. 28, 2020;Guptaet.al;pp. 1-10 (Year: 2020). [cited by examiner]
Katti A.R., et al., “Applying Sequence-to-Mask Models for Information Extraction from Invoices,” DAS 2018 Short Papers Booklet , 13th IAPR International Workshop on Document Analysis Systems, Vienna, Austria, Apr. 24-27… [cited by applicant]
Ma E., “3 Subword Algorithms Help to Improve Your NLP Model Performance,” Introduction to subword, May 18, 2019, pp. 1-7, Retrieved from URL: https://medium.com/@makcedward/how-subword-helps-on-your-nlp-model-83dd1b836f… [cited by applicant]
Mozharova V A., et al., “Investigation of Features for Extraction of Named Entities From Texts in Russian,” 2017, 2 Pages, Retrieved from URL: https://patents.google.com/scholar/16285686872319998042?q=text+field+entry+w… [cited by applicant]
Palm R.B., et al., “CloudScan—A Configuration-Free Invoice Analysis System Using Recurrent Neural Networks,” IEEE, 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), arXiv: 1708.07403v… [cited by applicant]
Raoui-Outach R., et al., “Deep Learning for Automatic Sale Receipt Understanding,” IEEE, arXiv: 1712.01606v1 [cs.CV], Dec. 5, 2017, 08 Pages. [cited by applicant]
Sandhan J., et al., “Revisiting the Role of Feature Engineering for Compound Type Identification in Sanskrit,” Indian Institute of Technology, Kanpur, UP, India, 17 Pages, Retrieved from URL: https://www.aclweb.org/anth… [cited by applicant]
Zuyev K., et al., “Text Field Detection Using Neural Networks,” U.S. Appl. No. 16/017,683, filed Jun. 25, 2018, 57 Pages. [cited by applicant]
Forsati R., et al., “Efficient Stochastic Algorithms for Document Clustering,” Information Sciences, 2012, 21 Pages. [cited by applicant]
Maher K., et al., “Effectiveness of Different Similarity Measures for Text Classification and Clustering,” International Journal of Computer Science and Information Technologies, 2016, vol. 7(4), pp. 1715-1720. [cited by applicant]
Rashad M. A., et al., “Document Classification Using Enhanced Grid Based Clustering Algorithm,” Springer International Publishing, 2015, pp. 207-215. [cited by applicant]