IP Library › Granted Patent US 12,597,281
Granted Patent B2
US 12,597,281 · App. 17/947,737 · Granted Apr 7, 2026

Image and semantic based table recognition

Inventors: Jiuxiang Gu (Greenbelt City, MD); Vlad Morariu (Potomac, MD); Tong Sun (San Ramon, CA); Jason wen yong Kuen (Santa Clara, CA); Ani Nenkova (Philadelphia, PA)
Assignee: Adobe Inc.
G06V30/412G06V30/262G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,281
App. No.
17/947,737
Granted
Apr 7, 2026
Kind
B2
Abstract

In various examples, a table recognition model receives an image of a table and generates, using a first encoder of the table recognition machine learning model, an image feature vector including features extracted from the image of the table; generates, using a first decoder of the table recognition machine learning model and the image feature vector, a set of coordinates within the image representing rows and columns associated with the table, and generates, using a second decoder of the table recognition machine learning model and the image feature vector, a set of bounding boxes and semantic features associated with cells the table, then determines, using a third decoder of the table recognition machine learning model, a table structure associated with the table using the image feature vector, the set of coordinates, the set of bounding boxes, and the semantic features.

Claims (38)

1 . A method comprising:

receiving an image of a table;

generating, using a first encoder of a table recognition machine learning model, an image feature vector including features extracted from the image of the table;

generating, using a first decoder of the table recognition machine learning model and the image feature vector, a set of coordinates within the image representing rows and columns associated with the table;

generating, using a second decoder of the table recognition machine learning model and the image feature vector, a set of bounding boxes and semantic features associated with cells the table, wherein the second decoder includes an optical character recognition (OCR) decoder; and

determining, using a third decoder of the table recognition machine learning model, a table structure associated with the table using the image feature vector, the set of coordinates, the set of bounding boxes, and the semantic features.

2 . The method of claim 1 , wherein the table structure further comprises a label associated with a cell of the table.

3 . The method of claim 1 , wherein the table recognition machine learning model includes a transformer network.

4 . The method of claim 1 , wherein the first decoder includes a self-attention layer.

5 . The method of claim 1 , wherein the method further comprises performing a task based on the table structure, the task including at least one of: table question answering, table fact verification, table formatting, table manipulation, and table captioning.

6 . The method of claim 1 , wherein determining the table structure further comprises determine a relationship between the cells of the table by at least performing a pair-wise comparison of features associated with the cells of the table.

7 . The method of claim 1 , wherein the method further comprises fine tuning the table recognition machine learning model by at least performing a bipartite matching between predictions generated by the table recognition machine learning model and a ground truth associated with the table.

8 . The method of claim 1 , wherein the method further comprises causing the table recognition machine learning model to execute a self-supervised pre-training task based on the table.

9 . The method of claim 8 , wherein the self-supervised pre-training task further comprises causing the table recognition machine learning model to predict semantic information based on a modified image of the table generated by at least masking a set of pixels of the image corresponding to a cell of the table.

10 . A non-transitory computer-readable medium storing executable instructions embodied thereon, which, when executed by a processing device, cause the processing device to perform operations comprising:

determining, using a table recognition machine learning model, a first set of features based on an input image depicting a table;

determining, based on the first set of features, a first set of predictions representing rows and columns associated with the table;

determining, based on the first set of features, semantic information included in the table and a set of bounding boxes corresponding to the semantic information;

determining, using the table recognition machine learning model, a set of labels associated with cells of the table based on the first set of feature, the first set of predictions, the semantic information, and the set bounding boxes; and

causing the table recognition machine learning model to execute a self-supervised pre-training task based on the table by at least causing the table recognition machine learning model to predict at least a portion of the semantic information based on a modified image of the table generated by at least masking a set of pixels of the image corresponding to a cell of the table.

11 . The medium of claim 10 , wherein determining the first set of features is performed by an encoder of the table recognition machine learning model.

12 . The medium of claim 10 , wherein determining the set of labels is performed by a decoder of the table recognition machine learning model.

13 . The medium of claim 10 , wherein the self-supervised pre-training task further comprises causing the table recognition machine learning model to predict a modification to a row or a column of the table.

14 . The medium of claim 13 , wherein the modification includes at least one of: merging rows of the table, adding rows to the table, merging columns of the table, adding columns to the table, merging spans of the table, and adding spans to the table.

15 . The medium of claim 10 , wherein the self-supervised pre-training task further comprises causing the table recognition machine learning model to predict a rotated angle associated with the table based on the input image.

16 . A system comprising:

a memory component; and

a processing device coupled to the memory component, the processing device to perform operations comprising:

pre-training a table recognition machine learning model by at least causing the table recognition machine learning model to perform a set of self-supervision tasks to generate a pre-trained table recognition machine learning model, wherein the set of self-supervision tasks includes at least:

predicting semantic values corresponding to a cell of a table based on a set of masked pixels included in an image of the table corresponding to the cell;

predicting a set of column and a set of row associated with the table based on modifications to rows and columns of the table; and

identifying a rotated angle associated with a table based on a rotated image of the table;

training the pre-trained table recognition machine learning model by at least fine tuning the pre-trained table recognition machine learning model to generate a trained table recognition machine learning model; and

using the trained table recognition machine learning model to perform a task.

17 . The system of claim 16 , wherein the table recognition machine learning model includes a transformer network.

18 . The system of claim 16 , wherein fine tuning the pre-trained table recognition machine learning model further comprises performing a bipartite matching between predictions generated by the pre-trained trained table recognition machine learning model and a ground truth associated with the table.

19 . The system of claim 16 , wherein fine tuning the pre-trained table recognition machine learning model further comprises performing binary classification based on a result of predicting the set of columns and the set of rows and a ground truth associated with the table.

20 . The system of claim 16 , wherein the task includes at least one of: table question answering, table fact verification, table formatting, table manipulation, and table captioning.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE THIRD INVENTOR'S NAME PREVIOUSLY RECORDED AT REEL: 061139 FRAME: 0625. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT . Recorded Sep 26, 2022
From: GU, JIUXIANG; MORARIU, VLAD; SUN, TONG; KUEN, JASON WEN YOUNG; NENKOVA, ANI
To: ADOBE INC.
Reel/Frame 061604/0513 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2022
From: GU, JIUXIANG; MORARIU, VLAD; SUNG, TONG; KUEN, JASON WEN YOUNG; NENKOVA, ANI
To: ADOBE INC.
Reel/Frame 061139/0625 →
Continuity (1)
Related Publication 20240104951A1 · Mar 28, 2024
References Cited (32)
US 20200097759A1 · Nadim · 2020 [cited by examiner]
Zhang, et al. (Split, Embed, and Merge: An accurate table structure recognizer). (Year: 2022). [cited by examiner]
Nassar, et al. (TableFormer: Table Structure Understanding with Transformers). (Year: 2022). [cited by examiner]
Iida, et al. (TABBIE: Pretrained Representations of Tabular data (Year: 2021). [cited by examiner]
Bodla, N., et al., “Soft-NMS—Improving Object Detection With One Line of Code”, In Proceedings of the IEEE International Conference on Computer Vision, pp. 5561-5569 (2017). [cited by applicant]
Chi, Z., et al., “Complicated Table Structure Recognition”, arXiv:1908.04729v2, pp. 1-9 (Aug. 28, 2019). [cited by applicant]
Coüasnon, B., and Lemaitre, A., “Recognition of Tables and Forms”, Handbook of Document Image Processing and Recognition, pp. 647-677 (2014). [cited by applicant]
Devlin, J., et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Proceedings of NAACL-HLT, Association for Computational Linguistics, pp. 4171-4186 (Jun. 2-7, 2019). [cited by applicant]
Erhan, D., et al., “Scalable Object Detection using Deep Neural Networks”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1-8 (2014). [cited by applicant]
Gao, L., et al., “ICDAR 2019 Competition on Table Detection and Recognition (cTDaR)”, International Conference on Document Analysis and Recognition (ICDAR), pp. 1510-1515 (2019). [cited by applicant]
Göbel, M., et al., “ICDAR 2013 Table Competition”, In 12th International Conference on Document Analysis and Recognition, pp. 1449-1453 (2013). [cited by applicant]
Hosang, J., et al., “Learning non-maximum suppression”, In Proceedings of the IEEE Conference On Computer Vision and Pattern Recognition, pp. 4507-4515 (2017). [cited by applicant]
Hu, H., et al., “Relation Networks for Object Detection”, In Proceedings of the IEEE Conference On Computer Vision and Pattern Recognition, pp. 3588-3597 (2018). [cited by applicant]
Hurst, M., “A Constraint-based Approach to Table Structure Derivation”, Proceedings of the Seventh International Conference on Document Analysis and Recognition (ICDAR), IEEE Computer Society, vol. 2, pp. 1-5 (2003). [cited by applicant]
Itonori, K., “Table Structure Recognition based on Textblock Arrangement and Ruled Line Position”, In Proceedings of 2nd International Conference on Document Analysis and Recognition (ICDAR), pp. 765-768 (1993). [cited by applicant]
Kieninger, T. G., “Table Structure Recognition Based On Robust Block Segmentation”, German Research Center for Artificial Intelligence, pp. 1-11 (1998). [cited by applicant]
Lewis, M., et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension”, arXiv:1910.13461v1, pp. 1-10 (Oct. 29, 2019). [cited by applicant]
Lin, T-Y., et al., “Focal Loss for Dense Object Detection”, In Proceedings of the IEEE International Conference on Computer Vision, pp. 2980-2988 (2017). [cited by applicant]
Long, J., et al., “Fully Convolutional Networks for Semantic Segmentation”, In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition, pp. 3431-3440 (2015). [cited by applicant]
Lüscher, C., et al., “RWTH ASR Systems for LibriSpeech: Hybrid vs Attention—w/o Data Augmentation”, arXiv:1905.03072v3, pp. 1-5 (Jul. 25, 2019). [cited by applicant]
Parmar, N., et al., “Image Transformer”, Proceedings of the 35th International Conference on Machine Learning (PMLR), pp. 1-10 (2018). [cited by applicant]
Pawlik, M., and Augsten, N., “Tree edit distance: Robust and memory-efficient”, Information Systems, pp. 1-17 (Aug. 22, 2015). [cited by applicant]
Radford, A., et al., “Language Models are Unsupervised Multitask Learners”, OpenAI Blog, vol. 1, No. 8, pp. 1-24 (2019). [cited by applicant]
Raja, S., et al., “Table Structure Recognition using Top-Down and Bottom-Up Cues”, arXiv:2010.04565v1, pp. 1-38 (Oct. 9, 2020). [cited by applicant]
Redmon, J., et al., “You Only Look Once: Unified, Real-Time Object Detection”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 779-788 (2016). [cited by applicant]
Ren, S., et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, Advances in Neural Information Processing Systems (NIPS), pp. 1-9 (2015). [cited by applicant]
Schreiber, S., et al., “DeepDeSRT: Deep Learning for Detection and Structure Recognition of Tables in Document Images”, In 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), vol. 1, pp. 1-6… [cited by applicant]
Siddiqui, S. A., et al., “Rethinking Semantic Segmentation for Table Structure Recognition in Documents”, In International Conference on Document Analysis and Recognition (ICDAR), pp. 1-6 (2019). [cited by applicant]
Tensmeyer, C., et al., “Deep Splitting and Merging for Table Structure Decomposition”, International Conference on Document Analysis and Recognition (ICDAR), pp. 114-121 (2019). [cited by applicant]
Zhang, S., et al., “Bridging the Gap Between Anchor-based and Anchor-free Detection via Adaptive Training Sample Selection”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9759… [cited by applicant]
Zhong, X., et al., “Image-based table recognition: data, model, and evaluation”, Computer Vision—ECCV, pp. 1-17 (2020). [cited by applicant]
Zhou, X., et al., “Objects as Points”, arXiv:1904.07850v2, pp. 1-12 (Apr. 25, 2019). [cited by applicant]