IP Library › Granted Patent US 12,354,396
Granted Patent B2
US 12,354,396 · App. 18/490,652 · Granted Jul 8, 2025

System for information extraction from form-like documents

Inventors: Sandeep Tata (San Francisco, CA); Bodhisattwa Prasad Majumder (La Jolla, CA); Qi Zhao (Santa Clara, CA); James Bradley Wendt (San Francisco, CA); Marc Najork (Palo Alto, CA); Navneet Potti (Sunnyvale, CA)
Assignee: GOOGLE LLC
G06V30/413G06F18/21G06F18/22G06N5/04G06N20/00G06T7/70G06V30/274G06V30/412G06V30/416G06T2207/30176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,396
App. No.
18/490,652
Granted
Jul 8, 2025
Kind
B2
Abstract

The present disclosure is directed to extracting text from form-like documents. In particular, a computing system can obtain an image of a document that contains a plurality of portions of text. The computing system can extract one or more candidate text portions for each field type included in a target schema. The computing system can generate a respective input feature vector for each candidate for the field type. The computing system can generate a respective candidate embedding for the candidate text portion. The computing system can determine a respective score for each candidate text portion for the field type based at least in part on the respective candidate embedding for the candidate text portion. The computing system can assign one or more of the candidate text portions to the field type based on the respective scores.

Claims (39)

1. A computer-implemented method for extracting information from images of structured documents, the method comprising:

obtaining, by a computing system comprising one or more computing devices, an image of a document, wherein the image of the document includes one or more portions of text;

determining, by the computing system, a document type associated with the document;

accessing, by the computing system, a target schema associated with the document type, the target schema including one or more field types;

providing, by the computing system, the image of the document as input to a machine-learned model; and

receiving, by the computing system and from the machine-learned model, model output, the model output associating a respective portion of the document with a respective field type in the one or more field types, wherein the model output includes, for a plurality of candidate portions in the one or more portions of text, a score assigned to a respective candidate portion based on a degree to which it is associated with the respective field type in the one or more field types.

2. The computer-implemented method of claim 1 , wherein the respective portion of the document is selected to be associated with the respective field type based on one or more scores associated with one or more candidate portions of the document.

3. The computer-implemented method of claim 1 , wherein the score for the respective candidate portion is determined based on a similarity metric between the respective candidate portion and one or more field characteristics associated with the field type.

4. The computer-implemented method of claim 3 , wherein the similarity metric comprises a cosine similarity metric.

5. The computer-implemented method of claim 1 , wherein providing, by the computing system, the image of the document as input to a machine-learned model further comprises:

extracting, by the computing system from the image of the document, one or more candidate portions for each of one or more field types included in a target schema; and

providing, by the computing system, the extracted one or more candidate portions to the machine-learned model.

6. The computer-implemented method of claim 5 , wherein providing, by the computing system, the extracted one or more candidate portions to the machine-learned model further comprises:

providing, by the computing system, for a respective candidate portion, data describing a respective position of one or more neighbor text portions that are proximate to the respective candidate portion and data describing a relative normalized position of the one or more neighbor text portions relative to the respective candidate portion.

7. The computer-implemented method of claim 6 , wherein providing, by the computing system, for a respective candidate portion, data describing the respective position of one or more neighbor portions that are proximate to the respective candidate portion comprises:

defining, by the computing system, a respective neighborhood zone for each candidate portion; and

identifying, by the computing system, the one or more neighbor portions for each candidate portion based at least in part on the respective neighborhood zone for each candidate portion and the respective positions of the one or more neighbor portions.

8. The computer-implemented method of claim 7 , wherein defining, by the computing system, the respective neighborhood zone for each candidate portion comprises, for each candidate portion:

defining, by the computing system, the respective neighborhood zone to extend from a position of the candidate portion leftwards to a margin of the document and to extend from the position of the candidate portion upwards a threshold amount of the document.

9. The computer-implemented method of claim 5 , wherein providing, by the computing system, the extracted one or more candidate portions to the machine-learned model further comprises:

providing, by the computing system, for a respective candidate portion, data describing an absolute position of the respective candidate portion.

10. The computer-implemented method of claim 5 , wherein providing, by the computing system, the extracted one or more candidate portions to the machine-learned model further comprises:

providing, by the computing system, for a respective candidate portion, data describing text contained in the respective candidate portion.

11. A computing system for extracting information from images of structured documents, the system comprising:

one or more processors; and

a non-transitory computer-readable memory that stores instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining an image of a document, wherein the image of the document includes one or more portions of text;

determining a document type associated with the document;

accessing a target schema associated with the document type, the target schema including one or more field types;

providing the image of the document as input to a machine-learned model; and

receiving, from the machine-learned model, model output, the model output associating a respective portion of the document with a respective field type in the one or more field types, wherein the model output includes, for a plurality of candidate portions in the one or more portions of text, a score assigned to a respective candidate portion based on a degree to which it is associated with the respective field type in the one or more field types.

12. The computing system of claim 11 , wherein the respective portion of the document is selected to be associated with the respective field type based on one or more scores associated with one or more candidate portions of the document.

13. A non-transitory computer-readable medium storing instruction that, when executed by one or more computing devices, cause the one or more computing devices to perform operations, the operations comprising:

obtaining an image of a document, wherein the image of the document includes one or more portions of text;

determining a document type associated with the document;

accessing a target schema associated with the document type, the target schema including one or more field types;

providing the image of the document as input to a machine-learned model; and

receiving, from the machine-learned model, model output, the model output associating a respective portion of the document with a respective field type in the one or more field types, wherein the model output includes, for a plurality of candidate portions in the one or more portions of text, a score assigned to a respective candidate portion based on a degree to which it is associated with the respective field type in the one or more field types.

14. The non-transitory computer-readable medium of claim 13 , wherein the respective portion of the document is selected to be associated with the respective field type based on one or more scores associated with one or more candidate portions of the document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2023
From: TATA, SANDEEP; MAJUMDER, BODHISATTWA PRASAD; ZHAO, QI; WENDT, JAMES BRADLEY; NAJORK, MARC; POTTI, NAVNEET
To: GOOGLE LLC
Reel/Frame 065297/0367 →
Continuity (3)
Continuation 17867300 · Jul 18, 2022
Continuation 16890287 · Jun 2, 2020
Related Publication 20240046684A1 · Feb 8, 2024
References Cited (35)
US 5933823A · Cullen · 1999 [cited by applicant]
US 10395772B1 · Lucas · 2019 [cited by examiner]
US 20140003721A1 · Saund · 2014 [cited by examiner]
US 20170068866A1 · Kostyukov · 2017 [cited by applicant]
US 20170109610A1 · Macciola et al. · 2017 [cited by applicant]
US 20190332658A1 · Heckel et al. · 2019 [cited by applicant]
US 20200126533A1 · Doyle · 2020 [cited by examiner]
US 20200142999A1 · Pedersen · 2020 [cited by examiner]
US 20200159820A1 · Rodriguez et al. · 2020 [cited by applicant]
US 20200394763A1 · Ma et al. · 2020 [cited by applicant]
US 20210012102A1 · Cristescu · 2021 [cited by examiner]
US 20210012199A1 · Zhang · 2021 [cited by examiner]
US 20210090694A1 · Colley · 2021 [cited by examiner]
Cai et al., “Block-based Web Search”, 27 [cited by applicant]
Cai et al., “Extracting Content Structure for Web Pages based on Visual Representation”, 5 [cited by applicant]
Chiticariu et al., “Rule-based Information Extraction is Dead! Long Live Rule-based Information Extraction Systems!”, 2013 Conference on Empirical Methods in Natural Language Processing, Oct. 18-21, 2013, Seattle, WA, p… [cited by applicant]
Chowdhury, “Template Mining for Information Extraction from Digital Documents”, Library Trends, vol. 48, No. 1, Summer 1999, pp. 182-208. [cited by applicant]
Dalvi et al., “Automatic Wrappers for Large Scale Web Extraction”, 37 [cited by applicant]
Denk et al., “BERTgrid: Contextualized Embedding for 2D Document Representation and Understanding”, arXiv:1909.04948v2, Oct. 14, 2019, 4 pages. [cited by applicant]
Doermann et al., “The Processing of Form Documents”, 2 [cited by applicant]
Duchi et al., “Adaptive Subgradient Methods for Online Learning and Stochastic Optimization”, Journal of Machine Learning Research, vol. 12, 2011, pp. 2121-2159. [cited by applicant]
Hobbs et al., “Information Extraction”, Handbook of Natural Language Processing, Chapman & Hall/CRC Press, 2010, pp. 511-532. [cited by applicant]
Katti et al., “Chargrid: Towards Understanding 2D Documents”, 2018 Conference on Empirical Methods in Natural Language Processing, Oct. 31,-Nov. 4, 2018, Brussels, Belgium, pp. 4459-4469. [cited by applicant]
Lample et al., “Neural Architectures for Named Entity Recognition”, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun. 12-17, 2016, San … [cited by applicant]
Liu et al., “Graph Convolution for Multimodal Information Extraction from Visually Rich Documents”, 2019 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Jun. 2-7, 2019, … [cited by applicant]
Palm et al., “CloudScan—A configuration-free invoice analysis system using recurrent neural networks”, arXiv:1708.07403v1, Aug. 24, 2017, 8 pages. [cited by applicant]
Peng et al., “Cross-Sentence N-ary Relation Extraction with Graph LSTMs”, Transactions of the Association for Computational Linguistics, vol. 5, Apr. 2017, pp. 101-115. [cited by applicant]
Sarawagi, “Information Extraction” Foundations and Trends in Databases, vol. 1, No. 3, 2008, pp. 261-377. [cited by applicant]
Schuster et al., “Intellix—End-User Trained Information Extraction for Document Archiving”, 12 [cited by applicant]
Vandermaaten et al., “Visualizing Data using t-SNE”, Journal of Machine Learning, vol. 9, No. 86, 2008, pp. 2579-2605. [cited by applicant]
Vaswani et al., “Attention Is All You Need”, 31 [cited by applicant]
Yu et al., “Improving Psuedo-Relevance Feedback in Web Information Retrieval Using Web Page Segmentation”, The Twelfth International World Wide Web Conference, May 20-24, 2003, Budapest, Hungary, 8 pages. [cited by applicant]
Zhao et al., “CUTIE: Learning to Understand Documents with Convolutional Universal Text Information Extractor”, arXIV:1903.12363v2, Apr. 4, 2019, 9 pages. [cited by applicant]
Zhu et al., “2D Conditional Random Fields for Web Information Extraction”, 22 [cited by applicant]
Zhu et al., “Simultaneous Record Detection and Attribute Labeling in Web Data Extraction”, Twelfth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Aug. 20-23, 2006, Philadelphia, PA, pp. 494-… [cited by applicant]