IP Library › Granted Patent US 12,530,537
Granted Patent B2
US 12,530,537 · App. 18/113,903 · Granted Jan 20, 2026

Self-attentive key-value extraction

Inventors: Eduardo Vellasques (Stuttgart, CA); Xiang Yu (Berlin, DE); Stefan Klaus Baur (Heidelberg, DE); Manuel Zeise (Karlsruhe, DE)
Assignee: SAP SE
G06F40/40G06F16/3347G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,537
App. No.
18/113,903
Granted
Jan 20, 2026
Kind
B2
Abstract

Systems and methods are provided for automated identification of key-value pairs in documents. A document including readable text is received. The document is processed to determine, from the readable text, a plurality of tokens. Pairs of vectors corresponding to the plurality of tokens are determined, each pair of vectors comprising a query vector and a key vector. Attention scores are determined for the plurality of tokens by using the pairs of vectors. The attention scores are normalized to generate normalized attention scores. Connected tokens are identified in the plurality of tokens using the normalized attention scores.

Claims (43)

1 . A computer-implemented method comprising:

receiving, by one or more processors of a server system, a document comprising readable text;

processing, by the one or more processors, the document to determine, from the readable text, a plurality of tokens;

determining, by the one or more processors, pairs of vectors corresponding to the plurality of tokens, each pair of vectors comprising a query vector and a key vector;

determining, by the one or more processors, attention scores for the plurality of tokens by using the pairs of vectors;

normalizing, by the one or more processors, the attention scores to generate normalized attention scores;

identifying, by the one or more processors, connected tokens in the plurality of tokens using the normalized attention scores; and

causing, by the one or more processors, the connected tokens to be processed by an application configured to perform scanned document processing by providing the connected tokens as input to the application via the server system.

2 . The computer-implemented method of claim 1 , wherein the document is formatted as an image.

3 . The computer-implemented method of claim 2 , further comprising:

processing the image, using an optical character recognition application, to generate the readable text.

4 . The computer-implemented method of claim 3 , wherein the readable text comprises keys and values.

5 . The computer-implemented method of claim 1 , wherein the vectors comprise key vectors and query vectors.

6 . The computer-implemented method of claim 1 , wherein the normalized attention scores are generated using 1) a first predictive model trained to predict links between words belonging to a same instance and configured to generate instance scores as output, and 2) a second predictive model trained to predict links between words belonging to a key-value pair and configured to generate key-value scores as output.

7 . A non-transitory computer-readable storage medium comprising programming code, which when executed by at least one data processor of a server system, causes operations comprising:

receiving a document comprising readable text;

processing the document to determine, from the readable text, a plurality of tokens;

determining pairs of vectors corresponding to the plurality of tokens, each pair of vectors comprising a query vector and a key vector;

determining attention scores for the plurality of tokens by using the pairs of vectors;

normalizing the attention scores to generate normalized attention scores;

identifying connected tokens in the plurality of tokens using the normalized attention scores; and

causing the connected tokens to be processed by an application configured to perform scanned document processing by providing the connected tokens as input to the application via the server system.

8 . The non-transitory computer-readable storage medium of claim 7 , wherein the document is formatted as an image.

9 . The non-transitory computer-readable storage medium of claim 8 , wherein the operations further comprise:

processing the image, using an optical character recognition application, to generate the readable text.

10 . The non-transitory computer-readable storage medium of claim 9 , wherein the readable text comprises keys and values.

11 . The non-transitory computer-readable storage medium of claim 7 , wherein the vectors comprise key vectors and query vectors.

12 . The non-transitory computer-readable storage medium of claim 7 , wherein the normalized attention scores are generated using 1) a first predictive model trained to predict links between words belonging to a same instance and configured to generate instance scores as output, and 2) a second predictive model trained to predict links between words belonging to a key-value pair and configured to generate key-value scores as output.

13 . A system comprising:

at least one data processor of a server system; and

at least one memory storing instructions, which when executed by the at least one data processor, cause operations comprising:

receiving a document comprising readable text;

processing the document to determine, from the readable text, a plurality of tokens;

determining pairs of vectors corresponding to the plurality of tokens, each pair of vectors comprising a query vector and a key vector;

determining attention scores for the plurality of tokens by using the pairs of vectors;

normalizing the attention scores to generate normalized attention scores;

identifying connected tokens in the plurality of tokens using the normalized attention scores; and

causing the connected tokens to be processed by an application configured to perform scanned document processing by providing the connected tokens as input to the application via the server system.

14 . The system of claim 13 , wherein the document is formatted as an image.

15 . The system of claim 14 , wherein the operations further comprise:

processing the image, using an optical character recognition application, to generate the readable text, wherein the readable text comprises keys and values.

16 . The system of claim 13 , wherein the vectors comprise key vectors and query vectors.

17 . The system of claim 13 , wherein the normalized attention scores are generated using 1) a first predictive model trained to predict links between words belonging to a same instance and configured to generate instance scores as output, and 2) a second predictive model trained to predict links between words belonging to a key value pair and configured to generate key-value scores as output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2023
From: VELLASQUES, EDUARDO; YU, XIANG; BAUR, STEFAN KLAUS; ZEISE, MANUEL
To: SAP SE
Reel/Frame 062805/0808 →
Continuity (1)
Related Publication 20240289557A1 · Aug 29, 2024
References Cited (14)
US 10824808B2 · Reisswig · 2020 [cited by examiner]
US 11423093B2 · Xiong · 2022 [cited by examiner]
US 11557283B2 · Moritz · 2023 [cited by examiner]
US 12013902B2 · Xiong · 2024 [cited by examiner]
US 20200159828A1 · Reisswig · 2020 [cited by examiner]
US 20210089594A1 · Xiong · 2021 [cited by examiner]
US 20220310070A1 · Moritz · 2022 [cited by examiner]
US 20220374479A1 · Xiong · 2022 [cited by examiner]
US 20230376676A1 · Mallinson · 2023 [cited by examiner]
US 20240289557A1 · Vellasques · 2024 [cited by examiner]
CN 117994671A · 2024 [cited by examiner]
Devlin, J. et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv preprint (2018): arXiv:1810.04805. [cited by applicant]
Luong, M.-T. et al., “Effective Approaches to Attention-based Neural Machine Translation,” arXiv preprint (2015): arXiv:1508.04025. [cited by applicant]
Vaswani, A. et al., “Attention Is All You Need,” arXiv preprint (2017): arxiv:1706.03762. [cited by applicant]