IP Library › Granted Patent US 12,586,404
Granted Patent B2
US 12,586,404 · App. 18/239,778 · Granted Mar 24, 2026

Method and system for relevant data extraction from a document

Inventors: Nirmal Ramesh Rayulu Vanapalli Venkata (Visakhapatnam, IN); Madhusudan Singh (Bangalore, IN); Tamilarasan Ellappan (Tiruvannamalai District, IN)
Assignee: L&T TECHNOLOGY SERVICES LIMITED
G06V30/414G06F40/169G06F40/186G06V10/945G06V20/62G06V30/19013G06V30/19147G06V30/1916
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,404
App. No.
18/239,778
Granted
Mar 24, 2026
Kind
B2
Abstract

A method and system for relevant data extraction from a document is disclosed. The method includes determining first positional information corresponding to a key from a plurality of predefined keys in the document image based on a deep learning model. Further, second positional information corresponding to the key is determined based on OCR of the document image and an NLP model. Final positional information is determined based on the first positional information and the second positional information, in case a difference between the first positional information and the second positional information is minimal. Relevant data is extracted for the key in the OCR document image based on the final positional information.

Claims (54)

1 . A method of extracting relevant data from a document image, the method comprising:

determining, by a processor, first positional information comprising at least one first bounding box, corresponding to at least one key from a plurality of predefined keys in the document image based on a deep learning model,

wherein the deep learning model is trained based on user-inputted predefined mapping information for a plurality of templates, and

wherein the first positional information is determined based on the user-inputted predefined mapping information corresponding to each of the plurality of predefined keys;

determining, by the processor, second positional information comprising one or more second bounding boxes, corresponding to the at least one key from the plurality of predefined keys based on an optical character recognition (OCR) of the document image and an NLP model,

wherein the NLP model is trained based on the plurality of templates corresponding to the plurality of predefined keys, and

wherein the second positional information is further processed based on a plurality of pre-defined rules comprising a predefined mapping information for each of the plurality of predefined keys in each of the plurality of templates;

determining, by the processor, final positional information for the at least one key based on comparison of the first positional information and the second positional information, in case a difference between the first positional information and the second positional information is minimal; and

extracting, by the processor, relevant data for the at least one key in the OCR document image based on the final positional information.

2 . The method of claim 1 , wherein the predefined mapping information comprises a plurality of attributes and a direction information corresponding to each of the plurality of predefined keys,

wherein the directional information comprises a direction in which a relevant data is present corresponding to each of the plurality of predefined keys in each of the plurality of templates, and

wherein the second positional information is determined based on determination of at least one of the plurality of attributes based on the OCR of the document image.

3 . The method of claim 2 , wherein the first positional information corresponding to the at least one key, comprises the at least one first bounding box determined for each of the plurality of templates.

4 . The method of claim 3 , wherein the user inputted predefined mapping information for each of the plurality of templates is determined based on user defined annotation in each of the of the plurality of templates for each of the plurality of predefined keys,

wherein the second positional information corresponding to the at least one key, comprises the one or more second bounding boxes determined for each of the plurality of templates, and

wherein the one or more second bounding boxes are determined based on the plurality of pre-defined rules.

5 . The method of claim 4 , wherein the final positional information for the at least one key is determined based on a comparison of the at least one first bounding box with each of the one or more second bounding boxes determined for each of the plurality of templates.

6 . The method of claim 5 , comprises: extracting, by the processor, a final data comprising the relevant text and the at least one attribute, from the OCR document image corresponding to the at least one key based on a validation of the relevant data, wherein the relevant data is validated based on pre-defined validation rules.

7 . A system for extracting relevant data from a document image, the system comprising:

a processor; and

a memory communicatively coupled to the processor, wherein the memory stores processor-executable instructions, which, on execution by the processor, cause the processor to:

determine first positional information comprising at least one first bounding box, corresponding to at least one key from a plurality of predefined keys in the document image based on a deep learning model,

wherein the deep learning model is trained based on user-inputted predefined mapping information for a plurality of templates, and

wherein the first positional information is determined based on the user-inputted predefined mapping information corresponding to each of the plurality of predefined keys;

determine second positional information comprising one or more second bounding boxes, corresponding to the at least one key from the plurality of predefined keys based on an optical character recognition (OCR) of the document image and an NLP model,

wherein the NLP model is trained based on the plurality of templates corresponding to the plurality of predefined keys, and

wherein the second positional information is further processed based on a plurality of pre-defined rules comprising a predefined mapping information for each of the plurality of predefined keys in each of the plurality of templates;

determine final positional information for the at least one key based on comparison of the first positional information and the second positional information, in case a difference between the first positional information and the second positional information is minimal; and

extract relevant data for the at least one key in the OCR document image based on the final positional information.

8 . The system of claim 7 , wherein the predefined mapping information comprises a plurality of attributes and a direction information corresponding to each of the plurality of predefined keys,

wherein the directional information comprises a direction in which a relevant data is present corresponding to each of the plurality of predefined keys in each of the plurality of templates, and

wherein the second positional information is determined based on determination of at least one of the plurality of attributes based on the OCR of the document image.

9 . The system of claim 8 , wherein the first positional information corresponding to the at least one key, comprises the at least one first bounding box determined for each of the plurality of templates,

wherein the second positional information corresponding to the at least one key, comprises the one or more second bounding boxes determined for each of the plurality of templates, wherein the one or more second bounding boxes are determined based on the plurality of pre-defined rules, and

wherein the final positional information for the at least one key is determined based on a comparison the at least one first bounding box with each of the one or more second bounding boxes determined for each of the plurality of templates.

10 . The system of claim 9 , wherein the processor is configured to extract a final data comprising the relevant text and the at least one attribute, from the OCR document image corresponding to the at least one key based on a validation of the relevant data, wherein the relevant data is validated based on pre-defined validation rules.

11 . A non-transitory computer-readable medium storing computer-executable instructions for extracting relevant data from a document image, the computer-executable instructions configured for:

determining first positional information comprising at least one first bounding box, corresponding to at least one key from a plurality of predefined keys in the document image based on a deep learning model,

wherein the deep learning model is trained based on user-inputted predefined mapping information for a plurality of templates, and

wherein the first positional information is determined based on the user-inputted predefined mapping information corresponding to each of the plurality of predefined keys;

determining second positional information comprising one or more second bounding boxes, corresponding to the at least one key from the plurality of predefined keys based on an optical character recognition (OCR) of the document image and an NLP model,

wherein the NLP model is trained based on the plurality of templates corresponding to the plurality of predefined keys, and

wherein the second positional information is further processed based on a plurality of pre-defined rules comprising a predefined mapping information for each of the plurality of predefined keys in each of the plurality of templates;

determining final positional information for the at least one key based on comparison of the first positional information and the second positional information, in case a difference between the first positional information and the second positional information is minimal; and

extracting relevant data for the at least one key in the OCR document image based on the final positional information.

12 . The non-transitory computer-readable medium of claim 11 , wherein the predefined mapping information comprises a plurality of attributes and a direction information corresponding to each of the plurality of predefined keys,

wherein the directional information comprises a direction in which a relevant data is present corresponding to each of the plurality of predefined keys in each of the plurality of templates, and

wherein the second positional information is determined based on determination of at least one of the plurality of attributes based on the OCR of the document image.

13 . The non-transitory computer-readable medium of claim 12 , wherein the first positional information corresponding to the at least one key, comprises the at least one first bounding box determined for each of the plurality of templates.

14 . The non-transitory computer readable medium of claim 13 , wherein the user inputted predefined mapping information for each of the plurality of templates is determined based on user defined annotation in each of the of the plurality of templates for each of the plurality of predefined keys,

wherein the second positional information corresponding to the at least one key, comprises the one or more second bounding boxes determined for each of the plurality of templates, and

wherein the one or more second bounding boxes are determined based on the plurality of pre-defined rules.

15 . The non-transitory computer readable medium of claim 14 , wherein the final positional information for the at least one key is determined based on a comparison of the at least one first bounding box with each of the one or more second bounding boxes determined for each of the plurality of templates.

16 . The non-transitory computer readable medium of claim 15 , comprises: extracting, by the processor, a final data comprising the relevant text and the at least one attribute, from the OCR document image corresponding to the at least one key based on a validation of the relevant data, wherein the relevant data is validated based on pre-defined validation rules.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 31, 2023
From: VANAPALLI VENKATA, NIRMAL RAMESH RAYULU; SINGH, MADHUSUDAN; ELLAPPAN, TAMILARASAN
To: L&T TECHNOLOGY SERVICES LIMITED
Reel/Frame 064761/0596 →
Priority Claims (1)
IN 202341028817 · Apr 20, 2023 · national
Continuity (1)
Related Publication 20240355136A1 · Oct 24, 2024
References Cited (5)
US 20220121821A1 · Yaramada · 2022 [cited by examiner]
US 20240242527A1 · Sathi · 2024 [cited by examiner]
US 20240312232A1 · Wang · 2024 [cited by examiner]
Dream Haddad; OCR Table Extraction Using Deep Learning (DL); Acodis website, Jul. 21, 2022. [cited by applicant]
Diego Leon; Extracting Information From PDF Invoices Using Deep Learning; Aug. 17, 2021; pp. 1-61; Stockholm; Sweden. [cited by applicant]