IP Library › Granted Patent US 12,602,713
Granted Patent B2
US 12,602,713 · App. 18/480,783 · Granted Apr 14, 2026

Deduction claim document parsing engine

Inventors: Debashish Sahu (Hyderabad, IN); Souranil De (Hyderabad, IN); Ritwik Upadhyay (Hyderabad, IN); Srinivas Rapaka (Hyderabad, IN); Susanta Kumar Sahoo (Hyderabad, IN); Lohit Vankina (Hyderabad, IN)
Assignee: HighRadius Corporation
G06Q30/04G06F40/174G06V30/412G06V30/413
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,713
App. No.
18/480,783
Granted
Apr 14, 2026
Kind
B2
Abstract

The present invention is related to data processing methods and systems thereof. According to an embodiment, the present invention provides a method of processing claim deduction documents using a machine learning model. The process begins by accessing data files and extracting information from them, which is subsequently stored. This document information, along with the machine learning model trained on various document formats, is used to classify the data files and generate tabular data. From this tabular data, data objects are created and included in an output data file. The information from the output file is then used to update the data of the machine learning model, optimizing it for improved future document processing. There are other embodiments as well.

Claims (65)

1 . A method for processing documents, the method comprising:

receiving a first document containing a deduction claim;

extracting document information from the first document, wherein the document information comprises at least one of tabular data or non-tabular data;

classifying the document information using a first machine learning model, the first machine learning model being trained using a plurality of document formats;

generating one or more data objects associated with the deduction claim based on the classification of the document information using the first machine learning model,

wherein generating the one or more data objects associated with the deduction claim based on the tabular data comprises:

processing the tabular data with two or more tabular data extraction modules, wherein the two or more tabular data extraction modules process the tabular data in parallel; and

generating the one or more data objects associated with the deduction claim based on the tabular data processed by the two or more tabular data extraction modules;

wherein generating the one or more data objects associated with the deduction claim based on the non-tabular data comprises:

processing the non-tabular data with two or more key-value pair data extraction modules, wherein the two or more key-value pair data extraction modules process the non-tabular data in parallel; and

generating the one or more data objects associated with the deduction claim based on the non-tabular data processed by the two or more key-value pair data extraction modules;

providing an output data file comprising the one or more data objects; and

updating the first machine learning model using at least the output data file.

2 . The method of claim 1 , wherein the document information comprises at least one of customer information, entity information, invoice information, deduction claim information, supporting information, contact information, signature information, or date information.

3 . The method of claim 1 , wherein classifying the document information using the first machine learning model comprises:

identifying within the document information at least one of tabular data or non-tabular data; and

classifying using the first machine learning model the document information based on an identification of at least one of tabular data or non-tabular data.

4 . The method of claim 3 , wherein the non-tabular data comprises at least one of key-value pairs or paragraph data.

5 . The method of claim 1 , wherein the document information comprises both tabular data and non-tabular data, wherein classifying the document information using the first machine learning model comprises:

identifying document information associated with the tabular data and the non-tabular data;

extracting using a first extraction module the tabular data; and

extracting using a second extraction module the non-tabular data.

6 . The method of claim 5 , further comprising generating a table using the one or more data objects, the table comprising one or more first data objects generated based on the extraction of the tabular data and one or more second data objects generated based on the extraction of the non-tabular data.

7 . The method of claim 1 , wherein, when the two or more tabular data extraction modules generate repeating or redundant first data objects, the one or more repeating or redundant first data objects are removed based on an order of preference associated with the two or more tabular data extraction modules and wherein, when the two or more key-value pair data extraction modules generate repeating or redundant second data objects, the one or more repeating or redundant second data objects are removed based on an order of preference associated with the two or more key-value pair data extraction modules.

8 . The method of claim 1 , wherein classifying using the first machine learning model the document information further comprises classifying the document information into at least one of one or more keys or one or more values;

wherein generating the one or more data objects further comprises defining at least one data object of the one or more data objects as at least one of a key or a value; and

wherein the method further comprises associating the at least one data object using a second machine learning model with one or more attributes of the data output file.

9 . The method of claim 8 , wherein associating the one or more data objects with the one or more attributes of the data output file further comprises:

associating the at least one data object with one or more attributes of the data output file based on whether the at least one data object is defined as at least one of the key or the value;

based on a determination that the at least one data object is defined as the key, matching the key with a first header of the data output file using the second machine learning algorithm; and

based on a determination that the at least one data object is defined as the value, determining a corresponding key using the second machine learning algorithm associated with the value from the document information or the one or more data objects, matching using the second machine learning algorithm the corresponding key with a corresponding header of the data output file, and adding using the second machine learning algorithm the value to the corresponding header of the data output file.

10 . The method of claim 1 , further comprising generating a table using the one or more data objects, the table comprising a header that is based at least on the document information.

11 . The method of claim 10 , wherein the header is at least one of customer information, entity information, invoice information, deduction claim information, supporting information, contact information, signature information, or date information.

12 . A method for processing documents, the method comprising:

extracting document information from one or more deduction claim documents by a data extraction module, wherein the document information comprises at least one of tabular data or non-tabular data;

classifying the one or more deduction claim documents or the document information using a machine learning model, the machine learning model being trained using a plurality of document formats;

generating one or more data objects associated with at least one corresponding deduction claim based on the classification of the document information using the machine learning model,

wherein generating the one or more data objects associated with the deduction claim based on the tabular data comprises:

processing the tabular data with two or more tabular data extraction modules, wherein the two or more tabular data extraction modules process the tabular data in parallel; and

generating the one or more data objects associated with the deduction claim based on the tabular data processed by the two or more tabular data extraction modules;

wherein generating the one or more data objects associated with the deduction claim based on the non-tabular data comprises:

processing the non-tabular data with two or more key-value pair data extraction modules, wherein the two or more key-value pair data extraction modules process the non-tabular data in parallel; and

generating the one or more data objects associated with the deduction claim based on the non-tabular data processed by the two or more key-value pair data extraction modules;

providing an output data file comprising the one or more data objects;

providing an accuracy assessment for the at least one corresponding deduction claim by comparing the output data file to reference data or ground truth data; and

modifying the machine learning model using at least the accuracy assessment.

13 . The method of claim 12 , further comprising identifying error patterns associated with the output data file, wherein the error patterns comprise at least one error pattern associated with at least one of customer information, entity information, invoice information, deduction claim information, supporting information, contact information, signature information, or date information for the at least one corresponding deduction claim.

14 . The method of claim 13 , further comprising modifying the machine learning model using at least the error patterns.

15 . The method of claim 13 , further comprising modifying the data extraction module using at least the error patterns.

16 . The method of claim 12 , wherein comparing the output data file to reference data or ground truth data comprises comparing one or more data objects in the output data file to at least one customer information, entity information, invoice information, deduction claim information, supporting information, contact information, signature information, or date information stored in a data storage and associated with the at least one corresponding deduction claim.

17 . The method of claim 12 , further comprising classifying the document information into tabular data, key-value pair data, value data, or paragraph data.

18 . A method for processing documents, the method comprising:

accessing one or more deduction claims;

extracting document information from the one or more deduction claims, wherein the document information comprises at least one of tabular data or non-tabular data;

classifying the document information using a first machine learning model, the first machine learning model being trained using a plurality of document formats;

generating one or more data objects associated with at least one corresponding deduction claim and based on the classification of the document information using the first machine learning model,

wherein generating the one or more data objects associated with the deduction claim based on the tabular data comprises:

processing the tabular data with two or more tabular data extraction modules, wherein the two or more tabular data extraction modules process the tabular data in parallel; and

generating the one or more data objects associated with the deduction claim based on the tabular data processed by the two or more tabular data extraction modules;

wherein generating the one or more data objects associated with the deduction claim based on the non-tabular data comprises:

processing the non-tabular data with two or more key-value pair data extraction modules, wherein the two or more key-value pair data extraction modules process the non-tabular data in parallel; and

generating the one or more data objects associated with the deduction claim based on the non-tabular data processed by the two or more key-value pair data extraction modules;

associating the one or more data objects using a second machine learning model with one or more attributes of an output data file; and

generating the output data file based on the second machine learning module comprising the one or more data objects associated with the one or more attributes of the output data file.

19 . The method of claim 18 , further comprising prioritizing or ordering two or more corresponding deduction claims in the output data file based on at least one of customer information, entity information, invoice information, deduction claim information, supporting information, contact information, signature information, or date information associated with the at least one corresponding deduction claim.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2024
From: SAHOO, SUSANTA KUMAR; UPADHYAY, RITWIK; DE, SOURANIL; SAHU, DEBASHISH; RAPAKA, SRINIVAS; VANKINA, LOHIT
To: HIGHRADIUS CORPORATION
Reel/Frame 067425/0540 →
Continuity (1)
Related Publication 20250117833A1 · Apr 10, 2025
References Cited (21)
US 10810420B2 · Bhatnagar · 2020 [cited by examiner]
US 12087068B2 · Rossi · 2024 [cited by examiner]
US 20080252924A1 · Gangai · 2008 [cited by examiner]
US 20150106257A1 · Maczuszenko · 2015 [cited by examiner]
US 20200074169A1 · Mukhopadhyay · 2020 [cited by examiner]
US 20200125592A1 · Teruya et al. · 2020 [cited by applicant]
US 20220309549A1 · Xu · 2022 [cited by examiner]
US 20220392047A1 · Wheaton · 2022 [cited by examiner]
US 20220414328A1 · Nguyen · 2022 [cited by examiner]
US 20230065915A1 · Berestovsky · 2023 [cited by examiner]
US 20230095673A1 · Dharmasiri et al. · 2023 [cited by applicant]
US 20230113578A1 · Kumar et al. · 2023 [cited by applicant]
US 20230267273A1 · Theriappan et al. · 2023 [cited by applicant]
US 20240419742A1 · Marcum · 2024 [cited by examiner]
US 20250181650A1 · Wang · 2025 [cited by examiner]
US 20250191308A1 · Kusber · 2025 [cited by examiner]
US 20250348478A1 · Papir · 2025 [cited by examiner]
US 20250370908A1 · McCourt · 2025 [cited by examiner]
CN 113420116B · 2022 [cited by applicant]
Extended European Search Report for Application No. 24204453.5 dated Dec. 16, 2024, 12 pages. [cited by applicant]
Liu Hao et al: “Show, Read and Reason 1-15, Table Structure Recognition with Flexible Context Aggregator”, Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval,… [cited by applicant]