IP Library Granted Patent US 12,339,908
Granted Patent B2
US 12,339,908 · App. 17/488,108 · Granted Jun 24, 2025

Systems and methods for machine learning-based data extraction

Inventors: Ankit Kumar Sinha (Dallas, TX); Hasan Kohadawala (Chennai, IN); Bhargava Reddy Karumuri (Dallas, TX); Saravanan Annamalai (Chennai, IN); Kishorekumar Torangallu (Dallas, TX); SaiNikitha Cheruku (Begaluru, IN)
Assignee: Nationstar Mortgage LLC
G06F16/90344G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,339,908
App. No.
17/488,108
Granted
Jun 24, 2025
Kind
B2
Abstract

In some aspects, the disclosure is directed to methods and systems for machine learning-based data extraction using multiple string searching models. String extraction logic may differ depending on the type of document received. For documents identified to contain line item structures, broader searching models are applied to the document to account for the increased variability of data in the document inherent in data organized in line item structures. For documents identifier to contain non-line item structures, stricter searching models are applied to the document to account for predictable data in the document associated with data organized in non-line item structures.

Claims (59)

1. A method for machine learning-based data extraction, comprising:

receiving, by a computing device, a document comprising one or more strings, the document associated with a document label separate from the one or more strings;

determining, by the computing device, whether the document label indicates a type of document structure;

receiving, by the computing device, a rule associated with the document label comprising a string sequence;

determining, by the computing device, whether the string sequence matches the one or more strings in the document by applying a regular expression parser to match the string sequence with the one or more strings in the document;

increasing, by the computing device, a confidence score in response to the string sequence matching the one or more strings in the document;

determining, by the computing device, a similarity score indicating a similarity between the string sequence and the one or more strings in the document;

increasing, by the computing device, the confidence score in response to the similarity score meeting or exceeding a first score value; and

displaying, by the computing device, on a display, the one or more strings in the document in response to the confidence score meeting or exceeding a second score value.

2. The method of claim 1 , wherein the type of document structure is a line item structure.

3. The method of claim 1 , further comprising:

receiving, by the computing device, a rule constraint associated with the rule;

discarding, by the computer, the one or more strings in response to the one or more strings not satisfying the rule constraint.

4. The method of claim 1 , wherein the similarity score is determined using a fuzzy model.

5. The method of claim 1 , further comprising:

executing, by the computing device, at least one of an optical character recognition model or a natural language processing model to determine the document label.

6. A method for machine learning-based data extraction comprising:

receiving, by a computing device, a document comprising one or more strings, the document associated with a document label separate from the one or more strings;

determining, by the computing device, whether the document label indicates a type of document structure;

applying, by the computing device, a plurality of searching models to the document to generate a corresponding plurality of arrays of one or more strings;

identifying, by the computing device, a first similarity score indicating a similarity between each of the plurality of arrays;

identifying, by the computing device, a second similarity score, different from the first similarity score, indicating a correctness of each of the one or more strings in each of the plurality of arrays, in response to the first similarity score meeting or exceeding a first score value, wherein the second similarity score indicating the correctness of each of the one or more strings in each of the plurality of arrays is determined using a fuzzy model; and

displaying, by the computing device, on a display, the one or more strings in the document in response to the second similarity score meeting or exceeding a second score value.

7. The method of claim 6 , wherein the first similarity score indicating the similarity between each of the plurality of arrays is determined using a similarity measure index.

8. The method of claim 6 , wherein the type of document structure is a non-line item structure.

9. The method of claim 6 , further comprising:

applying, by the computing device, the plurality of searching models to a second document to generate a second corresponding plurality of arrays of one or more strings;

identifying, by the computing device, that a third similarity score indicating a similarity between each of the second corresponding plurality of arrays does not meet or exceed the first score value; and

applying, by the computing device, a second, different plurality of searching models to the second document to generate a corresponding another plurality of arrays of one or more strings.

10. A system for machine learning-based data extraction comprising:

a computing device comprising processing circuitry and a network interface;

wherein the network interface is configured to receive a document comprising one or more strings, the document associated with a document label, separate from the one or more strings, indicating a type of document structure; and

wherein the processing circuitry is configured to:

determine whether the document label indicates a first type of document structure;

receive a rule associated with the document label comprising a string sequence;

determine whether the string sequence matches the one or more strings in the document by applying a regular expression parser to match the string sequence with the one or more strings in the document;

increase a confidence score in response to the string sequence matching the one or more strings in the document;

determine a first similarity score indicating a similarity between the string sequence and the one or more strings in the document;

increase the confidence score based on the similarity score meeting or exceeding a first score value;

display the one or more strings in the document in response to the confidence score meeting or exceeding a second score value.

11. The system of claim 10 , wherein the processing circuitry is further configured to:

determine whether the document label indicates a second type of document structure;

apply a plurality of searching models to the document to generate a corresponding plurality of arrays of one or more strings;

identify a second similarity score indicating a similarity between each of the plurality of arrays; and

identify a third similarity score indicating a correctness of each of the one or more strings in each of the one or more arrays, in response to the second similarity score meeting or exceeding a third score value.

12. The system of claim 10 , wherein the first type of document structure is a line item structure.

13. The system of claim 10 , wherein the processing circuitry is further configured to:

receive a rule constraint associated with the rule; and

discard the one or more strings in response to the one or more strings not satisfying the rule constraint.

14. The system of claim 10 , wherein the first similarity score is determined using a fuzzy model.

15. The system of claim 11 , wherein the second similarity score is determined using a similarity measure index.

16. The system of claim 11 , wherein the processing circuitry is further configured to:

determine whether another document label indicates the second type of document structure, the another document label associated with the document comprising one or more strings;

apply the plurality of searching models to the document to generate the corresponding plurality of arrays of one or more strings;

identify that the second similarity score does not meet or exceed the second score value; and

apply another plurality of searching models to the another document to generate a corresponding another plurality of arrays of one or more strings;

wherein the plurality of searching models comprises a first searching model and a second searching model; and

wherein the another plurality of searching models comprises the first searching model and a third searching model.

17. The system of claim 10 , wherein the processing circuitry is further configured to execute at least one or an optical character recognition model or a natural language processing model to determine the document label.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2021
From: SINHA, ANKIT KUMAR; KOHADAWALA, HASAN; KARUMURI, BHARGAVA REDDY; ANNAMALAI, SARAVANAN; TORANGALLU, KISHOREKUMAR; CHERUKU, SAINIKITHA
To: NATIONSTAR MORTGAGE LLC, D/B/A MR. COOPER
Reel/Frame 057715/0317 →
Continuity (1)
Related Publication 20230101817A1 · Mar 30, 2023
References Cited (21)
US 6738517B2 · Loce et al. · 2004 [cited by applicant]
US 7117432B1 · Shanahan · 2006 [cited by examiner]
US 8996350B1 · Dub et al. · 2015 [cited by applicant]
US 10354134B1 · Becker et al. · 2019 [cited by applicant]
US 20090092320A1 · Berard et al. · 2009 [cited by applicant]
US 20140072219A1 · Tian · 2014 [cited by applicant]
US 20140149105A1 · Lamba · 2014 [cited by examiner]
US 20170024466A1 · Bordawekar · 2017 [cited by examiner]
US 20190065576A1 · Peng · 2019 [cited by examiner]
US 20190147103A1 · Bhowan · 2019 [cited by examiner]
US 20200074169A1 · Mukhopadhyay et al. · 2020 [cited by applicant]
US 20200274990A1 · Berfanger et al. · 2020 [cited by applicant]
US 20220156298A1 · Mahmoud · 2022 [cited by examiner]
US 20220179896A1 · Farrell · 2022 [cited by examiner]
Eliza Grames; An automated approach to identifying search terms for systematic reviews using keyword co-occurrence networks; 2019; WileyOnlineLibrary; pp. 1645-1654. [cited by examiner]
Non-Final Office Action for U.S. Appl. No. 16/990,900 dated Aug. 11, 2021 (13 pages). [cited by applicant]
Adnan Kiran et al: “Limitations of information extraction methods and techniques for heterogeneous unstructured big data”, International Journal of Engineering, Business Management, vol. 11; http://journals.sagepub.com/… [cited by applicant]
Tanvir Qaisar; abstract, Retrieved from the Internet: URL:https://towardsdatascience.com/multi-page-document-classification-using-machine-learning-and-nlp-ba6151405c03, Aug. 7, 2021. [cited by applicant]
Doermann, David et al., Handbook of Document Image Processing and Recognition: excerpts; Techniques for Logical Labeling, Feb. 2, 2014, pp. 194-216, 223-225. [cited by applicant]
Liu Ying: “TableSeer: Automatic Table, Extraction, Search and Understanding”, PhD Dissertation, Pennsylvania State University, https://etda.libraries.psu.edu/files/final_submissions/6921; Dec. 31, 2009. [cited by applicant]
Baviskar Dipali et al: “Efficient Automated Processing of the Unstructured Documents Using Artificial Intelligence A Systematic Literature Review and Future Directions”, IEEE Access, IEEE, USA, vol. 9, Apr. 13, 2021. [cited by applicant]