IP Library Granted Patent US 12,561,113
Granted Patent B2
US 12,561,113 · App. 18/442,947 · Granted Feb 24, 2026

Systems and methods for unstructured and structured data-driven sorting

Inventors: Bryan Grant (Melissa, TX); Zach Rusk (Argyle, TX)
Assignee: National Mortgage LLC
G06F7/08G06F16/334G06F16/335G06F16/38G06V30/19007G06V30/416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,113
App. No.
18/442,947
Granted
Feb 24, 2026
Kind
B2
Abstract

In some aspects, the disclosure is directed to methods and systems for sorting, correlation, and/or matching of structured and unstructured data. In some implementations, metadata associated with structured data may be applied to unstructured data based on the results of the sort or correlation. Such sorting or matching may be performed through an efficient iterative process with progressive confidence scores, and use a priori knowledge from correlations between structured data values to identify potentially associated unstructured data values.

Claims (37)

1 . A method for matching data sets in different formats, comprising:

receiving, by one or more processors of a computing device, a first plurality of structured data sets, each structured data set comprising a plurality of values associated with a corresponding plurality of metadata;

receiving, by the one or more processors, a second plurality of unstructured data sets generated via optical character recognition (OCR) from a plurality of printed documents, each unstructured data set of the second plurality of data sets lacking associated metadata;

selecting, by the one or more processors, a first subset of the plurality of values of a first structured data set;

applying, by the one or more processors, a first comparison ruleset associated with a first confidence score to the selected first subset of the plurality of values of the first structured data set and values of each unstructured data set;

identifying a match, by the one or more processors, between the selected first subset of the plurality of values of the first structured data set and the values of a first unstructured data set;

responsive to identifying the match, adding, by the one or more processors, metadata of the first structured data set to the first unstructured data set;

selecting, by the one or more processors, a second subset of the plurality of values of the first structured data set;

applying, by the one or more processors, a second comparison ruleset associated with a second, lower confidence score to the selected second subset of the plurality of values of the first structured data set and values of each unstructured data set, excluding the first unstructured data set, wherein applying the second comparison ruleset requires more processing resources than the first comparison ruleset;

identifying a match, by the one or more processors, between the selected second subset of the plurality of values of the first structured data set and the values of a second unstructured data set; and

responsive to identifying the match, adding, by the one or more processors, metadata of the first structured data set to the second unstructured data set.

2 . The method of claim 1 , wherein a data set of the second plurality of data sets comprises a second plurality of values associated with the corresponding plurality of metadata, and a number of the second plurality of values is smaller than a number of the plurality of metadata.

3 . The method of claim 1 , wherein a first data set in the first plurality of data sets is matched to a second data set in the second plurality of data sets, and wherein the first data set is different than the second data set.

4 . The method of claim 3 , wherein the first data set comprises a first subset of values matching values of the second data set, and a second subset of values not matching values of the second data set.

5 . The method of claim 4 , wherein the first subset of values are associated with higher confidence matches than the second subset of values.

6 . The method of claim 4 , wherein the first subset of values are associated with lower confidence matches than the second subset of values.

7 . The method of claim 3 , wherein the first data set in the first plurality of data sets is matched to the second data set in the second plurality of data sets, responsive to a first value in the first data set having a similarity score to a second value in the second data set greater than a first threshold but less than a second threshold.

8 . The method of claim 1 , further comprising excluding, by the one or more processors, a subset of data sets from the second plurality of data sets prior to the iterative comparing, responsive to values of the subset of data sets matching a predetermined filter rule.

9 . A system for matching data sets in different formats, comprising:

one or more processors, and a memory device storing a first plurality of structured data sets, each structured data set comprising a plurality of values associated with a corresponding plurality of metadata; and

wherein the one or more processors are configured to:

receive a second plurality of unstructured data sets generated via optical character recognition (OCR) from a plurality of printed documents, each data set of the second unstructured plurality of data sets lacking associated metadata,

select a first subset of the plurality of values of a first structured data set;

apply a first comparison ruleset associated with a first confidence score to the selected first subset of the plurality of values of the first structured data set and values of each unstructured data set;

identify a match between the selected first subset of the plurality of values of the first structured data set and the values of a first unstructured data set;

responsive to identifying the match, add metadata of the first structured data set to the first unstructured data set;

select a second subset of the plurality of values of the first structured data set;

apply a second comparison ruleset associated with a second, lower confidence score to the selected second subset of the plurality of values of the first structured data set and values of each unstructured data set, excluding the first unstructured data set, wherein applying the second comparison ruleset requires more processing resources than the first comparison ruleset;

identify a match between the selected second subset of the plurality of values of the first structured data set and the values of a second unstructured data set; and

responsive to identifying the match, add metadata of the first structured data set to the second unstructured data set.

10 . The system of claim 9 , wherein a data set of the second plurality of data sets comprises a second plurality of values associated with the corresponding plurality of metadata, and a number of the second plurality of values is smaller than a number of the plurality of metadata.

11 . The system of claim 9 , wherein a first data set in the first plurality of data sets is matched to a second data set in the second plurality of data sets, and wherein the first data set is different than the second data set.

12 . The system of claim 11 , wherein the first data set comprises a first subset of values matching values of the second data set, and a second subset of values not matching values of the second data set.

13 . The system of claim 12 , wherein the first subset of values are associated with higher confidence matches than the second subset of values.

14 . The system of claim 12 , wherein the first subset of values are associated with lower confidence matches than the second subset of values.

15 . The system of claim 11 , wherein the first data set in the first plurality of data sets is matched to the second data set in the second plurality of data sets, responsive to a first value in the first data set having a similarity score to a second value in the second data set greater than a first threshold but less than a second threshold.

16 . The system of claim 9 , wherein the one or more processors are further configured to exclude a subset of data sets from the second plurality of data sets prior to the iterative comparing, responsive to values of the subset of data sets matching a predetermined filter rule.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2025
From: GRANT, BRYAN; RUSK, ZACH
To: NATIONSTAR MORTGAGE LLC, D/B/A/ MR. COOPER
Reel/Frame 070951/0611 →
Continuity (1)
Related Publication 20250265037A1 · Aug 21, 2025
References Cited (3)
US 11061874B1 · Funk · 2021 [cited by examiner]
US 20080247004A1 · Yeung · 2008 [cited by examiner]
US 20190147103A1 · Bhowan · 2019 [cited by examiner]