IP Library Granted Patent US 12,566,728
Granted Patent B2
US 12,566,728 · App. 18/517,187 · Granted Mar 3, 2026

Data remediation using an evolving model

Inventors: Michael Stoute (Gastonia, NC); Christopher Tate (Asheville, NC); Venkata Sudhakar Bulusu (Concord, NC); Robert Posert (Portland, OR); Carol Garcia (South Pasadena, CA); Stephen Karhnak (Matthews, NC); Siva Satram (Hyderabad, IN); Julian Stapleford (Glen Ridge, NJ); Doug G. Pewowaruk (Farmington, MN); Sitarama Raju Alluru (Woodbury, MN); Melissa Winnix (Middleburg, FL); Meghana Puligalla (Little Elm, TX)
Assignee: Wells Fargo Bank, N.A.
G06F16/162G06F18/241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,728
App. No.
18/517,187
Granted
Mar 3, 2026
Kind
B2
Abstract

This disclosure describes techniques for performing data remediation. In one example, this disclosure describes a method that includes identifying a plurality of stale files; applying a classification model to each of the plurality of stale files; identifying a plurality of unclassified files, wherein each of the unclassified files is one of the plurality of stale files that the classification model was not able to classify with a confidence level that exceeds a threshold confidence level; updating the classification model, over a period of time, to generate an evolved classification model; applying the evolved classification model to each of the unclassified files; identifying a subset of the unclassified files that the evolved classification model was not able to classify with a confidence level that exceeds the threshold confidence level; and deleting each of the files in the subset of the unclassified files.

Claims (67)

1 . A method, comprising:

identifying, by a computing system and for a plurality of files in a storage system, a plurality of stale files;

applying, by the computing system, a classification model to each of the plurality of stale files;

identifying, by the computing system and based on applying the classification model to each of the stale files, a plurality of unclassified files, wherein each of the unclassified files is one of the plurality of stale files that the classification model was not able to classify with an associated confidence level that exceeds a threshold confidence level, wherein the confidence level indicates accuracy of classification;

updating the classification model, by the computing system and over a period of time, to generate an evolved classification model, wherein the evolved classification model is trained using training samples developed over time to improve the evolved classification model;

applying, by the computing system, the evolved classification model to each of the unclassified files;

identifying, by the computing system and based on applying the evolved classification model to each of the unclassified files, a subset of the unclassified files that the evolved classification model was not able to classify with a confidence level that exceeds the threshold confidence level;

identifying, by the computing system and based on applying the classification model to each of the stale files, a plurality of classified files;

moving, by the computing system, a first subset of the plurality of classified files to a content management repository, wherein moving the first subset includes identifying each of the stale files in the first subset as an official record;

deleting, by the computing system, a second subset of the plurality of classified files, wherein deleting the second subset includes identifying each of the stale files in the second subset as not an official record; and

deleting, by the computing system, each of the files in the subset of the unclassified files.

2 . The method of claim 1 , wherein identifying the plurality of stale files includes:

identifying a plurality of files that have not been modified during a threshold period of time.

3 . The method of claim 1 , wherein identifying the plurality of stale files includes:

identifying a plurality of files that have not been accessed during a threshold period of time.

4 . The method of claim 1 , wherein identifying the plurality of stale files includes:

identifying a plurality of stale files that each represent unstructured data.

5 . The method of claim 1 ,

wherein each of the classified files of the plurality of classified files is one of the plurality of stale files that the classification model was able to classify with a confidence level that exceeds the threshold confidence level.

6 . The method of claim 1 , wherein identifying the subset of the unclassified files includes:

identifying a subset of the unclassified files that are stale when applying the evolved classification model to the unclassified files.

7 . The method of claim 1 , wherein the subset of the unclassified files is a first subset of the unclassified files, and wherein the method further comprises:

identifying, by the computing system and based on applying the evolved classification model to the unclassified files, a second subset of the unclassified files that the evolved classification model was able to classify with a confidence level that exceeds the threshold confidence level.

8 . The method of claim 7 , wherein identifying the second subset of the unclassified files includes:

identifying a first file in the second subset as an official document; and

identifying a second file in the second subset as not an official document.

9 . The method of claim 8 , further comprising:

moving, by the computing system, the first file to a content management repository; and

deleting, by the computing system, the second file.

10 . The method of claim 8 , wherein identifying the second subset of the unclassified files further includes:

identifying a second subset of the unclassified files as files that are stale when applying the evolved classification model to the unclassified files.

11 . The method of claim 1 , wherein updating the classification model to generate the evolved classification model includes:

updating the classification model over an approximately three year period of time.

12 . The method of claim 1 , wherein updating the classification model to generate the evolved classification model includes:

repeatedly updating the classification model over the period of time to generate a sequence of updated classification models.

13 . The method of claim 12 , wherein repeatedly updating the classification model over the period of time includes:

repeatedly updating the classification model at least six times over an at least three-year period of time.

14 . The method of claim 12 , wherein identifying the subset of the unclassified files includes:

identifying a subset of the unclassified files that none of the updated classification models in the sequence of updated classification models was able to classify with a confidence level that exceeds the threshold confidence level.

15 . A computing system comprising:

a storage device; and

processing circuitry, wherein the processing circuitry has access to the storage device and is configured to:

identify, for a plurality of files in a storage system, a plurality of stale files;

apply a classification model to each of the plurality of stale files;

identify, based on applying the classification model to each of the stale files, a plurality of unclassified files, wherein each of the unclassified files is one of the plurality of stale files that the classification model was not able to classify with an associated confidence level that exceeds a threshold confidence level, wherein the confidence level indicates accuracy of classification;

update the classification model over a period of time to generate an evolved classification model, wherein the evolved classification model is trained using training samples developed over time to improve the evolved classification model;

apply the evolved classification model to each of the unclassified files;

identify, based on applying the evolved classification model to each of the unclassified files, a subset of the unclassified files that the evolved classification model was not able to classify with a confidence level that exceeds the threshold confidence level;

identify, based on applying the classification model to each of the stale files, a plurality of classified files;

move a first subset of the plurality of classified files to a content management repository, wherein to move the first subset, the processing circuitry is further configured to identify each of the stale files in the first subset as an official record;

delete a second subset of the plurality of classified files, wherein to delete the second subset, the processing circuitry is further configured to identify each of the stale files in the second subset as not an official record; and

delete each of the files in the subset of the unclassified files.

16 . The computing system of claim 15 ,

wherein each of the classified files of the plurality of classified files is one of the plurality of stale files that the classification model was able to classify with a confidence level that exceeds the threshold confidence level.

17 . The computing system of claim 15 , wherein to identify the subset of the unclassified files, the processing circuitry is further configured to:

identify a subset of the unclassified files as files that are stale when applying the evolved classification model to the unclassified files.

18 . A non-transitory computer-readable medium comprising instructions that, when executed, configure processing circuitry of a computing system to:

identify, for a plurality of files in a storage system, a plurality of stale files;

apply a classification model to each of the plurality of stale files;

identify, based on applying the classification model to each of the stale files, a plurality of unclassified files, wherein each of the unclassified files is one of the plurality of stale files that the classification model was not able to classify with an associated confidence level that exceeds a threshold confidence level, wherein the confidence level indicates accuracy of classification;

update the classification model over a period of time to generate an evolved classification model, wherein the evolved classification model is trained using training sample developed over time to improve the evolved classification model;

apply the evolved classification model to each of the unclassified files;

identify, based on applying the evolved classification model to each of the unclassified files, a subset of the unclassified files that the evolved classification model was not able to classify with a confidence level that exceeds the threshold confidence level;

identify, based on applying the classification model to each of the stale files, a plurality of classified files;

move a first subset of the plurality of classified files to a content management repository, wherein to move the first subset, the instructions further cause the processing circuitry to identify each of the stale files in the first subset as an official record;

delete a second subset of the plurality of classified files, wherein to delete the second subset, the instructions further cause the processing circuitry to identify each of the stale files in the second subset as not an official record; and

delete each of the files in the subset of the unclassified files.

Assignments (2)
REQUEST FOR ADDRESS CHANGE Recorded Dec 5, 2025
From: WELLS FARGO BANK, N.A.
To: WELLS FARGO BANK, N.A.
Reel/Frame 073896/0195 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 8, 2024
From: STOUTE, MICHAEL; TATE, CHRISTOPHER; BULUSU, VENKATA SUDHAKAR; POSERT, ROBERT; GARCIA, CAROL; KARHNAK, STEPHEN; SATRAM, SIVA; STAPLEFORD, JULIAN; PEWOWARUK, DOUG G.; ALLURU, SITARAMA RAJU; WINNIX, MELISSA; PULIGALLA, MEGHANA
To: WELLS FARGO BANK, N.A.
Reel/Frame 066051/0666 →
Continuity (1)
Related Publication 20250165433A1 · May 22, 2025
References Cited (37)
US 6920450B2 · Aono et al. · 2005 [cited by applicant]
US 7398261B2 · Spivack et al. · 2008 [cited by applicant]
US 7630963B2 · Larimore et al. · 2009 [cited by applicant]
US 8666991B2 · Peters et al. · 2014 [cited by applicant]
US 9292507B2 · Calkowski et al. · 2016 [cited by applicant]
US 9311481B1 · Wawda · 2016 [cited by examiner]
US 9558218B2 · Dutta et al. · 2017 [cited by applicant]
US 9691027B1 · Sawant · 2017 [cited by examiner]
US 10489502B2 · Priestas et al. · 2019 [cited by applicant]
US 10579281B2 · Cherubini et al. · 2020 [cited by applicant]
US 11019088B2 · Pratt et al. · 2021 [cited by applicant]
US 11157652B2 · Basava et al. · 2021 [cited by applicant]
US 11425077B2 · Korotkikh · 2022 [cited by applicant]
US 11468022B2 · Ramaswamy et al. · 2022 [cited by applicant]
US 11575693B1 · Muddu et al. · 2023 [cited by applicant]
US 20120110259A1 · Mills et al. · 2012 [cited by applicant]
US 20130024637A1 · Hadley · 2013 [cited by applicant]
US 20170357807A1 · Harms · 2017 [cited by examiner]
US 20200257697A1 · Bhabesh et al. · 2020 [cited by applicant]
US 20200394396A1 · Yanamandra · 2020 [cited by examiner]
US 20210006472A1 · Khaspa et al. · 2021 [cited by applicant]
US 20210256420A1 · Elisha · 2021 [cited by examiner]
US 20220004523A1 · He et al. · 2022 [cited by applicant]
US 20220021652A1 · Moghe et al. · 2022 [cited by applicant]
US 20220156489A1 · Agarwal et al. · 2022 [cited by applicant]
US 20220245276A1 · Gupta · 2022 [cited by examiner]
CN 107193988A · 2017 [cited by applicant]
CN 113158246A · 2021 [cited by applicant]
CN 113380414A · 2021 [cited by applicant]
CN 114490625A · 2022 [cited by applicant]
CN 114996254A · 2022 [cited by applicant]
CN 115129703A · 2022 [cited by applicant]
CN 115185742A · 2022 [cited by applicant]
CN 115203182A · 2022 [cited by applicant]
DE 202022103462U1 · 2022 [cited by applicant]
KR 20200084460A · 2020 [cited by applicant]
Babu et al., “Retrieval of Large Data From Medical Lake Repository-Heart Note”, 2022 1st International Conference on Computational Science and Technology, IEEE, Nov. 9, 2022, pp. 518-522. [cited by applicant]