IP Library › Granted Patent US 12,321,824
Granted Patent B1
US 12,321,824 · App. 17/143,769 · Granted Jun 3, 2025

Pipelined machine learning frameworks

Inventors: Matthew Reeves (Boston, MA); Ben Thompson (Bellevue, WA)
Assignee: Liberty Mutual Insurance Company
G06N20/00G06F18/214G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,824
App. No.
17/143,769
Granted
Jun 3, 2025
Kind
B1
Abstract

In general, embodiments of the present invention provide methods, apparatus, systems, computing devices, computing entities, and/or the like for pipelined machine learning. Certain embodiments of the present invention utilize systems, methods, and computer program products that perform pipelined machine learning using at least one of error correction machine learning models and output augmentation machine learning models, for example using a pipelined implementation of error correction machine learning models followed by output augmentation machine learning models.

Claims (44)

1. A computer-implemented method for generating a predictive output based at least in part on an input data object, the computer-implemented method comprising:

generating, by one or more processors, one or more inference error correction engineered features based at least in part on the input data object, wherein the one or more inference error correction engineered features include an agent-based error likelihood value describing an estimated error likelihood of an input provider agent associated with the input data object;

processing, by the one or more processors, the one or more inference error correction engineered features, using a trained error correction machine learning model of a pipelined machine learning framework, to generate an error-corrected input data object in which one or more input data fields corresponding to the input data object are adjusted based at least in part on the one or more inference error correction engineered features; and

generating, by the one or more processors, the predictive output based at least in part on the error-corrected input data object.

2. The computer-implemented method of claim 1 , wherein the one or more inference error correction engineered features comprise a per-field-type error likelihood value for each selected input data field of one or more selected input fields of the input data object that describes an estimated error likelihood of an input data field type associated with the selected input data field.

3. The computer-implemented method of claim 1 , wherein generating the trained error correction machine learning model comprises:

identifying one or more training error correction data objects;

for each training error correction data object of the one or more training error correction data objects:

determining a set of training error correction engineered features, and

determining a finalized training error correction data object based at least in part on the set of training error correction engineered features; and

generating the trained error correction machine learning model based at least in part on each finalized training error correction data object for a training error correction data object of the one or more training error correction data objects.

4. The computer-implemented method of claim 3 , wherein each training error correction data object of the one or more training error correction data objects may describe a per-field error likelihood value for each selected input data field of one or more selected input fields of the input data object.

5. The computer-implemented method of claim 3 , wherein generating the trained error correction machine learning model based at least in part on each finalized training error correction data object for the training error correction data object comprises performing cross-validation with grid search using each finalized training error correction data object for the training error correction data object to generate the trained error correction machine learning model.

6. The computer-implemented method of claim 3 , wherein determining the finalized training error correction data object comprises performing mean target feature encoding with smoothing on the training error correction data object to determine the finalized training error correction data object.

7. The computer-implemented method of claim 1 , wherein the trained error correction machine learning model comprises a trained gradient boosting machine learning model.

8. A system for generating a predictive output based at least in part on an input data object, the system comprising one or more processors and memory including program code, the memory and the program code configured to, with the one or more processors, cause the system to at least:

generate one or more inference error correction engineered features based at least in part on the input data object, wherein the one or more inference error correction engineered features include an agent-based error likelihood value describing an estimated error likelihood of an input provider agent associated with the input data object;

process the one or more inference error correction engineered features using a trained error correction machine learning model of a pipelined machine learning framework to generate an error-corrected input data object in which one or more input data fields corresponding to the input data object are adjusted based at least in part on the one or more inference error correction engineered features; and

generate the predictive output based at least in part on the error-corrected input data object.

9. The system of claim 8 , wherein the one or more inference error correction engineered features comprise a per-field-type error likelihood value for each selected input data field of one or more selected input fields of the input data object that describes an estimated error likelihood of an input data field type associated with the selected input data field.

10. The system of claim 8 , wherein generating the trained error correction machine learning model comprises:

identifying one or more training error correction data objects;

for each training error correction data object of the one or more training error correction data objects:

determining a set of training error correction engineered features, and

determining a finalized training error correction data object based at least in part on the set of training error correction engineered features; and

generating the trained error correction machine learning model based at least in part on each finalized training error correction data object for a training error correction data object of the one or more training error correction data objects.

11. The system of claim 10 , wherein each training error correction data object of the one or more training error correction data objects may describe a per-field error likelihood value for each selected input data field of one or more selected input fields of the input data object.

12. The system of claim 10 , wherein generating the trained error correction machine learning model based at least in part on each finalized training error correction data object for the training error correction data object comprises performing cross-validation with grid search using each finalized training error correction data object for the training error correction data object to generate the trained error correction machine learning model.

13. The system of claim 10 , wherein determining the finalized training error correction data object comprises performing mean target feature encoding with smoothing on the training error correction data object to determine the finalized training error correction data object.

14. The system of claim 8 , wherein the trained error correction machine learning model comprises a trained gradient boosting machine learning model.

15. A computer program product for generating a predictive output based at least in part on an input data object, the computer program product comprising at least one non-transitory computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions configured to:

generate one or more inference error correction engineered features based at least in part on the input data object, wherein the one or more inference error correction engineered features include an agent-based error likelihood value describing an estimated error likelihood of an input provider agent associated with the input data object;

process the one or more inference error correction engineered features using a trained error correction machine learning model of a pipelined machine learning framework to generate an error-corrected input data object in which one or more input data fields corresponding to the input data object are adjusted based at least in part on the one or more inference error correction engineered features; and

generate the predictive output based at least in part on the error-corrected input data object.

16. The computer program product of claim 15 , wherein the one or more inference error correction engineered features comprise a per-field-type error likelihood value for each selected input data field of one or more selected input fields of the input data object that describes an estimated error likelihood of an input data field type associated with the selected input data field.

17. The computer program product of claim 15 , wherein generating the trained error correction machine learning model comprises:

identifying one or more training error correction data objects;

for each training error correction data object of the one or more training error correction data objects:

determining a set of training error correction engineered features, and

determining a finalized training error correction data object based at least in part on the set of training error correction engineered features; and

generating the trained error correction machine learning model based at least in part on each finalized training error correction data object for a training error correction data object of the one or more training error correction data objects.

18. The computer program product of claim 17 , wherein each training error correction data object of the one or more training error correction data objects may describe a per-field error likelihood value for each selected input data field of one or more selected input fields of the input data object.

19. The computer program product of claim 17 , wherein generating the trained error correction machine learning model based at least in part on each finalized training error correction data object for the training error correction data object comprises performing cross-validation with grid search using each finalized training error correction data object for the training error correction data object to generate the trained error correction machine learning model.

20. The computer program product of claim 17 , wherein determining the finalized training error correction data object comprises performing mean target feature encoding with smoothing on the training error correction data object to determine the finalized training error correction data object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 11, 2021
From: REEVES, MATTHEW; THOMPSON, BEN
To: LIBERTY MUTUAL INSURANCE COMPANY
Reel/Frame 054877/0154 →
Continuity (1)
Provisional Application 62958379 · Jan 8, 2020
References Cited (48)
US 4831526A · Luchs · 1989 [cited by examiner]
US 6117076A · Cassidy · 2000 [cited by examiner]
US 10713577B1 · Faruquie · 2020 [cited by examiner]
US 11042909B2 · Vij · 2021 [cited by examiner]
US 11062378B1 · Ross · 2021 [cited by examiner]
US 20160048766A1 · McMahon · 2016 [cited by examiner]
US 20160063636A1 · Feimster · 2016 [cited by examiner]
US 20170330099A1 · de Vial · 2017 [cited by examiner]
US 20180356807A1 · Honda · 2018 [cited by examiner]
US 20190087887A1 · Westphal · 2019 [cited by examiner]
US 20190164084A1 · Gulin · 2019 [cited by examiner]
US 20200005045A1 · Lloyd, II · 2020 [cited by examiner]
US 20200020098A1 · Odry · 2020 [cited by examiner]
US 20200034732A1 · Freed · 2020 [cited by examiner]
US 20200074997A1 · Jankowski, Jr. · 2020 [cited by examiner]
US 20200090056A1 · Singhal · 2020 [cited by examiner]
US 20200117523A1 · Morrison · 2020 [cited by examiner]
US 20200226503A1 · Subramanian · 2020 [cited by examiner]
US 20200279219A1 · Desai · 2020 [cited by examiner]
US 20200372561A1 · Sanghavi · 2020 [cited by examiner]
US 20210065048A1 · Salonidis · 2021 [cited by examiner]
US 20210117855A1 · Hoover · 2021 [cited by examiner]
US 20210374566A1 · Goyal · 2021 [cited by examiner]
US 20220150275A1 · McNee · 2022 [cited by examiner]
US 20220198156A1 · Rao · 2022 [cited by examiner]
US 20230060099A1 · Babich · 2023 [cited by examiner]
US 20230418654A1 · Zuccarelli · 2023 [cited by examiner]
Unfold Data Science, “Gradient Boost Machine Learning |How Gradient boost work in Machine Learning”, available online at [https://www.youtube.com/watch?v=j034-r302Cg], posted on Feb. 18, 2020. (Year: 2020). [cited by examiner]
Ronald Joseph, “Grid search for model tuning”, available online at <https://towardsdatascience.com/grid-search-for-model-tuning-3319b259367e>, published on Dec. 29, 2018 (Year: 2018). [cited by examiner]
“Chapter 9—Temporal-Difference Learning,” Stanford University, (18 pages). [Online]. [Retrieved from the Internet Feb. 17, 2021] <URL: https://web.stanford.edu/group/pdplab/pdphandbook/handbookch10.html>. [cited by applicant]
“Minimizing Real-Time Prediction Serving Latency In Machine Learning,” Google Cloud, Nov. 16, 2020, (21 pages). [Article, Online]. [Retrieved from the Internet Feb. 17, 2021] <URL: https://cloud.google.com/solutions/mac… [cited by applicant]
Dugas, Charles et al. “Statistical Learning Algorithms Applied To Automobile Insurance Ratemaking,” In The Casualty Actuarial Society Forum, Mar. 27, 2003, vol. 1, No. 1, pp. 179-214. [cited by applicant]
Karaman, Baris. “Market Response Models—Predicting Incremental Gains Of Promotional Campaigns,” Towards Data Science, Jul. 28, 2019, (17 pages). [Article, Online]. [Retrieved from the Internet Feb. 17, 2021] <URL: https… [cited by applicant]
Kim, Jin Kyu et al. “Smpframe: A Distributed Framework For Scheduled Model Parallel Machine Learning,” Parallel Data Laboratory, Carnegie Mellon University, May 2015, (26 pages). [cited by applicant]
Koen, Semi. “Architecting A Machine Learning Pipeline,” Towards Data Science, Apr. 5, 2019, (22 pages). [Article, Online]. [Retrieved from the Internet Feb. 17, 2021] <URL: https://towardsdatascience.com/architecting-a-… [cited by applicant]
Krishnan, Sanjay et al. “BoostClean: Automated Error Detection and Repair for Machine Learning,” arXiv:1711.01299v1, Nov. 3, 2017, (15 pages). [cited by applicant]
Sato, Danilo et al. “Continuous Delivery For Machine Learning,” Martin Fowler, Sep. 19, 2019, (34 pages). [Article, Online]. [Retrieved from the Internet Feb. 17, 2021] <URL: https://martinfowler.com/articles/cd4ml.html… [cited by applicant]
Soares, Carlos et al. “Machine Learning and Statistics To Detect Errors In Forms: Competition or Cooperation?”, In Proceedings of the European Conference on Machine Learning and Principles and Practice of Knowledge Disc… [cited by applicant]
Spedicato, Giorgio Alfredo et al. “Machine Learning Methods To Perform Pricing Optimization. A Comparison With Standard Generalized Linear Models,” Variance, vol. 12, Issue 1, Dec. 2018, pp. 69-89. [cited by applicant]
Sporleder, Caroline et al. “Spotting The ‘Odd-One-Out’: Data-Driven Error Detection and Correction In Textual Databases,” 11th Conference of the European Chapter of the Association of Computational Linguistics, In Proce… [cited by applicant]
(Nithila Jeyakumar, “Analysis of the Digital Direct-to-Customer channel in Insurance,” MIT (2016) [Thesis]) (Year: 2016). [cited by applicant]
Non-Final Rejection Mailed on Jun. 5, 2024 for U.S. Appl. No. 17/143,773, 37 page(s). [cited by applicant]
Robert E. Schapire, “The Boosting Approach to Machine Learning: An Overview,” (2001) (Year: 2001). [cited by applicant]
Non-Final Rejection Mailed on May 3, 2024 for U.S. Appl. No. 17/143,761, 80 page(s). [cited by applicant]
Final Rejection Mailed on Aug. 29, 2024 for U.S. Appl. No. 17/143,761, 41 page(s). [cited by applicant]
Advisory Action Mailed on Nov. 8, 2024 for U.S. Appl. No. 17/143,761, 3 page(s). [cited by applicant]
Final Rejection Mailed on Dec. 13, 2024 for U.S. Appl. No. 17/143,773, 33 page(s). [cited by applicant]
Lin et al., “A Recommender for Targeted Advertisement of Unsought Products in E-Commerce,” IEEE (2005) (Year: 2005). [cited by applicant]