IP Library › Granted Patent US 12,307,333
Granted Patent B2
US 12,307,333 · App. 17/135,920 · Granted May 20, 2025

Loss augmentation for predictive modeling

Inventors: Pavithra Harsha (White Plains, NY); Brian Leo Quanz (Yorktown Heights, NY); Shivaram Subramanian (Frisco, TX); Wei Sun (Tarrytown, NY); Max Biggs (Charlottesville, VA)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N20/00G06F18/21322G06F18/214G06F18/217G06F18/21326
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,333
App. No.
17/135,920
Filed
Dec 28, 2020
Granted
May 20, 2025
Kind
B2
Art Unit
2142
USPC
706/12
Abstract

A machine learning system that incorporates arbitrary constraints into deep learning model is provided. The machine learning system provides a set of penalty data points en a set of arbitrary constraints in addition to a set of original training data points. The machine learning system assigns a penalty to each penalty data point in the set of penalty data points. The machine learning system optimizes a machine learning model by solving an objective function based on an original loss function and a penalty loss function. The original loss function is evaluated over a set of original training data points and the penalty loss function is evaluated over the set of penalty data points. The machine learning system provides the optimized machine learning model based on a solution of the objective function.

Claims (35)

1. A computing device comprising:

a processor; and

a storage device storing a set of instructions, wherein an execution of the set of instructions by the processor configures the computing device to perform acts comprising:

providing a set of penalty data points enforcing a set of arbitrary constraints in addition to a set of original training data points for a deep learning model;

assigning a penalty to each penalty data point in the set of penalty data points;

optimizing the deep learning model by solving an objective function based on an original loss function and a penalty loss function, wherein the original loss function is evaluated over a set of original training data points and the penalty loss function is evaluated over the set of penalty data points; and

providing the optimized deep learning model based on a solution of the objective function.

2. The computing device of claim 1 , wherein optimizing the deep learning model further comprises maximizing the penalty loss function over additional data points in addition to the set of penalty data points.

3. The computing device of claim 2 , wherein the set of instructions further configures the computing device to perform acts comprising: upon determining that the maximized penalty loss function causes the objective function to be greater than a threshold, updating the set of penalty data points with the additional data points and continuing to optimize the deep learning model by solving the objective function.

4. The computing device of claim 2 , wherein the set of instructions further configures the computing device to perform an act comprising: upon determining that the maximized penalty loss function causes the objective function to be less than a threshold, terminating optimization of the deep learning model.

5. The computing device of claim 2 , wherein the additional data points are samples identified based on an earlier iteration of a stochastic gradient descent operation used to optimize the deep learning.

6. The computing device of claim 5 , wherein the additional data points are identified based on a violation of the arbitrary constraints during the earlier iteration of the stochastic gradient descent operation.

7. The computing device of claim 1 , wherein the original loss function with the penalty loss function are additive terms in the objective function.

8. The computing device of claim 1 , wherein the deep learning model comprises one or more intermediate layers.

9. A computer-implemented method comprising:

providing a set of penalty data points enforcing a set of arbitrary constraints in addition to a set of original training data points for a deep learning model;

assigning a penalty to each penalty data point in the set of penalty data points;

optimizing the deep learning model by solving an objective function based on an original loss function and a penalty loss function, wherein the original loss function is evaluated over a set of original training data points and the penalty loss function is evaluated over the set of penalty data points; and

providing the optimized deep learning model based on a solution of the objective function.

10. The computer-implemented method of claim 9 , wherein optimizing the deep learning model further comprises maximizing the penalty loss function over additional data points in addition to the set of penalty data points.

11. The computer-implemented method of claim 10 , further comprising, upon determining that the maximized penalty loss function causes the objective function to be greater than a threshold, updating the set of penalty data points with the additional data points and continuing to optimize the deep learning model by solving the objective function.

12. The computer-implemented method of claim 10 , further comprising, upon determining that the maximized penalty loss function causes the objective function to be less than a threshold, terminating optimization of the deep learning model.

13. The computer-implemented method of claim 10 , wherein the additional data points are samples identified based on an earlier iteration of a stochastic gradient descent operation used to optimize the deep learning.

14. The computer-implemented method of claim 13 , wherein the additional data points are identified based on a violation of the arbitrary constraints during the earlier iteration of the stochastic gradient descent operation.

15. The computer-implemented method of claim 9 , wherein the original loss function with the penalty loss function are additive terms in the objective function.

16. The computer-implemented method of claim 9 , wherein the deep learning model comprises one or more intermediate layers.

17. A computer program product comprising:

one or more non-transitory computer-readable storage devices and program instructions stored on at least one of the one or more non-transitory storage devices, the program instructions executable by a processor, the program instructions comprising sets of instructions for:

providing a set of penalty data points enforcing a set of arbitrary constraints in addition to a set of original training data points for a deep learning model;

assigning a penalty to each penalty data point in the set of penalty data points;

optimizing the deep learning model by solving an objective function based on an original loss function and a penalty loss function, wherein the original loss function is evaluated over a set of original training data points and the penalty loss function is evaluated over the set of penalty data points; and

providing the optimized deep learning model based on a solution of the objective function.

18. The computer program product of claim 17 , wherein optimizing the deep learning model further comprises maximizing the penalty loss function over additional data points in addition to the set of penalty data points.

19. The computer program product of claim 17 , wherein the original loss function with the penalty loss function are additive terms in the objective function.

20. The computer program product of claim 17 , wherein the deep learning model comprises one or more intermediate layers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2020
From: HARSHA, PAVITHRA; QUANZ, BRIAN LEO; SUBRAMANIAN, SHIVARAM; SUN, WEI; BIGGS, MAX
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054759/0508 →
Continuity (1)
Related Publication 20220207413A1 · Jun 30, 2022
References Cited (45)
US 5002313A · Salvatore · 1991 [cited by applicant]
US 7072852B1 · Kamille · 2006 [cited by applicant]
US 7451123B2 · Platt · 2008 [cited by examiner]
US 8280768B2 · Davis · 2012 [cited by applicant]
US 9965770B2 · Mason-Gugenheim et al. · 2018 [cited by applicant]
US 11281969B1 · Rangapuram · 2022 [cited by examiner]
US 20060224533A1 · Thaler · 2006 [cited by examiner]
US 20120158474A1 · Fahner et al. · 2012 [cited by applicant]
US 20120330867A1 · Gemulla · 2012 [cited by examiner]
US 20130006738A1 · Horvitz et al. · 2013 [cited by applicant]
US 20130006744A1 · Redford et al. · 2013 [cited by applicant]
US 20130006750A1 · Simmons, Jr. · 2013 [cited by applicant]
US 20130024261A1 · Main et al. · 2013 [cited by applicant]
US 20130041737A1 · Mishra et al. · 2013 [cited by applicant]
US 20130073372A1 · Novick et al. · 2013 [cited by applicant]
US 20150100402A1 · Gadotti · 2015 [cited by applicant]
US 20150205759A1 · Israel · 2015 [cited by examiner]
US 20170011416A1 · Vaysman · 2017 [cited by applicant]
US 20170161640A1 · Shamir · 2017 [cited by examiner]
US 20170300814A1 · Shaked et al. · 2017 [cited by applicant]
US 20180107925A1 · Choi · 2018 [cited by examiner]
US 20190130218A1 · Albright et al. · 2019 [cited by applicant]
US 20190171908A1 · Salavon · 2019 [cited by examiner]
US 20190172082A1 · Ganti Mahapatruni et al. · 2019 [cited by applicant]
US 20190180303A1 · Ventrice et al. · 2019 [cited by applicant]
US 20200026996A1 · Kolter · 2020 [cited by examiner]
US 20200234374A1 · Bawadhankar et al. · 2020 [cited by applicant]
US 20210097052A1 · Hans et al. · 2021 [cited by applicant]
US 20210295175A1 · Kennel et al. · 2021 [cited by applicant]
US 20210334664A1 · Li et al. · 2021 [cited by applicant]
US 20220207413A1 · Harsha · 2022 [cited by examiner]
US 20220230065A1 · Berthelot · 2022 [cited by applicant]
US 20220366218A1 · Parisotto et al. · 2022 [cited by applicant]
US 20230141655A1 · Gonzalez et al. · 2023 [cited by applicant]
US 20230154055A1 · Besenbruch · 2023 [cited by examiner]
US 20230385603A1 · Arikawa et al. · 2023 [cited by applicant]
Chen, J. T. et al., “Wide & Deep Learning for Recommender Systems”, In Proceedings of the 1st Workshop on Deep Learning for Recommender Systems (2016), pp. 7-10. [cited by applicant]
Bedi, J. et al., “Deep Learning Framework to Forecast Electricity Demand”; Applied Energy (2019); vol. 238; pp. 1312-1326. [cited by applicant]
Carbonneau, R. et al., “Application of Machine Learning Techniques for Supply Chain Demand Forecasting”; European Journal of Operational Research (2008); vol. 184; pp. 1140-1154. [cited by applicant]
Ke, J. et al., “Short-Term Forecasting of Passenger Demand under On-Demand Ride Services: A Spatio-Temporal Deep Learning Approach”; arXiv:1706.06279v1 [cs.LG] (2017); 39 pgs. [cited by applicant]
Qiu, X. et al., “Empirical Mode Decomposition based Ensemble Deep Learning for Load Demand Time Series Forecasting”; Applied Soft Computing (2017); 40 pgs. [cited by applicant]
Yao, H. et al., “Deep Multi-View Spatial-Temporal Network for Taxi Demand Prediction”; The Thirty-Second AAAI Conference on Artificial Intelligence (2018); pp. 2588-2595. [cited by applicant]
List of IBM Patents or Patent Applications Treated as Related, 2 pgs. [cited by applicant]
Kilroy, J. et al., “Method and System for Exchanging and Trading Online Coupons”; IP.Com (2011); 4 pgs. [cited by applicant]
Borghesi et al., “Improving Deep Learning Models via Constraint-Based Domain Knowledge: A Brief Survey,” arXiv:2005.10691v1 [cs.LG], May 19, 2020, 14 pages. [cited by applicant]