IP Library › Granted Patent US 11,586,919
Granted Patent B2
US 11,586,919 · App. 16/899,681 · Granted Feb 21, 2023

Task-oriented machine learning and a configurable tool thereof on a computing environment

Inventors: Yada Zhu (Westchester, NY); Di Chen (Ithaca, NY); Xiaodong Cui (Chappaqua, NY); Upendra Chitnis (Fairfield, CT); Kumar Bhaskaran (Englewood Cliffs, NJ); Wei Zhang (Harbin, CN)
Assignee: International Business Machines Corporation
G06N3/08G06N3/0445G06Q40/00G06Q40/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,586,919
App. No.
16/899,681
Granted
Feb 21, 2023
Kind
B2
Abstract

A task-based learning using task-directed prediction network can be provided. Training data can be received. Contextual information associated with a task-based criterion can be received. A machine learning model can be trained using the training data. A loss function computed during training of the machine learning model integrates the task-based criterion, and minimizing the loss function during training iterations includes minimizing the task-based criterion.

Claims (44)

1. A computer-implemented method, comprising:

receiving training data;

receiving contextual information associated with a task-based criterion; and

training a machine learning model using the training data, wherein a loss function computed during the training integrates the task-based criterion, and wherein minimizing the loss function during training iterations includes minimizing the task-based criterion,

the minimizing the task-based criterion including at least

generating true task-based loss using predictions of the machine learning model, ground truth labels associated with the predictions of the machine learning model, and the contextual information,

training a task-oriented loss estimator network that takes encodings of at least the predictions of the machine learning model, the ground truth labels associated with the predictions of the machine learning model, the contextual information, and that minimize a discrepancy between learned loss of the task-oriented loss estimator network and the true task-based loss,

wherein the training the machine learning model includes, responsive to determining that the discrepancy is less than a threshold value, updating parameters of the machine learning model based on the learned loss of the task-oriented loss estimator network, and responsive to determining that the discrepancy is not less than the threshold value, updating parameters of the machine learning model based on standard loss of the machine learning model.

2. The method of claim 1 , further including providing a tool for building and managing the machine leaning model on a computing environment, the computing environment allowing an on-demand network access to a shared pool of configurable computing resources, the configurable computing resources including at least one of networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services.

3. The method of claim 1 , wherein the machine learning model includes a neural network.

4. The method of claim 1 , further including learning a surrogate loss function that is differentiable, to guide the machine learning model.

5. The method of claim 4 , wherein the machine learning model is updated using gradients obtained from the learned surrogate loss function.

6. The method of claim 4 , wherein the surrogate loss function is learned via a neural network parameterized by a weight.

7. The method of claim 4 , wherein the surrogate loss function is initialized with a warm-up loss function.

8. The method of claim 4 , further including approximating a true task-based loss by minimizing a discrepancy between a surrogate loss from the learned surrogate loss function and the true task-based loss.

9. The method of claim 8 , wherein the approximating is performed by a neural network.

10. The method of claim 1 , further including extracting by a long short-term memory (LSTM) network features from the training data, wherein the machine learning model is trained based on the features.

11. A system comprising:

a hardware processor; and

a memory device coupled with the hardware processor,

the hardware processor configured to:

receive training data;

receive contextual information associated with a task-based criterion; and

train a machine learning model using the training data, wherein a loss function computed during training of the machine learning model integrates the task-based criterion, and wherein minimizing the loss function during training iterations includes minimizing the task-based criterion,

the minimizing the task-based criterion including at least

generating true task-based loss using predictions of the machine learning model, ground truth labels associated with the predictions of the machine learning model, and the contextual information,

training a task-oriented loss estimator network that takes encodings of at least the predictions of the machine learning model, the ground truth labels associated with the predictions of the machine learning model, the contextual information, and that minimize a discrepancy between learned loss of the task-oriented loss estimator network and the true task-based loss,

wherein the hardware processor is configured to train the machine learning model at least by, responsive to determining that the discrepancy is less than a threshold value, updating parameters of the machine learning model based on the learned loss of the task-oriented loss estimator network, and responsive to determining that the discrepancy is not less than the threshold value, updating parameters of the machine learning model based on standard loss of the machine learning model.

12. The system of claim 11 , wherein the machine learning model includes a neural network.

13. The system of claim 11 , wherein the hardware processor is further configured to learn a surrogate loss function that is differentiable, to guide the machine learning model.

14. The system of claim 13 , wherein the hardware processor is further configured to update the machine learning model using gradients obtained from the learned surrogate loss function.

15. The system of claim 13 , wherein the hardware processor is further configured to learn the surrogate loss function via a neural network parameterized by a weight.

16. The system of claim 13 , wherein the hardware processor is further configured to initialize the surrogate loss function with a warm-up loss function.

17. The system of claim 13 , wherein the hardware processor is further configured to approximate a true task-based loss by minimizing a discrepancy between a surrogate loss from the learned surrogate loss function and the true task-based loss.

18. The system of claim 17 wherein a neural network performs the minimizing a discrepancy between a surrogate loss from the learned surrogate loss function and the true task-based loss to approximate the true task-based loss.

19. A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions readable by a device to cause the device to:

receive training data;

receive contextual information associated with a task-based criterion; and

train a machine learning model using the training data, wherein a loss function computed during training of the machine learning model integrates the task-based criterion, and wherein minimizing the loss function during training iterations includes minimizing the task-based criterion,

the minimizing the task-based criterion including at least

generating true task-based loss using predictions of the machine learning model, ground truth labels associated with the predictions of the machine learning model, and the contextual information,

training a task-oriented loss estimator network that takes encodings of at least the predictions of the machine learning model, the ground truth labels associated with the predictions of the machine learning model, the contextual information, and that minimize a discrepancy between learned loss of the task-oriented loss estimator network and the true task-based loss,

wherein the device is caused to train the machine learning model at least by, responsive to determining that the discrepancy is less than a threshold value, updating parameters of the machine learning model based on the learned loss of the task-oriented loss estimator network, and responsive to determining that the discrepancy is not less than the threshold value, updating parameters of the machine learning model based on standard loss of the machine learning model.

20. The computer program product of claim 19 , wherein the machine learning model includes a neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 12, 2020
From: ZHU, YADA; CHEN, DI; CUI, XIAODONG; CHITNIS, UPENDRA; BHASKARAN, KUMAR; ZHANG, WEI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052922/0053 →
Continuity (1)
Related Publication 20210397941A1 · Dec 23, 2021
Cited By (1)
US 12,645,985