IP Library › Granted Patent US 11,605,028
Granted Patent B2
US 11,605,028 · App. 16/551,466 · Granted Mar 14, 2023

Methods and systems for sequential model inference

Inventors: Michele Gazzetti (Dublin, IE); Srikumar Venugopal (Dublin, IE); Christian Pinto (Dublin, IE)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N20/20G06N5/04G06N5/046H04L1/0017H04L67/61
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,605,028
App. No.
16/551,466
Granted
Mar 14, 2023
Kind
B2
Abstract

Embodiments for processing data with multiple machine learning models are provided. Input data is received. The input data is caused to be evaluated by a first machine learning model to generate a first inference result. The first inference result is compared to at least one quality of service (QoS) parameter. Based on the comparison of the first inference result to the at least one QoS parameter, the input data is caused to be evaluated by a second machine learning model to generate a second inference result.

Claims (40)

1. A method for processing data with multiple machine learning models, by a processor, comprising:

receiving input data;

causing the input data to be evaluated by a first machine learning model to generate a first inference result;

comparing the first inference result to at least one quality of service (QoS) parameter; and

based on said comparison of the first inference result to the at least one QoS parameter, causing the input data to be conditionally evaluated by a second machine learning model to generate a second inference result.

2. The method of claim 1 , wherein the input data comprises at least one data item, and wherein each of the first machine learning model and the second machine learning model is configured to perform the same prediction task.

3. The method of claim 1 , wherein the causing of the input data to be evaluated by the second machine learning model to generate a second inference result based on said comparison of the first inference result to the at least one QoS parameter includes providing the first inference result to a user if the first inference result meets the at least one QoS parameter.

4. The method of claim 3 , wherein the causing of the input data to be evaluated by the second machine learning model to generate a second inference result based on said comparison of the first inference result to the at least one QoS parameter further includes:

causing the input data to be evaluated by the second learning model if the first inference result does not meet the at least one QoS parameter; and

providing the second inference result to the user.

5. The method of claim 1 , further comprising receiving the at least one QoS parameter from a user, and wherein the at least one QoS parameter is associated with at least one of latency, precision, and recall.

6. The method of claim 1 , wherein the first machine learning model is implemented utilizing a first computing device, and the second machine learning model is implemented utilizing a second computing device.

7. The method of claim 6 , wherein the second computing device is remote from the first computing device and in operable communication with the first computing device via a communications network.

8. A system for processing data with multiple machine learning models comprising:

a processor executing instructions stored in a memory device, wherein the processor:

receives input data;

causes the input data to be evaluated by a first machine learning model to generate a first inference result;

compares the first inference result to at least one quality of service (QoS) parameter; and

based on said comparison of the first inference result to the at least one QoS parameter, causes the input data to be conditionally evaluated by a second machine learning model to generate a second inference result.

9. The system of claim 8 , wherein the input data comprises at least one data item, and wherein each of the first machine learning model and the second machine learning model is configured to perform the same prediction task.

10. The system of claim 8 , wherein the causing of the input data to be evaluated by the second machine learning model to generate a second inference result based on said comparison of the first inference result to the at least one QoS parameter includes providing the first inference result to a user if the first inference result meets the at least one QoS parameter.

11. The system of claim 10 , wherein the causing of the input data to be evaluated by the second machine learning model to generate a second inference result based on said comparison of the first inference result to the at least one QoS parameter further includes:

causing the input data to be evaluated by the second learning model if the first inference result does not meet the at least one QoS parameter; and

providing the second inference result to the user.

12. The system of claim 8 , wherein the processor further receives the at least one QoS parameter from a user, and wherein the at least one QoS parameter is associated with at least one of latency, precision, and recall.

13. The system of claim 8 , wherein the first machine learning model is implemented utilizing a first computing device, and the second machine learning model is implemented utilizing a second computing device.

14. The system of claim 13 , wherein the second computing device is remote from the first computing device and in operable communication with the first computing device via a communications network.

15. A computer program product for processing data with multiple machine learning models, by a processor, the computer program product embodied on a non-transitory computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions comprising:

an executable portion that receives input data;

an executable portion that causes the input data to be evaluated by a first machine learning model to generate a first inference result;

an executable portion that compares the first inference result to at least one quality of service (QoS) parameter; and

an executable portion that, based on said comparison of the first inference result to the at least one QoS parameter, causes the input data to be conditionally evaluated by a second machine learning model to generate a second inference result.

16. The computer program product of claim 15 , wherein the input data comprises at least one data item, and wherein each of the first machine learning model and the second machine learning model is configured to perform the same prediction task.

17. The computer program product of claim 15 , wherein the causing of the input data to be evaluated by the second machine learning model to generate a second inference result based on said comparison of the first inference result to the at least one QoS parameter includes providing the first inference result to a user if the first inference result meets the at least one QoS parameter.

18. The computer program product of claim 17 , wherein the causing the of input data to be evaluated by the second machine learning model to generate a second inference result based on said comparison of the first inference result to the at least one QoS parameter further includes:

causing the input data to be evaluated by the second learning model if the first inference result does not meet the at least one QoS parameter; and

providing the second inference result to the user.

19. The computer program product of claim 15 , wherein the computer-readable program code portions further include an executable portion that receives the at least one QoS parameter from a user, and wherein the at least one QoS parameter is associated with at least one of latency, precision, and recall.

20. The computer program product of claim 15 , wherein the first machine learning model is implemented utilizing a first computing device, and the second machine learning model is implemented utilizing a second computing device.

21. The computer program product of claim 20 , wherein the second computing device is remote from the first computing device and in operable communication with the first computing device via a communications network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 26, 2019
From: GAZZETTI, MICHELE; VENUGOPAL, SRIKUMAR; PINTO, CHRISTIAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 050173/0006 →
Continuity (1)
Related Publication 20210065063A1 · Mar 4, 2021
Cited By (1)
US 12,309,041