IP Library Granted Patent US 12,586,683
Granted Patent B2
US 12,586,683 · App. 17/381,141 · Granted Mar 24, 2026

Decision-making under selective labels

Inventor: Dennis Wei (Sunnyvale, CA)
Assignee: International Business Machines Corporation
G16H50/20G06N5/045G06N20/00G16H20/13
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,683
App. No.
17/381,141
Granted
Mar 24, 2026
Kind
B2
Abstract

A computer-implemented method of decision-making using selective labels, includes receiving a conditional success probability value of a feature associated with an entity. A confidence value of the received success probability value is received. A parameter value that is a trade-off between a short-term learning and a long-term utility is selected. A decision is rendered to accept or reject the feature associated with the entity according to a machine learning policy.

Claims (39)

1 . A computer-implemented method of automatic decision-making using selective labels, the method comprising:

receiving, at a success probability model of a computing device, a feature vector associated with an entity and generating, via the success probability model analyzing the feature vector, a conditional success probability value representing a likelihood of success of a feature associated with the entity if the entity is accepted;

receiving, at a confidence model of the computing device, a sample size associated with prior accepted entities having similar features and generating, via the confidence model analyzing the sample size, a confidence value associated with the conditional success probability value;

determining, at a discount factor model of the computing device, a tuning parameter value that is <1 and that represents a trade-off between a short-term learning cost and a long-term utility;

rendering a decision, using a policy model of the computing device, to accept the entity based on a threshold function of the conditional success probability value, the confidence value, and the tuning parameter value; and

iteratively updating the success probability model and the confidence model using an observed outcome that results from accepting the entity.

2 . The computer-implemented method of claim 1 , wherein:

the machine learning policy for rendering the decision comprises an optimal homogeneous policy; and

the feature associated with the entity is non-distinguishable from a population of other entities.

3 . The computer-implemented method of claim 1 , wherein the machine learning policy for rendering the decision comprises a homogeneous policy for rendering the decision to accept based on the feature associated with the entity.

4 . The computer-implemented method of claim 1 , wherein the machine learning policy for rendering the decision comprises a finite-domain case policy used for rendering the decision to accept based on the feature associated with the entity.

5 . The computer-implemented method of claim 1 , wherein the machine learning policy for rendering the decision comprises an infinite-domain case policy used for rendering the decision to accept the feature associated with the entity.

6 . The computer-implemented method of claim 1 , further comprising updating the policy model based on the observed outcome that results from accepting the entity.

7 . The computer-implemented method of claim 1 , further comprising training a machine learning model to render the decision to accept the feature associated with the entity.

8 . The computer-implemented method of claim 1 , wherein the entity includes multiple features, and the computer-implemented method further comprises training a machine learning model to render the decision to accept based on two or more of the multiple features associated with the entity.

9 . The computer-implemented method of claim 1 , wherein the machine learning policy is based on the conditional success probability value and the confidence value.

10 . A computing device configured for decision-making using selective labels, the computing device comprising:

a processor,

a memory coupled to the processor, the memory storing instructions to cause the processor to perform acts comprising:

receiving, at a success probability model of a computing device, a feature vector associated with an entity and generating, via the success probability model analyzing the feature vector, a conditional success probability value representing a likelihood of success of a feature associated with the entity if the entity is accepted;

receiving, at a confidence model of the computing device, a sample size associated with prior accepted entities having similar features and generating, via the confidence model analyzing the sample size, a confidence value associated with the conditional success probability value;

determining, at a discount factor model of the computing device, a tuning parameter value that is <1 and that represents a trade-off between a short-term learning cost and a long-term utility;

rendering a decision, using a policy model of the computing device, to accept the entity based on a threshold function of the conditional success probability value, the confidence value, and the tuning parameter value; and

iteratively updating the success probability model and the confidence model using an observed outcome that results from accepting the entity.

11 . The computing device of claim 10 , wherein the instructions cause the processor to perform an additional act of training a machine learning model to render the decision to accept the feature associated with the entity.

12 . The computing device of claim 10 , wherein the instructions cause the processor to perform an additional act of training a machine learning model to render the decision to accept based on two or more of the features associated with the entity.

13 . The computing device of claim 10 , wherein the instructions cause the processor to perform an additional act of rendering the decision to accept based on an optimal homogeneous policy.

14 . A computing device configured to perform decision-making using selective labels, the computing device comprising:

one or more processors including a dialog processor configured to process extracted text from a plurality of participants of a collaborative query;

a memory coupled to the one or more processors;

a plurality of models configured in the one or more processors, the plurality of models comprising:

a success probability model configured to provide a conditional success probability value representing an empirical success rate of a feature associated with an entity;

a confidence model configured to provide a confidence value of the conditional success probability value, based at least on a sample size of entities;

a discount factor model configured to determine a tuning parameter value that is <1 and that represents a trade-off between a short-term learning cost and a long-term utility; and

a policy model configured to render a decision to accept the feature associated with the entity based on a threshold function of the conditional success probability value, the confidence value, and the tuning parameter value; and

iteratively update the success probability model and the confidence model with a resultant outcome when the decision is to accept.

15 . The computing device according to claim 14 , wherein the success probability model and the confidence model each comprise a predictive model.

16 . The computing device according to claim 14 , wherein the success probability model is configured for automatically rendering decisions for dispensing a requested pharmaceutical or biological treatment.

17 . The computing device according to claim 14 , wherein the success probability model is configured for rendering decisions regarding suspending an operating license.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 20, 2021
From: WEI, DENNIS
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 056922/0975 →
Continuity (1)
Related Publication 20230034542A1 · Feb 2, 2023
References Cited (20)
US 8001001B2 · Brady et al. · 2011 [cited by applicant]
US 8682724B2 · Gonen · 2014 [cited by applicant]
US 10635707B2 · Perez et al. · 2020 [cited by applicant]
US 11551117B1 · Malhotra · 2023 [cited by examiner]
US 20150051973A1 · Li et al. · 2015 [cited by applicant]
US 20150095271A1 · Ioannidis · 2015 [cited by applicant]
US 20200273000A1 · Li et al. · 2020 [cited by applicant]
US 20200409983A1 · Miller et al. · 2020 [cited by applicant]
US 20210089959A1 · Ghosh et al. · 2021 [cited by applicant]
US 20220004863A1 · Park · 2022 [cited by examiner]
US 20220402522A1 · Tummala · 2022 [cited by examiner]
Kilbertus, Niki et al. “Fair Decisions Despite Imperfect Predictions.” arXiv (Cornell University) (2020): n. pag. Web (Year: 2019). [cited by examiner]
Bietti, A., Agarwal, A., & Langford, J. (2021). A contextual bandit bake-off. Ithaca: (Year: 2021). [cited by examiner]
Casimiro, Maria, et al. “Lynceus: Cost-efficient tuning and provisioning of data analytic jobs.” 2020 IEEE 40th International Conference on Distributed Computing Systems (ICDCS). IEEE, 2020. (Year: 2020). [cited by examiner]
Mell, P. et al., “Recommendations of the National Institute of Standards and Technology”; NIST Special Publication 800-145 (2011); 7 pgs. [cited by applicant]
Bietti, A. et al., “A Contextual Bandit Bake-off”; airXiv:1802.040645v5 [stat.ML] (2021); 49 pgs. [cited by applicant]
Eckles, D. et al., “Thompson Sampling with the Online Bootstrap”; arXiv:1410.4009v1 [cs.LG] (2014); 13 pgs. [cited by applicant]
Foster, D. et al., “Practical Contextual Bandits with Regression Oracles”; PMLR (2018); 2 pgs. [cited by applicant]
Kilbertus, N. et al., “Fair Decisions Despite Imperfect Predictions”; Proceedings of the 23rdInternational Conference on Artificial Intelligence and Statistics (AISTATS) 2020; 10 pgs. [cited by applicant]
Langford, J. et al., “The Epoch-Greedy Algorithm for Contextual Multi-armed Bandits”; Yahoo Research (2008); 8 pgs. [cited by applicant]