IP Library › Granted Patent US 12,394,508
Granted Patent B2
US 12,394,508 · App. 18/302,185 · Granted Aug 19, 2025

Systems and methods for training multi-armed bandit models

Inventors: Mohsen Afrasiabi (Madison, WI); Tanzeem Choudhury (New York, NY); Cecilia M. Livesey (Merion Station, PA); Jared Dustin Martin (Minneapolis, MN); Herk Anthony Confer (San Francisco, CA); Daniel Joseph Mulcahy (Evanston, IL); Rony Krell (Brooklyn, NY)
Assignee: UnitedHealth Group Incorporated
G16H20/00G06N3/092
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,508
App. No.
18/302,185
Granted
Aug 19, 2025
Kind
B2
Abstract

A method for determining a treatment recommendation using a multi-armed bandit (MAB) model can include receiving first patient information, determining, using the MAB model, the treatment recommendation based on the first patient information, wherein the MAB model is trained based on a MAB treatment recommendation determined by the MAB model using second patient information and a clinical treatment recommendation determined according to clinical guidelines based on the second patient information, and providing the treatment recommendation.

Claims (43)

1. A computer-implemented method for determining a treatment recommendation using a multi-armed bandit (MAB) model, the method comprising:

receiving, by one or more processors, first patient information;

determining, by the one or more processors and using the MAB model, the treatment recommendation based on the first patient information, wherein the MAB model is trained based on a MAB model treatment recommendation determined by the MAB model using second patient information and a clinical treatment recommendation determined according to clinical guidelines based on the second patient information; and

providing, by the one or more processors, the treatment recommendation, wherein the MAB model is trained by:

receiving training data including third patient information, selected treatment information, and selected treatment outcome information;

determining a reward probability distribution of treatment options of the MAB model using the training data and an artificial intelligence (AI) model;

receiving the second patient information;

determining the MAB model treatment recommendation using the MAB model configured with the reward probability distribution and a Thompson sampling technique, based on the second patient information;

determining the clinical treatment recommendation based on the second patient information;

determining a confidence score of the MAB model treatment recommendation;

determining a hybrid treatment recommendation based on the MAB model treatment recommendation, the clinical treatment recommendation, and the confidence score; and

training the MAB model based on the hybrid treatment recommendation.

2. The computer-implemented method of claim 1 , wherein the reward probability distribution indicates non-equal likelihoods of effectiveness of treatment options.

3. The computer-implemented method of claim 1 , wherein the receiving the first patient information comprises receiving the first patient information from a user device based on the first patient information being input via a graphical user interface of the user device.

4. A system for determining a treatment recommendation using a multi-armed bandit (MAB) model, the system comprising:

one or more non-transitory computer readable media storing processor-executable instructions; and

one or more processors configured to execute the processor-executable instructions to perform operations comprising:

receiving first patient information;

determining, using the MAB model, the treatment recommendation based on the first patient information, wherein the MAB model is trained based on a MAB model treatment recommendation determined by the MAB model using second patient information and a clinical treatment recommendation determined according to clinical guidelines based on the second patient information; and

providing the treatment recommendation, wherein the MAB model is trained by:

receiving training data including third patient information, selected treatment information, and selected treatment outcome information;

determining a reward probability distribution of treatment options of the MAB model using the training data and an artificial intelligence (AI) model;

receiving the second patient information;

determining the MAB model treatment recommendation using the MAB model configured with the reward probability distribution and a Thompson sampling technique, based on the second patient information;

determining the clinical treatment recommendation based on the second patient information;

determining a confidence score of the MAB model treatment recommendation;

determining a hybrid treatment recommendation based on the MAB model treatment recommendation, the clinical treatment recommendation, and the confidence score; and

training the MAB model based on the hybrid treatment recommendation.

5. The system of claim 4 , wherein the reward probability distribution indicates non-equal likelihoods of effectiveness of treatment options.

6. The system of claim 4 , wherein the receiving the first patient information comprises receiving the first patient information from a user device based on the first patient information being input via a graphical user interface of the user device.

7. One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors for determining a treatment recommendation using a multi-armed bandit (MAB) model, cause the one or more processors to perform operations comprising:

receiving first patient information;

determining, using the MAB model, the treatment recommendation based on the first patient information, wherein the MAB model is trained based on a MAB model treatment recommendation determined by the MAB model using second patient information and a clinical treatment recommendation determined according to clinical guidelines based on the second patient information; and

providing the treatment recommendation, wherein the MAB model is trained by:

receiving training data including third patient information, selected treatment information, and selected treatment outcome information;

determining a reward probability distribution of treatment options of the MAB model using the training data and an artificial intelligence (AI) model;

receiving the second patient information;

determining the MAB model treatment recommendation using the MAB model configured with the reward probability distribution and a Thompson sampling technique, based on the second patient information;

determining the clinical treatment recommendation based on the second patient information;

determining a confidence score of the MAB model treatment recommendation;

determining a hybrid treatment recommendation based on the MAB model treatment recommendation, the clinical treatment recommendation, and the confidence score; and

training the MAB model based on the hybrid treatment recommendation.

8. The one or more non-transitory computer-readable media of claim 7 , wherein the reward probability distribution indicates non-equal likelihoods of effectiveness of treatment options.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2023
From: AFRASIABI, MOHSEN; CHOUDHURY, TANZEEM; LIVESEY, CECILIA M.; MARTIN, JARED DUSTIN; CONFER, HERK ANTHONY; MULCAHY, DANIEL JOSEPH; KRELL, RONY
To: UNITEDHEALTH GROUP INCORPORATED
Reel/Frame 063369/0047 →
Continuity (2)
Provisional Application 63381392 · Oct 28, 2022
Related Publication 20240145057A1 · May 2, 2024
References Cited (10)
US 20150140527A1 · Gilad-Barach · 2015 [cited by examiner]
US 20210241873A1 · Kapaldo · 2021 [cited by examiner]
US 20220013230A1 · Wu · 2022 [cited by examiner]
US 20220415472A1 · Hakala · 2022 [cited by examiner]
US 20230103124A1 · Kano · 2023 [cited by examiner]
EP 3859741A2 · 2021 [cited by applicant]
WO 2022076221A1 · 2022 [cited by applicant]
Zhou et al., “Spoiled for Choice? Personalized Recommendation for Healthcare Decisions: A Multi-Armed Bandit Approach,” arXiv: 2009.06108, (Year: 2020). [cited by examiner]
Cortes, David, “Adapting multi-armed bandits policies to contextual bandits scenarios,” research paper, Nov. 23, 2019, arXiv preprint arXiv:1811.04383, accessible at: https://arxiv.org/pdf/1811.04383.pdf. [cited by applicant]
Cortes, David, “Contextual Bandits,” website, accessible at: https://github.com/david-cortes/contextualbandits. [cited by applicant]