IP Library › Granted Patent US 12,737,541
Granted Patent B2
US 12,737,541 · App. 18/493,472 · Granted Sep 15, 2026

Deep model integration techniques for machine learning entity interpretation

Inventors: Brian C. Potter (Carlsbad, CA); Ryan Michael Swan (Santa Monica, CA); Rafael Campos Do Amaral e Vasconcellos (Plymouth, MN); Fraser Ian Lockwood (Dacula, GA)
Assignee: Optum, Inc.
G06F40/279G06N20/00G06T11/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,541
App. No.
18/493,472
Filed
Oct 24, 2023
Granted
Sep 15, 2026
Kind
B2
Art Unit
2693
USPC
704/9
Abstract

Various embodiments of the present disclosure provide machine learning training techniques for implementing a multi-modal interpretation process to generate holistic outputs for an event. The techniques may include generating, using first layers of a multi-modal machine learning model, text-based intermediate representations for an entity based on textual input data. The techniques include generating, using second layers of the multi-modal machine learning model, image-based intermediate representations for the entity based on the text-based intermediate representations and input images for the entity. The techniques include generating, using one or more third layers of the multi-modal machine learning model, an entity representation summary based on the image-based intermediate representations and an image narrative summary for the input images. The techniques include initiating the performance of a prediction-based action based on the entity representation summary.

Claims (52)

1 . A computer-implemented method comprising:

generating, by one or more processors and using a first layer of a multi-modal machine learning model, a text-based intermediate representation for an entity based on textual input data, wherein the textual input data comprises one or more of a historical entity data record or a contextual image record corresponding to a plurality of input images for the entity;

generating, by the one or more processors and using a second layer of the multi-modal machine learning model, an image-based intermediate representation for the entity based on the text-based intermediate representation and the plurality of input images;

generating, by the one or more processors and using a third layer of the multi-modal machine learning model, an entity representation summary based on the image-based intermediate representation and an image narrative summary for the plurality of input images; and

initiating, by the one or more processors, performance of a prediction-based action based on the entity representation summary.

2 . The computer-implemented method of claim 1 , wherein:

the text-based intermediate representation comprises one or more of an initial structured textual output or an initial weight matrix for the entity, and

the image-based intermediate representation comprises one or more of an augmented structured textual output or an augmented weight matrix for the entity.

3 . The computer-implemented method of claim 2 , wherein the entity representation summary comprises a third structured textual output that is based on the augmented structured textual output and the augmented weight matrix.

4 . The computer-implemented method of claim 1 , wherein the first layer, the second layer, and the third layer of the multi-modal machine learning model are trained end-to-end using a labeled training dataset.

5 . The computer-implemented method of claim 1 , further comprising:

receiving a user summary for the plurality of input images; and

generating, using the third layer of the multi-modal machine learning model, the entity representation summary based on the image-based intermediate representation, the image narrative summary, and the user summary.

6 . The computer-implemented method of claim 5 , further comprising:

generating a plurality of performance insights based on a comparison between the entity representation summary and the user summary.

7 . The computer-implemented method of claim 6 , wherein the plurality of performance insights are indicative of a confidence score for the user summary or the entity representation summary.

8 . The computer-implemented method of claim 6 , wherein initiating the performance of the prediction-based action based on the entity representation summary comprises:

generating a performance alert based on the plurality of performance insights; and

providing the performance alert to a user associated with the user summary.

9 . The computer-implemented method of claim 1 , further comprising:

augmenting the textual input data with the entity representation summary.

10 . A system comprising:

one or more processors; and

one or more memories storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

generate, using a first layer of a multi-modal machine learning model, a text-based intermediate representation for an entity based on textual input data, wherein the textual input data comprises one or more of a historical entity data record or a contextual image record corresponding to a plurality of input images for the entity;

generate, using a second layer of the multi-modal machine learning model, an image-based intermediate representation for the entity based on the text-based intermediate representation and the plurality of input images for the entity;

generate, using a third layer of the multi-modal machine learning model, an entity representation summary based on the image-based intermediate representation and an image narrative summary for the plurality of input images; and

initiate performance of a prediction-based action based on the entity representation summary.

11 . The system of claim 10 , wherein:

the text-based intermediate representation comprises one or more of an initial structured textual output or an initial weight matrix for the entity, and

the image-based intermediate representation comprises one or more of an augmented structured textual output or an augmented weight matrix for the entity.

12 . The system of claim 11 , wherein the entity representation summary comprises a third structured textual output that is based on the augmented structured textual output and the augmented weight matrix.

13 . The system of claim 10 , wherein the first layer, the second layer, and the third layer of the multi-modal machine learning model are trained end-to-end using a labeled training dataset.

14 . The system of claim 10 , wherein the one or more processors are further configured to:

receive a user summary for the plurality of input images; and

generate, using the third layer of the multi-modal machine learning model, the entity representation summary based on the image-based intermediate representation, the image narrative summary, and the user summary.

15 . The system of claim 14 , further comprising:

generating a plurality of performance insights based on a comparison between the entity representation summary and the user summary.

16 . The system of claim 15 , wherein initiating the performance of the prediction-based action based on the entity representation summary comprises:

generating a performance alert based on the plurality of performance insights; and

providing the performance alert to a user associated with the user summary.

17 . One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

generating, using a first layer of a multi-modal machine learning model, a text-based intermediate representation for an entity based on textual input data, wherein the textual input data comprises one or more of a historical entity data record or a contextual image record corresponding to a plurality of input images for the entity;

generating, using a second layer of the multi-modal machine learning model, an image-based intermediate representation for the entity based on the text-based intermediate representation and the plurality of input images;

generating, using a third layer of the multi-modal machine learning model, an entity representation summary based on the image-based intermediate representation and an image narrative summary for the plurality of input images; and

initiating performance of a prediction-based action based on the entity representation summary.

18 . The one or more non-transitory computer-readable media of claim 17 , wherein the one or more processors are caused to perform operations further comprising:

receive a user summary for the plurality of input images; and

generate, using the third layer of the multi-modal machine learning model, the entity representation summary based on the image-based intermediate representation, the image narrative summary, and the user summary.

19 . The one or more non-transitory computer-readable media of claim 18 , wherein the one or more processors are caused to perform operations further comprising:

generate a plurality of performance insights based on a comparison between the entity representation summary and the user summary.

20 . The one or more non-transitory computer-readable media of claim 19 , wherein the plurality of performance insights are indicative of a confidence score for the user summary or the entity representation summary.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2023
From: POTTER, BRIAN C.; SWAN, RYAN MICHAEL; VASCONCELLOS, RAFAEL CAMPOS DO AMARAL E; LOCKWOOD, FRASER IAN
To: OPTUM, INC.
Reel/Frame 065327/0422 →
Continuity (1)
Related Publication 20250131196A1 · Apr 24, 2025
References Cited (13)
US 10395772B1 · Lucas et al. · 2019 [cited by applicant]
US 10593426B2 · Amarasingham et al. · 2020 [cited by applicant]
US 10943681B2 · Yao et al. · 2021 [cited by applicant]
US 11205516B2 · Lieberman · 2021 [cited by applicant]
US 11244755B1 · Syeda-Mahmood et al. · 2022 [cited by applicant]
US 20170098153A1 · Mao · 2017 [cited by examiner]
US 20170124432A1 · Chen · 2017 [cited by examiner]
US 20190290172A1 · Hadad et al. · 2019 [cited by applicant]
US 20210232773A1 · Wang · 2021 [cited by examiner]
US 20210343411A1 · Zhang et al. · 2021 [cited by applicant]
US 20240185588A1 · Kumari · 2024 [cited by examiner]
Bohr, et al., “The Rise of Artificial Intelligence in Healthcare Applications”, Artificial Intelligence in Healthcare, 37 Pages, (2020), doi.org/10.1016/B978-0-12-818438-7.00002-2. [cited by applicant]
Esteva, et al., “A Guide to Deep Learning in Healthcare”, Nature Medicine, vol. 25, pp. 24-29, Jan. 2019, DOI: 10.1038/s41591-018-0316-z. [cited by applicant]