IP Library › Granted Patent US 12,573,494
Granted Patent B1
US 12,573,494 · App. 19/223,888 · Granted Mar 10, 2026

Multi-dimensional reinforcement learning in an AI platform

Inventors: Luís Ungaro Pinto Coelho (Oporto, PT); Ana Clara Ferreira Matos (Oporto, PT); Filipe Daniel Martins Rodrigues (Oporto, PT); Virgílio António Ferro Bento (Oporto, PT); Ivo Emanuel Marques Gabriel (Oporto, PT); Pedro Henrique Oliveira Santos (Oporto, PT); Fabíola Maria Tavares Alves da Costa Moutinho (Oporto, PT); Fernando Emanuel Dias Correia (Oporto, PT)
Assignee: SWORD HEALTH, S.A.
G16H20/30G16H50/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,494
App. No.
19/223,888
Granted
Mar 10, 2026
Kind
B1
Abstract

Examples in the present disclosure relate to multi-dimensional reinforcement learning in a digital platform. The digital platform maintains and updates user data associated with tracked users. One or more machine learning models process the user data to generate recommendation data that includes automatically generated exploratory program elements for known activity programs of respective users of the digital platform. The exploratory program elements are validated by supervising users before implementation. The digital platform continuously updates the one or more machine learning models using dual reinforcement signals: one based on supervising user-validated recommendations and another based on activity program results. Further recommendation data is generated after updating the one or more machine learning models.

Claims (65)

1 . A system comprising:

a plurality of hardware devices associated with a plurality of tracked users of the system;

one or more hardware processors; and

memory storing instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations comprising:

for each tracked user of the plurality of tracked users:

maintaining, in at least one database, user data comprising at least one of user tracking data or user interaction data related to an activity program of the tracked user, the user tracking data based at least partially on sensor data from one or more sensors of a hardware device from among the plurality of hardware devices that is associated with the tracked user;

processing, by one or more generative machine learning models, the user data to obtain recommendation data comprising one or more automatically generated exploratory program elements of the activity program of the tracked user, the recommendation data being generated by the one or more generative machine learning models based on one or more prompts comprising the user data;

generating validated activity program data based on user input, received from a user device of a supervising user from among a plurality of supervising users of the system, in response to the recommendation data, the user input comprising at least one of an approval of the recommendation data or an update associated with the recommendation data; and

updating, in the at least one database, the user data to record results of the activity program performed according to the validated activity program data;

automatically modifying parameters of the one or more generative machine learning models by training the one or more generative machine learning models based on both the validated activity program data of the plurality of tracked users and the results of the plurality of tracked users; and

after modifying the parameters of the one or more generative machine learning models, generating further recommendation data for at least one of the plurality of tracked users, the further recommendation data being generated by the one or more generative machine learning models based on one or more further prompts comprising further user data.

2 . The system of claim 1 , wherein the training of the one or more generative machine learning models comprises performing at least one of supervised fine-tuning, reinforcement fine-tuning, or reinforcement learning with human feedback.

3 . The system of claim 1 , wherein the training of the one or more generative machine learning models comprises performing automated reinforcement learning by:

updating a first reinforcement learning signal based on the recommendation data as validated by one or more of the plurality of supervising users;

updating a second reinforcement learning signal based on the results of the activity programs; and

training the one or more generative machine learning models by applying both the first reinforcement learning signal and the second reinforcement learning signal to the one or more generative machine learning models.

4 . The system of claim 3 , wherein the first reinforcement learning signal and the second reinforcement learning signal are derived from paired data comprising a supervising user-approved recommendation and a corresponding measured outcome in relation to a particular activity program.

5 . The system of claim 3 , wherein generating the further recommendation data comprises, based on the automated reinforcement learning, automatically detecting and applying successful patterns for different user cohorts across the plurality of tracked users.

6 . The system of claim 1 , wherein the training of the one or more generative machine learning models is performed during a second stage using updated training data based on the validated activity program data of the plurality of tracked users and the results of the plurality of tracked users, the operations further comprising, in a first stage prior to the second stage:

training the one or more generative machine learning models using base training data.

7 . The system of claim 1 , wherein at least one of the one or more generative machine learning models has a transformer architecture, and the operations further comprise controlling generation of the one or more automatically generated exploratory program elements at least in part by adjusting a temperature parameter of the one or more generative machine learning models.

8 . The system of claim 1 , wherein the one or more automatically generated exploratory program elements comprise at least one of: a new activity not contained in the user data, an activity variation not contained in the user data, or a new activity sequence not contained in the user data.

9 . The system of claim 8 , wherein the one or more automatically generated exploratory program elements comprise the new activity sequence, and the recommendation data indicates a recommended change in an order of activities to be performed by a particular tracked user of the plurality of tracked users.

10 . The system of claim 1 , the operations further comprising:

transmitting the recommendation data to the user device of the supervising user; and

receiving the user input from the user device of the supervising user, the user input indicating approval of the one or more automatically generated exploratory program elements in the recommendation data, wherein the validated activity program data is automatically generated to include the one or more automatically generated exploratory program elements in response to receiving the user input.

11 . The system of claim 1 , the operations comprising:

transmitting the recommendation data to the user device of the supervising user; and

receiving the user input from the user device of the supervising user, the user input indicating an update to the one or more automatically generated exploratory program elements in the recommendation data, wherein the validated activity program data is generated by applying the update to the one or more automatically generated exploratory program elements in response to receiving the user input.

12 . The system of claim 1 ,

wherein the user tracking data comprises at least one of: real-time tracking data from one or more activity sessions, performance data, biomechanical data, physiological data, behavioral data, or environmental data, and

wherein the user interaction data comprises at least one of: clinical data, baseline assessment data, enrollment data, demographic data, activity session data, progress evaluation data, communication history data, user feedback data, or supervising user feedback data.

13 . The system of claim 1 , wherein the one or more sensors comprise one or more image sensors and the one or more automatically generated exploratory program elements comprise one or more new exercises, the operations further comprising, for each tracked user of the plurality of tracked users:

capturing, by the one or more image sensors, images of the tracked user while the tracked user performs the one or more new exercises;

processing the images to generate real-time tracking data;

updating the user tracking data based on the real-time tracking data; and

generating, by the one or more generative machine learning models, the further recommendation data based at least partially on the updated user tracking data.

14 . The system of claim 13 , the operations further comprising:

displaying a user interface comprising a digital video feed representing the images together with at least one of activity tracking output or instructions for performing the one or more new exercises.

15 . The system of claim 1 , wherein the hardware device comprises at least one of: a tablet computing device, a mobile phone, a wearable device, or an intra-body device.

16 . The system of claim 1 , wherein the one or more sensors comprise one or more tracking sensors and one or more environmental sensors.

17 . The system of claim 1 , wherein the one or more generative machine learning models comprise one or more recommender machine learning models, the operations further comprising:

executing one or more session analyzer artificial intelligence (AI) agents to process the user tracking data associated with one or more activity sessions of the activity program; and

executing one or more behavioral analyzer AI agents to process behavioral data associated with user behavior outside of the one or more activity sessions,

wherein the one or more recommender machine learning models process outputs of the one or more session analyzer AI agents and the one or more behavioral analyzer AI agents to generate the recommendation data.

18 . The system of claim 1 , the operations further comprising:

executing one or more user engagement AI agents to generate personalized interactions with the plurality of tracked users;

receiving, from a particular tracked user of the plurality of tracked users, feedback in response to a personalized interaction generated by the one or more user engagement AI agents; and

in response to receiving the feedback, automatically updating the activity program of the particular tracked user.

19 . A computer-implemented method comprising:

for each tracked user of a plurality of tracked users:

maintaining, in at least one database, user data comprising at least one of user tracking data or user interaction data related to an activity program of the tracked user, the user tracking data based at least partially on sensor data from one or more sensors of a hardware device associated with the tracked user;

processing, by one or more generative machine learning models, the user data to obtain recommendation data comprising one or more automatically generated exploratory program elements of the activity program of the tracked user, the recommendation data being generated by the one or more generative machine learning models based on one or more prompts comprising the user data;

generating validated activity program data based on user input, received from a user device of a supervising user, in response to the recommendation data, the user input comprising at least one of an approval of the recommendation data or an update associated with the recommendation data; and

updating, in the at least one database, the user data to record results of the activity program performed according to the validated activity program data;

automatically modifying parameters of the one or more generative machine learning models by training the one or more generative machine learning models based on both the validated activity program data of the plurality of tracked users and the results of the plurality of tracked users; and

after modifying the parameters of the one or more generative machine learning models, generating further recommendation data for at least one of the plurality of tracked users, the further recommendation data being generated by the one or more generative machine learning models based on one or more further prompts comprising further user data.

20 . One or more non-transitory machine-readable storage media including instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations comprising:

for each tracked user of a plurality of tracked users:

maintaining, in at least one database, user data comprising at least one of user tracking data or user interaction data related to an activity program of the tracked user, the user tracking data based at least partially on sensor data from one or more sensors of a hardware device associated with the tracked user;

processing, by one or more generative machine learning models, the user data to obtain recommendation data comprising one or more automatically generated exploratory program elements of the activity program of the tracked user, the recommendation data being generated by the one or more generative machine learning models based on one or more prompts comprising the user data;

generating validated activity program data based on user input, received from a user device of a supervising user, in response to the recommendation data, the user input comprising at least one of an approval of the recommendation data or an update associated with the recommendation data; and

updating, in the at least one database, the user data to record results of the activity program performed according to the validated activity program data;

automatically modifying parameters of the one or more generative machine learning models by training the one or more generative machine learning models based on both the validated activity program data of the plurality of tracked users and the results of the plurality of tracked users; and

after modifying the parameters of the one or more generative machine learning models, generating further recommendation data for at least one of the plurality of tracked users, the further recommendation data being generated by the one or more generative machine learning models based on one or more further prompts comprising further user data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2025
From: UNGARO PINTO, LUÍS UNGARO PINTO; MATOS, ANA CLARA FERREIRA; MARTINS RODRIGUES, FILIPE DANIEL; FERRO, VIRGÍLIO ANTÓNIO FERRO; SANTOS, PEDRO HENRIQUE OLIVEIRA; COSTA MOUTINHO, FABÍOLA MARIA TAVARES ALVES DA; DIAS CORREIA, FERNANDO EMANUEL; GABRIEL, IVO EMANUEL MARQUES
To: SWORD HEALTH, S.A.
Reel/Frame 071269/0776 →
References Cited (50)
US 9898789B2 · Ram et al. · 2018 [cited by applicant]
US 10130311B1 · De Sapio et al. · 2018 [cited by applicant]
US 10413238B1 · Cooper et al. · 2019 [cited by applicant]
US 11039763B2 · Ye et al. · 2021 [cited by applicant]
US 12397198B1 · Bento et al. · 2025 [cited by applicant]
US 20070179816A1 · Lemme · 2007 [cited by applicant]
US 20120290319A1 · Saria et al. · 2012 [cited by applicant]
US 20150038806A1 · Kaleal, III et al. · 2015 [cited by applicant]
US 20150324532A1 · Jones et al. · 2015 [cited by applicant]
US 20180330810A1 · Gamarnik · 2018 [cited by examiner]
US 20190328322A1 · Inada · 2019 [cited by applicant]
US 20200066390A1 · Svendrys et al. · 2020 [cited by applicant]
US 20200114207A1 · Weldemariam et al. · 2020 [cited by applicant]
US 20210202103A1 · Bostic et al. · 2021 [cited by applicant]
US 20220016484A1 · Bissonnette et al. · 2022 [cited by applicant]
US 20220076666A1 · Trehan · 2022 [cited by applicant]
US 20220208385A1 · Voschina et al. · 2022 [cited by applicant]
US 20220246268A1 · Hunter et al. · 2022 [cited by applicant]
US 20220392611A1 · Appelbaum · 2022 [cited by examiner]
US 20230071274A1 · Trehan · 2023 [cited by applicant]
US 20250273351A1 · Matos et al. · 2025 [cited by applicant]
CN 115023763 · 2022 [cited by applicant]
WO 2015103442 · 2015 [cited by applicant]
WO 2019010435 · 2019 [cited by applicant]
WO 2020198065 · 2020 [cited by applicant]
WO WO2022086454A1 · 2022 [cited by applicant]
Alowais et al. Revolutionizing healthcare: the role of artificial intelligence in clinical practice. BMC Med Educ. 2023;23(1):689. Published Sep. 2, 20232. doi: 10.1186/s12909-023-04698-z (Year: 2023). [cited by examiner]
Mosqueira-Rey et al. Human-in-the-loop machine learning: a state of the art. Artificial Intelligence Review 56, 3005-3054 (2023). https://doi.org/10.1007/s10462-022-10246-w (Year: 2023). [cited by examiner]
Mukherjee, Subhabrata, et al., “Polaris A Safety focused LLM Constellation Architecture for Healthcare”, arXiv:2403.13313v1 [cs.AI] Mar. 20, 2024, (Mar. 20, 2024), 53 pages. [cited by applicant]
Oberst, Michael, et al., “Pioneering the Science of AI Evaluation”, [Online]. Retrieved from the Internet: <https://www.abridge.com/ai/science-ai-evaluation>, (Accessed online May 21, 2025), 25 pages. [cited by applicant]
“U.S. Appl. No. 18/585,380, Non Final Office Action mailed Apr. 22, 2024”, 25 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,380, Response filed Jul. 22, 2024 to Non Final Office Action mailed Apr. 22, 2024”, 19 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,380, Examiner Interview Summary mailed Jul. 26, 2024”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,355, Non Final Office Action mailed Jul. 30, 2024”, 20 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,380, Final Office Action mailed Aug. 9, 2024”, 23 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,380, Response filed Sep. 30, 2024 to Final Office Action mailed Aug. 9, 2024”, 15 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,380, Advisory Action mailed Oct. 9, 2024”, 3 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,355, Response filed Oct. 29, 2024 to Non Final Office Action mailed Jul. 30, 2024”, 15 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,380, Response filed Nov. 8, 2024 to Advisory Action mailed Oct. 9, 2024”, 17 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,355, Final Office Action mailed Dec. 2, 2024”, 18 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,380, Non Final Office Action mailed Jan. 29, 2025”, 20 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,355, Response filed Jan. 31, 2025 to Final Office Action mailed Dec. 2, 2024”, 17 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,355, Advisory Action mailed Feb. 11, 2025”, 3 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,355, Notice of Allowance mailed Mar. 19, 2025”, 7 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,380, Response filed Apr. 24, 2025 to Non Final Office Action mailed Jan. 29, 2025”, 14 pgs. [cited by applicant]
“U.S. Appl. No. 18/585,380, Final Office Action mailed May 23, 2025”, 19 pgs. [cited by applicant]
U.S. Appl. No. 18/585,380, Examiner Interview Summary mailed Jul. 1, 2025, 2 pages. [cited by applicant]
U.S. Appl. No. 18/585,380, Notice of Allowance mailed Aug. 4, 2025, 13 pgs. [cited by applicant]
U.S. Appl. No. 18/585,380, Response filed Jul. 17, 2025 to Final Office Action mailed May 23, 2025, 13 pgs. [cited by applicant]
Kalakoti, Yogesh, et al., “TransDTI: Transformer-Based Language Models for Estimating DTIs and Building a Drug Recommendation Workflow”, ACS Publications; ACS Omega, 2706-2717, (2022), 12 pgs. [cited by applicant]