IP Library › Granted Patent US 12,470,465
Granted Patent B2
US 12,470,465 · App. 17/959,665 · Granted Nov 11, 2025

Continuously improving API service endpoint selections via adaptive reinforcement learning

Inventors: Rong Nickle Chang (Pleasantville, NY); Hongyi Bian (Ames, IA); Nitin Gaur (Round Rock, TX)
Assignee: International Business Machines Corporation
H04L41/16H04L41/0246H04L43/55
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,470,465
App. No.
17/959,665
Granted
Nov 11, 2025
Kind
B2
Abstract

A method includes: receiving, by a processor set, a request for a web-based service; generating, by the processor set, a feature vector including values based on parameters of the request; generating, by the processor set, an endpoint selection vector including plural probabilities corresponding to plural endpoints, wherein the endpoint selection vector is generated using the feature vector with a machine learning model; selecting, by the processor set, one of the plural endpoints based on the plural probabilities; and invoking, by the processor set, the selected endpoint.

Claims (62)

1 . A method comprising:

receiving, by a processor set, a request for a web-based service;

generating, by the processor set, a feature vector including values based on parameters of the request;

generating, by the processor set, an endpoint selection vector including plural probabilities corresponding to plural endpoints, wherein the endpoint selection vector is generated using the feature vector with a machine learning model;

selecting, by the processor set, one of the plural endpoints based on a random number and the plural probabilities;

invoking, by the processor set, the selected endpoint;

in response to the selecting one of the plural endpoints, providing data defining the selected one of the plural endpoints to an endpoint selection logger; and

in response to the invoking the selected endpoint, providing quality-of-experience data to the endpoint selection logger, wherein the quality-of-experience data comprises data that quantifies a quality-of-experience associated with the selected endpoint handling the request.

2 . The method of claim 1 , wherein the selected endpoint comprises one of an application program interface (API) endpoint and a web service endpoint.

3 . The method of claim 1 , wherein the machine learning model comprises a deep learning model.

4 . The method of claim 3 , wherein the deep learning model comprises a neural network.

5 . The method of claim 1 , wherein the generating the endpoint selection vector comprises:

inputting the feature vector to the machine learning model; and

receiving the endpoint selection vector as an output of the machine learning model.

6 . The method of claim 1 , wherein the selecting one of the plural endpoints comprises:

generating the random number;

comparing the random number to the plural probabilities; and

in response to the comparing, selecting the one of the plural endpoints based on the random number matching a probability range associated with the one of the plural endpoints.

7 . The method of claim 1 , further comprising training the machine learning model.

8 . The method of claim 7 , wherein the training the machine learning model comprises:

generating a stream of normalized feature vectors, wherein each respective one of the normalized feature vectors is associated with a respective one of plural logged requests;

generating a stream of probabilistic endpoint selection vectors;

determining loss values comprising a respective loss value for each one of the probabilistic endpoint selection vectors;

training the machine learning model using the loss values.

9 . The method of claim 8 , wherein the respective loss value for one of the probabilistic endpoint selection vectors comprises a difference between an expected reward generated by the machine learning model and an expected reward derived from log data.

10 . The method of claim 9 , wherein the expected reward derived from log data comprises a value that represents a quality-of-experience for a pair comprising a particular request feature vector and a particular endpoint.

11 . A computer program product comprising one or more computer readable storage media having program instructions collectively stored on the one or more computer readable storage media, the program instructions executable to:

receive a request for a web-based service;

generate a feature vector including values based on parameters of the request;

generate an endpoint selection vector including plural probabilities corresponding to plural endpoints, wherein the generating the endpoint selection vector comprises providing the feature vector as an input to a deep learning model;

select one of the plural endpoints based on a random number and the plural probabilities; and

invoke the selected endpoint,

wherein training the deep learning model comprises:

generating a stream of normalized feature vectors, wherein each respective one of the normalized feature vectors is associated with a respective one of plural logged requests;

generating a stream of probabilistic endpoint selection vectors;

determining loss values comprising a respective loss value for each one of the probabilistic endpoint selection vectors;

training the deep learning model using the loss values; and

wherein the respective loss value for one of the probabilistic endpoint selection vectors comprises a difference between an expected reward generated by the machine learning model and an expected reward derived from log data.

12 . The computer program product of claim 11 , wherein the selecting one of the plural endpoints comprises:

generating the random number;

comparing the random number to the plural probabilities; and

in response to the comparing, selecting the one of the plural endpoints based on the random number matching a probability range associated with the one of the plural endpoints.

13 . The computer program product of claim 11 , wherein the expected reward derived from log data comprises a value that represents a quality-of-experience for a pair comprising a particular request feature vector and a particular endpoint.

14 . The computer program product of claim 11 , wherein the program instructions are executable to:

in response to the selecting one of the plural endpoints, provide data defining the selected one of the plural endpoints to an endpoint selection logger; and

in response to the invoking the selected endpoint, provide quality-of-experience data to the endpoint selection logger, wherein the quality-of-experience data comprises data that quantifies a quality-of-experience associated with the selected endpoint handling the request.

15 . A system comprising:

a processor set, one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable to:

receive a request for a web-based service;

generate a feature vector including values based on parameters of the request;

generate an endpoint selection vector including plural probabilities corresponding to plural endpoints, wherein the generating the endpoint selection vector comprises providing the feature vector as an input to a deep learning model;

select one of the plural endpoints based on a random number and the plural probabilities; and

invoke the selected endpoint,

wherein the selecting one of the plural endpoints comprises:

generating the random number;

comparing the random number to the plural probabilities; and

in response to the comparing, selecting the one of the plural endpoints based on the random number matching a probability range associated with the one of the plural endpoints.

16 . The system of claim 15 , wherein training the deep learning model comprises:

generating a stream of normalized feature vectors, wherein each respective one of the normalized feature vectors is associated with a respective one of plural logged requests;

generating a stream of probabilistic endpoint selection vectors;

determining loss values comprising a respective loss value for each one of the probabilistic endpoint selection vectors, wherein the respective loss values are based on expected rewards;

training the deep learning model using the loss values.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2022
From: CHANG, RONG NICKLE; BIAN, HONGYI; GAUR, NITIN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061306/0325 →
Continuity (1)
Related Publication 20240113945A1 · Apr 4, 2024
References Cited (32)
US 10985987B2 · Imperia · 2021 [cited by examiner]
US 11095534B1 · Dunsmore · 2021 [cited by examiner]
US 11729071B1 · Kolar · 2023 [cited by examiner]
US 11776273B1 · Chen · 2023 [cited by examiner]
US 11882046B1 · Wang · 2024 [cited by examiner]
US 20120014388A1 · Shinohara · 2012 [cited by examiner]
US 20200076842A1 · Zhou · 2020 [cited by examiner]
US 20200162391A1 · Savalle · 2020 [cited by examiner]
US 20200371851A1 · Liu · 2020 [cited by examiner]
US 20210035025A1 · Kalluri · 2021 [cited by examiner]
US 20210055977A1 · Lisuk · 2021 [cited by examiner]
US 20210256391A1 · Karlinsky · 2021 [cited by examiner]
US 20220030086A1 · Meng et al. · 2022 [cited by applicant]
US 20220093246A1 · Karri · 2022 [cited by examiner]
US 20220245175A1 · Hawco · 2022 [cited by examiner]
US 20230031654A1 · Pandey · 2023 [cited by examiner]
US 20230206058A1 · Wellmann · 2023 [cited by examiner]
US 20240031292A1 · Reddy · 2024 [cited by examiner]
CN 110809306 · 2021 [cited by applicant]
CN 112561104 · 2021 [cited by applicant]
CN 112671865 · 2021 [cited by applicant]
CN 112272353 · 2021 [cited by applicant]
CN 114003387 · 2022 [cited by applicant]
CN 112506657 · 2022 [cited by applicant]
WO 2021208720 · 2021 [cited by applicant]
Wang et al., “Delay-Aware Microservice Coordination in Mobile Edge Computing: A Reinforcement Learning Approach”, IEEE Transactions on Mobile Computing, vol. 20, No. 3, doi: 10.1109/TMC.2019.2957804, Mar. 2021, 13 pages. [cited by applicant]
Magableh et al., “A Deep Recurrent Q network towards Self-adapting Distributed Microservice architecture”, Software: Practice and Experience 50.2, https://arxiv.org/pdf/1901.04011.pdf, 2020, 13 pages. [cited by applicant]
Fissaa et al., “An Intelligent Approach for Context-Aware Service Selection using Machine Learning”, In Proceedings of the International Conference on Learning and Optimization Algorithms: Theory and Applications, https… [cited by applicant]
Ren et al., “A Reinforcement Learning Method for Constraint-Satisfied Services Composition”, IEEE Transactions on Services Computing, vol. 13, No. 5, 2017, 15 pages. [cited by applicant]
Chandrakant, “Reinforcement Learning with Neural Network”, Baeldung on Computer Science, https://www.baeldung.com/cs/author/kumar-chandrakant, Jun. 20, 2022, 21 pages. [cited by applicant]
Anonymous, “Reinforcement learning”, Wikipedia, https://en.wikipedia.org/wiki/Reinforcement_learning, Sep. 13, 2022, 16 pages. [cited by applicant]
Anonymous, “Artificial neural network”, Wikipedia, https://en.wikipedia.org/wiki/Artificial_neural_network, Sep. 13, 2022, 28 pages. [cited by applicant]