Method and system for dynamic request for quotation pricing using reinforcement learning and symbolic regression
A method and a system for dynamic request for quotation (RFQ) pricing using reinforcement learning and symbolic regression in order to obtain a pseudo-optimal nonlinear state-feedback controller for tracking a metric with increased interpretability are provided. The method includes: receiving bid price information and ask price information that relates to a financial instrument; selecting a metric to be used in conjunction with a determination of an RFQ price with respect to the financial instrument; formulating a Markov Decision Process (MDP) that relates to the determination of the RFQ price; using the MDP to define an RFQ pricing controller with respect to the metric; applying a reinforcement learning algorithm in order to tune a behavior of the RFQ pricing controller with respect to the metric; and using the RFQ pricing controller, the first information, and the second information to determine the RFQ price.
1 . A method for performing a dynamic request for quotation (RFQ) pricing operation, the method being implemented by at least one processor, the method comprising: receiving, by the at least one processor, first information that relates to a bid price for a first instrument and second information that relates to an ask price for the first instrument; selecting a metric to be used in conjunction with a determination of an RFQ price with respect to the first instrument; formulating a Markov Decision Process (MDP) that relates to the determination of the RFQ price, wherein the MDP is designed to reduce a tracking error that corresponds to a difference between an observed value of the metric and a target value of the metric; using the MDP to define an RFQ pricing controller with respect to the metric, wherein the MDP is configured to minimize the tracking error, wherein the using of the MDP to define the RFQ pricing controller includes training and applying a neural network to a set of policy parameters relating to the metric, wherein the policy parameters include a proportional gain, an integral gain and a derivative gain, wherein the proportional gain controls a response of the RFQ pricing controller to an instantaneous tracking error for yielding higher control outputs as corrections to higher instantaneous tracking errors, and wherein the integral gain controls a response of the RFQ pricing controller to accumulated tracking error over past time, and the derivative gain controls a response of the RFQ pricing controller to a rate of change of the tracking error for yielding higher control output when the rate of change of the tracking error increases; selectively applying a reinforcement learning algorithm in order to tune a behavior of the RFQ pricing controller with respect to the metric, wherein the reinforcement learning algorithm is configured to be only applied when the first information or the second information affect a behavior of a dynamic system above a reference threshold; and using the RFQ pricing controller, the first information, and the second information to determine the RFQ price, wherein the using of the MDP to define the RFQ pricing controller includes: performing a symbolic regression process to obtain an algebraic approximation of the neural network, and performing the MDP within a Monte-Carlo simulation of the RFQ pricing operation; evaluating, via the Monte-Carlo simulation, robustness of the RFQ pricing controller against a prediction error based on a simulated performance of the RFQ pricing controller using the proportional gain, the integral gain and the derivative gain, wherein the Monte-Carlo simulation implements an ability to tune and evaluate the RFQ pricing controller for constraints in an actuation of the RFQ pricing controller by introducing an upper bound and a lower bound to a range of outputs of the RFQ pricing controller, wherein the evaluating of the RFQ pricing controller includes introducing, in the Monte-Carlo simulation and using an imperfect prediction model, random noise in outputs of the RFQ pricing controller; performing tuning of the RFQ pricing controller based on the evaluated robustness of the RFQ pricing controller in the Monte-Carlo simulation for reducing the prediction error of the RFQ pricing controller; and applying the tuned RFQ pricing controller on the first information and the second information for tracking the metric with reduced tracking error, wherein the reduced tracking error includes minimized tracking error.
2 . The method of claim 1 , wherein the metric includes at least one from among a hit ratio, a performance ratio, an expected profit, an expected market share, and a risk level.
3 . The method of claim 1 , wherein the first instrument includes at least one from among a stock that relates to a first entity, a bond that relates to the first entity, an option that relates to the first entity, and a derivative financial instrument that relates to the first entity.
4 . A computing apparatus for performing a dynamic request for quotation (RFQ) pricing operation, the computing apparatus comprising: a processor; a memory; and a communication interface coupled to each of the processor and the memory, wherein the processor is configured to: receive, via the communication interface, first information that relates to a bid price for a first instrument and second information that relates to an ask price for the first instrument; select a metric to be used in conjunction with a determination of an RFQ price with respect to the first instrument; formulate a Markov Decision Process (MDP) that relates to the determination of the RFQ price, wherein the MDP is designed to reduce a tracking error that corresponds to a difference between an observed value of the metric and a target value of the metric; use the MDP to define an RFQ pricing controller with respect to the metric, wherein the MDP is configured to minimize the tracking error, wherein the using of the MDP to define the RFQ pricing controller includes training and applying a neural network to a set of policy parameters relating to the metric, wherein the policy parameters include a proportional gain, an integral gain and a derivative gain, wherein the proportional gain controls a response of the RFQ pricing controller to an instantaneous tracking error for yielding higher control outputs as corrections to higher instantaneous tracking errors, and wherein the integral gain controls a response of the RFQ pricing controller to accumulated tracking error over past time, and the derivative gain controls a response of the RFQ pricing controller to a rate of change of the tracking error for yielding higher control output when the rate of change of the tracking error increases; selectively apply a reinforcement learning algorithm in order to tune a behavior of the RFQ pricing controller with respect to the metric, wherein the reinforcement learning algorithm is configured to be only applied when the first information or the second information affect a behavior of a dynamic system above a reference threshold; and use the RFQ pricing controller, the first information, and the second information to determine the RFQ price, wherein the using of the MDP to define the RFQ pricing controller includes: performing a symbolic regression process to obtain an algebraic approximation of the neural network, and performing the MDP within a Monte-Carlo simulation of the RFQ pricing operation, evaluating, via the Monte-Carlo simulation, robustness of the RFQ pricing controller against a prediction error based on a simulated performance of the RFQ pricing controller using the proportional gain, the integral gain and the derivative gain, wherein the Monte-Carlo simulation implements an ability to tune and evaluate the RFQ pricing controller for constraints in an actuation of the RFQ pricing controller by introducing an upper bound and a lower bound to a range of outputs of the RFQ pricing controller, wherein the evaluating of the RFQ pricing controller includes introducing, in the Monte-Carlo simulation and using an imperfect prediction model, random noise in outputs of the RFQ pricing controller, perform tuning of the RFQ pricing controller based on the evaluated robustness of the RFQ pricing controller in the Monte-Carlo simulation for reducing the prediction error of the RFQ pricing controller, and apply the tuned RFQ pricing controller on the first information and the second information for tracking the metric with reduced tracking error, wherein the reduced tracking error includes minimized tracking error.
5 . The computing apparatus of claim 4 , wherein the metric includes at least one from among a hit ratio, a performance ratio, an expected profit, an expected market share, and a risk level.
6 . The computing apparatus of claim 4 , wherein the first instrument includes at least one from among a stock that relates to a first entity, a bond that relates to the first entity, an option that relates to the first entity, and a derivative financial instrument that relates to the first entity.
7 . A non-transitory computer readable storage medium storing instructions for performing a dynamic request for quotation (RFQ) pricing operation, the storage medium comprising executable code which, when executed by a processor, causes the processor to: receive first information that relates to a bid price for a first instrument and second information that relates to an ask price for the first instrument; select a metric to be used in conjunction with a determination of an RFQ price with respect to the first instrument; formulate a Markov Decision Process (MDP) that relates to the determination of the RFQ price, wherein the MDP is designed to reduce a tracking error that corresponds to a difference between an observed value of the metric and a target value of the metric; use the MDP to define an RFQ pricing controller with respect to the metric, wherein the MDP is configured to minimize the tracking error, wherein the using of the MDP to define the RFQ pricing controller includes training and applying a neural network to a set of policy parameters relating to the metric, wherein the policy parameters include a proportional gain, an integral gain and a derivative gain, wherein the proportional gain controls a response of the RFQ pricing controller to an instantaneous tracking error for yielding higher control outputs as corrections to higher instantaneous tracking errors, and wherein the integral gain controls a response of the RFQ pricing controller to accumulated tracking error over past time, and the derivative gain controls a response of the RFQ pricing controller to a rate of change of the tracking error for yielding higher control output when the rate of change of the tracking error increases; selectively apply a reinforcement learning algorithm in order to tune a behavior of the RFQ pricing controller with respect to the metric, wherein the reinforcement learning algorithm is configured to be only applied when the first information or the second information affect a behavior of a dynamic system above a reference threshold; and use the RFQ pricing controller, the first information, and the second information to determine the RFQ price, wherein the using of the MDP to define the RFQ pricing controller includes: performing a symbolic regression process to obtain an algebraic approximation of the neural network, and performing the MDP within a Monte-Carlo simulation of the RFQ pricing operation, evaluate, via the Monte-Carlo simulation, robustness of the RFQ pricing controller against a prediction error based on a simulated performance of the RFQ pricing controller using the proportional gain, the integral gain and the derivative gain, wherein the Monte-Carlo simulation implements an ability to tune and evaluate the RFQ pricing controller for constraints in an actuation of the RFQ pricing controller by introducing an upper bound and a lower bound to a range of outputs of the RFQ pricing controller, wherein the evaluating of the RFQ pricing controller includes introducing, in the Monte-Carlo simulation and using an imperfect prediction model, random noise in outputs of the RFQ pricing controller, perform tuning of the RFQ pricing controller based on the evaluated robustness of the RFQ pricing controller in the Monte-Carlo simulation for reducing the prediction error of the RFQ pricing controller, and apply the tuned RFQ pricing controller on the first information and the second information for tracking the metric with reduced tracking error, wherein the reduced tracking error includes minimized tracking error.
8 . The storage medium of claim 7 , wherein the metric includes at least one from among a hit ratio, a performance ratio, an expected profit, an expected market share, and a risk level.