IP Library › Granted Patent US 11,250,313
Granted Patent B2
US 11,250,313 · App. 16/258,919 · Granted Feb 15, 2022

Autonomous trading with memory enabled neural network learning

Inventors: Sakyasingha Dasgupta (Shinagawa-ku, JP); Rudy R. Harry Putra (Yokohama, JP)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/0472G06F5/06G06F17/18G06N3/0454G06Q40/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,250,313
App. No.
16/258,919
Granted
Feb 15, 2022
Kind
B2
Abstract

A computer-implemented method is provided for autonomously making continuous trading decisions for assets using a first eligibility trace enabled Neural Network (NN). The method includes pretraining the first eligibility trace enabled NN, using asset price time series data, to generation predictions of future asset price time series data. The method further includes initializing a second eligibility trace enabled NN for reinforcement learning using learned parameters of the first eligibility trace enabled NN. The method also includes augmenting state information of the second eligibility trace enabled NN for reinforcement learning using an output from the first eligibility trace enabled NN. The method additionally includes performing continuous actions for trading assets at each of multiple time points.

Claims (34)

1. A computer-implemented method for autonomously making continuous trading decisions for assets using a first eligibility trace enabled Neural Network (NN), comprising:

pretraining the first eligibility trace enabled NN, using asset price time series data, to generation predictions of future asset price time series data;

initializing a second eligibility trace enabled NN for reinforcement learning using learned parameters of the first eligibility trace enabled NN;

augmenting state information of the second eligibility trace enabled NN for reinforcement learning using an output from the first eligibility trace enabled NN; and

performing continuous actions for trading assets at each of multiple time points.

2. The computer-implemented method of claim 1 , further comprising training a DyBM with eligibility traces and FIFO queues for online learning on the asset price time series data based on maximizing a likelihood of the asset price time series data at each time-step.

3. The computer-implemented method of claim 1 , further comprising initializing another DyBM with eligibility traces and FIFO queues for reinforcement learning using the parameters of the trained DyBM.

4. The computer-implemented method of claim 1 , wherein the first eligibility trace enabled NN comprises an ensemble of eligibility trace enabled neural networks.

5. The computer-implemented method of claim 1 , wherein nodes of the second eligibility trace enabled NN, corresponding to representative neurons, are divided into two groups, wherein a first group of the two groups represents actions and a second group of the two groups represents observations.

6. The computer-implemented method of claim 5 , wherein the observations comprise the asset price time series data and the predictions of future asset price time series data.

7. The computer-implemented method of claim 1 , wherein the second eligibility trace enabled NN is a Gaussian eligibility trace enabled NN.

8. The computer-implemented method of claim 1 , further comprising updating parameters of the second eligibility trace enabled NN using a temporal-difference error.

9. The computer-implemented method of claim 1 , wherein the continuous actions are determined based on a Gaussian exploration policy for maximizing a future reward.

10. The computer implemented method of claim 1 , wherein trading the assets comprises performing an action selected from the group consisting of buying and selling.

11. A computer program product for autonomously making continuous trading decisions for assets using a first eligibility trace enabled Neural Network (NN), the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

pretraining the first eligibility trace enabled NN, using asset price time series data, to generation predictions of future asset price time series data;

initializing a second eligibility trace enabled NN for reinforcement learning using learned parameters of the first eligibility trace enabled NN;

augmenting state information of the second eligibility trace enabled NN for reinforcement learning using an output from the first eligibility trace enabled NN; and

performing continuous actions for trading assets at each of multiple time points.

12. The computer program product of claim 11 , wherein the method further comprises training a DyBM with eligibility traces and FIFO queues for online learning on the asset price time series data based on maximizing a likelihood of the asset price time series data at each time-step.

13. The computer program product of claim 11 , wherein the method further comprises initializing another DyBM with eligibility traces and FIFO queues for reinforcement learning using the parameters of the trained DyBM.

14. The computer program product of claim 11 , wherein the first eligibility trace enabled NN comprises an ensemble of eligibility trace enabled neural networks.

15. The computer program product of claim 11 , wherein nodes of the second eligibility trace enabled NN, corresponding to representative neurons, are divided into two groups, wherein a first group of the two groups represents actions and a second group of the two groups represents observations.

16. The computer program product of claim 15 , wherein the observations comprise the asset price time series data and the predictions of future asset price time series data.

17. The computer program product of claim 11 , wherein the second eligibility trace enabled NN is a Gaussian eligibility trace enabled NN.

18. The computer program product of claim 11 , wherein the method further comprises updating parameters of the second eligibility trace enabled NN using a temporal-difference error.

19. The computer program product of claim 11 , wherein the continuous actions are determined based on a Gaussian exploration policy for maximizing a future reward.

20. A computer processing system for autonomously making continuous trading decisions for assets using a first eligibility trace enabled Neural Network (NN), comprising:

a memory for storing program code; and

a processor device for running the program code to

pretrain the first eligibility trace enabled NN, using asset price time series data, to generation predictions of future asset price time series data;

initialize a second eligibility trace enabled NN for reinforcement learning using learned parameters of the first eligibility trace enabled NN;

augment state information of the second eligibility trace enabled NN for reinforcement learning using an output from the first eligibility trace enabled NN; and

perform continuous actions for trading assets at each of multiple time points.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 7, 2019
From: HARRY PUTRA, RUDY R.; DASGUPTA, SAKYASINGHA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 049990/0701 →
Continuity (1)
Related Publication 20200242449A1 · Jul 30, 2020
Cited By (1)
US 12,596,594