IP Library Granted Patent US 12,367,350
Granted Patent B2
US 12,367,350 · App. 16/946,586 · Granted Jul 22, 2025

Random action replay for reinforcement learning

Inventors: Wei Zhang (Littleton, MA); Murray Scott Campbell (Yorktown Heights, NY); Yang Yu (Acton, MA); Sadhana Kumaravel (White Plains, NY)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F40/35G06F18/214G06F40/284G06F40/56G06N3/082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,350
App. No.
16/946,586
Granted
Jul 22, 2025
Kind
B2
Abstract

An artificial intelligence (AI) platform to support random action replay for natural language (NL) learning. A NL conversation is subject to exploration to train a neural network. One or more tuples are leveraged for the training, with each tuple representing an input action, a vector, an output action, and a reward value. An action is sampled from the vector, with the sampling configured to assess a corresponding first gradient. The first gradient is applied to selectively adjust the neural network. As NL input is received and applied to the selectively adjusted neural network, an output corresponding to the NL input is identified and a corresponding action is subject to be executed.

Claims (40)

1. A computer system comprising:

a processing unit operatively coupled to memory;

an artificial intelligence (AI) platform operatively coupled to the processing unit, the AI platform configured with one or more tools to support random action replay for natural language (NL) learning, the one or more tools comprising:

a training manager configured to train a neural network, the training further comprising the training manager to:

explore a NL conversation, the exploration to leverage one or more tuples associated with the NL conversation, each tuple representing at least an input action, an output action, a policy vector, and a reward value;

select a tuple and sample a first action, from a distribution of actions, associated with the selected tuple;

assess the sampled first action, including generate output associated with the assessment, compare the generated output to a value of the sampled first action corresponding to the policy vector and, based on the comparison, calculate a first gradient representing a distance of the generated output from the sampled first action in the selected tuple associated with the NL conversation; and

apply the first gradient to selectively adjust the neural network;

a language manager operatively coupled to the training manager, the language manager configured to receive and apply NL input to the selectively adjusted neural network, and generate a NL output corresponding to the received NL input; and

the language manager configured to execute an identified action corresponding to the identified output.

2. The computer system of claim 1 , further comprising an interaction manager operatively coupled to the training manager, the interaction manager configured to create the one or more tuples in an interactive environment with corresponding first and second agents, the interactive environment to identify one or more actions from the distribution of actions as a response to receipt of the input action.

3. The computer system of claim 1 , further comprising the training manager configured to re-train the neural network and incorporate a sampled second action from the distribution of actions, calculate a second gradient representing a distance of the sampled second action from the input action, and apply the second gradient to selectively adjust the neural network.

4. The computer system of claim 3 , further comprising the training manager configured to assess the first and second gradients, and responsive to identification of a convergence of the first and second gradients the training manager further configured to terminate training of the neural network.

5. The computer system of claim 1 , further comprising the training manager configured to utilize a random choice function to select the first action from the distribution of actions for sampling.

6. The computer system of claim 1 , wherein the trained neural network is configured to evaluate the received NL input and to determine one or more NL components of the evaluated NL input.

7. The computer system of claim 6 , further comprising the trained neural network configured to evaluate the determined one or more NL components and determine an action corresponding to the received NL input.

8. A computer program product comprising a computer readable storage medium having program code embodied therewith, the program code executable by a processor to:

train a neural network, the training further comprising the program code to:

explore a natural language (NL) conversation, the exploration to leverage one or more tuples associated with the NL conversation, each tuple representing at least an input action, an output action, a policy vector, and a reward value;

select a tuple and sample a first action, from a distribution of actions, associated with the selected tuple;

assess the sampled first action, including generate output associated with the assessment, compare the generated output to a value of the sampled first action corresponding to the policy vector and, based on the comparison, calculate a first gradient representing a distance of the generated output from the sampled first action in the selected tuple associated with the NL conversation; and

apply the first gradient to selectively adjust the neural network;

receive and apply NL input to the selectively adjusted neural network, and generate a NL output corresponding to the received NL input; and

execute an identified action corresponding to the identified output.

9. The computer program product of claim 8 , further comprising the program code executable by the processor to create the one or more tuples in an interactive environment with corresponding first and second agents, the interactive environment to identify one or more actions from the distribution of actions as a response to receipt of the input action.

10. He computer program product of claim 8 , further comprising the program code executable by the processor to re-train the neural network and incorporate a sampled second action from the distribution of actions, calculate a second gradient representing a distance of the sampled second action from the input action; and apply the second gradient to selectively adjust the neural network.

11. The computer program product of claim 10 , further comprising the program code executable by the processor to assess the first and second gradients, and responsive to identification of a convergence of the first and second gradients terminate training of the neural network.

12. The computer program product of claim 8 , further comprising the program code executable by the processor to utilize a random choice function to select the first action from the distribution of actions for sampling.

13. A computer implemented method comprising:

training a neural network, the training further comprising:

exploring a natural language (NL) conversation, the exploration to leverage one or more tuples associated with the NL conversation, each tuple representing an input action, an output action, a policy vector, and a reward value;

selecting a tuple and sampling a first action, from a distribution of actions, associated with the selected tuple;

assessing the sampled first action, including generate output associated with the assessment, compare the generated output to a value of the sampled first action corresponding to the policy vector and, based on the comparison, calculate a first gradient representing a distance of the generated output from the sampled first action in the selected tuple associated with the NL conversation; and

applying the first gradient to selectively adjust the neural network;

receiving and applying NL input to the selectively adjusted neural network, and generating a NL output corresponding to received NL input; and

executing an identified action corresponding to the identified output.

14. The method of claim 13 , further comprising creating the one or more tuples in an interactive environment with corresponding first and second agents, the interactive environment to identify one or more actions from the distribution of actions as a response to receipt of the input action.

15. The method of claim 13 , further comprising re-training the neural network and incorporating a sampled second action from the distribution of actions, calculating a second gradient representing a distance of the sampled second action from the input action, and applying the second gradient to selectively adjust the neural network.

16. The method of claim 15 , further comprising assessing the first and second gradients, and responsive to identification of a convergence of the first and second gradients terminating training of the neural network.

17. The method of claim 13 , further comprising utilizing a random choice function to select the first action from the distribution of actions for sampling.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2020
From: ZHANG, WEI; CAMPBELL, MURRAY SCOTT; YU, YANG; KUMARAVEL, SADHANA
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 053072/0394 →
Continuity (1)
Related Publication 20210406689A1 · Dec 30, 2021
References Cited (67)
US 9480908B2 · Brown et al. · 2016 [cited by applicant]
US 9679258B2 · Mnih · 2017 [cited by examiner]
US 10204097B2 · Lipton et al. · 2019 [cited by applicant]
US 10510000B1 · Commons · 2019 [cited by examiner]
US 11715042B1 · Liu · 2023 [cited by examiner]
US 20170330556A1 · Fatemi Booshehri · 2017 [cited by examiner]
US 20180300317A1 · Bradbury · 2018 [cited by applicant]
US 20190115027A1 · Shah · 2019 [cited by examiner]
US 20190362074A1 · Wang et al. · 2019 [cited by applicant]
US 20200012953A1 · Sun et al. · 2020 [cited by applicant]
US 20200143247A1 · Jonnalagadda · 2020 [cited by examiner]
US 20210019642A1 · O'Malia · 2021 [cited by examiner]
US 20210232922A1 · Zhang · 2021 [cited by examiner]
US 20210272559A1 · Medalion · 2021 [cited by examiner]
US 20220036884A1 · Ramachandran · 2022 [cited by examiner]
US 20220103891A1 · Xu · 2022 [cited by examiner]
EP 2381393 · 2011 [cited by applicant]
Wen, Tsung-Hsien, et al. “Multi-domain neural network language generation for spoken dialogue systems.” arXiv preprint arXiv: 1603.01232 (2016) (Year: 2016). [cited by examiner]
Cerisara, Christophe, Pavel Kral, and Ladislav Lenc. “On the effects of using word2vec representations in neural networks for dialogue act recognition.” Computer Speech & Language 47 (2018): 175-193 (Year: 2018). [cited by examiner]
Liu, Bing, and Ian Lane. “Iterative policy learning in end-to-end trainable task-oriented neural dialog models.” 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2017 (Year: 2017). [cited by examiner]
Prollochs, N., et al., “Reinforcement Learning in R”, University of Oxford, Oct. 31, 2018. [cited by applicant]
Mnih, V., et al., “Playing Atari with Deep Reinforcement Learning”, arXiv: 1312.5602v1, Dec. 19, 2013. [cited by applicant]
Schaul, T., et al., “Prioritized Experience Replay”, ICLR 2016, arXiv: 1511.05952v4, Feb. 25, 2016. [cited by applicant]
Wang, Z., et al., “Sample Efficieny Actor-Critic With Experience Replay”, ICLR 2017, arXiv: 1611.01224v2, Jul. 10, 2017. [cited by applicant]
Foerster, J., et al., “Stabilising Experience Replay for Deep Multi-Agent Reinforcement Learning”, Proceedings of the 34th International Conference on Machine Learning, PMLR 70, 2017, arXiv: 1702.08887v3, May 21, 2018. [cited by applicant]
Liang, C., et al.. “Memory Agumented Policy Optimization for Program Synthesis and Semantic Parsing”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), arXiv: 1807.02322v5, Jan. 13, 2019. [cited by applicant]
Abadi, M., et al., “TensorFlow: A System for Large-Scale Machine Learning”, 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI '16), pp. 265-283, Nov. 2-4, 2016. [cited by applicant]
Andrews, M., et al., “Integrating Experiential and Distributional Data to Learn Semantic Representations”, Psychological Review 2009, vol. 116, No. 3, pp. 463-498. [cited by applicant]
Busoniu, L., et al., “A comprehensive survey of multi-agent reinforcement learning”, IEEE Transactions on Systems, Man, and Cybernetics, Part C: Applications and Reviews, vol. 38, No. 2, pp. 156-172, Mar. 2008. [cited by applicant]
Church, K. W., et al., “Word Association Norms, Mutual Information and Lexicography”, Computational Linguistics, 16(1), pp. 22-29, Mar. 1990. [cited by applicant]
Das, A., et al., “Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning”, IEEE International Conference on Computer Vision (ICCV), pp. 2951-2960, 2017. [cited by applicant]
De Deyne, S., et al., “Predicting Human Similarity Judgments with Distributional Models: The Value of Word Associations”, Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence (IJCAI-… [cited by applicant]
De Vries, H., et al., “GuessWhat?! Visual object discovery through multi-modal dialogue”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5503-5512, 2017. [cited by applicant]
Foerster, J. N., et al., “Learning to Communicate with Deep Multi-Agent Reinforcement Learning”, Advances in Neural Information Processing Systems, pp. 2137-2145, 2016. [cited by applicant]
Foerster, J. N., et al., “Counterfactual Multi-Agent Policy Gradients”, 32nd AAAI Conference on Artificial Intelligence, 2018. [cited by applicant]
Frank, M. C., “Predicting Pragmatic Reasoning in Language Games”, Science, vol. 336, pp. 998-998, May 25, 2012. [cited by applicant]
Ghazvininejad, M., et al., “A Knowledge-Grounded Neural Conversation Model”, 32nd AAAI Confernce on Artificial Intelligence (AAAI-18), pp. 5510-5517, 2018. [cited by applicant]
Havrylov, S., et al., “Emergence of Language with Multi-agent Games: Learning to Communicate with Sequences of Symbols”, Advances in Neural Information Processing Systems, pp. 2149-2159, 2017. [cited by applicant]
He, H., et al., “Learning Symmetric Collaborative Dialogue Agents with Dynamic Knowledge Graph Embeddings”, arXiv: 1704.0713v1, Apr. 24, 2017. [cited by applicant]
Hinton, G., et al., “Neural Networks for Machine Learning”, Lecture 6a, Overview of mini-batch gradient descent, Feb. 5, 2016. [cited by applicant]
Hu, H., et al., “Playing 20 Question Game with Policy-Based Reinforcement Learning”, arXiv:1808.07645v3, Jun. 24, 2019. [cited by applicant]
Lazardou, A., et al., “Multi-Agent Cooperation and the Emergence of (Natural) Language”, arXiv: 1612.01782v2, Mar. 5, 2017. [cited by applicant]
Levin, J. A., et al. “Dialogue-games: Metacommunication Structures for Natural Language Interaction”, Cognitive Science, 1(4), pp. 395-420, 1977. [cited by applicant]
Li, J., et al., “Deep Reinforcement Learning for Dialogue Generation”, Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pp. 1192-1202, arXiv: 1606.01541v4, Sep. 29, 2016. [cited by applicant]
Lowe, R., et al., “Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments”, 31st Conference on Neural Information Processing Systems (NIPS 2017). [cited by applicant]
Lowe, R., et al., “Training End-to-End Dialogue Systems with the Ubuntu Dialogue Corpus”, Dialogue & Discourse 8(1), pp. 31-65, 2017. [cited by applicant]
Mordatch, I., et al., “Emergence of Grounded Compositional Language in Multi-Agent Populations”, 32nd AAAI Conference on Artificial Intelligence, arXiv: 1703.04908v2, Jul. 24, 2018. [cited by applicant]
Narasimhan, K., et al., “Language Understanding for Text-based Games using Deep Reinforcement Learning”, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Sep. 2015. [cited by applicant]
Nelson, D. L., et al., “The University of South Florida Word Association, Rhyme and Word Fragment Norms”, Behavior Research Methods, Instruments & Computers, 36(3), pp. 40-407, 1998. [cited by applicant]
Omidshafiei, S., et al., “Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability”, Proceedings of the 34th International Conferene on Machine Learning, 2017, arXiv: 1703.06182v4, Ju… [cited by applicant]
OpenAI Five, https://openai.com/blog/openai-five/, Jun. 25, 2018. [cited by applicant]
Pincus, E., et al., “Towards Automatic Identification of Effective Clues for Team Word-Guessing Games”, Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), European Language… [cited by applicant]
Rubin, D. C., et al., “Predicting which words get recalled: Measures of free recall, availability, goodness, emotionality, and pronunciability for 925 nouns”, Memory & Recognition, 14(1), pp. 79-94, 1986. [cited by applicant]
Serban, J. V., et al., “Building end-to-end dialogue systems using generative hierarchical neural network models”, AAAI '16: Proceedings of the 30th AAAI Conference on Artificial Intelligence, pp. 3776-3783, Feb. 2016. [cited by applicant]
Silver, D., et al., “Mastering the game of Go with deep neural networks and tree search”, Nature 529, pp. 484-489, 2016. [cited by applicant]
Speer, R., et al., “ConceptNet 5.5: An Open Multilingual Graph of General Knowledge”, 31st AAAI Conference on Artificial Intelligence, arXiv:1612.03975v2, Dec. 11, 2018. [cited by applicant]
Sukhbaatar, S., et al., “Learning multiagent communication with backpropagation” NIPS'16: Proceedings of the 30 International Conference on Neural Information Processing Systems, pp. 2252-2260, Dec. 2016. [cited by applicant]
Sutton, R.S., et al., “Policy Gradient Methods for Reinforcement Learning with Function Approximation”, Advances in Neural Information Processing Systems, pp. 1057-1063, 2000. [cited by applicant]
https://www.hasbro.com/common/instruct/Taboo(2000).pdf, Accessed: Mar. 10, 2019. [cited by applicant]
Vinyals, O., et al., “A Neural Conversation Model”, Proceedings of the 31st International Conference on Machine Learning, arXiv: 1506.05869v3, Jul. 22, 2015. [cited by applicant]
Von Ahn, L., et al., “Designing games with a purpose”, Communications of the ACM, vol. 51, Issue 8, Aug. 2008. [cited by applicant]
Wen, T-H. , et al., “A Network-based End-to-End Trainable Task-oriented Dialogue System”, arXiv:1604.04562v3, Apr. 24, 2017. [cited by applicant]
Williams, Ronald, J., “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning”, Machine Learning, 8, pp. 229-256, 1992. [cited by applicant]
Wu, Chien-Sheng, et al., “Global-to-Local Memory Pointer Networks for Task-Oriented Dialogue”, arXiv:1901.04713v2, Mar. 29, 2019. [cited by applicant]
Young, S., et al., “POMDP-based Statistical Spoken Dialogue Systems: a Review”, Proceedings of the IEEE, vol. 101, Issue 5, pp. 1160-1179, May 2013. [cited by applicant]
Young, T., et al., “Augmenting End-to-End Dialogue Systems with Commonsense Knowledge”, arXiv: 1709.05453v3, Feb. 12, 2018. [cited by applicant]
Zhou, H., et al., “Commonsense Knowledge Aware Conversation Generation with Graph Attention”, Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence (IJCAI-18), pp. 4623-4629, 2018. [cited by applicant]