IP Library › Granted Patent US 12,260,328
Granted Patent B2
US 12,260,328 · App. 17/494,055 · Granted Mar 25, 2025

Neuro-symbolic reinforcement learning with first-order logic

Inventors: Daiki Kimura (Kanagawa, JP); Masaki Ono (Tokyo, JP); Subhajit Chaudhury (Kawasaki, JP); Michiaki Tatsubori (Oiso, JP)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/08G06F40/205G06F40/289G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,328
App. No.
17/494,055
Granted
Mar 25, 2025
Kind
B2
Abstract

A computer-implemented method for reinforcement learning with Logical Neural Networks (LNNs) is provided including receiving a plurality of observation text sentences from a target environment, extracting one or more propositional logic values from the plurality of observation text sentences, finding a class for each propositional logic value by using external knowledge, converting each propositional logic value into a first-order logic by replacing a part in the propositional logic value with a variable word, the part indicating the class, selecting a LNN based on the class among LNNs prepared in advance for each class, each LNN receiving the one or more propositional logic values as a status input and outputting an action with a score indicating a degree of preference for taking the action, and performing a highest score action to the target environment to obtain a next state of the target environment and a reward for the highest score action.

Claims (40)

1. A computer-implemented method for reinforcement learning (RL) with Logical Neural Networks (LNNs), the method comprising:

receiving a plurality of observation text sentences from a target environment;

extracting one or more propositional logic values from the plurality of observation text sentences;

finding a class for each propositional logic value by using external knowledge;

converting, by an FOL converter, each propositional logic value into a first-order logic (FOL) by replacing a part in the propositional logic value with a variable word, the part indicating the class;

selecting a LNN based on the class among LNNs prepared in advance for each class, each LNN receiving the one or more propositional logic values as a status input and outputting an action with a score indicating a degree of preference for taking the action; and

performing a highest score action to the target environment to obtain a next state of the target environment and a reward for the highest score action.

2. The computer-implemented method of claim 1 , further comprising storing the current state, the highest score action, the next state, and the reward in a replay buffer to update each selected LNN.

3. The computer-implemented method of claim 1 , further comprising generating the FOL based on a history of the states of the target environment.

4. The computer-implemented method of claim 1 , wherein the FOL converter uses a semantic parser or natural language processing method.

5. The computer-implemented method of claim 1 , wherein the LNN includes one or more conjunction gates, one or more disjunction gates, one or more negation nodes, one or more universal quantifiers, and one or more existential quantifiers.

6. The computer-implemented method of claim 1 , wherein an action policy is trained in the LNN from the FOL.

7. The computer-implemented method of claim 1 , wherein the reward for the highest score action is used for adding one or more new nodes and for updating a weight value for each connection.

8. A computer program product for reinforcement learning (RL) with Logical Neural Networks (LNNs), the computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to:

receive a plurality of observation text sentences from a target environment;

extract one or more propositional logic values from the plurality of observation text sentences;

find a class for each propositional logic value by using external knowledge;

convert, by an FOL converter, each propositional logic value into a first-order logic (FOL) by replacing a part in the propositional logic value with a variable word, the part indicating the class;

select a LNN based on the class among LNNs prepared in advance for each class, each LNN receiving the one or more propositional logic values as a status input and outputting an action with a score indicating a degree of preference for taking the action; and

perform a highest score action to the target environment to obtain a next state of the target environment and a reward for the highest score action.

9. The computer program product of claim 8 , wherein the current state, the highest score action, the next state, and the reward are stored in a replay buffer to update each selected LNN.

10. The computer program product of claim 8 , wherein the FOL is generated based on a history of the states of the target environment.

11. The computer program product of claim 8 , wherein the FOL converter uses a semantic parser or natural language processing method.

12. The computer program product of claim 8 , wherein the LNN includes one or more conjunction gates, one or more disjunction gates, one or more negation nodes, one or more universal quantifiers, and one or more existential quantifiers.

13. The computer program product of claim 8 , wherein an action policy is trained in the LNN from the FOL.

14. The computer program product of claim 8 , wherein the reward for the highest score action is used for adding one or more new nodes and for updating a weight value for each connection.

15. A system for reinforcement learning (RL) with Logical Neural Networks (LNNs), the system comprising:

a memory; and

one or more processors in communication with the memory configured to:

receive a plurality of observation text sentences from a target environment;

extract one or more propositional logic values from the plurality of observation text sentences;

find a class for each propositional logic value by using external knowledge;

convert, by an FOL converter, each propositional logic value into a first-order logic (FOL) by replacing a part in the propositional logic value with a variable word, the part indicating the class;

select a LNN based on the class among LNNs prepared in advance for each class, each LNN receiving the one or more propositional logic values as a status input and outputting an action with a score indicating a degree of preference for taking the action; and

perform a highest score action to the target environment to obtain a next state of the target environment and a reward for the highest score action.

16. The system of claim 15 , wherein the current state, the highest score action, the next state, and the reward are stored in a replay buffer to update each selected LNN.

17. The system of claim 15 , wherein the FOL is generated based on a history of the states of the target environment.

18. The system of claim 15 , wherein the FOL converter uses a semantic parser or natural language processing method.

19. The system of claim 15 , wherein the LNN includes one or more conjunction gates, one or more disjunction gates, one or more negation nodes, one or more universal quantifiers, and one or more existential quantifiers.

20. The system of claim 15 , wherein an action policy is trained in the LNN from the FOL.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 5, 2021
From: KIMURA, DAIKI; ONO, MASAKI; CHAUDHURY, SUBHAJIT; TATSUBORI, MICHIAKI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057700/0754 →
Continuity (1)
Related Publication 20230108135A1 · Apr 6, 2023
References Cited (44)
US 20180157977A1 · Saikia · 2018 [cited by examiner]
US 20180349256A1 · Fong et al. · 2018 [cited by applicant]
US 20190108448A1 · O'Malia · 2019 [cited by examiner]
US 20200320435A1 · Sequeira · 2020 [cited by examiner]
US 20210019642A1 · O'Malia · 2021 [cited by examiner]
US 20210232915A1 · Dalli · 2021 [cited by examiner]
US 20210256377A1 · Dalli · 2021 [cited by examiner]
US 20210271817A1 · Sen · 2021 [cited by examiner]
US 20210357738A1 · Luus · 2021 [cited by examiner]
US 20210365817A1 · Riegel · 2021 [cited by examiner]
US 20210406669A1 · Yu · 2021 [cited by examiner]
US 20220036180A1 · Archuleta · 2022 [cited by examiner]
US 20220067520A1 · Dalli · 2022 [cited by examiner]
US 20220114369A1 · Debnath · 2022 [cited by examiner]
US 20220114417A1 · Dalli · 2022 [cited by examiner]
US 20220147876A1 · Dalli · 2022 [cited by examiner]
US 20220172050A1 · Dalli · 2022 [cited by examiner]
US 20220180166A1 · Wachi · 2022 [cited by examiner]
US 20220198254A1 · Dalli · 2022 [cited by examiner]
US 20220269858A1 · Sen · 2022 [cited by examiner]
US 20220277217A1 · Francis · 2022 [cited by examiner]
US 20220300799A1 · Jiang · 2022 [cited by examiner]
US 20230048764A1 · Zheng · 2023 [cited by examiner]
US 20230108135A1 · Kimura · 2023 [cited by examiner]
US 20230143937A1 · Wachi · 2023 [cited by examiner]
US 20230156556A1 · Choi · 2023 [cited by examiner]
US 20230179489A1 · Latapie · 2023 [cited by examiner]
US 20230196063A1 · Latapie · 2023 [cited by examiner]
US 20230259766A1 · Kasioumis · 2023 [cited by examiner]
US 20230274137A1 · Makariou · 2023 [cited by examiner]
US 20230409872A1 · Makondo · 2023 [cited by examiner]
US 20240320503A1 · Kimura · 2024 [cited by examiner]
CN 109858630A · 2019 [cited by applicant]
CN 111382253A · 2020 [cited by applicant]
WO 2020068877A1 · 2020 [cited by applicant]
Anonymous Authors, “Interpretable Reinforcement Learning With Neural Symbolic Logic”, ICLR 2021 Conference, Mar. 5, 2021, pp. 1-11 (Year: 2021). [cited by examiner]
Yuan et al., “Counting to Explore and Generalize in Text-based Games”, 35th International Conference on Machine Learning, arXiv:1806.11525v2 [cs.CL] Mar. 7, 2019, pp. 1-12. [cited by applicant]
Dong et al., “Neural Logic Machines”, 1 arXiv:1904.11694v1 [cs.AI] Apr. 26, 2019, pp. 1-22. [cited by applicant]
Riegel et al., “Logical Neural Networks”, arXiv:2006.13155v1 [cs.AI] Jun. 23, 2020, pp. 1-48. [cited by applicant]
Roukos et al., “Getting AI to Reason: Using Neuro-Symbolic AI for Knowledge-Based Question Answering”, IBM Research Blog, https://research.ibm.com/blog/ai-neurosymbolic-common-sense, Dec. 4, 2020, pp. 1-9. [cited by applicant]
Anonymous Authors, “Interpretable Reinforcement Learning With Neural Symbolic Logic”, ICLR 2021 Conference, Mar. 5, 2021, pp. 1-11. [cited by applicant]
Kimura et al., “Reinforcement Learning with External Knowledge by using Logical Neural Networks” arXiv:2103.02363v1 [cs.AI] Mar. 3, 2021, pp. 1-4. [cited by applicant]
Authors et al., Disclosed Anonymously, “Reinforcement Learning with Logical Neural Networks”, IP.com No. IPCOM000265954D, May 28, 2021, pp. 1-6. [cited by applicant]
International Search Report from PCT/CN2022/117992 dated Dec. 5, 2022. (10 pages). [cited by applicant]
Cited By (1)
US 12,572,674