IP Library Granted Patent US 12,362,049
Granted Patent B2
US 12,362,049 · App. 17/661,946 · Granted Jul 15, 2025

Differentiable multi-agent actor-critic for multi-step radiology report summarization system

Inventors: Sanjeev Kumar Karn (Plainsboro, NJ); Oladimeji Farri (Upper Saddle River, NJ)
Assignee: Siemens Healthineers AG
G16H15/00G06N3/084G06N3/092G06F40/216
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,362,049
App. No.
17/661,946
Granted
Jul 15, 2025
Kind
B2
Abstract

Systems and methods for using a differentiable multi-agent Actor-Critic (DiMAC) for multi-step radiology report summarization. The tasks of extracting salient sentences and phrases are divided across two collaborating agents that are trained end-to-end using reinforcement learning (RL).

Claims (33)

1. A system for multi-step radiology report summarization, the system comprising:

a word extraction network configured to extract one or more words from a radiology report;

a sentence extraction network configured to extract one or more salient sentences from the radiology report based at least in part on the extracted one or more words; and

an abstractor network configured to condense one or more of the one or more extracted salient sentences into a concise summary of the radiology report;

wherein the word extraction network, the sentence extraction network, and the abstractor network are trained end to end using at least a critic and a communication channel, wherein the critic is configured to estimate a value function that is used to compute a policy gradient for the word extraction network and the sentence extraction network, and the communication channel is configured to pass gradient information between the word extraction network and the sentence extraction network.

2. The system of claim 1 , wherein the word extraction network and the sentence extraction network comprise a bi-directional LSTM based word encoder and a bi-directional LSTM sentence encoder respectively; wherein the bi-directional LSTM based word encoder is configured to obtain word representations from a FINDINGS portion of the radiology report and the bi-directional LSTM sentence encoder is configured to obtain sentence representations.

3. The system of claim 1 , wherein the abstractor network comprises a pointer generator network.

4. The system of claim 1 , wherein during training, the word extraction network and sentence extraction network are configured using multiple agent reinforcement learning.

5. The system of claim 1 , wherein during training, the word extraction network and sentence extraction network actions at a previous step are input into the communication channel that generates a sigmoidal, mt, that is input into the sentence extraction network, wherein 1−mt is input into the word extraction network.

6. The system of claim 1 , wherein during training, when the word extraction network or the sentence extraction network selects an action, a reward is provided to the word extraction network or the sentence extraction network to it based on an individual reward function computed specifically for either the word extraction network or the sentence extraction network.

7. The system of claim 6 , wherein the reward for the sentence extraction network is computed using ROUGE L recall by comparing extracted salient sentences and condensed salient sentences with ground truth data.

8. The system of claim 6 , wherein the reward for the word extraction network is computed by checking if an extracted word is in a given set of keywords.

9. The system of claim 1 , further comprising an output interface configured to provide the concise summary to a user.

10. A method for configuring a system for multi-step radiology report summarization, wherein the system comprises at least a sentence extraction network, a word extraction network, a critic network, and a communication channel, the method comprising:

inputting training data from a FINDINGS section of a radiology report to the sentence extraction network and the word extraction network;

computing states for the sentence extraction network, the word extraction network, and the critic network;

sampling actions by the sentence extraction network and the word extraction network;

computing rewards based on the actions and the states; and

updating the sentence extraction network, the word extraction network, and the critic network based on the computed rewards.

11. The method of claim 10 , wherein inputting, computing, sampling, computing, and updating is iteratively performed for a plurality of iterations.

12. The method of claim 11 , wherein when sampling actions by the sentence extraction network and the word extraction network, only one of the sentence extraction network or the word extraction network is active while the other is paused.

13. The method of claim 10 , wherein each of the sentence extraction network and the word extraction network selects one of its actions and communicates with the other respective network at every step.

14. The method of claim 10 , wherein the communication channel is configured to generate a sigmoidal, mt, that is input into the sentence extraction network wherein 1−mt is input into the word extraction network.

15. The method of claim 10 , wherein the word extraction network is configured to extract one or more words from the FINDINGS section based on a list of keywords.

16. The method of claim 15 , wherein the sentence extraction network is configured to extract one or more sentences from the FINDINGS section based in part on the extracted one or more words.

17. A method for automatic summarization of a FINDINGS section of a radiology report, the method comprising:

acquiring a radiology report that describes results of a medical procedure;

inputting the FINDINGS section of the radiology report into an automatic summarization system comprising at least a word extraction network, a sentence extraction network, a communications channel, and an abstractor, wherein the automatic summarization system is trained using differential Multi-agent Actor-Critic reinforcement learning;

outputting, by the automatic summarization system, an IMPRESSIONS section for the radiology report; and

providing the IMPRESSIONS section to a user.

18. The method of claim 17 , wherein the word extraction network is configured to extract one or more words from the FINDINGS section based on a list of keywords.

19. The method of claim 18 , wherein the sentence extraction network is configured to extract one or more sentences from the FINDINGS section based in part on the extracted one or more words.

20. The method of claim 17 , wherein the communications channel is configured to pass information between the word extraction network and the sentence extraction network.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: SIEMENS HEALTHCARE GMBH
To: SIEMENS HEALTHINEERS AG
Reel/Frame 066267/0346 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2022
From: SIEMENS MEDICAL SOLUTIONS USA, INC.
To: SIEMENS HEALTHCARE GMBH
Reel/Frame 059887/0634 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2022
From: KUMAR KARN, SANJEEV; FARRI, OLADIMEJI
To: SIEMENS MEDICAL SOLUTIONS USA, INC.
Reel/Frame 059868/0064 →
Continuity (2)
Provisional Application 63186246 · May 10, 2021
Related Publication 20220374584A1 · Nov 24, 2022
References Cited (35)
US 11197036B2 · Chao · 2021 [cited by examiner]
US 20220375210A1 · Ngo · 2022 [cited by examiner]
US 20240347156A1 · Paulett · 2024 [cited by examiner]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. [cited by applicant]
N Reed Dunnick and Curtis P Langlotz. 2008. The radiology report of the future: a summary of the 2007 Intersociety conference. Journal of the American College of Radiology, 5(5):626-629. [cited by applicant]
Günes Erkan and Dragomir R. Radev. 2011. Lexrank: Graph-based lexical centrality as salience in text summarization. CoRR, abs/1109.2128. [cited by applicant]
Jakob Foerster, Gregory Farquhar, Triantafyllos Afouras, Nantas Nardelli, and Shimon Whiteson. 2018. Counterfactual multi-agent policy gradients. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32. [cited by applicant]
Jakob N Foerster, Yannis M Assael, Nando De Freitas, and Shimon Whiteson. 2016. Learning to communicate with deep multi-agent reinforcement learning. arXiv preprint arXiv:1605.06676. [cited by applicant]
Kilem Li Gwet. 2008. Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology, 61(1):29-48. [cited by applicant]
Eduard Hovy, Chin-Yew Lin, et al. 1999. Automated text summarization in summarist. Advances in automatic text summarization, 14:81-94. [cited by applicant]
Vijay R Konda and John N Tsitsiklis. 2000. Actorcritic algorithms. In Advances in neural information processing systems, pp. 1008-1014. Citeseer. [cited by applicant]
Landon Kraemer and Bikramjit Banerjee. 2016. Multiagent reinforcement learning as a rehearsal for decentralized planning. Neurocomputing, 190:82-94. [cited by applicant]
Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pp. 74-81, Barcelona, Spain. Association for Computational Linguistics. [cited by applicant]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781. [cited by applicant]
Igor Mordatch and Pieter Abbeel. 2018. Emergence of grounded compositional language in multi-agent populations. In Thirty-second AAAI conference on artificial intelligence. [cited by applicant]
Jingbo Shang, Jialu Liu, Meng Jiang, Xiang Ren, Clare R Voss, and Jiawei Han. 2018. Automated phrase mining from massive text corpora. IEEE Transactions on Knowledge and Data Engineering, 30(10):1825-1837. [cited by applicant]
Sainbayar Sukhbaatar, Rob Fergus, et al. 2016. Learning multiagent communication with backpropagation. Advances in neural information processing systems, 29:2244-2252. [cited by applicant]
Richard S Sutton, David A McAllester, Satinder P Singh, Yishay Mansour, et al. 1999. Policy gradient methods for reinforcement learning with function approximation. In NIPs, vol. 99, pp. 1057-1063. Citeseer. [cited by applicant]
Ming Tan. 1993. Multi-agent reinforcement learning: Independent vs. cooperative agents. In Proceedings of the tenth international conference on machine learning, pp. 330-337. [cited by applicant]
Zheng Tian, Shihao Zou, Ian Davies, Tim Warr, Lisheng Wu, Haitham Bou Ammar, and Jun Wang. 2020. Learning to communicate implicitly by actions. In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, … [cited by applicant]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. arXiv preprint arXiv:1706.03762. [cited by applicant]
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015. Pointer networks. In Advances in Neural Information Processing Systems, vol. 28. Curran Associates, Inc. [cited by applicant]
A Wallis and P McCoubrie. 2011. The radiology report—are we getting the message across? Clinical radiology, 66(11):1015-1022. [cited by applicant]
Manzil Zaheer, Guru Guruganesh, Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontanon, Philip Pham, Anirudh Ravula, QifanWang, Li Yang, et al. 2020. Big bird: Transformers for longer sequences. arXiv preprint a… [cited by applicant]
Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Peter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In International Conference on Machine Learning, pp. 11328-11339. PMLR. [cited by applicant]
Yuhao Zhang, Derek Merck, Emily Bao Tsai, Christopher D. Manning, and Curtis P. Langlotz. 2019. Optimizing the factual correctness of a summary: A study of summarizing radiology reports. [cited by applicant]
Emily Alsentzer, John Murphy, William Boag, Wei-Hung Weng, Di Jin, Tristan Naumann, and Matthew McDermott. 2019. Publicly available clinical BERT embeddings. In Proceedings of the 2nd Clinical Natural Language Processin… [cited by applicant]
Yen-Chun Chen and Mohit Bansal. 2018. Fast abstractive summarization with reinforce-selected sentence rewriting. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long P… [cited by applicant]
Wan-Ting Hsu, Chieh-Kai Lin, Ming-Ying Lee, Kerui Min, Jing Tang, and Min Sun. 2018. A unified model for extractive and abstractive summarization using inconsistency loss. In Proceedings of the 56th Annual Meeting of th… [cited by applicant]
Aishwarya Jadhav and Vaibhav Rajan. 2018. Extractive summarization with Swap-Net: Sentences and words from alternating pointer networks. In Proceedings of the 56th Annual Meeting of the Association for Computational Lin… [cited by applicant]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer' 2019. Bart: Denoising sequence-to-sequence pretraining for natural language generation, … [cited by applicant]
Yang Liu and Mirella Lapata. 2019. Text summarization with pretrained encoders. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Nat… [cited by applicant]
Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014. The Stanford CoreNLP natural language processing toolkit. In Association for Computational Linguistics (ACL)… [cited by applicant]
Abigail See, Peter J. Liu, and Christopher D. Manning. 2017. Get to the point: Summarization with pointergenerator networks. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (vo… [cited by applicant]
Yuhao Zhang, Daisy Yi Ding, Tianpei Qian, Christopher D. Manning, and Curtis P. Langlotz. 2018. Learning to summarize radiology findings. In Proceedings of the Ninth International Workshop on Health Text Mining and Info… [cited by applicant]