IP Library › Granted Patent US 12,380,344
Granted Patent B2
US 12,380,344 · App. 17/139,214 · Granted Aug 5, 2025

Generating summary and next actions in real-time for multiple users from interaction records in natural language

Inventors: Yufang Hou (Dublin, IE); Akihiro Kishimoto (Setagaya, JP); Beat Buesser (Ashtown, IE); Bei Chen (Blanchardstown, IE)
Assignee: International Business Machines Corporation
G06N5/04G06F40/174G06F40/186G06F40/279G06N20/00G10L15/22G10L15/26G06F40/205G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,344
App. No.
17/139,214
Granted
Aug 5, 2025
Kind
B2
Abstract

A system receives messaging, video and/or audio input streams including dialogue spoken by users at a group meeting. From these inputs, the system obtains single or multiple interaction records including natural language text memorializing content spoken by each speaker at a meeting, analyzes the content, and identifies single or multiple action item tasks in the interaction records. The system then generates summaries indicating the action item tasks for the users. From the dialogue content, the system further detects whether each action item is addressed, and whether the action item for a user has a solution, or not. The system further detects whether one action item is a precondition for resolving another action item by the user or in conjunction with another user. Using a pre-configured template, the system generates action item summaries, any associated solution, and any relationship or precondition between action items and presents the summary to a user.

Claims (61)

1. A computer-implemented method comprising:

receiving, at one or more processors, multiple interaction records of natural language text relating to a group of users attending a meeting;

analyzing, using one or more classifier models, said text of said multiple interaction records to identify one or more action items from said text, an action item relating to a respective task to be performed;

detecting, from said interaction records, a respective user assigned to perform the respective task;

detecting, from said interaction records, whether each action item from the identified one or more action items is addressed with a solution for carrying out said task or not, said detecting, from said interaction records, whether each action item from the identified one or more action items is addressed with a solution for carrying out said task comprising:

converting a detected action item into a question form;

forming a question-answering task; and

running a trained question-answering model for extracting spans from said interaction records, and determining whether an answer to the question exists from an extracted span;

identifying, using a natural language processing (NLP) entailment model, whether any action item from the identified one or more action items is dependent upon another action item from the identified one or more action items, or whether an action item relating to a task is a precondition to performing another action item task from the identified one or more action items; and

generating a text summary of a solution that addresses each action item for the user based upon the identifying and the determining that the answer to the question exists from the extracted span.

2. The computer-implemented method as claimed in claim 1 , further comprising:

receiving a real-time audio signal stream or audio/video signal stream capturing users spoken content at said meeting; and

converting said real-time audio signal stream or audio/video signal stream into a textual representation to form said multiple interaction records.

3. The computer-implemented method as claimed in claim 1 , further comprising:

receiving, at said one or more processors, one of: an activation signal to initiate said receiving of interaction records, said analyzing, said action item detecting and said summary generating, at a beginning of or during said meeting, or a deactivation signal to terminate said receiving of said interaction records, said analyzing, said action item detecting and said summary generating during said meeting.

4. The computer-implemented method as claimed in claim 1 , wherein said generating a summary of a solution that addresses each action item for the user further comprises:

obtaining a template having pre-defined fields for receiving information related to said action items detected in said interaction records; and

populating said template with action item information including a corresponding user owning that action item, and for each action item: specifying in said template information comprising: the solution to said action item and/or any precondition action item upon which an action item depends.

5. The computer-implemented method as claimed in claim 4 , further comprising:

receiving, from a user, based on said populated template summary information, a feedback information relating to an action item task, its detected corresponding user, a solution to said action item; and

using said feedback information to refine tuning one or more of: said one or more classifier models, the trained question-answering model, and said NLP entailment model.

6. A computer-implemented system comprising:

a memory storage device for storing a computer-readable program, and at least one processor adapted to run said computer-readable program to configure the at least one processor to:

receive multiple interaction records of natural language text relating to a group of users attending a meeting;

analyze, using one or more classifier models, said text of said multiple interaction records to identify one or more action items from said text, an action item relating to a respective task to be performed;

detect, from said interaction records, a respective user assigned to perform the respective task;

detect, from said interaction records, whether each action item from the identified one or more action items is addressed with a solution for carrying out said task or not, wherein to detect, from said interaction records, whether each action item from the identified one or more action items is addressed with a solution for carrying out said task, said at least one processor is further configured to:

convert a detected action item into a question form;

form a question-answering task; and

run a trained question-answering model for extracting spans from said interaction records, and determine whether an answer to the question exists from an extracted span;

identify, using a natural language processing (NLP) entailment model, whether any action item is dependent upon another action item or whether an action item relating to a task is a precondition to performing another action item task; and

generate a text summary of a solution that addresses each action item for the user based upon the identifying and the determining that the answer to the question exists from the extracted span.

7. The computer-implemented system as claimed in claim 6 , wherein said at least one processor is further configured to:

receive a real-time audio signal stream or audio/video signal stream capturing users spoken content at said meeting; and

convert said real-time audio signal stream or audio/video signal stream into a textual representation to form said multiple interaction records.

8. The computer-implemented system as claimed in claim 6 , wherein said at least one processor is further configured to:

receive one of: an activation signal to initiate said receiving of interaction records, said analyzing, said action item detecting and said summary generating, at a beginning of or during said meeting, or a deactivation signal to terminate said receiving of said interaction records, said analyzing, said action item detecting and said summary generating during said meeting.

9. The computer-implemented system as claimed in claim 6 , wherein to generate a summary of a solution that addresses each action item for the user, said at least one processor is further configured to:

obtain a template having pre-defined fields for receiving information related to said action items detected in said interaction records; and

populate said template with action item information including a corresponding user owning that action item, and for each action item: specifying in said template information comprising: the solution to said action item and/or any precondition action item upon which an action item depends.

10. The computer-implemented system as claimed in claim 9 , wherein said at least one processor is further configured to:

receive, from a user, based on said populated template summary information, a feedback information relating to an action item task, its detected corresponding user, a solution to said action item; and

use said feedback information to refine tuning of one or more of: said one or more classifier models, the trained question-answering model, and said NLP entailment model.

11. A computer program product, the computer program product comprising a computer-readable storage medium having a computer-readable program stored therein, wherein the computer-readable program, when executed on a computer including at least one processor, causes the at least one processor to:

receive multiple interaction records of natural language text relating to a group of users attending a meeting;

analyze, using one or more classifier models, said text of said multiple interaction records to identify one or more action items from said text, an action item relating to a respective task to be performed;

detect, from said interaction records, a respective user assigned to perform the respective task;

detect, from said interaction records, whether each action item from the identified one or more action items is addressed with a solution for carrying out said task or not, wherein to detect, from said interaction records, whether each action item from the identified one or more action items is addressed with a solution for carrying out said task, the computer-readable program further configuring the at least one processor to:

convert a detected action item into a question form;

form a question-answering task; and

run a trained question-answering model for extracting spans from said interaction records, and determine whether an answer to the question exists from an extracted span;

identify, using a natural language processing (NLP) entailment model, whether any action item is dependent upon another action item or whether an action item relating to a task is a precondition to performing another action item task; and

generate a text summary of a solution that addresses each action item for the user based upon the identifying and the determining that the answer to the question exists from the extracted span.

12. The computer program product as claimed in claim 11 , wherein the computer-readable program further configures the at least one processor to:

receive one of: an activation signal to initiate said receiving of interaction records, said analyzing, said action item detecting and said summary generating, at a beginning of or during said meeting, or a deactivation signal to terminate said receiving of said interaction records, said analyzing, said action item detecting and said summary generating during said meeting.

13. The computer program product as claimed in claim 11 , wherein to generate a summary of a solution that addresses each action item for the user, the computer-readable program further configures the at least one processor to:

obtain a template having pre-defined fields for receiving information related to said action items detected in said interaction records; and

populate said template with action item information including a corresponding user owning that action item, and for each action item: specifying in said template information comprising: the solution to said action item and/or any precondition action item upon which an action item depends.

14. The computer program product as claimed in claim 13 , wherein the computer-readable program further configures the at least one processor to:

receive, from a user, based on said populated template summary information, a feedback information relating to an action item task, its detected corresponding user, a solution to said action item; and

use said feedback information to refine tuning of one or more of: said one or more classifier models, the trained question-answering model, and said NLP entailment model, use the explanation or reason as a recommendation to prospectively change a condition of an existing policy in anticipation of a future occurring event.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 31, 2020
From: HOU, YUFAN; KISHIMOTO, AKIHIRO; BUESSER, BEAT; CHEN, BEI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054785/0423 →
Continuity (1)
Related Publication 20220207392A1 · Jun 30, 2022
References Cited (82)
US 9094361B2 · Camacho et al. · 2015 [cited by applicant]
US 10102198B2 · Astigarraga et al. · 2018 [cited by applicant]
US 10157618B2 · Higbie · 2018 [cited by examiner]
US 10181326B2 · Raanani et al. · 2019 [cited by applicant]
US 10262654B2 · Hakkani-Tur et al. · 2019 [cited by applicant]
US 10318636B2 · Chatterjee et al. · 2019 [cited by applicant]
US 10373588B2 · Sergott et al. · 2019 [cited by applicant]
US 10984193B1 · Shalev · 2021 [cited by examiner]
US 11200885B1 · Mandal · 2021 [cited by examiner]
US 11520813B2 · Boguraev · 2022 [cited by examiner]
US 20060235694A1 · Cross · 2006 [cited by examiner]
US 20070033086A1 · Christensen et al. · 2007 [cited by applicant]
US 20070112926A1 · Brett et al. · 2007 [cited by applicant]
US 20090259718A1 · O'Sullivan et al. · 2009 [cited by applicant]
US 20130138746A1 · Tardelli · 2013 [cited by examiner]
US 20150100356A1 · Bessler · 2015 [cited by examiner]
US 20150234806A1 · Bhagwan et al. · 2015 [cited by applicant]
US 20150348538A1 · Donaldson · 2015 [cited by applicant]
US 20170032260A1 · Bhagwat · 2017 [cited by examiner]
US 20170134446A1 · Kitada et al. · 2017 [cited by applicant]
US 20170161258A1 · Astigarraga et al. · 2017 [cited by applicant]
US 20170316777A1 · Perez · 2017 [cited by examiner]
US 20180131904A1 · Segal · 2018 [cited by applicant]
US 20180232436A1 · Elson · 2018 [cited by examiner]
US 20180350395A1 · Simko · 2018 [cited by examiner]
US 20190130355A1 · Gupta et al. · 2019 [cited by applicant]
US 20190132265A1 · Nowak-Przygodzki et al. · 2019 [cited by applicant]
US 20190156366A1 · Wilson et al. · 2019 [cited by applicant]
US 20190189117A1 · Kumar · 2019 [cited by applicant]
US 20190294261A1 · Lohse · 2019 [cited by examiner]
US 20190339820A1 · Wu et al. · 2019 [cited by applicant]
US 20190341050A1 · Diamant et al. · 2019 [cited by applicant]
US 20190392066A1 · Kim · 2019 [cited by examiner]
US 20200133745A1 · Dugan · 2020 [cited by examiner]
US 20200184155A1 · Galitsky · 2020 [cited by examiner]
US 20200201496A1 · Wong · 2020 [cited by applicant]
US 20200243095A1 · Adlersberg et al. · 2020 [cited by applicant]
US 20200409519A1 · Faulkner · 2020 [cited by examiner]
US 20210110895A1 · Shriberg · 2021 [cited by examiner]
US 20210174023A1 · Gao · 2021 [cited by examiner]
US 20210342551A1 · Yang · 2021 [cited by examiner]
US 20210342689A1 · Schuff · 2021 [cited by examiner]
US 20210342785A1 · Mann · 2021 [cited by examiner]
US 20220013235A1 · Al-Sinan · 2022 [cited by examiner]
US 20220067285A1 · Sellam · 2022 [cited by examiner]
CN 106685916A · 2017 [cited by applicant]
CN 110741601A · 2020 [cited by applicant]
CN 112075075A · 2020 [cited by applicant]
CN 114691857A · 2022 [cited by applicant]
DE 102020106066A1 · 2021 [cited by examiner]
EP 3364295B1 · 2020 [cited by examiner]
GB 2569440A · 2019 [cited by applicant]
GB 2603842A · 2022 [cited by applicant]
JP 2012053628A · 2012 [cited by applicant]
JP 2019061594A · 2019 [cited by applicant]
JP 2019191276A · 2019 [cited by applicant]
JP 2022105273A · 2022 [cited by applicant]
KR 102623727B1 · 2024 [cited by examiner]
WO WO2019093239A1 · 2019 [cited by examiner]
Punyakanok, Vasin, Dan Roth, and Wen-tau Yih. “Natural language inference via dependency tree mapping: An application to question answering.” (Year: 2004). [cited by examiner]
Herrera, Jesus, Anselmo Penas, and Felisa Verdejo. “Textual entailment recognition based on dependency analysis and WordNet.” Machine Learning Challenges Workshop. Berlin, Heidelberg: Springer Berlin Heidelberg. (Year: … [cited by examiner]
Sammons, Mark, VG Vinod Vydiswaran, and Dan Roth. “Ask not what textual entailment can do for you . . . ” Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics. (Year: 2010). [cited by examiner]
Poliak, Adam. “A survey on recognizing textual entailment as an NLP evaluation.” arXiv preprint arXiv:2010.03061. (Year: 2020). [cited by examiner]
Nielsen, Rodney D., Wayne Ward, and James H. Martin. “Recognizing entailment in intelligent tutoring systems.” Natural Language Engineering 15.4 (2009): 479-501. (Year: 2009). [cited by examiner]
Iftene, Adrian. “Textual Entailment.” Eurolan 2007 Summer School Alexandru Ioan Cuza University of Iaşi (2009): 99. (Year: 2009). [cited by examiner]
Sharma, Nidhi, Richa Sharma, and Kanad K. Biswas. “Recognizing textual entailment using dependency analysis and machine learning.” Proceedings of the 2015 Conference of the North American Chapter of the Association for … [cited by examiner]
He, Luheng, Mike Lewis, and Luke Zettlemoyer. “Question-answer driven semantic role labeling: Using natural language to annotate natural language.” Proceedings of the 2015 conference on empirical methods in natural lang… [cited by examiner]
Examination Report dated Jul. 22, 2022 from related application GB2117761.3. [cited by applicant]
Chen et al.; “Detecting Actionable Items in Meetings by Convolutional Deep Structured Semantic Models”, ASRU IEEE Workshop On, pp. 375-382, Dec. 13-17, 2015. [cited by applicant]
Purver et al.; “Detecting Action Items in Multi-Party Meetings: Annotation and Initial Experiments”, MLMI'06 3rd ACM International Workshop On, pp. 200-211, May 1-4, 2006. [cited by applicant]
Moore et al.; “ActiveNavigator: Toward Real-Time . . . In Design Workspace*”, International Journal of Engineering Education, vol. 34, No. 2(B), pp. 723-733, Jan. 12, 2018. [cited by applicant]
Chen, et al., “AIMU: Actionable Items For Meeting Understanding”, LREC 10th International Conference On, pp. 739-743, Jun. 24, 2016, https://www.microsoft.com/en-us/research/wp-content/uploads/2016/06/LREC16_AIMU.pdf. [cited by applicant]
Whittaker et al.; “Design and Evaluation of Systems to Support Interaction Capture and Retrieval”, Personal and Ubiquitous Computing, vol. 12,No. 3,pp. 197-221,Jan. 2008. [cited by applicant]
Li et al.; “Keep Meeting Summaries on Topic: Abstractive Multi-Modal Meeting Summarization”, Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp. 2190-2196, Florence, Italy, Jul. … [cited by applicant]
Wang et al.; “ ”, Proceedings of the 13th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SIGDIAL), pp. 304-313, Seoul, South Korea, Jul. 5-6, 2012. 2012 Association for Computational Linguistics… [cited by applicant]
Shang et al.; “Unsupervised Abstractive Meeting Summarization with Multi-Sentence Compression and Budgeted Submodular Maximization”, Proceedings of the 56th Annual Meeting of the Association for Computational Linguistic… [cited by applicant]
Purver et al.; “Detecting and Summarizing Action Items in Multi-Party Dialogue”, SIGDIAL 2007. http://www.eecs.qmul.ac.uk/˜mpurver/papers/purver-et-al07sigdial.pdf; https://pdfs.semanticscholar.org/e381/e352f9a5e2673ccf… [cited by applicant]
Office Action dated Jun. 10, 2022 received from the United Kingdom Patent Office in related GB 2117761.3. [cited by applicant]
Response to Office Action dated Jun. 15, 2022 filed in GB 2117761.3. [cited by applicant]
“The Stanford Question Answering Dataset”, SQuAD 2.0, downloaded from the Internet on Oct. 7, 2020, 38 pages, https://rajpurkar.github.io/SQuAD-explorer/. [cited by applicant]
Japan Patent Office, “Notice of Reasons for Refusal” May 20, 2025, 06 Pages, JP Application No. 2021-201458. [cited by applicant]
Keizer et al., “Detecting and Summarizing Action Items in Multi-Party Dialogue”, Proceedings of the 8th SIGdial Workshop on Discourse and Dialogue, Dec. 31, 2007-Sep. 2, 2007, 16 pages. [cited by applicant]