IP Library Granted Patent US 12,647,532
Granted Patent B2
US 12,647,532 · App. 18/441,548 · Granted Jun 2, 2026

Online meeting summarization for videoconferencing

Inventors: Davide Giovanardi (San Jose, CA); Bilung Lee (Irvine, CA); Felix Schneider (Baden-Wurttemberg, DE); Marco Turchi (Trento, IT); Alexander Waibel (Sammamish, WA); Yun Zhang (Pittsburgh, PA)
Assignee: Zoom Communications, Inc.
H04N7/155H04N7/152
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,647,532
App. No.
18/441,548
Granted
Jun 2, 2026
Kind
B2
Abstract

Systems and methods for generating online meeting summaries for videoconferencing are provided. For example, a computing device establishes a video conference for a plurality of participants. While the video conference is in progress, the computing device receives a first portion of a transcript of the video conference and generates a first meeting summary based on the first portion of the transcript. The computing device causes the first meeting summary to be presented in a user interface accessible by a client computing device associated with at least one of the plurality of participants. The computing device receives a second portion of the transcript of the video conference and generates a second meeting summary based on the first portion and the second portion of the transcript. The computing device causes the second meeting summary to be presented in the user interface.

Claims (79)

1 . A method performed by one or more computing devices, the method comprising:

establishing a video conference for a plurality of participants;

while the video conference is in progress,

receiving a first portion of a transcript of the video conference;

generating a first meeting summary based on the first portion of the transcript comprising:

determining different sets of chunks based on the first portion of the transcript, each set of chunks comprising a different subset of one or more chunks of the first portion of the transcript,

generating a plurality of candidate meeting summaries comprising applying a meeting summarization model to the different sets of chunks, and

selecting a candidate meeting summary from the plurality of candidate meeting summaries that has a highest quality score as the first meeting summary,

causing the first meeting summary to be presented in a user interface accessible by a client computing device associated with at least one of the plurality of participants;

receiving a second portion of the transcript of the video conference;

generating a second meeting summary based on the first portion and the second portion of the transcript; and

causing the second meeting summary to be presented in the user interface.

2 . The method of claim 1 , wherein a quality score comprises one or more of a confidence score output by the meeting summarization model or a similarity metric calculated based on embeddings of a corresponding candidate meeting summary and the corresponding set of chunks.

3 . The method of claim 1 , wherein generating the plurality of candidate meeting summaries comprises:

generating a first candidate meeting summary by providing a first chunk into the meeting summarization model;

generating a second candidate meeting summary by providing the first chunk and a second chunk into the meeting summary model; and

generating a third candidate meeting summary by providing the one or more chunks into the meeting summary model.

4 . The method of claim 1 , wherein generating the second meeting summary based on the first portion and the second portion of the transcript comprises:

generating a second plurality of candidate meeting summaries by applying the meeting summarization model to subsets of chunks in the second portion of the transcript and one or more of the chunks in the first portion of the transcript that are not used in generating the selected candidate meeting summary.

5 . The method of claim 1 , wherein the second meeting summary comprises the first meeting summary and an additional meeting summary, and wherein:

generating the first meeting summary comprises applying the meeting summarization model to the first portion of the transcript;

generating the second meeting summary based on the first portion and the second portion of the transcript comprises applying the meeting summary model to a first concatenation of the first portion of the transcript and the second portion of the transcript by forcing the meeting summary model to output the first meeting summary before the additional meeting summary in the second meeting summary is output; and

presenting the second meeting summary in the user interface comprises presenting the additional meeting summary.

6 . The method of claim 5 , further comprising:

receiving a third portion of the transcript of the video conference;

generating a third meeting summary based on the second portion and the third portion of the transcript by applying the meeting summarization model to a concatenation of the second portion of the transcript and the third portion of the transcript by forcing the meeting summary model to output the second meeting summary before other meeting summary in the third meeting summary is output; and

presenting the other meeting summary in the third meeting summary in the user interface.

7 . The method of claim 5 , further comprising:

receiving an edit to the first meeting summary, wherein generating the second meeting summary based on the first portion and the second portion of the transcript comprises applying the meeting summarization model to a second concatenation of the first portion of the transcript and the second portion of the transcript by forcing the meeting summary model to output the edited first meeting summary before the additional meeting summary in the second meeting summary is output.

8 . A computing device, comprising:

a non-transitory computer-readable medium; and

a processor communicatively coupled to the non-transitory computer-readable medium, the processor configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

establish a video conference for a plurality of participants;

while the video conference is in progress,

receive a first portion of a transcript of the video conference;

generate a first meeting summary based on the first portion of the transcript, comprising executing processor-executable instructions stored in the non-transitory computer-readable medium to:

determine different sets of chunks based on the first portion of the transcript, each set of chunks comprising a different subset of one or more chunks of the first portion of the transcript,

apply a meeting summarization model to the different sets of chunks to generate a plurality of candidate meeting summaries, and

select a candidate meeting summary from the plurality of candidate meeting summaries that has a highest quality score as the first meeting summary;

cause the first meeting summary to be presented in a user interface accessible by a client computing device associated with at least one of the plurality of participants;

receive a second portion of the transcript of the video conference;

generate a second meeting summary based on the first portion and the second portion of the transcript; and

cause the second meeting summary to be presented in the user interface.

9 . The computing device of claim 8 , wherein a quality score comprises one or more of a confidence score output by the meeting summarization model or a similarity metric calculated based on embeddings of a corresponding candidate meeting summary and the corresponding set of chunks.

10 . The computing device of claim 8 , wherein generating the plurality of candidate meeting summaries comprises:

generating a first candidate meeting summary by providing a first chunk into the meeting summarization model;

generating a second candidate meeting summary by providing the first chunk and a second chunk into the meeting summary model; and

generating a third candidate meeting summary by providing the one or more chunks into the meeting summary model.

11 . The computing device of claim 8 , wherein generating the second meeting summary based on the first portion and the second portion of the transcript comprises:

generating a second plurality of candidate meeting summaries by applying the meeting summarization model to subsets of chunks in the second portion of the transcript and one or more of the chunks in the first portion of the transcript that are not used in generating the selected candidate meeting summary.

12 . The computing device of claim 8 , wherein the second meeting summary comprises the first meeting summary and an additional meeting summary, and wherein:

generating the first meeting summary comprises applying the meeting summarization model to the first portion of the transcript;

generating the second meeting summary based on the first portion and the second portion of the transcript comprises applying the meeting summary model to a concatenation of the first portion of the transcript and the second portion of the transcript by forcing the meeting summary model to output the first meeting summary before the additional meeting summary in the second meeting summary is output; and

presenting the second meeting summary in the user interface comprises presenting the additional meeting summary.

13 . The computing device of claim 12 , wherein the processor is configured to execute the processor-executable instructions stored in the non-transitory computer-readable medium to further:

receive a third portion of the transcript of the video conference;

generate a third meeting summary based on the second portion and the third portion of the transcript by applying the meeting summarization model to a concatenation of the second portion of the transcript and the third portion of the transcript by forcing the meeting summary model to output the second meeting summary before other meeting summary in the third meeting summary is output; and

present the other meeting summary in the third meeting summary in the user interface.

14 . A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:

establish a video conference for a plurality of participants;

while the video conference is in progress,

receive a first portion of a transcript of the video conference;

generate a first meeting summary based on the first portion of the transcript, comprising processor-executable instructions stored in the non-transitory computer-readable medium to cause the one or more processors to:

determine different sets of chunks based on the first portion of the transcript, each set of chunks comprising a different subset of one or more chunks of the first portion of the transcript,

apply a meeting summarization model to the different sets of chunks to generate a plurality of candidate meeting summaries, and

select a candidate meeting summary from the plurality of candidate meeting summaries that has a highest quality score as the first meeting summary;

cause the first meeting summary to be presented in a user interface accessible by a client computing device associated with at least one of the plurality of participants;

receive a second portion of the transcript of the video conference;

generate a second meeting summary based on the first portion and the second portion of the transcript; and

cause the second meeting summary to be presented in the user interface.

15 . The non-transitory computer-readable medium of claim 14 , wherein a quality score comprises one or more of a confidence score output by the meeting summarization model or a similarity metric calculated based on embeddings of a corresponding candidate meeting summary and the corresponding set of chunks.

16 . The non-transitory computer-readable medium of claim 14 , wherein generating the plurality of candidate meeting summaries comprises:

generating a first candidate meeting summary by providing a first chunk into the meeting summarization model;

generating a second candidate meeting summary by providing the first chunk and a second chunk into the meeting summary model; and

generating a third candidate meeting summary by providing the one or more chunks into the meeting summary model.

17 . The non-transitory computer-readable medium of claim 14 , wherein the second meeting summary comprises the first meeting summary and an additional meeting summary, and wherein:

generating the first meeting summary comprises applying the meeting summarization model to the first portion of the transcript;

generating the second meeting summary based on the first portion and the second portion of the transcript comprises applying the meeting summary model to a concatenation of the first portion of the transcript and the second portion of the transcript by forcing the meeting summary model to output the first meeting summary before the additional meeting summary in the second meeting summary is output; and

presenting the second meeting summary in the user interface comprises presenting the additional meeting summary.

Assignments (2)
CHANGE OF NAME Recorded Apr 24, 2026
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 075471/0811 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2024
From: GIOVANARDI, DAVIDE; LEE, BILUNG; SCHNEIDER, FELIX; TURCHI, MARCO; WAIBEL, ALEXANDER; ZHANG, YUN
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 066529/0684 →
Continuity (1)
Related Publication 20250260790A1 · Aug 14, 2025
References Cited (42)
US 20220353470A1 · Huang et al. · 2022 [cited by applicant]
US 20230096782A1 · Crumley et al. · 2023 [cited by applicant]
US 20230153530A1 · Savkov · 2023 [cited by examiner]
US 20240086461A1 · Varakin · 2024 [cited by examiner]
US 20240176960A1 · Maurer · 2024 [cited by examiner]
US 20240340193A1 · Zhu · 2024 [cited by examiner]
US 20250005289A1 · Wang · 2025 [cited by examiner]
US 20250140246A1 · Lee · 2025 [cited by examiner]
US 20250232141A1 · Bahirwani · 2025 [cited by examiner]
“GPT-4 Technical Report”, Available online at: https://cdn.openai.com/papers/gpt-4.pdf, Dec. 19, 2023, pp. 1-100. [cited by applicant]
Arivazhagan, et al., “Re-Translation Strategies for Long Form, Simultaneous, Spoken Language Translation”, Institute of Electrical and Electronics Engineers International Conference on Acoustics, Speech and Signal Proce… [cited by applicant]
Asi, et al., “An End-to-End Dialogue Summarization System for Sales Calls”, North American Chapter of the Association for Computational Linguistics, Available online at: https://aclanthology.org/2022.naacl-industry.6.pd… [cited by applicant]
Fabbri, et al., “ConvoSumm: Conversation Summarization Benchmark and Improved Abstractive Summarization with Argument Mining”, Association for Computational Linguistics, vol. 1, Jun. 1, 2021, 15 pages. [cited by applicant]
Fabbri, et al., “SummEval: Re-evaluating Summarization Evaluation”, Transactions of the Association for Computational Linguistics, vol. 9, Apr. 26, 2021, pp. 391-409. [cited by applicant]
Garg, et al., “Clusterrank: A Graph Based Method for Meeting Summarization”, IDIAP Research Report, Jun. 2009, 6 pages. [cited by applicant]
Ghosal, et al., “Overview of the Second Shared Task on Automatic Minuting (AutoMin) at INLG 2023”, Proceedings of the 16th International Natural Language Generation Conference: Generation Challenges, Available online at… [cited by applicant]
Gliwa, et al., “SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization”, Association for Computational Linguistics, Proceedings of the 2nd Workshop on New Frontiers in Summarization, Nov. 29, 20… [cited by applicant]
Janin, et al., “The ICSI Meeting Corpus”, Available online at: https://www1.icsi.berkeley.edu/ftp/global/pub/speech/papers/icassp03-janin.pdf, Apr. 6-10, 2003, 4 pages. [cited by applicant]
Lewis, et al., “BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension”, Annual Meeting of the Association for Computational Linguistics, Oct. 29, 2019, 10 page… [cited by applicant]
Lin, “ROUGE: A Package for Automatic Evaluation of Summaries”, Text Summarization Branches Out, Association for Computational Linguistics, Jul. 25, 2004, 8 pages. [cited by applicant]
Ma, et al., “SIMULEVAL: An Evaluation Toolkit for Simultaneous Translation”, Conference on Empirical Methods in Natural Language Processing, Jul. 31, 2020, 7 pages. [cited by applicant]
Murray, et al., “Generating and Validating Abstracts of Meeting Conversations: a User Study”, Proceedings of the 6th International Natural Language Generation Conference, Jun. 2010, 9 pages. [cited by applicant]
Nedoluzhko, et al., “ELITR Minuting Corpus: A Novel Dataset for Automatic Minuting from Multi-Party Meetings in English and Czech”, Proceedings of the 13th Conference on Language Resources and Evaluation (LREC 2022), Ju… [cited by applicant]
Niehues, et al., “Low-Latency Neural Speech Translation”, Available online at: https://www.researchgate.net/publication/326799383_Low-Latency_Neural_Speech_Translation, Aug. 1, 2018, 5 pages. [cited by applicant]
Olariu, “Efficient Online Summarization of Microblogging Streams”, Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, Apr. 26-30, 2014, pp. 236-240. [cited by applicant]
Papi, et al., “Over-Generation Cannot Be Rewarded: Length-Adaptive Average Lagging for Simultaneous Speech Translation”, Proceedings of the Third Workshop on Automatic Simultaneous Translation, Jun. 20, 2022, 6 pages. [cited by applicant]
Pham, et al., “Select, Prompt, Filter: Distilling Large Language Models for Summarizing Conversations”, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Dec. 6-10, 2023, pp. 12257-… [cited by applicant]
Schneider, et al., “Team Zoom@ AutoMin 2023: Utilizing Topic Segmentation and LLM Data Augmentation for Long-Form Meeting Summarization”, Proceedings of the 16th International Natural Language Generation Conference: Gen… [cited by applicant]
Sequiera, et al., “Overview of the TREC 2018 Real-Time Summarization Track”, Available online at: https://cs.uwaterloo.ca/˜jimmylin/publications/Sequiera_etal_TREC2018.pdf, 2018, 5 pages. [cited by applicant]
Tixier, et al., “Combining Graph Degeneracy and Submodularity for Unsupervised Extractive Summarization”, Proceedings of the Workshop on New Frontiers in Summarization, Sep. 7, 2017, pp. 48-58. [cited by applicant]
Tur, et al., “The Calo Meeting Speech Recognition and Understanding System”, Conference: Spoken Language Technology Workshop, Dec. 2008, pp. 69-72. [cited by applicant]
Vaswani, et al., “Attention Is All You Need”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Jun. 2017, 11 pages. [cited by applicant]
Waibel, et al., “Meeting Browser: Tracking and Summarizing Meetings”, Available online at: https://isl.anthropomatik.kit.edu/downloads/_TRACKING.pdf, Aug. 2000, 6 pages. [cited by applicant]
Zechner, “Automatic Summarization of Open-Domain Multiparty Dialogues Diverse Genres”, Association for Computational Linguistics, vol. 28, No. 4, Dec. 1, 2002, pp. 447-485. [cited by applicant]
Zhang, et al., “SummN: A Multi-Stage Summarization Framework for Long Input Dialogues and Documents”, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, vol. 1, Apr. 13, 2022, 13 pa… [cited by applicant]
Zhong, et al., “DIALOGLM: Pre-trained Model for Long Dialogue Understanding and Summarization”, The Thirty-Sixth AAAI Conference on Artificial Intelligence, vol. 44, No. 4, Sep. 6, 2021, pp. 11765-11773. [cited by applicant]
Zhong, et al., “Extractive Summarization as Text Matching”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Apr. 19, 2020, 12 pages. [cited by applicant]
Zhong, et al., “QMSum: A New Benchmark for Query-based Multi-domain Meeting Summarization”, North American Chapter of the Association for Computational Linguistics, Apr. 13, 2021, 17 pages. [cited by applicant]
Zhou, et al., “Hierarchical Recurrent Aggregative Generation for Few-Shot NLG”, Findings of the Association for Computational Linguistics, May 22-27, 2022, pp. 2167-2181. [cited by applicant]
Zhu, et al., “A Hierarchical Network for Abstractive Meeting Summarization with Cross-Domain Pre training”, Findings of the Association for Computational Linguistics: EMNLP 2020, Sep. 20, 2020, 14 pages. [cited by applicant]
McCowan, et al., “The AMI meeting corpus”, Int'l. Conf. on Methods and Techniques in Behavioral Research, 2005, 4 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/US2025/013733 mailed Apr. 14, 2025. [cited by applicant]