IP Library Granted Patent US 12,323,264
Granted Patent B2
US 12,323,264 · App. 18/538,184 · Granted Jun 3, 2025

Dynamic communication session topic generation

Inventors: Davide Giovanardi (Stanford, CA); Helgi Hilmarsson (Stanford, CA); Stephen Muchovej (Bishop, CA); Mengxiao Qian (Santa Clara, CA); Xiaoli Song (Sunnyvale, CA); Min Xiao-Devins (San Jose, CA)
Assignee: Zoom Communications, Inc.
H04L12/1831G10L15/04G10L15/08G10L15/1815G10L15/26H04L12/1818G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,323,264
App. No.
18/538,184
Granted
Jun 3, 2025
Kind
B2
Abstract

Methods and systems provide for dynamically generated topic segments for a communication session. In one embodiment, the system connects to a communication session with a number of participants; receives a list of topics; receives a transcript of a conversation between the participants produced during the communication session, the transcript including timestamps for a number of utterances associated with speaking participants; for each topic in the list of topics, segments the utterances into one or more topic segments based on the topic; for each of the segments, classifies whether the topic segment is related to the topic, and transmits, to one or more client devices, a list of the topic segments for the communication session.

Claims (41)

1. A method, comprising:

receiving a transcript that includes one or more transcriptions of utterances associated with a list of topics;

for each topic in the list of topics, segmenting the utterances into one or more topic segments based on a determination of an utterance boundary based on a lexical score that is a vector product associated with an adjacent pair of text blocks, wherein a vector contains a number of times a lexical item occurs within a corresponding text block;

for each of the one or more topic segments, determining whether a respective topic segment is related to a topic using one or more zero-shot text classification techniques; and

transmitting, to one or more client devices, a list of topic segments that are related to the topic.

2. The method of claim 1 , further comprising:

generating a title for the respective topic segment based on the topic.

3. The method of claim 1 , wherein the list of topics is received from a client device of the one or more client devices.

4. The method of claim 1 , wherein the segmenting is performed via one or more text tiling techniques.

5. The method of claim 1 , wherein the segmenting comprises:

shifting a window over the utterances in the transcript one word at a time with a pre-specified window size to generate two blocks of utterances per each shift of the window;

at each shift of the window, comparing the two blocks of the utterances to determine whether the two blocks are semantically similar; and

defining a boundary between two topic segments when two blocks of utterances are semantically different.

6. The method of claim 1 , wherein at least a subset of the topic segments overlap with one or more of other topic segments.

7. The method of claim 1 , wherein the one or more topic segments comprise a span of the transcript comprising one or more lines or utterances.

8. The method of claim 1 , further comprising:

classifying whether the one or more topic segments are related to the topic based on one or more language models.

9. A system comprising:

one or more processors configured to:

receive a transcript that includes one or more transcriptions of utterances associated with a list of topics;

for each topic in the list of topics, segment the utterances into one or more topic segments based on a determination of an utterance boundary that is based on a lexical score that is a vector product associated with an adjacent pair of text blocks, wherein a vector contains a number of times a lexical item occurs within a corresponding text block;

for each of the one or more topic segments, determine whether a respective topic segment is related to a topic via one or more zero-shot text classification techniques; and

transmit, to one or more client devices a list of topic segments that are related to the topic.

10. The system of claim 9 , wherein the one or more processors are further configured to classify whether the respective topic segment is related to the topic based on a relatedness threshold.

11. The system of claim 9 , wherein the one or more processors are configured to transmit a starting timestamp and ending timestamp for each of the one or more topic segments.

12. The system of claim 9 , wherein the one or more processors are configured to segment the utterances via linear segmentation for each of the one or more topics.

13. The system of claim 9 , wherein the one or more processors are configured to classify whether the one or more topic segments are related to the topic based on one or more language models.

14. The system of claim 9 , wherein the one or more processors are configured to classify whether the one or more topic segments are related to the topic based on one or more keywords.

15. The system of claim 9 , wherein the one or more processors are further configured to:

transmit, to one or more client devices, a topic summary for a topic, the topic summary comprising one or more utterances from topic segments related to the topic.

16. A non-transitory computer-readable medium comprising instructions that when executed by one or more processors, causes the one or more processors to perform operations comprising:

receiving a transcript that includes one or more transcriptions of utterances associated with a list of topics;

for each topic in the list of topics, segmenting the utterances into one or more topic segments based on a determination of an utterance boundary based on a lexical score that is a vector product associated with an adjacent pair of text blocks, wherein a vector contains a number of times a lexical item occurs within a corresponding text block;

for each of the one or more topic segments, determining whether a respective topic segment is related to a topic using one or more zero-shot text classification techniques; and

transmitting, to one or more client devices, a list of topic segments that are related to the topic.

17. The non-transitory computer-readable medium of claim 16 , wherein the one or more processors are further configured to perform operations comprising:

transmitting, to one or more client devices, one or more utterance results based on a search for the topic within a communication session.

18. The non-transitory computer-readable medium of claim 16 , wherein the one or more processors are further configured to perform operations comprising:

transmitting, to one or more client devices, analytics data related to one or more topics within a communication session.

19. The non-transitory computer-readable medium of claim 16 , wherein the segmenting comprises one or more topic matching techniques.

20. The non-transitory computer-readable medium of claim 16 , wherein the segmenting comprises one or more utterance boundary detection techniques.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2023
From: GIOVANARDI, DAVIDE; HILMARSSON, HELGI; MUCHOVEJ, STEPHEN; QIAN, MENGXIAO; SONG, XIAOLI; XIAO-DEVINS, MIN
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 065857/0295 →
Continuity (2)
Continuation 17734038 · Apr 30, 2022
Related Publication 20240113906A1 · Apr 4, 2024
References Cited (21)
US 10374816B1 · Leblang et al. · 2019 [cited by applicant]
US 20120284016A1 · Tamura et al. · 2012 [cited by applicant]
US 20120296914A1 · Romanov et al. · 2012 [cited by applicant]
US 20120323575A1 · Gibbon et al. · 2012 [cited by applicant]
US 20150332168A1 · Bhagwat et al. · 2015 [cited by applicant]
US 20160072862A1 · Bader-Natal et al. · 2016 [cited by applicant]
US 20160352912A1 · Dhawan et al. · 2016 [cited by applicant]
US 20170062010A1 · Pappu · 2017 [cited by examiner]
US 20170263265A1 · Ashikawa et al. · 2017 [cited by applicant]
US 20180191912A1 · Cartwright et al. · 2018 [cited by applicant]
US 20180239822A1 · Reshef et al. · 2018 [cited by applicant]
US 20200065379A1 · Shires et al. · 2020 [cited by applicant]
US 20200342182A1 · Johnson Premkumar et al. · 2020 [cited by applicant]
US 20210027783A1 · Szymanski et al. · 2021 [cited by applicant]
US 20210210097A1 · Diamant et al. · 2021 [cited by applicant]
US 20220231873A1 · Werfelli et al. · 2022 [cited by applicant]
US 20220383865A1 · McDermid · 2022 [cited by examiner]
International Search Report and Written Opinion mailed on Jul. 19, 2023 in corresponding PCT Application No. PCT/US2023/019594. [cited by applicant]
Marti A Hearst: “TextTiling”, Computational Linguistics, M I T Press, US, vol. 23, No. 1, Mar. 1, 1997 (Mar. 1, 1997), pp. 33-64, XP058176768, ISSN: 0891-2017 section 5.2.1 in particular. [cited by applicant]
Bernadette Sharp et al: “Text segmentation of spoken meeting transcripts”, International Journal of Speech Technology, Kluwer Academic Publishers, BO, vol. 11, No. 3-4, Nov. 12, 2009 (Nov. 12, 2009), pp. 157-165, XP0197… [cited by applicant]
Gupta Vidhi et al: “A Comparative Study of the Performance of Unsupervised Text Segmentation Techniques on Dialogue Transcripts”, 2020 Systems and Information Engineering Design Symposium (SIEDS), IEEE, Apr. 24, 2020 (A… [cited by applicant]