IP Library › Granted Patent US 12,547,835
Granted Patent B2
US 12,547,835 · App. 18/065,758 · Granted Feb 10, 2026

Automatic extraction of semantically similar question topics

Inventors: Marek Šuppa (Bratislava, SK); Daniela Jašš (Bratislava, SK); Katarína Kozáková (Bratislava, SK); Daniel Skala (Groningen, NL)
Assignee: CISCO TECHNOLOGY, INC.
G06F40/30G06F16/35G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,547,835
App. No.
18/065,758
Granted
Feb 10, 2026
Kind
B2
Abstract

A method, computer system, and computer program product are provided for automatically extracting semantically-similar question topics. A set of documents, wherein each document includes a plurality of words. One or more clusters of documents are identified in the set of documents based on a presence of common words in the documents of the one or more clusters. The one or more clusters are adjusted based on semantic similarity by adding or removing one or more documents from the one or more clusters. A topic is extracted from each adjusted cluster of documents.

Claims (50)

1 . A computer-implemented method comprising:

obtaining, via at least one processor, a set of documents submitted by participants from a communication session conducted over a network, wherein each document is in electronic form and includes a plurality of words;

identifying, via the at least one processor, one or more clusters of documents in the set of documents based on a presence of a threshold number of common words in documents of the one or more clusters of documents;

adjusting the one or more clusters of documents, via the at least one processor, based on semantic similarity by adding or removing one or more documents from a cluster based on the one or more documents being within a threshold distance from the cluster, wherein the threshold distance is based on a plurality of adjustable values;

extracting, via the at least one processor, a topic from each adjusted cluster of documents;

determining, via the at least one processor, a document from each adjusted cluster of documents containing the topic;

sending, via the at least one processor, the document from each adjusted cluster of documents to a participant of the communication session to produce a response pertaining to the topic from each adjusted cluster of documents, wherein sending the document from each adjusted cluster enables the response to be produced for remaining documents in each adjusted cluster rather than sending individual documents of each adjusted cluster;

presenting, via the at least one processor, the response to the participants in the communication session, and

adjusting, via the at least one processor, the threshold distance for adjustment of document clusters by modifying the plurality of adjustable values based on user feedback for the topic from each adjusted cluster of documents.

2 . The computer-implemented method of claim 1 , wherein each document is a question submitted by a participant of a question-and-answer session.

3 . The computer-implemented method of claim 1 , wherein extracting the topic from a particular cluster includes identifying a particular document having a longest sequence of stemmed words in common with a majority of other documents of the particular cluster.

4 . The computer-implemented method of claim 1 , wherein adding or removing the one or more documents from a cluster comprises generating a representation of the set of documents in a vector space, identifying a centroid for the cluster in the vector space, and adding or removing the one or more documents based on the one or more documents being within the threshold distance from the centroid.

5 . The computer-implemented method of claim 1 , further comprising:

obtaining the user feedback in response to presenting the topic to one or more users.

6 . The computer-implemented method of claim 1 , further comprising presenting the topic to one or more users via a display.

7 . The computer-implemented method of claim 1 , wherein the plurality of words in each document are tokenized prior to identifying the one or more clusters of documents.

8 . A computer system comprising:

one or more computer processors;

one or more computer readable storage media; and

program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising instructions to perform operations including:

obtaining a set of documents submitted by participants from communication session conducted over a network, wherein each document is in electronic form and includes a plurality of words;

identifying one or more clusters of documents in the set of documents based on a presence of a threshold number of common words in documents of the one or more clusters of documents;

adjusting the one or more clusters of documents based on semantic similarity by adding or removing one or more documents from a cluster based on the one or more documents being within a threshold distance from the cluster, wherein the threshold distance is based on a plurality of adjustable values;

extracting a topic from each adjusted cluster of documents;

determining a document from each adjusted cluster of documents containing the topic:

sending the document from each adjusted cluster of documents to a participant of the communication session to produce a response pertaining to the topic from each adjusted cluster of documents, wherein sending the document from each adjusted cluster enables the response to be produced for remaining documents in each adjusted cluster rather than sending individual documents of each adjusted cluster:

presenting the response to the participants in the communication session, and

adjusting the threshold distance for adjustment of document clusters by modifying the plurality of adjustable values based on user feedback for the topic from each adjusted cluster of documents.

9 . The computer system of claim 8 , wherein each document is a question submitted by a participant of a question-and-answer session.

10 . The computer system of claim 8 , wherein the am instructions to extract the topic from a particular cluster include program instructions to perform further operations including a particular document having a longest sequence of stemmed words in common with a majority of other documents of the particular cluster.

11 . The computer system of claim 8 , wherein adding or removing the one or more documents from a cluster comprises generating a representation of the set of documents in a vector space, identifying a centroid for the cluster, and adding or removing the one or more documents based on the one or more documents being within the threshold distance from the centroid.

12 . The computer system of claim 8 , wherein the program instructions further comprise program instructions to perform further operations including:

obtaining the user feedback in response to presenting the topic to one or more users.

13 . The computer system of claim 8 , further comprising program instructions to perform further operations including presenting the topic to one or more users via a display.

14 . The computer system of claim 8 , wherein the plurality of words in each document are tokenized prior to identifying the one or more clusters of documents.

15 . One or more non-transitory computer readable storage media encoded with program instructions that when executed by a computer, cause the computer to perform operations including:

obtaining a set of documents submitted by participants from a communication session conducted over a network, wherein each document is in electronic form and includes a plurality of words;

identifying one or more clusters of documents in the set of documents based on a presence of a threshold number of common words in documents of the one or more clusters of documents;

adjusting the one or more clusters of documents based on semantic similarity by adding or removing one or more documents from a cluster based on the one or more documents being within a threshold distance from the cluster, wherein the threshold distance is based on a plurality of adjustable values;

extracting a topic from each adjusted cluster of documents;

determining document from each adjusted cluster of documents containing the topic;

sending the document from each adjusted cluster of documents to a participant of the communication session to produce a response pertaining to the topic from each adjusted cluster of documents, wherein sending the document from each adjusted cluster enables the response to be produced for remaining documents in each adjusted cluster rather than sending individual documents of each adjusted cluster:

presenting the response to the participants in the communication session; and

adjusting the threshold distance for adjustment of document clusters by modifying the plurality of adjustable values based on user feedback for the topic from each adjusted cluster of documents.

16 . The one or more non-transitory computer readable storage media of claim 15 , wherein each document is a question submitted by a participant of a question-and-answer session.

17 . The computer one or more non-transitory computer readable storage media of claim 15 , wherein the program instructions to extract the topic from a particular cluster further cause the computer to perform further operations including identifying a particular document having a longest sequence of stemmed words in common with a majority of other documents of the particular cluster.

18 . The one or more non-transitory computer readable storage media of claim 15 , wherein adding or removing the one or more documents from a cluster comprises generating a representation of the set of documents in a vector space, identifying a centroid for the cluster, and adding or removing the one or more documents based on the one or more documents being within the threshold distance from the centroid.

19 . The one or more non-transitory computer readable storage media of claim 5 , wherein the program instructions further cause the computer to perform further operations including:

obtaining the user feedback in response to presenting the topic to one or more users.

20 . The one or more non-transitory computer readable storage media of claim 15 , wherein the program instructions further cause the computer to perform further operations including presenting the topic to one or more users via a display.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: ¿UPPA, MAREK; JA¿¿, DANIELA; KOZÁKOVÁ, KATARÍNA; SKALA, DANIEL
To: CISCO TECHNOLOGY, INC.
Reel/Frame 062087/0079 →
Continuity (1)
Related Publication 20240202448A1 · Jun 20, 2024
References Cited (21)
US 8713021B2 · Bellegarda · 2014 [cited by examiner]
US 20130085745A1 · Koister et al. · 2013 [cited by applicant]
US 20160232221A1 · McCloskey · 2016 [cited by examiner]
US 20160314200A1 · Markman et al. · 2016 [cited by applicant]
US 20160371277A1 · Allen et al. · 2016 [cited by applicant]
US 20180329882A1 · Bennett · 2018 [cited by examiner]
US 20190005127A1 · Alkov et al. · 2019 [cited by applicant]
US 20190079938A1 · Agrawal · 2019 [cited by examiner]
US 20210365524A9 · Anderson · 2021 [cited by examiner]
US 20220004715A1 · Patel · 2022 [cited by examiner]
“Semantic based Document Clustering: A Detailed Review”; International Journal of Computer Applications (0975-8887) vol. 52—No. 5, Aug. 2012, (Shah et al.) (Year: 2012). [cited by examiner]
“Efficient Phrase-Based Document Similarity for Clustering”; IEEE Transactions on Knowledge and Data Engineering, vol. 20, No. 9, Sep. 2008, (Chim et al.) (Year: 2008). [cited by examiner]
“Centroid-based document classification: Analysis and experimental results.” European conference on principles of data mining and knowledge discovery. Berlin, Heidelberg: Springer Berlin Heidelberg, 2000, (Han et al.) (… [cited by examiner]
Velden, et al., “Comparison of Topic Extraction Approaches and Their Results,” https://pure.knaw.nl/ws/portalfiles/portal/4247535/preprint_Velden_comparison.pdf, Feb. 2017, 66 pages. [cited by applicant]
Dong, et al., “Topic Extraction from Online Reviews for Classification and Recommendation,” Proceedings of the Twenty-Third International Joint Conference on Artificial Intelligence (IJCAI 13), https://researchrepositor… [cited by applicant]
Mihalcea, et al., “TextRank: Bringing Order into Texts,” Proceedings of the 2004 Conference on Empirical Methods In Natural Language Processing, https://aclanthology.org/W04-3252.pdf, Jul. 2004, 8 pages. [cited by applicant]
Bougouin, et al., “TopicRank: Graph-Based Topic Ranking for Keyphrase Extraction,” International Joint Conference on Natural Language Processing (IJCNLP), https://hal.archives-ouvertes.fr/hal-00917969/document, Oct. 201… [cited by applicant]
Slido, “Teacher Convention,” https://app.sli.do/event/9EiTA1Nqz1h3sSU4PetrvT/live/questions, Jun. 2022, 2 pages. [cited by applicant]
Suppa, et al., “Cost-effective Deployment of BERT Models in Serverless Environment,” Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech… [cited by applicant]
Meng, et al., “Detecting topics and overlapping communities in Question and Answer sites,” Social Network Analysis and Mining, vol. 5, Dec. 2015, 27 pages. [cited by applicant]
Post Image, “Question-topics-bid,” retrieved from https://postimg.cc/NyNPCK6R on Jul. 17, 2025, 1 page. [cited by applicant]