IP Library › Granted Patent US 12,524,618
Granted Patent B2
US 12,524,618 · App. 17/933,391 · Granted Jan 13, 2026

Database systems and methods of representing conversations

Inventors: Lidiya Murakhovs'ka (San Francisco, CA); Chien-Sheng Wu (San Francisco, CA); Yixin Mao (San Francisco, CA)
G06F40/35G06F16/3329G06F16/345G06F16/358G06F16/383G10L15/083G10L15/1815G10L15/20G10L15/22G10L15/26H04L51/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,618
App. No.
17/933,391
Granted
Jan 13, 2026
Kind
B2
Abstract

Database systems and methods are provided for assigning structural metadata to records and creating automations using the structural metadata. One method of assigning structural metadata to a record associated with a conversation involves obtaining a plurality of utterances associated with the conversation, the plurality of utterances including at least a first set of utterances by a first actor and a second set of utterances corresponding to a second actor, obtaining a summarization of semantic content of the conversation based at least in part on an initial subset of the plurality of utterances using a summarization model, identifying, from among the first set of utterances corresponding to the first actor, a representative utterance that is closest to the summarization of the semantic content of the conversation, and automatically updating the record associated with the conversation at a database system to include metadata identifying the representative utterance by the first actor.

Claims (54)

1 . A method of assigning structural metadata to a record associated with a conversation at a database system to assign the conversation to a cluster group of one or more semantically similar conversations at the database system from among a plurality of different cluster groups of semantically similar conversations at the database system that are distinct from one another, the method comprising:

obtaining a plurality of utterances associated with the conversation, the plurality of utterances including at least a first set of utterances by a first actor and a second set of utterances corresponding to a second actor;

automatically generating conversation summary text comprising a summarization of semantic content of the conversation based at least in part on an initial subset of the plurality of utterances using a summarization model;

converting the conversation summary text to a numerical conversation summary vector represents semantic meaning of the conversation summary text in a numerical form using an encoder model configured to convert input text into a numerical vector that correlates to semantic meaning of the input text;

identifying, from among the first set of utterances, a representative utterance comprising an individual utterance corresponding to the first actor indicative of the semantic content of the conversation having a corresponding numerical representation using the encoder model that is closest to the numerical conversation summary vector, the representative utterance comprising a verb and noun indicative of an intent of the first actor that initiated the conversation;

automatically updating, at the database system, a field of the record associated with the conversation at the database system to include metadata identifying the representative utterance by the first actor;

automatically assigning, at the database system, the record associated with the conversation to an assigned cluster group of one or more semantically similar conversations among the plurality of different cluster groups of semantically similar conversations based at least in part on the corresponding numerical representation of the representative utterance; and

providing, by the database system, a graphical user interface (GUI) display comprising a graphical representation of the representative utterance in association with the assigned cluster group of one or more semantically similar conversations.

2 . The method of claim 1 , wherein identifying the representative utterance comprises:

for each utterance by the first actor in the initial subset of the plurality of utterances, converting the respective utterance by the first actor into a respective corresponding numerical representation using the encoder model; and

identifying the representative utterance as the respective utterance by the first actor in the initial subset of the plurality of utterances with the respective corresponding numerical representation closest to a numerical vector representation of the conversation summary text.

3 . The method of claim 2 , wherein:

identifying the representative utterance comprises identifying the respective utterance by the first actor in the initial subset of the plurality of utterances with the respective corresponding numerical representation having greatest cosine similarity with respect to the numerical conversation summary vector.

4 . The method of claim 1 , wherein obtaining the plurality of utterances comprises obtaining a transcript of audio of the conversation, the transcript including the plurality of utterances.

5 . The method of claim 4 , further comprising removing one or more semantically insignificant terms from the transcript, resulting in an augmented transcript, wherein automatically generating the conversation summary text comprises:

selecting the initial subset of the plurality of utterances from the augmented transcript; and

inputting the initial subset of the plurality of utterances to the summarization model to obtain the conversation summary text.

6 . The method of claim 5 , wherein removing the one or more semantically insignificant terms from the transcript comprises removing one or more utterances from the transcript using a reference vocabulary of terms or phrases for removal based on cosine similarity between a respective utterance of the one or more utterances and one of a reference set of numerical vectors corresponding to the reference vocabulary.

7 . The method of claim 6 , wherein removing the one or more utterances comprises removing one or more standard phrases by the second actor from the transcript.

8 . The method of claim 6 , wherein removing the one or more utterances comprises removing one or more noise utterances from the transcript.

9 . The method of claim 5 , further comprising concatenating successive utterances by the first actor after removing the one or more semantically insignificant terms from the transcript prior to selecting the initial subset of the plurality of utterances from the augmented transcript.

10 . The method of claim 5 , wherein removing the one or more semantically insignificant terms from the transcript comprises removing personally identifiable information from the transcript.

11 . The method of claim 1 , further comprising verifying the plurality of utterances includes a start of the conversation prior to automatically updating the record associated with the conversation to include the metadata identifying the representative utterance by the first actor.

12 . The method of claim 1 , further comprising inputting the representative utterance to the summarization model to obtain a shortened version of the representative utterance when a length of the representative utterance is greater than a threshold, wherein automatically updating the record comprises automatically updating the record associated with the conversation to include the metadata identifying the shortened version of the representative utterance by the first actor.

13 . The method of claim 1 , wherein automatically updating the record associated with the conversation at the database system comprises assigning the representative utterance as a contact reason associated with the first actor initiating the conversation.

14 . At least one non-transitory machine-readable storage medium that provides instructions that, when executed by at least one processor, are configurable to cause the at least one processor to perform operations comprising:

obtaining a plurality of utterances associated with a conversation, the plurality of utterances including at least a first set of utterances by a first actor and a second set of utterances corresponding to a second actor;

automatically generating conversation summary text comprising a summarization of semantic content of the conversation based at least in part on an initial subset of the plurality of utterances using a summarization model;

converting the conversation summary text to a numerical conversation summary vector represents semantic meaning of the conversation summary text in a numerical form using an encoder model configured to convert input text into a numerical vector that correlates to semantic meaning of the input text;

identifying, from among the first set of utterances, a representative utterance comprising an individual utterance corresponding to the first actor indicative of the semantic content of the conversation having a corresponding numerical representation using the encoder model that is closest to the numerical conversation summary vector using the summarization model, the representative utterance comprising a verb and noun indicative of an intent of the first actor that initiated the conversation;

automatically updating a field of a record associated with the conversation at a database system to include metadata identifying the representative utterance by the first actor;

automatically assigning, at the database system, the record associated with the conversation to an assigned cluster group of one or more semantically similar conversations among a plurality of different cluster groups of semantically similar conversations based at least in part on the corresponding numerical representation of the representative utterance; and

providing, by the database system, a graphical user interface (GUI) display comprising a graphical representation of the representative utterance in association with the assigned cluster group of one or more semantically similar conversations.

15 . The at least one non-transitory machine-readable storage medium of claim 14 , wherein the instructions are configurable to cause the at least one processor to:

for each utterance by the first actor in the initial subset of the plurality of utterances, convert the respective utterance by the first actor into a respective corresponding numerical representation using the encoder model; and

identify the representative utterance as the respective utterance by the first actor in the initial subset of the plurality of utterances with the respective corresponding numerical representation closest to the numerical conversation summary vector.

16 . The at least one non-transitory machine-readable storage medium of claim 14 , wherein the instructions are configurable to cause the at least one processor to:

obtain a transcript of audio of the conversation, the transcript including the plurality of utterances; and

remove one or more semantically insignificant terms from the transcript, resulting in an augmented transcript, wherein automatically generating the conversation summary text comprises:

selecting the initial subset of the plurality of utterances from the augmented transcript; and

inputting the initial subset of the plurality of utterances to the summarization model to obtain the conversation summary text.

17 . The at least one non-transitory machine-readable storage medium of claim 16 , wherein the one or more semantically insignificant terms comprise one or more terms in a removal reference vocabulary.

18 . The at least one non-transitory machine-readable storage medium of claim 16 , wherein the one or more semantically insignificant terms comprise one or more noise utterances.

19 . The at least one non-transitory machine-readable storage medium of claim 16 , wherein the instructions are configurable to cause the at least one processor to concatenate successive utterances by the first actor after removing the one or more semantically insignificant terms from the transcript prior to selecting the initial subset of the plurality of utterances from the transcript.

20 . A computing system comprising:

at least one non-transitory machine-readable storage medium that stores software; and

at least one processor, coupled to the at least one non-transitory machine-readable storage medium, to execute the software that implements a conversation mining service and that is configurable to perform operations comprising:

obtaining a plurality of utterances associated with a conversation, the plurality of utterances including at least a first set of utterances by a first actor and a second set of utterances corresponding to a second actor;

automatically generating conversation summary text comprising a summarization of semantic content of the conversation based at least in part on an initial subset of the plurality of utterances using a summarization model;

converting the conversation summary text to a numerical conversation summary vector represents semantic meaning of the conversation summary text in a numerical form using an encoder model configured to convert input text into a numerical vector that correlates to semantic meaning of the input text;

identifying, from among the first set of utterances, a representative utterance comprising an individual utterance corresponding to the first actor indicative of the semantic content of the conversation having a corresponding numerical representation using the encoder model that is closest to the numerical conversation summary vector using the summarization model, the representative utterance comprising a verb and noun indicative of an intent of the first actor that initiated the conversation;

automatically updating a field of a record associated with the conversation at a database system to include metadata identifying the representative utterance by the first actor;

automatically assigning the record associated with the conversation to an assigned cluster group of one or more semantically similar conversations among a plurality of different cluster groups of semantically similar conversations based at least in part on the representative utterance; and

providing a graphical user interface (GUI) display comprising a graphical representation of the representative utterance in association with the assigned cluster group of one or more semantically similar conversations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2022
From: MURAKHOVS'KA, LIDIYA; WU, CHIEN-SHENG; MAO, YIXIN
To: SALESFORCE, INC.
Reel/Frame 061634/0992 →
Continuity (2)
Provisional Application 63261397 · Sep 20, 2021
Related Publication 20230086668A1 · Mar 23, 2023
References Cited (191)
US 5577188A · Zhu · 1996 [cited by applicant]
US 5608872A · Schwartz et al. · 1997 [cited by applicant]
US 5649104A · Carleton et al. · 1997 [cited by applicant]
US 5715450A · Ambrose et al. · 1998 [cited by applicant]
US 5761419A · Schwartz et al. · 1998 [cited by applicant]
US 5819038A · Carleton et al. · 1998 [cited by applicant]
US 5821937A · Tonelli et al. · 1998 [cited by applicant]
US 5831610A · Tonelli et al. · 1998 [cited by applicant]
US 5873096A · Lim et al. · 1999 [cited by applicant]
US 5918159A · Fomukong et al. · 1999 [cited by applicant]
US 5963953A · Cram et al. · 1999 [cited by applicant]
US 6092083A · Brodersen et al. · 2000 [cited by applicant]
US 6161149A · Achacoso et al. · 2000 [cited by applicant]
US 6169534B1 · Raffel et al. · 2001 [cited by applicant]
US 6178425B1 · Brodersen et al. · 2001 [cited by applicant]
US 6189011B1 · Lim et al. · 2001 [cited by applicant]
US 6216135B1 · Brodersen et al. · 2001 [cited by applicant]
US 6233617B1 · Rothwein et al. · 2001 [cited by applicant]
US 6266669B1 · Brodersen et al. · 2001 [cited by applicant]
US 6295530B1 · Ritchie et al. · 2001 [cited by applicant]
US 6324568B1 · Diec et al. · 2001 [cited by applicant]
US 6324693B1 · Brodersen et al. · 2001 [cited by applicant]
US 6336137B1 · Lee et al. · 2002 [cited by applicant]
US D454139S · Feldcamp et al. · 2002 [cited by applicant]
US 6367077B1 · Brodersen et al. · 2002 [cited by applicant]
US 6393605B1 · Loomans · 2002 [cited by applicant]
US 6405220B1 · Brodersen et al. · 2002 [cited by applicant]
US 6434550B1 · Warner et al. · 2002 [cited by applicant]
US 6446089B1 · Brodersen et al. · 2002 [cited by applicant]
US 6535909B1 · Rust · 2003 [cited by applicant]
US 6549908B1 · Oomans · 2003 [cited by applicant]
US 6553563B2 · Ambrose et al. · 2003 [cited by applicant]
US 6560461B1 · Fomukong et al. · 2003 [cited by applicant]
US 6574635B2 · Stauber et al. · 2003 [cited by applicant]
US 6577726B1 · Huang et al. · 2003 [cited by applicant]
US 6601087B1 · Zhu et al. · 2003 [cited by applicant]
US 6604117B2 · Im et al. · 2003 [cited by applicant]
US 6604128B2 · Diec · 2003 [cited by applicant]
US 6609150B2 · Lee et al. · 2003 [cited by applicant]
US 6621834B1 · Scherpbier et al. · 2003 [cited by applicant]
US 6654032B1 · Zhu et al. · 2003 [cited by applicant]
US 6665648B2 · Brodersen et al. · 2003 [cited by applicant]
US 6665655B1 · Warner et al. · 2003 [cited by applicant]
US 6684438B2 · Brodersen et al. · 2004 [cited by applicant]
US 6711565B1 · Subramaniam et al. · 2004 [cited by applicant]
US 6724399B1 · Katchour et al. · 2004 [cited by applicant]
US 6728702B1 · Subramaniam et al. · 2004 [cited by applicant]
US 6728960B1 · Loomans et al. · 2004 [cited by applicant]
US 6732095B1 · Warshavsky et al. · 2004 [cited by applicant]
US 6732100B1 · Brodersen et al. · 2004 [cited by applicant]
US 6732111B2 · Brodersen et al. · 2004 [cited by applicant]
US 6754681B2 · Brodersen et al. · 2004 [cited by applicant]
US 6763351B1 · Subramaniam et al. · 2004 [cited by applicant]
US 6763501B1 · Zhu et al. · 2004 [cited by applicant]
US 6768904B2 · Kim · 2004 [cited by applicant]
US 6772229B1 · Achacoso et al. · 2004 [cited by applicant]
US 6782383B2 · Subramaniam et al. · 2004 [cited by applicant]
US 6804330B1 · Jones et al. · 2004 [cited by applicant]
US 6826565B2 · Ritchie et al. · 2004 [cited by applicant]
US 6826582B1 · Chatterjee et al. · 2004 [cited by applicant]
US 6826745B2 · Coker · 2004 [cited by applicant]
US 6829655B1 · Huang et al. · 2004 [cited by applicant]
US 6842748B1 · Warner et al. · 2005 [cited by applicant]
US 6850895B2 · Brodersen et al. · 2005 [cited by applicant]
US 6850949B2 · Warner et al. · 2005 [cited by applicant]
US 7062502B1 · Kesler · 2006 [cited by applicant]
US 7069231B1 · Cinarkaya et al. · 2006 [cited by applicant]
US 7181758B1 · Chan · 2007 [cited by applicant]
US 7289976B2 · Kihneman et al. · 2007 [cited by applicant]
US 7340411B2 · Cook · 2008 [cited by applicant]
US 7356482B2 · Frankland et al. · 2008 [cited by applicant]
US 7401094B1 · Kesler · 2008 [cited by applicant]
US 7412455B2 · Dillon · 2008 [cited by applicant]
US 7508789B2 · Chan · 2009 [cited by applicant]
US 7620655B2 · Larsson et al. · 2009 [cited by applicant]
US 7698160B2 · Beaven et al. · 2010 [cited by applicant]
US 7730478B2 · Weissman · 2010 [cited by applicant]
US 7779475B2 · Jakobson et al. · 2010 [cited by applicant]
US 8014943B2 · Jakobson · 2011 [cited by applicant]
US 8015495B2 · Achacoso et al. · 2011 [cited by applicant]
US 8032297B2 · Jakobson · 2011 [cited by applicant]
US 8082301B2 · Ahlgren et al. · 2011 [cited by applicant]
US 8095413B1 · Beaven · 2012 [cited by applicant]
US 8095594B2 · Beaven et al. · 2012 [cited by applicant]
US 8209308B2 · Rueben et al. · 2012 [cited by applicant]
US 8275836B2 · Beaven et al. · 2012 [cited by applicant]
US 8380511B2 · Cave · 2013 [cited by applicant]
US 8457545B2 · Chan · 2013 [cited by applicant]
US 8484111B2 · Frankland et al. · 2013 [cited by applicant]
US 8490025B2 · Jakobson et al. · 2013 [cited by applicant]
US 8504945B2 · Jakobson et al. · 2013 [cited by applicant]
US 8510045B2 · Rueben et al. · 2013 [cited by applicant]
US 8510664B2 · Rueben et al. · 2013 [cited by applicant]
US 8566301B2 · Rueben et al. · 2013 [cited by applicant]
US 8646103B2 · Jakobson et al. · 2014 [cited by applicant]
US 9430463B2 · Futrell · 2016 [cited by applicant]
US 10930272B1 · Orkin · 2021 [cited by applicant]
US 11507756B2 · Lima · 2022 [cited by applicant]
US 11551677B2 · Roy · 2023 [cited by applicant]
US 11568856B2 · Ho · 2023 [cited by applicant]
US 20010044791A1 · Richter et al. · 2001 [cited by applicant]
US 20020072951A1 · Lee et al. · 2002 [cited by applicant]
US 20020082892A1 · Raffel · 2002 [cited by applicant]
US 20020129352A1 · Brodersen et al. · 2002 [cited by applicant]
US 20020140731A1 · Subramanian et al. · 2002 [cited by applicant]
US 20020143997A1 · Huang et al. · 2002 [cited by applicant]
US 20020162090A1 · Parnell et al. · 2002 [cited by applicant]
US 20020165742A1 · Robbins · 2002 [cited by applicant]
US 20030004971A1 · Gong · 2003 [cited by applicant]
US 20030018705A1 · Chen et al. · 2003 [cited by applicant]
US 20030018830A1 · Chen et al. · 2003 [cited by applicant]
US 20030066031A1 · Laane et al. · 2003 [cited by applicant]
US 20030066032A1 · Ramachandran et al. · 2003 [cited by applicant]
US 20030069936A1 · Warner · 2003 [cited by applicant]
US 20030070000A1 · Coker et al. · 2003 [cited by applicant]
US 20030070004A1 · Mukundan et al. · 2003 [cited by applicant]
US 20030070005A1 · Mukundan et al. · 2003 [cited by applicant]
US 20030074418A1 · Coker et al. · 2003 [cited by applicant]
US 20030120675A1 · Stauber et al. · 2003 [cited by applicant]
US 20030151633A1 · George et al. · 2003 [cited by applicant]
US 20030159136A1 · Huang et al. · 2003 [cited by applicant]
US 20030187921A1 · Diec et al. · 2003 [cited by applicant]
US 20030189600A1 · Gune et al. · 2003 [cited by applicant]
US 20030204427A1 · Gune et al. · 2003 [cited by applicant]
US 20030206192A1 · Chen et al. · 2003 [cited by applicant]
US 20030225730A1 · Warner et al. · 2003 [cited by applicant]
US 20040001092A1 · Rothwein et al. · 2004 [cited by applicant]
US 20040010489A1 · Rio et al. · 2004 [cited by applicant]
US 20040015981A1 · Coker et al. · 2004 [cited by applicant]
US 20040027388A1 · Berg et al. · 2004 [cited by applicant]
US 20040128001A1 · Levin et al. · 2004 [cited by applicant]
US 20040186860A1 · Lee et al. · 2004 [cited by applicant]
US 20040193510A1 · Catahan et al. · 2004 [cited by applicant]
US 20040199489A1 · Barnes-Leon et al. · 2004 [cited by applicant]
US 20040199536A1 · Barnes-Leon et al. · 2004 [cited by applicant]
US 20040199543A1 · Braud et al. · 2004 [cited by applicant]
US 20040249854A1 · Barnes-Leon et al. · 2004 [cited by applicant]
US 20040260534A1 · Pak et al. · 2004 [cited by applicant]
US 20040260659A1 · Chan et al. · 2004 [cited by applicant]
US 20040268299A1 · Lei et al. · 2004 [cited by applicant]
US 20050050555A1 · Exley et al. · 2005 [cited by applicant]
US 20050091098A1 · Brodersen et al. · 2005 [cited by applicant]
US 20050251383A1 · Murray · 2005 [cited by applicant]
US 20060021019A1 · Hinton et al. · 2006 [cited by applicant]
US 20080249972A1 · Dillon · 2008 [cited by applicant]
US 20090063414A1 · White et al. · 2009 [cited by applicant]
US 20090100342A1 · Jakobson · 2009 [cited by applicant]
US 20090177744A1 · Marlow et al. · 2009 [cited by applicant]
US 20110238408A1 · Larcheveque · 2011 [cited by applicant]
US 20110247051A1 · Bulumulla et al. · 2011 [cited by applicant]
US 20120042218A1 · Cinarkaya et al. · 2012 [cited by applicant]
US 20120218958A1 · Rangaiah · 2012 [cited by applicant]
US 20120233137A1 · Jakobson et al. · 2012 [cited by applicant]
US 20130212497A1 · Zelenko et al. · 2013 [cited by applicant]
US 20130218948A1 · Jakobson · 2013 [cited by applicant]
US 20130218949A1 · Jakobson · 2013 [cited by applicant]
US 20130218966A1 · Jakobson · 2013 [cited by applicant]
US 20130247216A1 · Cinarkaya et al. · 2013 [cited by applicant]
US 20150347393A1 · Futrell · 2015 [cited by applicant]
US 20160012818A1 · Faizakof · 2016 [cited by applicant]
US 20180174600A1 · Chaudhuri · 2018 [cited by applicant]
US 20180240015A1 · Martin · 2018 [cited by applicant]
US 20180247549A1 · Martin · 2018 [cited by applicant]
US 20200014326A1 · Miller · 2020 [cited by applicant]
US 20200152183A1 · Wang · 2020 [cited by applicant]
US 20210256534A1 · An · 2021 [cited by applicant]
US 20210342554A1 · Martin · 2021 [cited by examiner]
US 20210390127A1 · Fox · 2021 [cited by examiner]
US 20220058343A1 · Sapugay · 2022 [cited by applicant]
US 20220156296A1 · de Oliveira · 2022 [cited by examiner]
US 20220156460A1 · Láinez Rodrigo · 2022 [cited by examiner]
US 20220200936A1 · Higgins · 2022 [cited by applicant]
US 20220222437A1 · Lauber · 2022 [cited by examiner]
US 20230054726A1 · Roy · 2023 [cited by examiner]
WO 2016142933A1 · 2016 [cited by applicant]
Cambridge University Press, Dropping common terms: stop words, Apr. 7, 2009. [cited by applicant]
Wikipedia, Hierarchical clustering, https://en.wikipedia.org/wiki/Hierarchical_clustering (accessed Sep. 19, 2022). [cited by applicant]
Wikipedia, Lemmatisation, https://en.wikipedia.org/wiki/Lemmatisation (accessed Sep. 19, 2022). [cited by applicant]
Wikipedia, Text normalization, https://en.wikipedia.org/wiki/Text_normalization (accessed Sep. 19, 2022). [cited by applicant]
Wikipedia, tf-idf, https://en.wikipedia.org/wiki/Tf%E2%80%93idf (accessed Sep. 19, 2022). [cited by applicant]
Nils Reimers, et al., Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks, Ubiquitous Knowledge Processing Lab (UKP-TUDA) Department of Computer Science, Technische Universitat Darmstadt, Aug. 27, 2019. [cited by applicant]
Mike Lewis, et al., BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation and Comprehension, Jul. 2020. [cited by applicant]
CYLNLP/DIALOGSUM, DialogSum: A Real-life Scenerio Dialogue Summarization Dataset—Findings of ACL 2021. [cited by applicant]
Scikit-Learn Developers (Bsd License), 2.1. Gaussian mixture models, https://scikit-learn.org/stable/modules/ mixture.html, 2007. [cited by applicant]
Cambridge University Press, Dropping common terms: stop words, https://nlp.stanford.edu/IR-book/html/ htmledition/dropping-common-terms-stop-words-1.html. [cited by applicant]
Wikipedia, Hierarchical clustering, https://en.wikipedia.org/wiki/Hierarchical_clustering. [cited by applicant]
Wikipedia, k-nearest neighbors algorithm, https://en.wikipedia.org/wiki/K-nearest_neighbors_algorithm. [cited by applicant]
Wikipedia, Lemmatisation, https://en.wikipedia.org/wiki/Lemmatisation. [cited by applicant]
Sklearn.metrics.silhouette_score, https://scikit-learn.org/stable/modules/generated/sklearn.metrics.silhouette_score. html. [cited by applicant]
Wikipedia, Text normalization, https://en.wikipedia.org/wiki/Text_normalization. [cited by applicant]
Wikipedia, tf-idf, https://en.wikipedia.org/wiki/Tf%E2%80%93idf. [cited by applicant]