Identification of symbol drift in written discourse
An embodiment establishes a text corpus database based at least in part on text data received from a written discourse. The embodiment extracts a first set of terms from the text corpus database and a first set of collocations for each term of the first set of terms. The embodiment constructs a summation model based at least in part on the first set of terms and the first set of collocations for each term of the first set of terms. The embodiment inputs additional text data into the summation model to determine a probability score that defines the probability that a term stored on the text corpus database will change meaning. The embodiment displays the probability score on a user interface.
1 . A computer-implemented method comprising:
integrating a drift analyzer graphical user interface (GUI) with a chat application graphical user interface (GUI) of a chat application;
establishing a text corpus database, the text corpus database comprising text data received from a written discourse, wherein the written discourse comprises a real-time chat discourse received from the chat application;
extracting a first set of terms from the text corpus database and a first set of collocations for each term of the first set of terms;
constructing a summation model based at least in part on the first set of terms and the first set of collocations for each term of the first set of terms;
inputting additional text data into the summation model to determine a probability score that defines the probability that a term stored on the text corpus database will change meaning;
overlaying computer-generated graphics on the chat application GUI upon detecting that the term stored on the text corpus database will change meaning; and
displaying the probability score on on the drift analyzer GUI.
2 . The computer-implemented method of claim 1 , wherein the summation model comprises a discrete Markov model.
3 . The computer-implemented method of claim 1 , further comprising generating temporal step data for each term stored on the text corpus database, wherein the temporal step data tracks the symbolic drift of each term.
4 . The computer-implemented method of claim 1 , wherein the summation model is trained for a first domain, and wherein the method further comprises re-training the summation model for a second domain that is different than the first domain.
5 . The computer-implemented method of claim 1 , further comprising identifying a synonym term within the additional text data corresponding to an existing term stored on the text corpus database.
6 . A computer program product comprising one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by a processor to cause the processor to perform operations comprising:
integrating a drift analyzer graphical user interface (GUI) with a chat application graphical user interface (GUI) of a chat application;
establishing a text corpus database, the text corpus database comprising text data received from a written discourse, wherein the written discourse comprises a real-time chat discourse received from the chat application;
extracting a first set of terms from the text corpus database and a first set of collocations for each term of the first set of terms;
constructing a summation model based at least in part on the first set of terms and the first set of collocations for each term of the first set of terms;
inputting additional text data into the summation model to determine a probability score that defines the probability that a term stored on the text corpus database will change meaning;
overlaying computer-generated graphics on the chat application GUI upon detecting that the term stored on the text corpus database will change meaning; and
displaying the probability score on on the drift analyzer GUI.
7 . The computer program product of claim 6 , wherein the stored program instructions are stored in a computer readable storage device in a data processing system, and wherein the stored program instructions are transferred over a network from a remote data processing system.
8 . The computer program product of claim 6 , wherein the stored program instructions are stored in a computer readable storage device in a server data processing system, and wherein the stored program instructions are downloaded in response to a request over a network to a remote data processing system for use in a computer readable storage device associated with the remote data processing system, further comprising:
program instructions to meter use of the program instructions associated with the request; and
program instructions to generate an invoice based on the metered use.
9 . The computer program product of claim 6 , wherein the summation model comprises a discrete Markov model.
10 . The computer program product of claim 6 , further comprising generating temporal step data for each term stored on the text corpus database, wherein the temporal step data tracks the symbolic drift of each term.
11 . The computer program product of claim 6 , wherein the summation model is trained for a first domain, and wherein the method further comprises re-training the summation model for a second domain that is different than the first domain.
12 . The computer program product of claim 6 , further comprising identifying a synonym term within the additional text data corresponding to an existing term stored on the text corpus database.
13 . A computer system comprising a processor and one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by the processor to cause the processor to perform operations comprising:
integrating a drift analyzer graphical user interface (GUI) with a chat application graphical user interface (GUI) of a chat application;
establishing a text corpus database, the text corpus database comprising text data received from a written discourse, wherein the written discourse comprises a real-time chat discourse received from the chat application;
extracting a first set of terms from the text corpus database and a first set of collocations for each term of the first set of terms;
constructing a summation model based at least in part on the first set of terms and the first set of collocations for each term of the first set of terms;
inputting additional text data into the summation model to determine a probability score that defines the probability that a term stored on the text corpus database will change meaning;
overlaying computer-generated graphics on the chat application GUI upon detecting that the term stored on the text corpus database will change meaning; and
displaying the probability score on on the drift analyzer GUI.
14 . The computer system of claim 13 , wherein the summation model comprises a discrete Markov model.
15 . The computer system of claim 13 , further comprising generating temporal step data for each term stored on the text corpus database, wherein the temporal step data tracks the symbolic drift of each term.
16 . The computer system of claim 13 , wherein the summation model is trained for a first domain, and wherein the system further comprises re-training the summation model for a second domain that is different than the first domain.
17 . The computer system of claim 13 , further comprising identifying a synonym term within the additional text data corresponding to an existing term stored on the text corpus database.
18 . The computer system of claim 17 , further comprising generating a modified representation of the additional text data, wherein the modified representation of the additional text data comprises the identified synonym term replaced with the existing term.