Intelligently identifying freshness of terms in documentation
The freshness of one or more terms in a documentation, indicative of a currency of the one or more terms is computed. Each term includes one or more constituent words, and the terms are visually marked as current or out-of-date based on the computed freshness. Upon marking a term as out-of-date, a latest term for the out-of-date term is retrieved or a most possible latest term for the out-of-date term is predicted.
1 . A computer-implemented method, comprising:
receiving, from a documentation including a plurality of sentences that form a corpora, a term comprising one or more constituent words;
computing an active year distribution space comprising a three-dimensional space including information about a most active year of words from the documentation;
computing, for the term, using the active year distribution space, a sequence of freshness distribution vectors representative of a freshness distribution of the term;
obtaining a target year of the term;
computing, using a trained self-supervised deep learning model, a freshness of the term based on a target year vector of the target year and the sequence of freshness distribution vectors, the freshness being indicative of a currency of the term;
visually marking the term as current, in a case where the computing of the freshness indicates that the freshness is a fresh freshness;
visually marking the term as out-of-date, in a case where the computing of the freshness indicates that the freshness is a non-fresh freshness; and
based on the marking of the term as out-of-date, automatically retrieving, for the term, a latest term from a term change history database, or automatically predicting a most possible latest term for the term.
2 . The computer-implemented method of claim 1 , wherein the active year distribution space is computed by:
tokenizing the plurality of sentences from the corpora of the documentation to obtain word embeddings that form a multi-dimensional vector space;
converting, by dimensionality reduction, the multi-dimensional vector space into a two-dimensional plane space;
adding a third dimension representing the most active year of words to the two-dimensional plane space; and
projecting words in the two-dimensional plane space onto the third dimension based on the most active year of words.
3 . The computer-implemented method of claim 2 , wherein the dimensionality reduction comprises principal component analysis.
4 . The computer-implemented method of claim 1 , wherein the most active year of words is provided by a word frequency history.
5 . The computer-implemented method of claim 1 , wherein:
the sequence of freshness distribution vectors comprises a freshness distribution vector for each constituent word of the one or more constituent words of the term, and the freshness distribution vector is obtained by:
generating a sphere in the active year distribution space centered on a constituent word of the one or more constituent words;
identifying a plurality of nearest n words as a plurality of most frequently cooccurred words, n being a predefined number;
computing a first span between a most active year of the constituent word and the target year;
computing, for each nearest word of the plurality of nearest n words, a second span between a most active year for each nearest word of the plurality of nearest n words and the target year; and
computing the freshness distribution vector by forming a vector with a size of n+1 from the first span and the second span.
6 . The computer-implemented method of claim 1 , wherein based on one of computing the term as non-fresh or receiving a term previously marked as non-fresh, the most possible latest term is automatically predicted and displayed by:
tokenizing the plurality of sentences from the corpora of the documentation to obtain word embeddings that form a multi-dimensional vector space;
for a target term identified as non-fresh, selecting a specific term that meets a predefined candidate term condition from the multi-dimensional vector space as a candidate term;
for the specific term selected as the candidate term, providing as input to the trained self-supervised deep learning model a sequence of gradually increased target years and the candidate term, to predict a freshness of the candidate term;
based on the predicting of a fresh freshness for the candidate term, storing the candidate term along with a greatest target year value of the gradually increased target years to obtain one or more stored candidate term and greatest target year pairs; and
computing the most possible latest term by selecting a stored candidate term and greatest target year pair that has a highest value of the greatest target year value among the one or more stored candidate term and greatest target year pairs.
7 . The computer-implemented method of claim 6 , wherein the predefined candidate term condition comprises computing a cosine value “K distance” between the word embeddings of the candidate term and the target term that exceeds a preset threshold.
8 . The computer-implemented method of claim 7 , further comprising selecting an instance having a shortest “K distance” based on obtaining of multiple instances of the most possible latest term.
9 . The computer-implemented method of claim 6 , wherein the word embeddings are computed through a BoW (Bag-of-words) model to constitute the multi-dimensional vector space.
10 . The computer-implemented method of claim 1 , further comprising:
configuring the trained self-supervised deep learning model, to compute the freshness of the term, based on a training term change history database and a first deep learning model by:
randomly selecting, from the training term change history database, a training original term and a training target year for the training original term;
computing, for the training original term, a training sequence of freshness distribution vectors based on the training target year;
computing, for the training target year, a training target year vector; and
providing the training sequence of freshness distribution vectors and the training target year vector as an input set to the first deep learning model to generate a corresponding processed training output representative of the freshness of the training original term.
11 . The computer-implemented method of claim 10 , further comprising:
performing a back propagation, based on an accuracy of the corresponding processed training output, to update parameters of the first deep learning model.
12 . The computer-implemented method of claim 1 , further comprising:
replacing, in the documentation, based on the marking of the term as out-of-date, the term with the latest term or the most possible latest term.
13 . The computer-implemented method of claim 1 , further comprising:
displaying, based on the marking of the term as out-of-date, the latest term, or the most possible latest term.
14 . A computer program product, comprising:
one or more computer-readable storage devices and program instructions stored on at least one of the one or more computer-readable storage devices, the program instructions executable by a processor cause the processor to:
receive, from a documentation including a plurality of sentences that form a corpora, a term comprising one or more constituent words;
compute an active year distribution space comprising a three-dimensional space including information about a most active year of words from the documentation;
compute, for the term, using the active year distribution space, a sequence of freshness distribution vectors representative of a freshness distribution of the term;
obtain a target year of the term;
compute, using a trained self-supervised deep learning model, a freshness of the term based on a target year vector of the target year and the sequence of freshness distribution vectors, the freshness being indicative currency of the term;
visually mark the term as current, in a case where the computation of the freshness indicates that the freshness is a fresh freshness;
visually mark the term as out-of-date, in a case where the computation of the freshness indicates that the freshness is a non-fresh freshness; and
based on the marking of the term as out-of-date, automatically retrieve, for the term, a latest term from a term change history database, or automatically predict a most possible latest term for the term.
15 . The computer program product of claim 14 , wherein in a case where the term is computed as non-fresh or a term previously marked as non-fresh is received, the most possible latest term is automatically predicted and displayed based on the program instructions that further cause the processor to:
tokenize the plurality of sentences from the corpora of the documentation to obtain word embeddings that form a multi-dimensional vector space;
for a target term identified as non-fresh, select a specific term that meets a predefined candidate term condition from the multi-dimensional vector space as a candidate term;
for the specific term selected as the candidate term, provide as input to the trained self-supervised deep learning model a sequence of gradually increased target years and the candidate term, to predict a freshness of the candidate term;
based on the prediction of a fresh freshness for the candidate term, store the candidate term along with a greatest target year value of the gradually increased target years to obtain one or more stored candidate term and greatest target year pairs; and
compute the most possible latest term by selecting the stored candidate term having a highest greatest target year value among the one or more stored candidate term and greatest target year pairs.
16 . A non-transitory computer readable storage medium tangibly embodying a computer readable program code having computer readable instructions that, when executed, causes a processor to carry out a method comprising:
receiving, from a documentation including a plurality of sentences that form a corpora, a term comprising one or more constituent words;
computing an active year distribution space comprising a three-dimensional space including information about a most active year of words from the documentation;
computing, for the term, using the active year distribution space, a sequence of freshness distribution vectors representative of a freshness distribution of the term;
obtaining a target year of the term;
computing, using a trained self-supervised deep learning model, a freshness of the term based on a target year vector of the target year and the sequence of freshness distribution vectors, the freshness being indicative of a currency of the term;
visually marking the term as current, in a case where the computing of the freshness indicates that the freshness is a fresh freshness;
visually marking the term as out-of-date, in a case where the computing of the freshness indicates that the freshness is a non-fresh freshness; and
based on the marking the term as out-of-date, automatically retrieving, for the term, a latest term from a term change history database, or automatically predicting a most possible latest term for the term.
17 . The non-transitory computer readable storage medium of claim 16 , wherein the method further comprises, based on one of computing the term as non-fresh or receiving a term previously marked as non-fresh, automatically predicting and displaying the most possible latest term by:
tokenizing the plurality of sentences from the corpora of the documentation to obtain word embeddings that form a multi-dimensional vector space;
for a target term identified as non-fresh, selecting a specific term that meets a predefined candidate term condition from the multi-dimensional vector space as a candidate term;
for the specific term selected as the candidate term, providing as input to the trained self-supervised deep learning model a sequence of gradually increased target years and the candidate term, to predict a freshness of candidate term;
based on the predicting of a fresh freshness for the candidate term, storing the candidate term along with a greatest target year value of the gradually increased target years to obtain one or more stored candidate term and greatest target year pairs; and
computing the most possible latest term by selecting the stored candidate term having a highest greatest target year value among the one or more stored candidate term and greatest target year pairs.