IP Library › Granted Patent US 12,651,112
Granted Patent B2
US 12,651,112 · App. 17/903,996 · Granted Jun 9, 2026

Intelligently identifying freshness of terms in documentation

Inventors: Dan Zhang (Shanghai, CN); Jun Qian Zhou (Shanghai, CN); Yuan Jie Song (Shanghai, CN); Meng Chai (Shanghai, CN); Zhen Ma (Shanghai, CN); Xiao Feng Ji (Shanghai, CN)
Assignee: International Business Machines Corporation
G06F40/166G06F16/3334G06F40/237G06F40/247G06F40/253G06F40/284G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,112
App. No.
17/903,996
Granted
Jun 9, 2026
Kind
B2
Abstract

The freshness of one or more terms in a documentation, indicative of a currency of the one or more terms is computed. Each term includes one or more constituent words, and the terms are visually marked as current or out-of-date based on the computed freshness. Upon marking a term as out-of-date, a latest term for the out-of-date term is retrieved or a most possible latest term for the out-of-date term is predicted.

Claims (75)

1 . A computer-implemented method, comprising:

receiving, from a documentation including a plurality of sentences that form a corpora, a term comprising one or more constituent words;

computing an active year distribution space comprising a three-dimensional space including information about a most active year of words from the documentation;

computing, for the term, using the active year distribution space, a sequence of freshness distribution vectors representative of a freshness distribution of the term;

obtaining a target year of the term;

computing, using a trained self-supervised deep learning model, a freshness of the term based on a target year vector of the target year and the sequence of freshness distribution vectors, the freshness being indicative of a currency of the term;

visually marking the term as current, in a case where the computing of the freshness indicates that the freshness is a fresh freshness;

visually marking the term as out-of-date, in a case where the computing of the freshness indicates that the freshness is a non-fresh freshness; and

based on the marking of the term as out-of-date, automatically retrieving, for the term, a latest term from a term change history database, or automatically predicting a most possible latest term for the term.

2 . The computer-implemented method of claim 1 , wherein the active year distribution space is computed by:

tokenizing the plurality of sentences from the corpora of the documentation to obtain word embeddings that form a multi-dimensional vector space;

converting, by dimensionality reduction, the multi-dimensional vector space into a two-dimensional plane space;

adding a third dimension representing the most active year of words to the two-dimensional plane space; and

projecting words in the two-dimensional plane space onto the third dimension based on the most active year of words.

3 . The computer-implemented method of claim 2 , wherein the dimensionality reduction comprises principal component analysis.

4 . The computer-implemented method of claim 1 , wherein the most active year of words is provided by a word frequency history.

5 . The computer-implemented method of claim 1 , wherein:

the sequence of freshness distribution vectors comprises a freshness distribution vector for each constituent word of the one or more constituent words of the term, and the freshness distribution vector is obtained by:

generating a sphere in the active year distribution space centered on a constituent word of the one or more constituent words;

identifying a plurality of nearest n words as a plurality of most frequently cooccurred words, n being a predefined number;

computing a first span between a most active year of the constituent word and the target year;

computing, for each nearest word of the plurality of nearest n words, a second span between a most active year for each nearest word of the plurality of nearest n words and the target year; and

computing the freshness distribution vector by forming a vector with a size of n+1 from the first span and the second span.

6 . The computer-implemented method of claim 1 , wherein based on one of computing the term as non-fresh or receiving a term previously marked as non-fresh, the most possible latest term is automatically predicted and displayed by:

tokenizing the plurality of sentences from the corpora of the documentation to obtain word embeddings that form a multi-dimensional vector space;

for a target term identified as non-fresh, selecting a specific term that meets a predefined candidate term condition from the multi-dimensional vector space as a candidate term;

for the specific term selected as the candidate term, providing as input to the trained self-supervised deep learning model a sequence of gradually increased target years and the candidate term, to predict a freshness of the candidate term;

based on the predicting of a fresh freshness for the candidate term, storing the candidate term along with a greatest target year value of the gradually increased target years to obtain one or more stored candidate term and greatest target year pairs; and

computing the most possible latest term by selecting a stored candidate term and greatest target year pair that has a highest value of the greatest target year value among the one or more stored candidate term and greatest target year pairs.

7 . The computer-implemented method of claim 6 , wherein the predefined candidate term condition comprises computing a cosine value “K distance” between the word embeddings of the candidate term and the target term that exceeds a preset threshold.

8 . The computer-implemented method of claim 7 , further comprising selecting an instance having a shortest “K distance” based on obtaining of multiple instances of the most possible latest term.

9 . The computer-implemented method of claim 6 , wherein the word embeddings are computed through a BoW (Bag-of-words) model to constitute the multi-dimensional vector space.

10 . The computer-implemented method of claim 1 , further comprising:

configuring the trained self-supervised deep learning model, to compute the freshness of the term, based on a training term change history database and a first deep learning model by:

randomly selecting, from the training term change history database, a training original term and a training target year for the training original term;

computing, for the training original term, a training sequence of freshness distribution vectors based on the training target year;

computing, for the training target year, a training target year vector; and

providing the training sequence of freshness distribution vectors and the training target year vector as an input set to the first deep learning model to generate a corresponding processed training output representative of the freshness of the training original term.

11 . The computer-implemented method of claim 10 , further comprising:

performing a back propagation, based on an accuracy of the corresponding processed training output, to update parameters of the first deep learning model.

12 . The computer-implemented method of claim 1 , further comprising:

replacing, in the documentation, based on the marking of the term as out-of-date, the term with the latest term or the most possible latest term.

13 . The computer-implemented method of claim 1 , further comprising:

displaying, based on the marking of the term as out-of-date, the latest term, or the most possible latest term.

14 . A computer program product, comprising:

one or more computer-readable storage devices and program instructions stored on at least one of the one or more computer-readable storage devices, the program instructions executable by a processor cause the processor to:

receive, from a documentation including a plurality of sentences that form a corpora, a term comprising one or more constituent words;

compute an active year distribution space comprising a three-dimensional space including information about a most active year of words from the documentation;

compute, for the term, using the active year distribution space, a sequence of freshness distribution vectors representative of a freshness distribution of the term;

obtain a target year of the term;

compute, using a trained self-supervised deep learning model, a freshness of the term based on a target year vector of the target year and the sequence of freshness distribution vectors, the freshness being indicative currency of the term;

visually mark the term as current, in a case where the computation of the freshness indicates that the freshness is a fresh freshness;

visually mark the term as out-of-date, in a case where the computation of the freshness indicates that the freshness is a non-fresh freshness; and

based on the marking of the term as out-of-date, automatically retrieve, for the term, a latest term from a term change history database, or automatically predict a most possible latest term for the term.

15 . The computer program product of claim 14 , wherein in a case where the term is computed as non-fresh or a term previously marked as non-fresh is received, the most possible latest term is automatically predicted and displayed based on the program instructions that further cause the processor to:

tokenize the plurality of sentences from the corpora of the documentation to obtain word embeddings that form a multi-dimensional vector space;

for a target term identified as non-fresh, select a specific term that meets a predefined candidate term condition from the multi-dimensional vector space as a candidate term;

for the specific term selected as the candidate term, provide as input to the trained self-supervised deep learning model a sequence of gradually increased target years and the candidate term, to predict a freshness of the candidate term;

based on the prediction of a fresh freshness for the candidate term, store the candidate term along with a greatest target year value of the gradually increased target years to obtain one or more stored candidate term and greatest target year pairs; and

compute the most possible latest term by selecting the stored candidate term having a highest greatest target year value among the one or more stored candidate term and greatest target year pairs.

16 . A non-transitory computer readable storage medium tangibly embodying a computer readable program code having computer readable instructions that, when executed, causes a processor to carry out a method comprising:

receiving, from a documentation including a plurality of sentences that form a corpora, a term comprising one or more constituent words;

computing an active year distribution space comprising a three-dimensional space including information about a most active year of words from the documentation;

computing, for the term, using the active year distribution space, a sequence of freshness distribution vectors representative of a freshness distribution of the term;

obtaining a target year of the term;

computing, using a trained self-supervised deep learning model, a freshness of the term based on a target year vector of the target year and the sequence of freshness distribution vectors, the freshness being indicative of a currency of the term;

visually marking the term as current, in a case where the computing of the freshness indicates that the freshness is a fresh freshness;

visually marking the term as out-of-date, in a case where the computing of the freshness indicates that the freshness is a non-fresh freshness; and

based on the marking the term as out-of-date, automatically retrieving, for the term, a latest term from a term change history database, or automatically predicting a most possible latest term for the term.

17 . The non-transitory computer readable storage medium of claim 16 , wherein the method further comprises, based on one of computing the term as non-fresh or receiving a term previously marked as non-fresh, automatically predicting and displaying the most possible latest term by:

tokenizing the plurality of sentences from the corpora of the documentation to obtain word embeddings that form a multi-dimensional vector space;

for a target term identified as non-fresh, selecting a specific term that meets a predefined candidate term condition from the multi-dimensional vector space as a candidate term;

for the specific term selected as the candidate term, providing as input to the trained self-supervised deep learning model a sequence of gradually increased target years and the candidate term, to predict a freshness of candidate term;

based on the predicting of a fresh freshness for the candidate term, storing the candidate term along with a greatest target year value of the gradually increased target years to obtain one or more stored candidate term and greatest target year pairs; and

computing the most possible latest term by selecting the stored candidate term having a highest greatest target year value among the one or more stored candidate term and greatest target year pairs.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2022
From: ZHANG, DAN; ZHOU, JUN QIAN; SONG, YUAN JIE; CHAI, MENG; MA, ZHEN; JI, XIAO FENG
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061007/0373 →
Continuity (1)
Related Publication 20240078372A1 · Mar 7, 2024
References Cited (19)
US 8001462B1 · Kupke et al. · 2011 [cited by applicant]
US 8762130B1 · Diaconescu · 2014 [cited by examiner]
US 10169427B2 · Aaron · 2019 [cited by examiner]
US 10902201B2 · Hewitt et al. · 2021 [cited by applicant]
US 20050108426A1 · Klein · 2005 [cited by examiner]
US 20150309989A1 · Brav · 2015 [cited by examiner]
US 20170111465A1 · Yellin et al. · 2017 [cited by applicant]
US 20200356626A1 · Cogley · 2020 [cited by examiner]
US 20220019429A1 · Farivar et al. · 2022 [cited by applicant]
US 20220245348A1 · Chen · 2022 [cited by examiner]
CN 102253998A · 2011 [cited by applicant]
CN 106484380A · 2017 [cited by applicant]
Li, Hongqiao, Chang-Ning Huang, Jianfeng Gao, and Xiaozhong Fan, “The Use of SVM for Chinese New Word Identification”, Mar. 2004, Proceedings of the First International Joint Conference on Natural Language Processing (I… [cited by examiner]
Szymanski, Terrence, “Temporal Word Analogies: Identifying Lexical Replacement with Diachronic Word Embeddings”, Aug. 2017, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (vol. 2… [cited by examiner]
Wang, Meng, Lanfen Lin, and Feng Wang, “New Word Identification in Social Network Text Based on Time Series Information”, May 2014, Proceedings of the 2014 IEEE 18th International Conference on Computer Supported Cooper… [cited by examiner]
Mell, P. et al., “Recommendations of the National Institute of Standards and Technology”; NIST Special Publication 800-145 (2011); 7 pgs. [cited by applicant]
Hong, C. et al., “Automatic Extraction of New Words Based on Google News Corpora for Supporting Lexicon-Based Chinese Word Segmentation Systems”; published online 2009; downloaded Jun. 7, 2022, 4 pgs. [cited by applicant]
No Author, “WolframAlpha”, Wolfram Language, Aug. 31, 2022, 2 Pages. [cited by applicant]
Pennington, et al., “GloVe: Global Vectors for Word Representation”, Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, Jan. 2014, 12 Pages. [cited by applicant]