IP Library Granted Patent US 12,596,735
Granted Patent B2
US 12,596,735 · App. 18/638,365 · Granted Apr 7, 2026

Semantic text analysis for glossary maintenance

Inventors: Toshihiro Takahashi (Nakano-ku, JP); Takaaki Tateishi (Yamato, JP)
Assignee: International Business Machines Corporation
G06F16/3344G06F16/367G06F40/242G06F40/40G06N5/022G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,735
App. No.
18/638,365
Granted
Apr 7, 2026
Kind
B2
Abstract

A computer-implemented method (CIM), according to one embodiment, includes causing a first search to be performed on a first knowledge base for a first descriptive name, and extracting sentences from results of the first search. The method further includes running at least one predetermined deep learning model on the results of the first search for determining similarity scores for each of the extracted sentences, where each of the similarity scores defines a similarity score for an associated one of the extracted sentences and a first glossary. A first of the determined similarity scores is used to map terms of a second glossary to terms of the first glossary to enhance context provided in search results generated using the glossaries.

Claims (38)

1 . A computer-implemented method (CIM), the CIM comprising:

causing a first search to be performed on a first knowledge base for a first descriptive name;

extracting sentences from results of the first search,

wherein the extracted sentences include labels associated with the first descriptive name and descriptions of the labels;

running at least one predetermined deep learning model on the results of the first search for determining similarity scores for the extracted sentences, wherein each of the similarity scores defines a similarity score for an associated one of the extracted sentences and a first glossary;

determining a first of the determined similarity scores, wherein the first determined similarity score has a relatively greater similarity score than the other determined similarity scores; and

using the first determined similarity score to map terms of a second glossary to terms of the first glossary to enhance context provided in search results generated using the glossaries.

2 . The CIM of claim 1 , wherein using the first determined similarity score to map terms of the second glossary to terms of the first glossary includes: mapping the label of the extracted sentence associated with the first determined similarity score to at least one of the terms of the second glossary, and comprising: in response to receiving, from a first device, a request for a second search to be performed on the at least one of the terms of the second glossary, causing the label of the extracted sentence associated with the first determined similarity score to be returned to the first device.

3 . The CIM of claim 2 , wherein the first glossary is a financial industry business ontology.

4 . The CIM of claim 1 , wherein the predetermined deep learning model run on the results of the first search is a Prompt-based Contrastive Learning for Sentence Embeddings (PromCSE) model.

5 . The CIM of claim 4 , wherein a second predetermined deep learning model is run on the results of the first search for determining the similarity scores for the extracted sentences, wherein the second predetermined deep learning model is a Term Frequency-Inverse Document Frequency (TF-IDF) model.

6 . The CIM of claim 1 , wherein the first descriptive name is selected from the group consisting of: a column name in a database, a header in a comma-separated values (CSV) formatted text file, and a categorical variable in a CSV formatted text file.

7 . The CIM of claim 1 , wherein the first knowledge base is an open and editable knowledge base accessible on the Internet.

8 . A computer program product (CPP), the CPP comprising:

a set of one or more computer-readable storage media; and

program instructions, collectively stored in the set of one or more computer-readable storage media, for causing a processor set to perform the following computer operations:

cause a first search to be performed on a first knowledge base for a first descriptive name;

extract sentences from results of the first search,

wherein the extracted sentences include labels associated with the first descriptive name and descriptions of the labels;

run at least one predetermined deep learning model on the results of the first search for determining similarity scores for the extracted sentences, wherein each of the similarity scores defines a similarity score for an associated one of the extracted sentences and a first glossary;

determine a first of the determined similarity scores, wherein the first determined similarity score has a relatively greater similarity score than the other determined similarity scores; and

use the first determined similarity score to map terms of a second glossary to terms of the first glossary to enhance context provided in search results generated using the glossaries.

9 . The CPP of claim 8 , wherein using the first determined similarity score to map terms of the second glossary to terms of the first glossary includes: mapping the label of the extracted sentence associated with the first determined similarity score to at least one of the terms of the second glossary, and the CPP comprising: program instructions, collectively stored in the set of one or more storage media, for causing the processor set to perform the following computer operations: in response to receiving, from a first device, a request for a second search to be performed on the at least one of the terms of the second glossary, cause the label of the extracted sentence associated with the first determined similarity score to be returned to the first device.

10 . The CPP of claim 9 , wherein the first glossary is a financial industry business ontology.

11 . The CPP of claim 8 , wherein the predetermined deep learning model run on the results of the first search is a Prompt-based Contrastive Learning for Sentence Embeddings (PromCSE) model.

12 . The CPP of claim 11 , wherein a second predetermined deep learning model is run on the results of the first search for determining the similarity scores for the extracted sentences, wherein the second predetermined deep learning model is a Term Frequency-Inverse Document Frequency (TF-IDF) model.

13 . The CPP of claim 8 , wherein the first descriptive name is selected from the group consisting of: a column name in a database, a header in a comma-separated values (CSV) formatted text file, and a categorical variable in a CSV formatted text file.

14 . The CPP of claim 8 , wherein the first knowledge base is an open and editable knowledge base accessible on the Internet.

15 . A computer system (CS), the CS comprising:

a processor set;

a set of one or more computer-readable storage media;

program instructions, collectively stored in the set of one or more storage media, for causing the processor set to perform the following computer operations:

cause a first search to be performed on a first knowledge base for a first descriptive name;

extract sentences from results of the first search,

wherein the extracted sentences include labels associated with the first descriptive name and descriptions of the labels;

run at least one predetermined deep learning model on the results of the first search for determining similarity scores for the extracted sentences, wherein each of the similarity scores defines a similarity score for an associated one of the extracted sentences and a first glossary;

determine a first of the determined similarity scores, wherein the first determined similarity score has a relatively greater similarity score than the other determined similarity scores; and

use the first determined similarity score to map terms of a second glossary to terms of the first glossary to enhance context provided in search results generated using the glossaries.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: TAKAHASHI, TOSHIHIRO; TATEISHI, TAKAAKI
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 067191/0774 →
Continuity (1)
Related Publication 20250328564A1 · Oct 23, 2025
References Cited (33)
US 10198491B1 · Semturs et al. · 2019 [cited by applicant]
US 11366966B1 · Ramsey · 2022 [cited by examiner]
US 11487708B1 · Dangi · 2022 [cited by examiner]
US 11769015B2 · Travis · 2023 [cited by examiner]
US 20120239677A1 · Neale · 2012 [cited by examiner]
US 20150186361A1 · Su · 2015 [cited by examiner]
US 20190295158A1 · Wu · 2019 [cited by examiner]
US 20200134757A1 · Raphael · 2020 [cited by examiner]
US 20210117617A1 · Blaya · 2021 [cited by examiner]
US 20220020288A1 · Naber · 2022 [cited by examiner]
US 20220230227A1 · Fan · 2022 [cited by examiner]
US 20230004729A1 · Cushman, II · 2023 [cited by examiner]
US 20240046038A1 · Yoshida · 2024 [cited by examiner]
US 20250095351A1 · Shin · 2025 [cited by examiner]
US 20250217388A1 · Melbouci · 2025 [cited by examiner]
US 20250328564A1 · Takahashi · 2025 [cited by examiner]
US 20260023783A1 · Saha · 2026 [cited by examiner]
CN 109271626B · 2023 [cited by applicant]
Modeling Score Distributions for Combining the Outputs of Search Engines (Year: 2001). [cited by examiner]
A Learning-Based Approach for Automatic Construction of Domain Glossary from Source Code and Documentation (Year: 1998). [cited by examiner]
Matching of Descriptive Labels to Glossary Descriptions (Year: 2023). [cited by examiner]
Takahashi et al., “Matching of Descriptive Labels to Glossary Descriptions,” arXiv, 2023, 10 pages, retrieved from https://arxiv.org/abs/2310.18385. [cited by applicant]
Tsekouras et al., “A Graph-based Text Similarity Measure That Employs Named Entity Information,” Proceedings of Recent Advances in Natural Language Processing, Sep. 2017, pp. 765-771. [cited by applicant]
Lobo et al., “Matching Table Metadata with Business Glossaries Using Large Language Models,” arXiv, 2023, 13 pages, retrieved from https://arxiv.org/abs/2309.11506. [cited by applicant]
Anonymous, “Semantic Text Similarity Using Name Enrichment for Glossary Maintenance,” Proceedings of 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2023, pp. 1-9. [cited by applicant]
Jiang et al., “Improved Universal Sentence Embeddings with Prompt-based Contrastive Learning and Energy-based Learning,” arXiv, 2022, 15 pages, retrieved from https://paperswithcode.com/paper/deep-continuous-prompt-for-… [cited by applicant]
Wikidata, “Wikidata:Introduction,” Wikidata, Jan. 2024, 3 pages, retrieved from https://www.wikidata.org/wiki/Wikidata:Introduction. [cited by applicant]
Wikipedia, “DBpedia,” Wikipedia, 2024, 7 pages, retrieved from https://en.wikipedia.org/wiki/Dbpedia. [cited by applicant]
Coursera, “What Is Kaggle and What Is It Used For?” Coursera, Nov. 29, 2023, 11 pages, retrieved from https://www.coursera.org/articles/kaggle. [cited by applicant]
Ontotext, “What is a Knowledge Graph?” Ontotext Knowledge Hub, 2024, 8 pages, retrieved from https://www.ontotext.com/knowledgehub/fundamentals/what-is-a-knowledge-graph/. [cited by applicant]
Jiang, Y., “PromCSE: Improved Universal Sentence Embeddings with Prompt-based Contrastive Learning and Energy-based Learning,” GitHub, 2022, 7 pages, retrieved from https://github.com/YJiangcm/PromCSE?tab=readme-ov-file… [cited by applicant]
Wikipedia, “Semantics,” Wikipedia, 2024, 38 pages, retrieved from https://en.wikipedia.org/wiki/Semantics. [cited by applicant]
Indoc.Pro, “Create a well crafted glossary for software documentation,” indoc.pro, 2023, 4 pages, retrieved from https://indoc.pro/documentation-types/glossary/#What_is_a_glossary. [cited by applicant]