IP Library Granted Patent US 12,411,825
Granted Patent B2
US 12,411,825 · App. 18/537,879 · Granted Sep 9, 2025

Expanding database column names

Inventors: Boris Rozenberg (Ramat Gan, IL); Yehoshua Sagron (Haifa, IL); Ariel Farkash (Shimshit, IL); Igor Gokhman (Haifa, IL); Micha Gideon Moffie (Zichron Yaakov, IL)
Assignee: International Business Machines Corporation
G06F16/221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,411,825
App. No.
18/537,879
Granted
Sep 9, 2025
Kind
B2
Abstract

A method, a structure, and a computer system for expanding database column names. Exemplary embodiments may include identifying glossary terms that sufficiently syntactically match a column name, extracting partial expansions from the glossary terms, selectively combining a set of the partial expansions, and storing the combined set of partial expansions in association with the column name.

Claims (46)

1. A computer-implemented method for expanding database column names, the method comprising:

identifying glossary terms that sufficiently syntactically match a column name based on dividing a number of matching tokens between at least one glossary term of the glossary terms and the column name by a number of tokens within the column name;

extracting partial expansions from the at least one glossary term;

selectively combining one or more of the partial expansions into a set of the partial expansions; and

storing the combined set of partial expansions in association with the column name.

2. The computer-implemented method of claim 1 , wherein the selectively combining the set of partial expansions further comprises:

scoring candidate combined sets of the partial expansions based on a matching score, a fragmentation score, and a conflict score; and

selectively combining at least one of the candidate combined sets of partial expansions having a best scoring.

3. The computer-implemented method of claim 2 , wherein the matching score is based on an approximate string match between the column name and the candidate combined sets of the partial expansions.

4. The computer-implemented method of claim 2 , wherein the fragmentation score is a count of the partial expansions used to create each set of the candidate combined sets of the partial expansions.

5. The computer-implemented method of claim 2 , wherein the conflict score is a measure of conflicting expansions within each of the candidate combined sets of the partial expansions.

6. The computer-implemented method of claim 2 , wherein the scoring the candidate combined sets of the partial expansions is further based on at least one of:

the candidate combined sets of the partial expansions sufficiently matching at least one of a name of a corresponding table and another column name within the table; and

the candidate combined sets of the partial expansions having a maximum diversity.

7. The computer-implemented method of claim 1 , wherein the identifying the glossary terms that sufficiently syntactically match the column name is based on an input relation score.

8. A computer program product for expanding database column names, the computer program product comprising:

one or more non-transitory computer-readable storage media and program instructions stored on the one or more non-transitory computer-readable storage media capable of performing a method, the method comprising:

identifying glossary terms that sufficiently syntactically match a column name based on dividing a number of matching tokens between at least one glossary term of the glossary terms and the column name by a number of tokens within the column name;

extracting partial expansions from the at least one glossary term;

selectively combining one or more of the partial expansions into a set of the partial expansions; and

storing the combined set of partial expansions in association with the column name.

9. The computer program product of claim 8 , wherein the selectively combining the set of partial expansions further comprises:

scoring candidate combined sets of the partial expansions based on a matching score, a fragmentation score, and a conflict score; and

selectively combining at least one of the candidate combined sets of partial expansions having a best scoring.

10. The computer program product of claim 9 , wherein the matching score is based on an approximate string match between the column name and the candidate combined sets of the partial expansions.

11. The computer program product of claim 9 , wherein the fragmentation score is a count of the partial expansions used to create each set of the candidate combined sets of the partial expansions.

12. The computer program product of claim 9 , wherein the conflict score is a measure of conflicting expansions within each of the candidate combined sets of the partial expansions.

13. The computer program product of claim 9 , wherein the scoring the candidate combined sets of the partial expansions is further based on at least one of:

the candidate combined sets of the partial expansions sufficiently matching at least one of a name of a corresponding table and another column name within the table; and

the candidate combined sets of the partial expansions having a maximum diversity.

14. The computer program product of claim 8 , wherein the identifying the glossary terms that sufficiently syntactically match the column name is based on an input relation score.

15. A computer system for expanding database column names, the system comprising:

one or more computer processors, one or more computer-readable storage media, and program instructions stored on the one or more of the computer-readable storage media for execution by at least one of the one or more processors capable of performing a method, the method comprising:

identifying glossary terms that sufficiently syntactically match a column name based on dividing a number of matching tokens between at least one glossary term of the glossary terms and the column name by a number of tokens within the column name;

extracting partial expansions from the at least one glossary term;

selectively combining one or more of the partial expansions into a set of the partial expansions; and

storing the combined set of partial expansions in association with the column name.

16. The computer system of claim 15 , wherein the selectively combining the set of partial expansions further comprises:

scoring candidate combined sets of the partial expansions based on a matching score, a fragmentation score, and a conflict score; and

selectively combining at least one of the candidate combined sets of partial expansions having a best scoring.

17. The computer system of claim 16 , wherein the matching score is based on an approximate string match between the column name and the candidate combined sets of the partial expansions.

18. The computer system of claim 16 , wherein the fragmentation score is a count of the partial expansions used to create each set of the candidate combined sets of the partial expansions.

19. The computer system of claim 16 , wherein the conflict score is a measure of conflicting expansions within each of the candidate combined sets of the partial expansions.

20. The computer system of claim 16 , wherein the scoring the candidate combined sets of the partial expansions is further based on at least one of:

the candidate combined sets of the partial expansions sufficiently matching at least one of a name of a corresponding table and another column name within the table; and

the candidate combined sets of the partial expansions having a maximum diversity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2023
From: ROZENBERG, BORIS; SAGRON, YEHOSHUA; FARKASH, ARIEL; GOKHMAN, IGOR; MOFFIE, MICHA GIDEON
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 065852/0659 →
Continuity (1)
Related Publication 20250200017A1 · Jun 19, 2025
References Cited (23)
US 7809715B2 · Wei · 2010 [cited by applicant]
US 9020805B2 · Boguraev · 2015 [cited by examiner]
US 9031832B2 · Boguraev · 2015 [cited by applicant]
US 9390081B2 · Anders · 2016 [cited by applicant]
US 10650808B1 · Platt · 2020 [cited by examiner]
US 10929455B2 · Jetley · 2021 [cited by applicant]
US 11068653B2 · Gahlot · 2021 [cited by applicant]
US 11152084B2 · Kondadadi · 2021 [cited by applicant]
US 11461687B1 · Shaikh · 2022 [cited by examiner]
US 20080086297A1 · Li · 2008 [cited by examiner]
US 20120084076A1 · Boguraev · 2012 [cited by examiner]
US 20160140116A1 · Li · 2016 [cited by examiner]
US 20180210821A1 · Raghavan · 2018 [cited by examiner]
US 20200409999A1 · Chen · 2020 [cited by examiner]
US 20210303786A1 · Veyseh · 2021 [cited by applicant]
US 20220350810A1 · Majumdar · 2022 [cited by examiner]
US 20230087421A1 · Chikoti · 2023 [cited by applicant]
US 20230169050A1 · Chandrahasan · 2023 [cited by examiner]
US 20230214275A1 · Abhyankar · 2023 [cited by examiner]
CN 114925698A · 2022 [cited by applicant]
CN 115293168A · 2022 [cited by applicant]
Zhang et al. “NameGuess: Column Name Expansion for Tabular Data”. Oct. 7, 2023. <https://openreview.net/forum?id=AfEowGM3qG> (Year: 2023). [cited by examiner]
Ifergan et al., “Fuzzy Matching of Obscure Texts With Meaningful Terms Included in a Glossary”, U.S. Appl. No. 17/892,169, filed Aug. 22, 2022, 34 pages. (Specs + Drawings). [cited by applicant]