IP Library Granted Patent US 12,561,613
Granted Patent B2
US 12,561,613 · App. 18/051,900 · Granted Feb 24, 2026

Data augmentation using semantic transforms

Inventors: Horst Cornelius Samulowitz (Yorktown Heights, NY); Udayan Khurana (Yorktown Heights, NY); Kavitha Srinivas (Yorktown Heights, NY); Takaaki Tateishi (Tokyo, JP); Ibrahim Abdelaziz (Yorktown Heights, NY); Julian Timothy Dolby (Yorktown Heights, NY)
Assignee: International Business Machines Corporation
G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,613
App. No.
18/051,900
Granted
Feb 24, 2026
Kind
B2
Abstract

A method of data augmentation includes receiving, by a processor, a set of data including a plurality of variables, mapping each variable to one or more target concepts associated with a name of each variable, and acquiring a set of semantic transforms, each semantic transform including a function applied to one or more concepts mapped to a respective variable. The method also includes comparing the one or more target concepts to the one or more concepts of each semantic transform, selecting at least one semantic transform based on the comparing, generating an expression for each selected semantic transform, each expression configured to apply a function of a selected semantic transform to at least one of the plurality of variables, and augmenting the set of data for use in an application by adding each expression to the set of data.

Claims (40)

1 . A method of data augmentation, comprising:

receiving, by a processor, a set of data including a plurality of variables;

mapping each variable to one or more target concepts associated with a name of each variable;

acquiring a set of semantic transforms, each semantic transform including a function applied to one or more concepts mapped to a respective variable;

comparing the one or more target concepts to the one or more concepts of each semantic transform, and selecting at least one semantic transform based on the comparing;

generating an expression for each selected semantic transform, each expression configured to apply a function of a selected semantic transform to at least one of the plurality of variables; and

augmenting the set of data for use in an application by adding each expression to the set of data.

2 . The method of claim 1 , wherein the application includes a machine learning application.

3 . The method of claim 1 , wherein the set of semantic transforms is stored in a transform repository, each semantic transform including a function indexed to one or more concepts.

4 . The method of claim 3 , wherein acquiring the set of semantic transforms includes selecting the set of transforms from the repository based on a level of similarity between a stored semantic transform and a target concept.

5 . The method of claim 1 , wherein the set of data includes tabular data having a column associated with each variable, and augmenting the set of data includes creating a new column for each expression.

6 . The method of claim 5 , wherein the new column is indexed based on the name of at least one variable and includes values calculated based on the expression.

7 . The method of claim 1 , wherein generating the expression includes replacing each concept in a selected semantic transform with a variable name associated with the concept.

8 . A system comprising:

a memory device; and

one or more processing units coupled with the memory device, the one or more processing units are configured to perform a method of data augmentation, the method comprising:

receiving a set of data including a plurality of variables;

mapping each variable to one or more target concepts associated with a name of each variable;

acquiring a set of semantic transforms, each semantic transform including a function applied to one or more concepts mapped to a respective variable;

comparing the one or more target concepts to the one or more concepts of each semantic transform, and selecting at least one semantic transform based on the comparing;

generating an expression for each selected semantic transform, each expression configured to apply a function of a selected semantic transform to at least one of the plurality of variables; and

augmenting the set of data for use in an application by adding each expression to the set of data.

9 . The system of claim 8 , wherein the set of semantic transforms is stored in a transform repository, each semantic transform including a function indexed to one or more concepts.

10 . The system of claim 9 , wherein acquiring the set of semantic transforms includes selecting the set of transforms from the repository based on a level of similarity between a stored semantic transform and a target concept.

11 . The system of claim 8 , wherein the set of data includes tabular data having a column associated with each variable.

12 . The system of claim 11 , wherein augmenting the set of data includes creating a new column for each expression.

13 . The system of claim 12 , wherein the new column is indexed based on the name of at least one variable and includes values calculated based on the expression.

14 . The system of claim 8 , wherein generating the expression includes replacing each concept in a selected semantic transform with a variable name associated with the concept.

15 . A computer program product comprising a computer-readable memory that has computer-executable instructions stored thereupon, the computer-executable instructions when executed by a processor cause the processor to perform operations comprising:

receiving a set of data including a plurality of variables;

mapping each variable to one or more target concepts associated with a name of each variable;

acquiring a set of semantic transforms, each semantic transform including a function applied to one or more concepts mapped to a respective variable;

comparing the one or more target concepts to the one or more concepts of each semantic transform, and selecting at least one semantic transform based on the comparing;

generating an expression for each selected semantic transform, each expression configured to apply a function of a selected semantic transform to at least one of the plurality of variables; and

augmenting the set of data for use in an application by adding each expression to the set of data.

16 . The computer program product of claim 15 , wherein the set of semantic transforms is stored in a transform repository, each semantic transform including a function indexed to one or more concepts.

17 . The method of claim 16 , wherein acquiring the set of semantic transforms includes selecting the set of transforms from the repository based on a level of similarity between a stored semantic transform and a target concept.

18 . The computer program product of claim 15 , wherein the set of data includes tabular data having a column associated with each variable.

19 . The computer program product of claim 18 , wherein augmenting the set of data includes creating a new column for each expression, and the new column is indexed based on the name of at least one variable and includes values calculated based on the expression.

20 . The computer program product of claim 15 , wherein generating the expression includes replacing each concept in a selected semantic transform with a variable name associated with the concept.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2022
From: SAMULOWITZ, HORST CORNELIUS; KHURANA, UDAYAN; SRINIVAS, KAVITHA; TATEISHI, TAKAAKI; ABDELAZIZ, IBRAHIM; DOLBY, JULIAN TIMOTHY
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061626/0822 →
Continuity (1)
Related Publication 20240144084A1 · May 2, 2024
References Cited (112)
US 6263339B1 · Hirsch · 2001 [cited by applicant]
US 6741974B1 · Harrison et al. · 2004 [cited by applicant]
US 6965902B1 · Ghatate · 2005 [cited by applicant]
US 7530054B2 · Reimer et al. · 2009 [cited by applicant]
US 7703085B2 · Poznanovic et al. · 2010 [cited by applicant]
US 7984422B2 · Graham · 2011 [cited by applicant]
US 8843884B1 · Koerner · 2014 [cited by applicant]
US 9639335B2 · Hoban et al. · 2017 [cited by applicant]
US 9658839B2 · Hale et al. · 2017 [cited by applicant]
US 9959326B2 · Duan et al. · 2018 [cited by applicant]
US 10073763B1 · Raman et al. · 2018 [cited by applicant]
US 10229200B2 · Bornea et al. · 2019 [cited by applicant]
US 10303448B2 · Steven et al. · 2019 [cited by applicant]
US 10402175B2 · Mcfarland · 2019 [cited by applicant]
US 10606885B2 · Brundage et al. · 2020 [cited by applicant]
US 11003994B2 · Liang et al. · 2021 [cited by applicant]
US 11954424B2 · Samulowitz et al. · 2024 [cited by applicant]
US 12124822B2 · Dolby et al. · 2024 [cited by applicant]
US 20020062463A1 · Hines · 2002 [cited by applicant]
US 20060294499A1 · Shim · 2006 [cited by applicant]
US 20070088697A1 · Charlebois · 2007 [cited by examiner]
US 20090119095A1 · Beggelman · 2009 [cited by examiner]
US 20090234640A1 · Boegl · 2009 [cited by examiner]
US 20100175049A1 · Ramsey et al. · 2010 [cited by applicant]
US 20100287214A1 · Narasayya et al. · 2010 [cited by applicant]
US 20110202559A1 · Todd · 2011 [cited by applicant]
US 20120086547A1 · Foster et al. · 2012 [cited by applicant]
US 20120233188A1 · Majumdar · 2012 [cited by examiner]
US 20130086547A1 · Said et al. · 2013 [cited by applicant]
US 20160315960A1 · Teilhet et al. · 2016 [cited by applicant]
US 20170109933A1 · Voorhees et al. · 2017 [cited by applicant]
US 20170221153A1 · Albright · 2017 [cited by applicant]
US 20170255536A1 · Weissinger et al. · 2017 [cited by applicant]
US 20190005163A1 · Farrell et al. · 2019 [cited by applicant]
US 20190213185A1 · Arroyo · 2019 [cited by examiner]
US 20200110746A1 · Lecue et al. · 2020 [cited by applicant]
US 20200143243A1 · Liang et al. · 2020 [cited by applicant]
US 20200175163A1 · Hassanshahi et al. · 2020 [cited by applicant]
US 20200210478A1 · Wada et al. · 2020 [cited by applicant]
US 20200233889A1 · Nassar · 2020 [cited by applicant]
US 20210064672A1 · Mahadi et al. · 2021 [cited by applicant]
US 20210173641A1 · Dolby et al. · 2021 [cited by applicant]
US 20210326312A1 · White · 2021 [cited by applicant]
US 20210342723A1 · Rao · 2021 [cited by applicant]
US 20230401467A1 · Ferrucci · 2023 [cited by examiner]
CN 108334321A · 2018 [cited by applicant]
CN 110134848A · 2019 [cited by applicant]
CN 110362596A · 2019 [cited by applicant]
CN 111091883A · 2020 [cited by applicant]
CN 111353005A · 2020 [cited by applicant]
CN 112287679A · 2021 [cited by applicant]
KR 101505546B1 · 2017 [cited by applicant]
KR 101762670B1 · 2017 [cited by applicant]
White, et al, “Deep Learning Code Fragments for Code Clone Detection”, In Proceedings of the 31st IEEE/ACM International Conference on Automated Software Engineering, ACM, 2016, 12 pages. [cited by applicant]
Agesen, et al, “Type Inference of SELF Analysis of Objects with Dynamic and Multiple Inheritance”, Journal of Software: Practice & Experience, vol. 25, 1995, 26 pages. [cited by applicant]
Allamanis, et al, “A Survey of Machine Learning for Big Code and Naturalness”, Abstract Only, ACM Computing Surveys (CSUR), vol. 51, https://doi.org/10.1145/3212695, (Retrieved: Jan. 22, 2020), 2018, 5 pages. [cited by applicant]
Allamanis, et al, “Learning to Represent Programs with Graphs”, arXiv preprint, 2018, 17 pages. [cited by applicant]
Alon, et al, “A General Path-Based Representation for Predicting Program Properties”, arXiv preprint, 2018, 16 pages. [cited by applicant]
Alon, et al, “CODE2SEQ: Generating Sequences from Structured Representations of Code”, arXiv preprint, 2019, 22 pages. [cited by applicant]
Bichsel, et al, “Statistical Deobfuscation of Android Applications”, In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security (CCS), 2016, 13 pages. [cited by applicant]
Bruch, et al, “Learning from Examples to Improve Code Completion Systems”, In Proceedings of the 7th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on The Foundations of Soft… [cited by applicant]
Chae, et al., “Automatically Generating Features for Learning Program Analysis Heuristics for C-Like Languages”, Proc. ACM Program. Lang., vol. 1, 2017, 25 pages. [cited by applicant]
Chambers, et al, “Iterative Type Analysis and Extended Message Splitting: Optimizing Dynamically-Typed Object-Oriented Programs”, In Proceedings of the ACM SIGPLAN 1990 Conference on Programming Language Design and Impl… [cited by applicant]
Dolby et al., “Mining Code Expressions for Data Analysis,” U.S. Appl. No. 17/895,881, filed Aug. 25, 2022. [cited by applicant]
Feldthaus, et al, “Efficient Construction of Approximate Call Graphs for JavaScript IDE Services”, In 35th International Conference on Software Engineering (ICSE), 2013, 10 pages. [cited by applicant]
Fernandes, et al., “Structured Neural Summarization”, arXiv preprint, 2019, 18 pages. [cited by applicant]
Feurer, et al, “Using Meta-Learning to Initialize Bayesian Optimization of Hyperparameters”, In Proceedings of the 2014 International Conference on Meta-learning and Algorithm Selection, vol. 1201, 2014, 8 pages. [cited by applicant]
Hsiao, et al, “Reducing MapReduce Abstraction Costs for Text-Centric Applications”, In 43rd International Conference on Parallel Processing (ICPP), 2014, 10 pages. [cited by applicant]
Hu, et al, “CodeSum: Translate Program Language to Natural Language”, arXiv preprint, http://www.arxiv-vanity.com/papers/1708.01837/, (Retrieved: Nov. 11, 2022), 19 pages. [cited by applicant]
Li, et al, “SySeVR: A Framework for Using Deep Learning to Detect Software Vulnerabilities”, arXiv preprint, 2018, 13 pages. [cited by applicant]
List of IBM Patents or Patent Applications Treated as Related; Date Filed: Feb. 16, 2023, 2 pages. [cited by applicant]
Nguyen, et al, “Graph-based Mining of Multiple Object Usage Patterns”, Abstract Only, In Proceedings of the 7th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on the Foundati… [cited by applicant]
Nguyen, et al., “Graph-based Statistical Language Model for Code,” In Proceedings of the 37th International Conference on Software Engineering (ICSE), vol. 1, 2015, 11 pgs. [cited by applicant]
Olson, et al., “Evaluation of a Tree-based Pipeline Optimization Tool for Automating Data Science”, arXiv preprint, 2016, 8 pgs. [cited by applicant]
Proksch, et al., “Intelligent Code Completion with Bayesian Networks”, Abstract Only, ACM Transactions on Software Engineering and Methodology (TOSEM), https://doi.org/10.1145/2744200, (Retrieved: Jan. 22, 2020), 2015, … [cited by applicant]
Samulowitz et al., “Automatic Domain Annotation of Structured Data,” U.S. Appl. No. 17/661,619,238, filed May 2, 2022. [cited by applicant]
Shivers, “Control-Flow Analysis of Higher-Order Languages”, Carnegie Mellon University, School of Computer Science, 1991, 200 pages. [cited by applicant]
“H20 AI Feature Store,” Downloaded from the internet on Oct. 26, 2002; from h20.ai/platform/ai-cloud/make/feature-store/; 5 pages. [cited by applicant]
Cambronero et al., “wranglesearch: Mining Data: Wrangling Functions from Python Programs [Under submission],” May 21, 2022; pp. 1-9. [cited by applicant]
Chen et al., Learning Semantic Annotations for Tabular Data. IJCAI 2019; 7 pages. [cited by applicant]
Chen et al., Colnet: Embedding the semantics of web tables for column type prediction. AAAI 2019; 8 pages. [cited by applicant]
Cremaschi et al. MantisTable: an Automatic Approach for the Semantic Table Interpretation. In SemTab@ ISWC. 2019; pp. 15-24. [cited by applicant]
Galhotra et al.; “Semantic Search over Structured Data.” ACM CIKM; 2020; 4 pages. [cited by applicant]
Galhotra et al; “Automated Feature Enhancement for Predictive Modeling using External Knowledge;” IEEE ICDM; 2019; 4 pages. [cited by applicant]
Grigoriu, et al., “SIENA: Semi-automatic Semantic Enhancement of Datasets Using Concept Recognition”, Journal of Biomedical Semantics, vol. 12, Article No. 5, Mar. 24, 2021, 12 pgs., <https://doi.org/10.1186/s13326-021-… [cited by applicant]
Huynh, et al., “Dagobah: Enhanced Scoring Algorithms for Scalable Annotations of Tabular Data”, SEMTAB 2020, Semantic Web Challenge on Tabular Data to Knowledge Graph Matching (SemTab 2020), co-located with the 19th Int… [cited by applicant]
Jiménez-Ruiz et al; “SemTab2019: Resources to Benchmark Tabular Data to Knowledge Graph Matching Systems;”European Semantic Web Conference. Springer 2020; pp. 514-530. [cited by applicant]
Kanter et al., “Deep feature synthesis: Towards automating data science endeavors.” 2015 IEEE international conference on data science and advanced analytics (DSAA). IEEE, 2015; 10 pages. [cited by applicant]
Katz et al., “Explorekit: Automatic feature generation and selection.” 2016 IEEE 16th International Conference on Data Mining (ICDM). IEEE, 2016; 6 pages. [cited by applicant]
Khurana et al.; “Semantic Annotation for Tabular Data.” ACM CIKM, 2021; 9 pages. [cited by applicant]
Khurana et al.; Feature Engineering for Predictive Modeling Using Reinforcement Learning; AAAI 2018; 8 pages. [cited by applicant]
Khurana, et al., “Semantic Annotation for Tabular Data”, Dec. 15, 2020, 9 pgs., DOI: 10.48550/arxiv.2012.08594. [cited by applicant]
Limaye et al.; Annotating and searching web tables using entities, types and relationships; VLDB 2010; 10 pages. [cited by applicant]
M Hulsebos et all. Sherlock: A Deep Learning Approach to Semantic Data Type Detection. ACM SIGKDD 2019; 9 pages. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, National Institute of Standards and Technology, Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Morikawa; 2019. Semantic Table Interpretation using LOD4ALL.. In SemTab@ ISWC. 49-56. [cited by applicant]
Namaki et al.; “Vamsa: Automated Provenance Tracking in Data Scient Scripts;” KDD 2020; Aug. 23-27, Virtual Event USA; 10 pages. [cited by applicant]
Neumaier et al; “Multi-level semantic labelling of numerical values;” In ICWS; 2016; 16 pages. [cited by applicant]
Nguyen et al; MTab: matching tabular data to knowledge graph using probability models arXiv preprint arXiv:1910.00246 (2019); 8 pages. [cited by applicant]
Oliveira et al.; ADOG-Annotating Data with Ontologies and Graphs.. In SemTab@ ISWC. 2019; pp. 6. [cited by applicant]
Ota et al.; Data Driven Domain Discovery for Structured Datasets. PVLDB (2020), pp. 953-965. [cited by applicant]
Ritze et al.; Matching html tables to dbpedia; In Proceedings of the 5th International Conference on Web Intelligence, Mining and Semantics, 2015; 6 pages. [cited by applicant]
Song; Autofe: efficient and robust automated feature engineering; Massachusetts Institute of Technology, 2018; 61 pages. [cited by applicant]
Srinivas et al.; Semantic Feature Discovery with Code Mining and Semantic Type Detection; AAAI; 2022; 3 pages. [cited by applicant]
Suhara, et al., “Annotating Columns with Pre-trained Language Models”, Mar. 1, 2022, 15 pgs., arXiv:2104.01785. [cited by applicant]
Yan et al; “Synthesising Type-Detection Logic for Rich Semantic Data Types using Open-Source Code;” SIGMOD 18;Dated Jun. 10-15, 2018; 16 pages. [cited by applicant]
Yu et al.; “Deep Code Curator—Technical Report on Code2Graph,” Apr. 2019, CECS Technical Report; Uniiverty of California, Irvine, 33 pages; 2019. [cited by applicant]
Zhang et al.; “Sato: Contextual Semantic Type Detection in Tables;” VLDB 2020; 14 pages. [cited by applicant]
Uren, Victoria, et al. “Semantic annotation for knowledge management: Requirements and a survey of the state of the art.” Journal of Web Semantics 4.1 (2006): 14-28 (Year: 2006). [cited by applicant]
Karim Ali et al., “A Study of Call Graph Construction for JVM-Hosted Languages”, Dec. 2021 IEEE, [Retrieved on Apr. 23, 2024], Retrieved from the internet: <URL: https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=… [cited by applicant]
Khurana et al., “Semantic concept annotation for tabular data.” Proceedings of the 30th ACM International Conference on Information & Knowledge Management. 2021 (Year: 2021), 10 pages. [cited by applicant]
Taheriyan, Mohsen, et al. “Learning the semantics of structured data sources.” Journal of Web Semantics 37 (2016): 152-169 (Year:2016). [cited by applicant]