IP Library Granted Patent US 12,524,629
Granted Patent B2
US 12,524,629 · App. 17/364,842 · Granted Jan 13, 2026

Universal data language translator

Inventors: James B. Cushman, II (Longboat Key, FL); Aurko Joshi (West Bloomfield, MI); Satyender Goel (Chicago, IL)
Assignee: Collibra Belgium BV
G06F40/47G06F40/242G06F40/49G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,629
App. No.
17/364,842
Granted
Jan 13, 2026
Kind
B2
Abstract

The present disclosure is directed to a universal data language (UDL) translator. Specifically, the systems and methods disclosed enable input data from a variety of sources to be translated into a UDL that can be consistently analyzed and compared against other sources of data. For example, an entity may upload input data that has a plurality of data terms and definitions (e.g., header column in a spreadsheet). These terms may be duplicative and/or inaccurate with respect to the underlying data. If the entity wishes to compare and transact data within a data marketplace, the entity may not fully comprehend what data it is missing and/or what data another entity may have to offer for trade. To remedy this problem of business semantic management, the present invention discloses steps for creating a UDL and a UDL translator so that any input data can be translated to UDL.

Claims (44)

1 . A system for translating input data into a universal data language, comprising:

a memory configured to store non-transitory computer readable instructions; and

a processor communicatively coupled to the memory, wherein the processor, when

executing the non-transitory computer readable instructions, is configured to:

receive input data from at least one trusted source, wherein the input data comprises at least one data term;

compare the at least one data term to at least one universal data language (UDL) library, wherein the at least one UDL library comprises a UDL entry for the at least one data term, wherein the at least one UDL library is comprised of a plurality of definitions from the at least one trusted source;

train at least one machine learning algorithm on the at least one UDL library;

process the at least one data term using the at least one machine learning algorithm to calculate a similarity score based on i) a similarity of context characteristics between the at least one data term and at least one UDL term in the at least one UDL library, and ii) a source of the at least one data term; and

if the similarity score is above a similarity threshold, translate the at least one data term to the at least one UDL term in the at least one UDL library.

2 . The system of claim 1 , wherein the processor is further configured to:

determine if the at least one data term is a duplicate of an already-existing UDL term in the at least one UDL library; and

based on the at least one data term being determined to be the duplicate, map the at least one data term to the already-exiting UDL term.

3 . The system of claim 2 , wherein determining if the at least one data term is the duplicate of the already-existing UDL term comprises comparing each word or each character in the at least one data term with each word or each character with the already-existing UDL term.

4 . The system of claim 2 , wherein determining if the at least one data term is the duplicate of the already-existing UDL term comprises comparing a definition of the at least one data term to a definition to the already-existing UDL term.

5 . The system of claim 4 , wherein the definition of the at least one data term is at least one equation.

6 . The system of claim 1 , wherein the at least one trusted source is at least one of: a publisher of a business glossary, a publisher of a dictionary, and a publisher of industry-specific ontology.

7 . The system of claim 1 , wherein the processor is further configured to: format the at least one data term to conform to a formatting for the at least one UDL term.

8 . The system of claim 1 , wherein processing the at least one data term further comprises evaluating a level of accuracy of the at least one data term.

9 . The system of claim 8 , wherein the level of accuracy is determined based on at least one of: an inactivity score and an outdatedness score.

10 . A method of creating a universal data language (UDL) translator, comprising:

receiving input data from at least one trusted source, wherein the input data comprises a plurality of data terms from at least one business glossary;

analyzing lexical features of each data term in the plurality of data terms;

analyzing contextual features of each data term in the plurality of data terms;

extracting at least one semantic ontology for each data term in the plurality of data terms;

creating a UDL library, wherein the UDL library comprises a UDL entry for each data term in the plurality of data terms;

training at least one machine learning algorithm on the UDL library;

receiving client-specific input data, wherein the client-specific input data comprises a newly-received data term;

processing the newly-received data term using the at least one machine learning algorithm by calculating a similarity score based on i) a similarity of context characteristics between the newly-received data term and at least one UDL term in the UDL library, and ii) a source of the newly-received data term; and

mapping the newly-received data term to at least one UDL term in the UDL library.

11 . The method of claim 10 , wherein the at least one trusted source is at least one of: a publisher of a business glossary, a publisher of an industry-specific glossary, and a publisher of a business ontology.

12 . The method of claim 10 , wherein the similarity score is calculated by comparing each character or each word in the newly-received data term with each character or each word in the at least one UDL term.

13 . The method of claim 10 , wherein the similarity score is calculated by comparing a definition of the newly-received data term to a definition of the at least one UDL term, wherein the definition of the newly-received data term is an equation.

14 . The method of claim 10 , wherein the similarity score is calculated by comparing at least one domain classification of the newly-received data term with at least one domain classification of the at least one UDL term.

15 . The method of claim 10 , further comprising determining if the newly-received data term is a duplicate of a preexisting UDL term.

16 . The method of claim 15 , wherein determining if the newly-received data term is a duplicate of a preexisting UDL term comprises evaluating the similarity score.

17 . The method of claim 10 , wherein the UDL library is comprised of at least one general business ontology, at least one financial ontology, and at least one life sciences ontology.

18 . A non-transitory computer-readable media storing computer executable instructions that when executed cause a computer system to perform a method for translating input data into a universal data language (UDL), comprising:

receiving input data from at least one trusted source, wherein the input data comprises at least one data term;

comparing the at least one data term to at least one universal data language (UDL) library, wherein the at least one UDL library comprises a UDL entry for the at least one data term;

training at least one machine learning algorithm on the at least one UDL library;

processing the at least one data term using the at least one machine learning algorithm to calculate a similarity score based on i) a similarity of context characteristics between the at least one data term and at least one UDL term in the at least one UDL library, and ii) a source of the at least one data term; and

if the similarity score is above a similarity threshold, translating the at least one data term to the at least one UDL term in the at least one UDL library;

determine if the at least one data term is a duplicate of an already-existing UDL term in the at least one UDL library; and

based on the at least one data term being determined to be the duplicate, map the at least one data term to the already-exiting UDL term.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 15, 2023
From: COLLIBRA NV; CNV NEWCO B.V.; COLLIBRA B.V.
To: COLLIBRA BELGIUM BV
Reel/Frame 062989/0023 →
SECURITY INTEREST Recorded Jan 4, 2023
From: COLLIBRA BELGIUM BV
To: SILICON VALLEY BANK UK LIMITED
Reel/Frame 062269/0955 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2021
From: CUSHMAN, JAMES B., II; JOSHI, AURKO; GOEL, SATYENDER
To: COLLIBRA NV
Reel/Frame 056843/0426 →
Continuity (1)
Related Publication 20230004729A1 · Jan 5, 2023
References Cited (5)
US 9390132B1 · Kapoor · 2016 [cited by examiner]
US 20060106824A1 · Stuhec · 2006 [cited by examiner]
US 20180032497A1 · Mukherjee · 2018 [cited by examiner]
WO WO2007029240A2 · 2007 [cited by examiner]
International Search Report and Written Opinion of the International Search Authority for International Application No. PCT/EP2022/068963 dated Oct. 7, 2022, 14 pages. [cited by applicant]