IP Library › Granted Patent US 12,748,734
Granted Patent B2
US 12,748,734 · App. 18/331,530 · Granted Sep 29, 2026

Canonical transformations using machine learning language model

Inventors: Sanjay Kumar Singh (Bengaluru, IN); Subhasis Jethy (Bangalore, IN); Udit Saini (Uttarakhand, IN); Carlos W. Morato (Sammamish, WA); Rahul Bhotika (Bellevue, WA); Ranju Das (Seattle, WA); Vasant Manohar (Apex, NC)
Assignee: UNITEDHEALTH GROUP INCORPORATED
G06F16/211G06F16/258G06N3/044
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,734
App. No.
18/331,530
Granted
Sep 29, 2026
Kind
B2
Abstract

Various embodiments of the present disclosure provide machine learning techniques for transforming disparate, third-party datasets to canonical representations. The techniques include generating, using a machine learning prediction model, a canonical representation for an input dataset. The machine learning prediction model is previously trained using permutative input embeddings for a training dataset based on canonical data entity features, such that each permutative input embedding corresponds to a different sequence of the canonical data entity features. The permutative input embeddings are leveraged to generate a latent representation for the training dataset. The latent representation is combined with a canonical data map to generate an alignment vector, which is refined to generate an output vector for the input dataset. The machine learning prediction model is trained using a model loss generated based on a comparison of the output vector with a corresponding labeled vector.

Claims (52)

1 . A computer-implemented method comprising:

generating, by one or more processors and using a machine learning prediction model, a canonical representation for an input dataset, wherein the machine learning prediction model is previously trained by:

generating a plurality of permutative input embeddings for a training dataset based on a plurality of canonical data entity features, wherein each permutative input embedding of the plurality of permutative input embeddings corresponds to a different sequence of the plurality of canonical data entity features;

generating a latent representation based on the plurality of permutative input embeddings;

generating an alignment vector representation for the training dataset based on a comparison between the latent representation and a canonical data map, wherein the canonical data map comprises a canonical labeled vector comprising a label corresponding to a data type of the input dataset;

generating an output vector for the training dataset based on the alignment vector representation;

generating, using a loss function, a model loss for the machine learning prediction model based on the output vector and a labeled vector for the training dataset; and

updating one or more parameters of the machine learning prediction model based on the model loss.

2 . The computer-implemented method of claim 1 further comprising:

receiving the input dataset from a third-party data source, wherein the input dataset comprises one or more data fields associated with inconsistent metadata indicative of one or more field descriptions or one or more column values specific to the third-party data source.

3 . The computer-implemented method of claim 1 , wherein the latent representation is generated using one or more neural network layers of the machine learning prediction model, and the latent representation is indicative of a plurality of feature weights for each of the plurality of canonical data entity features.

4 . The computer-implemented method of claim 3 , wherein the training dataset comprises a plurality of data fields and the plurality of feature weights comprises one or more feature weights between each of the plurality of data fields and each of the plurality of canonical data entity features.

5 . The computer-implemented method of claim 3 , wherein the one or more neural network layers of the machine learning prediction model comprise a bidirectional recurrent neural network.

6 . The computer-implemented method of claim 1 , wherein the alignment vector representation is based on a dot product between the latent representation and the canonical data map.

7 . The computer-implemented method of claim 1 , wherein generating the output vector for the training dataset comprises:

generating, using a sigmoid function, a hidden state output for the alignment vector representation,

generating, using an activation function, a refined hidden state output, and

generating the output vector based on the refined hidden state output.

8 . The computer-implemented method of claim 7 , wherein the activation function comprises a softmax function.

9 . The computer-implemented method of claim 7 , wherein the output vector for the training dataset comprises a dot product between the refined hidden state output and the canonical data map.

10 . The computer-implemented method of claim 1 , wherein the loss function comprises a cross-entropy loss function and the model loss comprises a cross-entropy loss between the output vector and the labeled vector.

11 . The computer-implemented method of claim 1 , wherein the output vector comprises a two dimensional vector indicative of a canonical table from a canonical model that corresponds to each of a plurality of data fields of the training dataset.

12 . The computer-implemented method of claim 1 , wherein the canonical labeled vector is generated using exhaustive raw data representing all data types of the input dataset.

13 . A system comprising:

one or more processors; and

one or more memories storing processor-executable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

generating, using a machine learning prediction model, a canonical representation for an input dataset, wherein the machine learning prediction model is previously trained by:

generating a plurality of permutative input embeddings for a training dataset based on a plurality of canonical data entity features, wherein each permutative input embedding of the plurality of permutative input embeddings corresponds to a different sequence of the plurality of canonical data entity features;

generating a latent representation based on the plurality of permutative input embeddings;

generating an alignment vector representation for the training dataset based on a comparison between the latent representation and a canonical data map, wherein the canonical data map comprises a canonical labeled vector comprising a label corresponding to a data type of the input dataset;

generating an output vector for the training dataset based on the alignment vector representation;

generating, using a loss function, a model loss for the machine learning prediction model based on the output vector and a labeled vector for the training dataset; and

updating one or more parameters of the machine learning prediction model based on the model loss.

14 . The system of claim 13 , wherein the one or more processors are further configured to:

receive the input dataset from a third-party data source, wherein the input dataset comprises one or more data fields associated with inconsistent metadata indicative of one or more field descriptions or one or more column values specific to the third-party data source.

15 . The system of claim 13 , wherein the latent representation is generated using one or more neural network layers of the machine learning prediction model, and the latent representation is indicative of a plurality of feature weights for each of the plurality of canonical data entity features.

16 . The system of claim 15 , wherein the training dataset comprises a plurality of data fields and the plurality of feature weights comprises one or more feature weights between each of the plurality of data fields and each of the plurality of canonical data entity features.

17 . The system of claim 15 , wherein the one or more neural network layers of the machine learning prediction model comprise a bidirectional recurrent neural network.

18 . The system of claim 13 , wherein the alignment vector representation is based on a dot product between the latent representation and the canonical data map.

19 . One or more non-transitory computer-readable media storing processor-executable instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

generating, using a machine learning prediction model, a canonical representation for an input dataset, wherein the machine learning prediction model is previously trained by:

generating a plurality of permutative input embeddings for a training dataset based on a plurality of canonical data entity features, wherein each permutative input embedding of the plurality of permutative input embeddings corresponds to a different sequence of the plurality of canonical data entity features;

generating a latent representation based on the plurality of permutative input embeddings;

generating an alignment vector representation for the training dataset based on a comparison between the latent representation and a canonical data map, wherein the canonical data map comprises a canonical labeled vector comprising a label corresponding to a data type of the input dataset;

generating an output vector for the training dataset based on the alignment vector representation;

generating, using a loss function, a model loss for the machine learning prediction model based on the output vector and a labeled vector for the training dataset; and

updating one or more parameters of the machine learning prediction model based on the model loss.

20 . The one or more non-transitory computer-readable media of claim 19 , wherein generating the output vector for the training dataset comprises:

generating, using a sigmoid function, a hidden state output for the alignment vector representation,

generating, using an activation function, a refined hidden state output, and

generating the output vector based on the refined hidden state output; and wherein

the activation function comprises a softmax function, and the output vector for the training dataset comprises a dot product between the refined hidden state output and the canonical data map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2023
From: SINGH, SANJAY KUMAR; JETHY, SUBHASIS; SAINI, UDIT; MORATO, CARLOS W.; BHOTIKA, RAHUL; DAS, RANJU; MANOHAR, VASANT
To: UNITEDHEALTH GROUP INCORPORATED
Reel/Frame 063897/0044 →
Continuity (1)
Related Publication 20240411732A1 · Dec 12, 2024
References Cited (50)
US 7822768B2 · Maymir-Ducharme et al. · 2010 [cited by applicant]
US 9626451B2 · Dietrich et al. · 2017 [cited by applicant]
US 10191969B2 · Yan · 2019 [cited by applicant]
US 10726025B2 · Guo et al. · 2020 [cited by applicant]
US 11170166B2 · Bellegarda · 2021 [cited by examiner]
US 11354582B1 · Prat · 2022 [cited by examiner]
US 11366966B1 · Ramsey et al. · 2022 [cited by applicant]
US 12242945B1 · Everest · 2025 [cited by examiner]
US 20170011107A1 · Shet et al. · 2017 [cited by applicant]
US 20180137404A1 · Fauceglia et al. · 2018 [cited by applicant]
US 20180150808A1 · Haley · 2018 [cited by examiner]
US 20180189265A1 · Chen et al. · 2018 [cited by applicant]
US 20180253669A1 · Thunoli et al. · 2018 [cited by applicant]
US 20190279101A1 · Habti et al. · 2019 [cited by applicant]
US 20200110809A1 · DeFelice · 2020 [cited by examiner]
US 20200201834A1 · Ye et al. · 2020 [cited by applicant]
US 20200349129A1 · Bracholdt et al. · 2020 [cited by applicant]
US 20200380212A1 · Butler et al. · 2020 [cited by applicant]
US 20210183484A1 · Shaib · 2021 [cited by examiner]
US 20210295822A1 · Tomkins et al. · 2021 [cited by applicant]
US 20220198286A1 · Prat · 2022 [cited by examiner]
US 20220207343A1 · Lei et al. · 2022 [cited by applicant]
US 20220269663A1 · Raphael et al. · 2022 [cited by applicant]
US 20220350810A1 · Majumdar et al. · 2022 [cited by applicant]
US 20220358373A1 · Bucher · 2022 [cited by examiner]
US 20220374735A1 · Rathod et al. · 2022 [cited by applicant]
US 20230038256A1 · Tal · 2023 [cited by examiner]
US 20230290114A1 · Prat · 2023 [cited by examiner]
US 20230325366A1 · Erett et al. · 2023 [cited by applicant]
CN 113505590A · 2021 [cited by applicant]
WO 2019203863A1 · 2019 [cited by applicant]
WO 2019227073A1 · 2019 [cited by applicant]
WO 2019229769A1 · 2019 [cited by applicant]
Mena et al., (Learning Latent Permutations With Gumbelsinkhorn Networks, Iclr 2018, 22 pages) (Year: 2018). [cited by examiner]
Imani et al., (Representation Alignment in Neural Networks, arXiv, Sep. 17, 2022, 26 pages) (Year: 2022). [cited by examiner]
Chen, Harr et al. “Content Modeling Using Latent Permutations,” Journal of Artificial Intelligence Research, vol. 36, Oct. 2009, pp. 129-163, DOI: 10.1613/jair.2830. [cited by applicant]
Chrupala, Grzegorz. “Normalizing Tweets With Edit Scripts and Recurrent Neural Embeddings,” In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, vol. 2: Short Papers, Jun. 2014, pp… [cited by applicant]
Duarte, José Marcio et al. “Deep Analysis of Word Sense Disambiguation via Semi-Supervised Learning and Neural Word Representations,” Information Sciences, vol. 570, Sep. 2021, pp. 278-297, DOI: 10.1016/j.ins.2021.04.00… [cited by applicant]
Huang, Hongzhao et al. “Leveraging Deep Neural Networks and Knowledge Graphs for Entity Disambiguation,” arXiv preprint arXiv:1504.07678v1 [cs.CL], Apr. 28, 2015, (10 pages), available online: https://arxiv.org/pdf/1504… [cited by applicant]
Lourentzou, Ismini et al. “Adapting Sequence to Sequence Models for Text Normalization in Social Media,” In Proceedings of the international AAAI Conference on Web and Social Media, vol. 13, Jul. 6, 2019, pp. 335-345. [cited by applicant]
Raiman, Jonathan et al. “DeepType: Multilingual Entity Linking by Neural Type System Evolution,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, No. 1, Apr. 27, 2018, pp. 5406-5413, DOI: 10.1609/… [cited by applicant]
Shen, Wei et al. “Entity Linking Meets Deep Learning: Techniques and Solutions,” arXiv preprint arXiv:2109.12520v1 [cs.CL], Sep. 26, 2201, pp. 1-20, available online: https://arxiv.org/pdf/2109.12520.pdf. [cited by applicant]
Sinha, Koustuv et al. “UnNatural Language Inference,” arXiv preprint arXiv:2101.00010v2 [cs.CL], Jun. 11, 2021, (18 pages), available online: https://arxiv.org/pdf/2101.00010.pdf. [cited by applicant]
Vretinaris, Alina et al. “Medical Entity Disambiguation Using Graph Neural Networks,” arXiv preprint arXiv:2104.01488v1 [cs.IR], Apr. 3, 2021, (10 pages), available online: https://arxiv.org/pdf/2104.01488.pdf. [cited by applicant]
Xie, Jie et al. “Joint Entity Linking for Web Tables With Hybrid Semantic Matching,” In Computational Science—ICCS 2020: 20th International Conference, Proceedings, Part II 20 2020, Amsterdam, The Netherlands, Jun. 15, … [cited by applicant]
Fan, Grace et al., “Semantics-aware Dataset Discovery from Data Lakes with Contextualized Column-based Representation Learning”, submitted Oct. 4, 2022, to Cornell University Olin Online Library Archive, available on th… [cited by applicant]
Luo, Yuan, et al., “Implementing a Portable Clinical NLP System with a Common Data Model—a Lisp Perspective”, Proceedings of 2018 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Dec. 3-6, 2018, p… [cited by applicant]
Non-Final Rejection Mailed on Feb. 25, 2026 for U.S. Appl. No. 18/322,091, 57 page(s). [cited by applicant]
Trabelsi, Mohamed et al., “Semantic Labeling Using a Deep Contextualized Language Model”, submitted to Cornell University Olin Online Library Archive on Oct. 30, 2020, available on the Internet at https://arxiv.org/pdf/… [cited by applicant]
Notice of Allowance and Fees Due (PTOL-85) Mailed on Jul. 6, 2026 for U.S. Appl. No. 18/322,091, 50 page(s). [cited by applicant]