IP Library Granted Patent US 12,645,726
Granted Patent B2
US 12,645,726 · App. 18/132,224 · Granted Jun 2, 2026

Latent concept analysis method

Inventors: Hassan Sajjad (Doha, QA); Fahim Dalvi (Doha, QA); Firoj Alam (Doha, QA); Nadir Durrani (Doha, QA); Abdul Rafae Khan (Doha, QA); Jia Xu (Doha, QA)
Assignee: HAMAD BIN KHALIFA UNIVERSITY
G06F16/35G06F16/38
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,726
App. No.
18/132,224
Granted
Jun 2, 2026
Kind
B2
Abstract

A method of constructing a dataset for identifying a plurality of latent concepts in a Natural Language Processing model is provided. The method includes executing a clustering process on a first dataset, preparing a second dataset, defining a hierarchical concept tag-set from the second dataset, and annotating the hierarchical concept tag-set.

Claims (35)

1 . A method of constructing a dataset for identifying a plurality of latent concepts in a Natural Language Processing model, comprising:

executing a clustering process on a first dataset, wherein the clustering process outputs a plurality of clusters;

preparing a second dataset, separate from the first dataset, wherein the second dataset is a portion of the first dataset which comprises less than a totality of the first dataset;

tokenizing the second dataset into a plurality of tokens and extracting contextualized representations for each of the plurality of tokens, wherein input tokens of the plurality of tokens which are split into at least one of sub-words and multiple tokens are mean pooled to create an embedding for each input token;

defining a hierarchical concept tag-set from the second dataset using at least one core language property, wherein the at least one core language property includes syntax; and

annotating the hierarchical concept tag-set comprising identifying a meaning of each of the plurality of clusters based on linguistic relation.

2 . The method of claim 1 , wherein the clustering process is an agglomerative hierarchical clustering method.

3 . The method of claim 1 , wherein a minimum variance criterion is used in the clustering process.

4 . The method of claim 1 , wherein the clustering process outputs a pre-defined number of clusters.

5 . The method of claim 1 , wherein the second dataset is randomly selected from the first dataset.

6 . The method of claim 1 , wherein preparing the second dataset comprises:

extracting a plurality of contextualized representations; and

tokenizing a plurality of sentences.

7 . The method of claim 1 , wherein a plurality of core language properties are used to define the hierarchical concept tag-set.

8 . The method of claim 1 , wherein annotating the hierarchical concept tag-set further comprises:

combining a portion of the plurality of clusters based on the meaning identified.

9 . A method of constructing a dataset for identifying a plurality of latent concepts in a Natural Language Processing model, comprising:

executing a clustering process on a first dataset, wherein the clustering process outputs a plurality of clusters;

preparing a second dataset, separate from the first dataset, wherein the second dataset is a portion of the first dataset which comprises less than a totality of the first dataset;

tokenizing the second dataset into a plurality of tokens and extracting contextualized representations for each of the plurality of tokens, wherein input tokens of the plurality of tokens which are split into at least one of sub-words and multiple tokens are mean pooled to create an embedding for each input token;

defining a hierarchical concept tag-set from the second dataset using at least one core language property, wherein the at least one core language property includes syntax;

annotating the hierarchical concept tag-set comprising identifying a meaning of each of the plurality of clusters based on linguistic relation, outputting an annotated dataset based on the annotated hierarchical concept tag-set; and

executing a logistic regression classifier on the annotated dataset.

10 . The method of claim 9 , wherein a threshold is used on the logistic regression classifier.

11 . The method of claim 9 , wherein the logistic regression classifier is trained on the first dataset, outputting a final dataset.

12 . The method of claim 9 , wherein the clustering process is an agglomerative hierarchical clustering method.

13 . The method of claim 9 , wherein a minimum variance criterion is used in the clustering process.

14 . The method of claim 9 , wherein the clustering process outputs a pre-defined number of clusters.

15 . The method of claim 9 , wherein the second dataset is randomly selected from the first dataset.

16 . The method of claim 9 , wherein preparing the second dataset comprises:

extracting a plurality of contextualized representations; and

tokenizing a plurality of sentences.

17 . The method of claim 9 , wherein a plurality of core language properties are used to define the hierarchical concept tag-set.

18 . The method of claim 9 , wherein annotating the hierarchical concept tag-set further comprises:

combining a portion of the plurality of clusters based on the meaning identified.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2025
From: QATAR FOUNDATION FOR EDUCATION, SCIENCE & COMMUNITY DEVELOPMENT
To: HAMAD BIN KHALIFA UNIVERSITY
Reel/Frame 069936/0656 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2024
From: SAJJAD, HASSAN; DALVI, FAHIM; ALAM, FIROJ; DURRANI, NADIR; KHAN, ABDUL RAFAE; XU, JIA
To: QATAR FOUNDATION FOR EDUCATION, SCIENCE AND COMMUNITY DEVELOPMENT
Reel/Frame 068470/0339 →
Continuity (2)
Provisional Application 63328498 · Apr 7, 2022
Related Publication 20230325426A1 · Oct 12, 2023
References Cited (19)
US 6625585B1 · MacCuish · 2003 [cited by examiner]
US 6658624B1 · Savitzky · 2003 [cited by examiner]
US 6781609B1 · Barker · 2004 [cited by examiner]
US 20030074369A1 · Schuetze · 2003 [cited by examiner]
US 20110225159A1 · Murray · 2011 [cited by examiner]
US 20120117484A1 · Convertino · 2012 [cited by examiner]
US 20160321617A1 · Shastri · 2016 [cited by examiner]
US 20170091692A1 · Guo · 2017 [cited by examiner]
US 20170235820A1 · Conrad · 2017 [cited by examiner]
US 20210343411A1 · Zhang · 2021 [cited by examiner]
US 20220108195A1 · Kehler · 2022 [cited by examiner]
US 20230315993A1 · Nieborowski · 2023 [cited by examiner]
Xu et al.: “A Survey On Multi-Output Learning” (Google Scholar ⋅ arxiv.org ⋅ Xu D ⋅ IEEE transactions on neural networks and learning systems Published 2019, hereafter “Xu”). (Year: 2019). [cited by examiner]
Anonymous; “Discovering Latent Concepts Learned in BERT”; conference paper at ICLR; 2022; (25 pages). [cited by applicant]
Xu, et al.; “A Survey on Multi-output Learning”; Cornell University; Oct. 2019; (21 pages). [cited by applicant]
Michael, et al.; “Asking without Telling: Exploring Latent Ontologies in Contextual Representations” ACL Anthology; Nov. 2020; (21 pages). [cited by applicant]
Mamou, et al.; “Emergence of Separable Manifolds in Deep Language Representations”; Proceedings of Machine Learning Research; vol. 119; Feb. 2023; (11 pages). [cited by applicant]
Jones, et al.; “Semantics-Based Machine Translation with Hyperedge Replacement Grammars” ACL Anthology; Dec. 2012; (18 pages). [cited by applicant]
Gowda, et al.; “Agglomerative clustering using the concept of mutual nearest neighbourhood”; ScienceDirect; vol. 10, Issue 2; 1978; (3 pages). [cited by applicant]