IP Library Granted Patent US 12,602,946
Granted Patent B2
US 12,602,946 · App. 18/464,620 · Granted Apr 14, 2026

Document classification using unsupervised text analysis with concept extraction

Inventor: Mikael Hillborg (Umeå, SE)
Assignee: International Business Machines Corporation
G06V30/413G06F40/258G06V30/19093
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,602,946
App. No.
18/464,620
Granted
Apr 14, 2026
Kind
B2
Abstract

An embodiment for classifying documents using unsupervised text analysis with concept extraction. The embodiment may obtain a generic description for a document by: extracting concepts from headings and descriptions of a target document using natural language processing, matching the extracted concepts to generic abstracts obtained from an abstract database, where the abstract database is independent and separate from the target document, and processing the generic abstracts to lemmatize words and remove common words to form the generic description. The embodiment may process an available series of technical classifications. The embodiment may perform bidirectional analysis to determine a most relevant technical classification for the target document based on the generic description for the target document. The embodiment may assign the most relevant technical classification to the target document.

Claims (56)

1 . A computer-based method for document classification, the method comprising:

obtaining a generic description for a target document, wherein obtaining the generic description comprises:

extracting concepts from headings and descriptions of the target document using natural language processing;

matching the extracted concepts to generic abstracts obtained from an abstract database, where the abstract database is independent and separate from the target document;

processing the generic abstracts to lemmatize words and remove common words to form the generic description;

processing an available series of technical classifications;

performing bidirectional analysis to determine a most relevant technical classification for the target document based on the generic description for the target document by summing occurrences of corresponding words within the generic description for the target document and aggregate descriptions of each of the technical classifications in the available series of technical classifications and calculating scores for each of the technical classifications in the available series of technical classifications based on the sum of occurrences; and

assigning the most relevant technical classification to the target document, wherein the most relevant technical classification corresponds to the technical classification with a highest calculated score, wherein the highest calculated score corresponds to a highest volume of word matches and a highest association between level descriptions of the technical classifications in the available series of technical classifications as compared to the generic description of the target document.

2 . The computer-based method of claim 1 , further comprising:

generating a word frequency table associated with the obtained generic description of the target document.

3 . The computer-based method of claim 1 , further comprising:

concatenating the level descriptions within a tree structure corresponding to each of the technical classifications in the available series of technical classifications;

further processing each of the technical classifications in the available series of technical classifications by lemmatizing the words and removing common non-technical words to obtain aggregate descriptions; and

generating word tables for each of the aggregate descriptions of each of the technical classifications in the available series of technical classifications.

4 . The computer-based method of claim 1 , wherein the extracted concepts from headings and descriptions of the target document are obtained using multiple large language models, and wherein at least a portion of the extracted concepts are abstract and non-technical in substance.

5 . The computer-based method of claim 1 , wherein the available series of technical classifications is derived from an available set of generic technical classifications.

6 . A computer system, the computer system comprising:

one or more processors, one or more computer-readable memories, one or more computer-readable tangible storage medium, and program instructions stored on at least one of the one or more computer-readable tangible storage medium for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories, wherein the computer system is capable of performing a method comprising:

obtaining a generic description for a target document, wherein obtaining the generic description comprises:

extracting concepts from headings and descriptions of the target document using natural language processing;

matching the extracted concepts to generic abstracts obtained from an abstract database, where the abstract database is independent and separate from the target document;

processing the generic abstracts to lemmatize words and remove common words to form the generic description;

processing an available series of technical classifications;

performing bidirectional analysis to determine a most relevant technical classification for the target document based on the generic description for the target document by summing occurrences of corresponding words within the generic description for the target document and aggregate descriptions of each of the technical classifications in the available series of technical classifications and calculating scores for each of the technical classifications in the available series of technical classifications based on the sum of occurrences; and

assigning the most relevant technical classification to the target document, wherein the most relevant technical classification corresponds to the technical classification with a highest calculated score, wherein the highest calculated score corresponds to a highest volume of word matches and a highest association between level descriptions of the technical classifications in the available series of technical classifications as compared to the generic description of the target document.

7 . The computer system of claim 6 , further comprising:

generating a word frequency table associated with the obtained generic description of the target document.

8 . The computer system of claim 6 , further comprising:

concatenating the level descriptions within a tree structure corresponding to each of the technical classifications in the available series of technical classifications;

further processing each of the technical classifications in the available series of technical classifications by lemmatizing the words and removing common non-technical words to obtain aggregate descriptions; and

generating word frequency tables for each of the aggregate descriptions of each of the technical classifications in the available series of technical classifications.

9 . The computer system of claim 6 , wherein the extracted concepts from headings and descriptions of the target document are obtained using multiple large language models, and wherein at least a portion of the extracted concepts are abstract and non-technical in substance.

10 . The computer system of claim 6 , wherein the available series of technical classifications is derived from an available set of generic technical classifications.

11 . A computer program product comprising:

one or more computer-readable storage media;

program instructions stored on the one or more computer-readable storage media to perform operations comprising:

obtaining a generic description for a target document, wherein obtaining the generic description comprises:

extracting concepts from headings and descriptions of the target document using natural language processing;

matching the extracted concepts to generic abstracts obtained from an abstract database, where the abstract database is independent and separate from the target document;

processing the generic abstracts to lemmatize words and remove common words to form the generic description;

processing an available series of technical classifications;

performing bidirectional analysis to determine a most relevant technical classification for the target document based on the generic description for the target document by summing occurrences of corresponding words within the generic description for the target document and aggregate descriptions of each of the technical classifications in the available series of technical classifications and calculating scores for each of the technical classifications in the available series of technical classifications based on the sum of occurrences; and

assigning the most relevant technical classification to the target document, wherein the most relevant technical classification corresponds to the technical classification with a highest calculated score, wherein the highest calculated score corresponds to a highest volume of word matches and a highest association between level descriptions of the technical classifications in the available series of technical classifications as compared to the generic description of the target document.

12 . The computer program product of claim 11 , further comprising:

generating a word frequency table associated with the generic description of the target document.

13 . The computer program product of claim 11 , further comprising:

concatenating the level descriptions within a tree structure corresponding to each of the technical classifications in the available series of technical classifications;

further processing each of the technical classifications in the available series of technical classifications by lemmatizing the words and removing common non-technical words to obtain aggregate descriptions; and

generating word frequency tables for each of the aggregate descriptions of each of the technical classifications in the available series of technical classifications.

14 . The computer program product of claim 11 , wherein the extracted concepts from headings and descriptions of the target document are obtained using multiple large language models, and wherein at least a portion of the extracted concepts are abstract and non-technical in substance.

15 . The computer program product of claim 14 , wherein the available series of technical classifications is derived from an available set of generic technical classifications.

16 . The computer-based method of claim 5 , wherein vector representations of the technical classifications are leveraged to extract the concepts.

17 . The computer-based method of claim 16 , wherein the extracted concepts are utilized in obtaining independent abstracts.

18 . The computer system of claim 10 , wherein vector representations of the technical classifications are leveraged to extract the concepts.

19 . The computer system of claim 18 , wherein the extracted concepts are utilized in obtaining independent abstracts.

20 . The computer program product of claim 15 , wherein vector representations of the technical classifications are leveraged to extract the concepts, and wherein the extracted concepts are utilized in obtaining independent abstracts.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2023
From: HILLBORG, MIKAEL
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064861/0107 →
Continuity (1)
Related Publication 20250087009A1 · Mar 13, 2025
References Cited (31)
US 7065514B2 · Yang-Stephens · 2006 [cited by examiner]
US 8132103B1 · Chowdhury · 2012 [cited by examiner]
US 8645389B2 · Oliver · 2014 [cited by examiner]
US 8843536B1 · Elbaz · 2014 [cited by examiner]
US 8903825B2 · Parker · 2014 [cited by examiner]
US 9201957B2 · Turdakov · 2015 [cited by examiner]
US 9235812B2 · Scholtes · 2016 [cited by examiner]
US 10896292B1 · Norton · 2021 [cited by examiner]
US 11887731B1 · Gallagher · 2024 [cited by examiner]
US 20020138529A1 · Yang-Stephens · 2002 [cited by examiner]
US 20040019601A1 · Gates · 2004 [cited by examiner]
US 20070198506A1 · Attaran Rezaei · 2007 [cited by examiner]
US 20080154875A1 · Morscher · 2008 [cited by examiner]
US 20080208840A1 · Zhang · 2008 [cited by examiner]
US 20080273220A1 · Couchman · 2008 [cited by examiner]
US 20130013603A1 · Parker · 2013 [cited by examiner]
US 20140108005A1 · Kassis · 2014 [cited by examiner]
US 20160132648A1 · Shah · 2016 [cited by examiner]
US 20160217126A1 · Kannan · 2016 [cited by examiner]
US 20180182381A1 · Singh · 2018 [cited by examiner]
US 20220182253A1 · Pawar · 2022 [cited by examiner]
US 20240289407A1 · Rofouei · 2024 [cited by examiner]
KR 101983752B1 · 2019 [cited by applicant]
KR 20220108924A · 2022 [cited by applicant]
WO 2016099019A1 · 2016 [cited by applicant]
Suzanne “using multi-terminology indexing for the assignment of MeSH descriptors to health resources in a French online catalogue” (Year: 2008). [cited by examiner]
Gayathri, et al., “Ontology based Concept Extraction and Classification of Ayurvedic Documents”, ELSEVIER, ScienceDirect Procedia Computer Science 172, 2020, pp. 511-516. [cited by applicant]
Rabiger, et al., “Context-based Extraction of Concepts from Unstructured Textual Documents”, Elsevier, 2021, pp. 1-23. https://www.sciencedirect.com/science/article/abs/pii/S0020025521012779?via%3Dihub. [cited by applicant]
Li, et al., “Bag-of-Concepts representation for document classification based on automatic knowledge acquisition from probabilistic knowledge base”, Elsevier, 2020, 47 Pages. [cited by applicant]
Litvak, et al., “Classification of Web Documents Using Concept Extraction from Ontologies”, Springer-Verlag Berlin Hedelberg, 2007, pp. 287-292. [cited by applicant]
Gu, et al., “Hierarchical document classification based on concept and context”, International Journal of Digital Content Technology and its Applications, Jan. 2011, 2 Pages (Abstract only). [cited by applicant]