IP Library › Granted Patent US 12,505,142
Granted Patent B2
US 12,505,142 · App. 18/040,788 · Granted Dec 23, 2025

Decarbonizing BERT with topics for efficient document classification

Inventors: Pankaj Gupta (Munich, DE); Thomas Runkler (Munich, DE); Khushbu Saxena (Zürich, CH)
Assignee: DRIMCO GMBH
G06F16/35G06F40/284G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,142
App. No.
18/040,788
Granted
Dec 23, 2025
Kind
B2
Abstract

Various embodiments of the teachings herein include a computer-implemented method of fine-tuning Natural Language Processing (NLP) models. Some examples include: providing a training data set including a multitude of training text documents; providing a NLP model including a Neural Network (NN) based Topic Model (TM) having scalable TM parameters and a parallel large-scale pre-trained Language Model (LM) having scalable LM parameters; and fine-tuning the NLP model by jointly training the NN-based TM and the parallel large-scale pre-trained LM using a projected vector comprising a combination and projection of a document topic proportion generated by the NN-based TM based on the scalable TM parameters from an input training text document of the multitude of training text documents, and of a contextualized document representation generated by the large-scale pre-trained LM based on the scalable LM parameters from the same input training text document.

Claims (19)

1 . A method of fine-tuning Natural Language Processing (NLP) models, the method comprising:

providing a training data set including a multitude of training text documents;

providing a NLP model including a Neural Network (NN) based Topic Model (TM) having scalable TM parameters and a parallel large-scale pre-trained Language Model (LM) having scalable LM parameters;

fine-tuning the NLP model by jointly training the NN-based TM and the parallel large-scale pre-trained LM using a projected vector comprising a combination and projection of a document topic proportion generated by the NN-based TM based on the scalable TM parameters from an input training text document of the multitude of training text documents, and of a contextualized document representation generated by the large-scale pre-trained LM based on the scalable LM parameters from the same input training text document

wherein the provided NLP model includes a downstream processing layer with scalable processing parameters; and

fine-tuning the NN-based TM, the parallel large-scale pre-trained LM, and the downstream processing layer includes joint training, with the projected vector as input to the processing layer;

fine-tuning comprises, for each training text document of at least a sub-set of the provided at least one training data set, iteratively:

putting in one training text document of the multitude of training text documents to the NN-based TM and to the parallel large-scale pre-trained LM;

generating a document topic proportion and a TM output vector based on the document topic proportion from the input training text document by the NN-based TM using the scalable TM parameters;

generating a contextualized document representation from the same input training text document or a decreased fraction thereof by the large-scale pre-trained LM using the scalable LM parameters;

combining and projecting the generated document topic proportion and the generated contextualized document representation into the projected vector;

generating a processed output vector from the projected vector by the processing layer using the scalable processing parameters;

combining an TM objective function based on the TM output vector of the NN-based TM, and an LM objective function, using the processed output vector of the processing layer, into a joint objective function; and

updating the scalable TM parameters of the NN-based TM, the scalable LM parameters of the large-scale pre-trained LM, and the scalable processing parameters of the processing layer using the joint objective function.

2 . A method according to claim 1 , wherein the contextualized document representation is generated by the large-scale pre-trained LM from an input decreased sequence of the same training text document.

3 . A method according to claim 1 , wherein:

the NN-based TM comprises a Neural Variational Document Model, NVDM; and/or

the large-scale pre-trained LM comprises a Bidirectional Encoder Representations from Transformers, BERT, model; and/or

the processing layer comprises a classification layer.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2026
From: SIEMENS AKTIENGESELLSCHAFT
To: DRIMCO GMBH
Reel/Frame 073814/0585 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 20, 2025
From: SIEMENS AKTIENGESELLSCHAFT
To: DRIMCO GMBH
Reel/Frame 072980/0378 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2025
From: SAXENA, KHUSHBU
To: SIEMENS AKTIENGESELLSCHAFT
Reel/Frame 070649/0581 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2024
From: GUPTA, PANKAJ; RUNKLER, THOMAS
To: SIEMENS AKTIENGESELLSCHAFT
Reel/Frame 069243/0992 →
Continuity (1)
Related Publication 20230306050A1 · Sep 28, 2023
References Cited (21)
US 11042700B1 · Walters · 2021 [cited by examiner]
US 20190236464A1 · Feinson · 2019 [cited by examiner]
US 20190325066A1 · Krishna · 2019 [cited by examiner]
US 20190370331A1 · Büttmer · 2019 [cited by applicant]
US 20200184339A1 · Li · 2020 [cited by applicant]
US 20210248192A1 · Lu · 2021 [cited by examiner]
US 20210326523A1 · Walters · 2021 [cited by examiner]
US 20220129724A1 · Feinson · 2022 [cited by examiner]
Search Report for International Application No. PCT/EP20210/072039, 12 pages, Apr. 26, 2021 [cited by applicant]
Miao Yishu, Lei Yu and Phil Blunsom. “Neural Variational Inference for Text Processing.” ICML, 2016. [cited by applicant]
Ashish Vaswani et al. “Attention is all you need.” In Advances in Neural Information Processing Systems, 2017. [cited by applicant]
Yatin, Chaudhary et al: “TopicBERT for Energy Efficient Document Classification”; arxiv.org, Cornell University Library; 201 Olin Library Cornell University Ithaca; NY 14853; XP081803199; Oct. 15, 2020. [cited by applicant]
CNN: Kim et al. (Yoon Kim. 2014. Convolutional neural net-works for sentence classification. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP, Doha, Qatar, A meeting of SI… [cited by applicant]
Chi Sun et al: “How to Fine-Tune BERT for Text Classification?”; ARXIV.Org, Cornell University Library; 201; Olin Library Cornell University Ithaca; NY; 1485; XP081592483; May 14, 2019. [cited by applicant]
Federico, Bianchi et al: “Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence”; ARXIV.Org; Cornell University Library; 201 Olin Library Cornell University Ithaca; NY; 14853; XP0816402… [cited by applicant]
Jacob Devlin et al, “BERT: Pre-training of Deep Didirectional Transformers for Language Understanding” NAACL 2019; Proceedings of NAACL-HLT 2019, pp. 4171-4186 Minneapolis, Minnesota. Association for Computational Lingu… [cited by applicant]
Zhou, Yujun et al: “Pre-trained Contextualized Representation for Chinese Conversation Topic Classification”; 2019 IEEE International Conference on Intelligence and Security Informatics (ISI); pp. 122-127; XP033611985; … [cited by applicant]
Rogers et al. (Anna Rogers, Olga Kovaleva, Anna Rumshisky. A Primer in BERTology: What we know about how BERT works), 2020. [cited by applicant]
Arman, Cohan et al: “Document-level Representation Learning using Citation-informed Transformers”; arxiv.org; Cornell University Library; 201 Olin Library Cornell University Ithaca; NY 14853; XP081644943; Apr. 15, 2020. [cited by applicant]
Alexandre Lacoste et al. Quantifying the carbon emissions of machine learning. arXiv preprint arXiv:1910.09700, 2019. [cited by applicant]
Peinelt Nicole et al: “tBERT: Topic Models and BERT Joining Forces for Semantic Similarity Detection”; Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics; pp. 7047-7055; Jul. 1, 2020. [cited by applicant]