Decarbonizing BERT with topics for efficient document classification
Various embodiments of the teachings herein include a computer-implemented method of fine-tuning Natural Language Processing (NLP) models. Some examples include: providing a training data set including a multitude of training text documents; providing a NLP model including a Neural Network (NN) based Topic Model (TM) having scalable TM parameters and a parallel large-scale pre-trained Language Model (LM) having scalable LM parameters; and fine-tuning the NLP model by jointly training the NN-based TM and the parallel large-scale pre-trained LM using a projected vector comprising a combination and projection of a document topic proportion generated by the NN-based TM based on the scalable TM parameters from an input training text document of the multitude of training text documents, and of a contextualized document representation generated by the large-scale pre-trained LM based on the scalable LM parameters from the same input training text document.
1 . A method of fine-tuning Natural Language Processing (NLP) models, the method comprising:
providing a training data set including a multitude of training text documents;
providing a NLP model including a Neural Network (NN) based Topic Model (TM) having scalable TM parameters and a parallel large-scale pre-trained Language Model (LM) having scalable LM parameters;
fine-tuning the NLP model by jointly training the NN-based TM and the parallel large-scale pre-trained LM using a projected vector comprising a combination and projection of a document topic proportion generated by the NN-based TM based on the scalable TM parameters from an input training text document of the multitude of training text documents, and of a contextualized document representation generated by the large-scale pre-trained LM based on the scalable LM parameters from the same input training text document
wherein the provided NLP model includes a downstream processing layer with scalable processing parameters; and
fine-tuning the NN-based TM, the parallel large-scale pre-trained LM, and the downstream processing layer includes joint training, with the projected vector as input to the processing layer;
fine-tuning comprises, for each training text document of at least a sub-set of the provided at least one training data set, iteratively:
putting in one training text document of the multitude of training text documents to the NN-based TM and to the parallel large-scale pre-trained LM;
generating a document topic proportion and a TM output vector based on the document topic proportion from the input training text document by the NN-based TM using the scalable TM parameters;
generating a contextualized document representation from the same input training text document or a decreased fraction thereof by the large-scale pre-trained LM using the scalable LM parameters;
combining and projecting the generated document topic proportion and the generated contextualized document representation into the projected vector;
generating a processed output vector from the projected vector by the processing layer using the scalable processing parameters;
combining an TM objective function based on the TM output vector of the NN-based TM, and an LM objective function, using the processed output vector of the processing layer, into a joint objective function; and
updating the scalable TM parameters of the NN-based TM, the scalable LM parameters of the large-scale pre-trained LM, and the scalable processing parameters of the processing layer using the joint objective function.
2 . A method according to claim 1 , wherein the contextualized document representation is generated by the large-scale pre-trained LM from an input decreased sequence of the same training text document.
3 . A method according to claim 1 , wherein:
the NN-based TM comprises a Neural Variational Document Model, NVDM; and/or
the large-scale pre-trained LM comprises a Bidirectional Encoder Representations from Transformers, BERT, model; and/or
the processing layer comprises a classification layer.