IP Library › Granted Patent US 12,511,476
Granted Patent B2
US 12,511,476 · App. 18/359,113 · Granted Dec 30, 2025

Concept-conditioned and pretrained language models based on time series to free-form text description generation

Inventors: Yuncong Chen (Plainsboro, NJ); Yanchi Liu (Monmouth Junction, NJ); Wenchao Yu (Plainsboro, NJ); Haifeng Chen (West Windsor, NJ)
Assignee: NEC Corporation
G06F40/242G06F40/205G06F40/284G06F40/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,476
App. No.
18/359,113
Granted
Dec 30, 2025
Kind
B2
Abstract

A computer-implemented method for employing a time-series-to-text generation model to generate accurate description texts is provided. The method includes passing time series data through a time series encoder and a multilayer perceptron (MLP) classifier to obtain predicted concept labels, converting the predicted concept labels, by a serializer, to a text token sequence by concatenating an aspect term and an option term of every aspect, inputting the text token sequence into a pretrained language model including a bidirectional encoder and an autoregressive decoder, and using adapter layers to fine-tune the pretrained language model to generate description texts.

Claims (37)

1 . A computer-implemented method for employing a time-series-to-text generation model to generate accurate description texts, comprising:

passing time series data through a time series encoder and a multilayer perceptron (MLP) classifier to obtain predicted concept labels;

converting the predicted concept labels, by a serializer, to a text token sequence by concatenating an aspect term and an option term of every aspect;

inputting the text token sequence into a pretrained language model (PLM) including a bidirectional encoder and an autoregressive decoder;

iteratively training the PLM updated with adapter lavers by updating parameters of the adapter layers by minimizing a sum of a cross-entropy loss of the predicted concept labels and a text token maximum likelihood estimation (MLE) loss of the text token sequence; and

using the adapter layers to fine-tune the PLM to generate description texts.

2 . The computer-implemented method of claim 1 , wherein a concept word parser is used to mine phrases related to the aspects from a training text of the times series data.

3 . The computer-implemented method of claim 1 , wherein concepts serve as a ground truth for training concept classifiers and are aggregated to form a domain concept dictionary.

4 . The computer-implemented method of claim 1 , wherein iteratively training the PLM further comprises computing the cross-entropy loss by aggregating the cross-entropy loss between the predicted concept label and a ground truth label for every aspect.

5 . The computer-implemented method of claim 1 , wherein the MLP classifier takes a last hidden vector of the time series encoder to obtain the predicted concept labels based on a domain concept dictionary.

6 . The computer-implemented method of claim 1 , wherein iteratively training the PLM further comprises computing the text token MLE loss by aggregating a cross-entropy loss between a probability distribution output of the PLM and a ground-truth token at every position.

7 . The computer-implemented method of claim 1 , wherein the PLM is fine-tuned with schedule sampling using a combination of teacher-forcing and free-running modes that determines an input at a next position as a true token at a position in a training text for the teacher forcing mode and as a token sampled from a vocabulary based on an output distribution for the free-running mode.

8 . A computer program product for employing a time-series-to-text generation model to generate accurate description texts, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

passing time series data through a time series encoder and a multilayer perceptron (MLP) classifier to obtain predicted concept labels;

converting the predicted concept labels, by a serializer, to a text token sequence by concatenating an aspect term and an option term of every aspect;

inputting the text token sequence into a pretrained language model (PLM) including a bidirectional encoder and an autoregressive decoder;

iteratively training the PLM updated with adapter layers by updating parameters of the adapter layers by minimizing a sum of a cross-entropy loss of the predicted concept labels and a text token maximum likelihood estimation (MLE) loss of the text token sequence; and

using the adapter layers to fine-tune the PLM to generate description texts.

9 . The computer program product of claim 8 , wherein a concept word parser is used to mine phrases related to the aspects from a training text of the times series data.

10 . The computer program product of claim 8 , wherein concepts serve as a ground truth for training concept classifiers and are aggregated to form a domain concept dictionary.

11 . The computer program product of claim 8 , wherein iteratively training the PLM further comprises computing the cross-entropy loss by aggregating the cross-entropy loss between the predicted concept label and a ground truth label for every aspect.

12 . The computer program product of claim 8 , wherein the MLP classifier takes a last hidden vector of the time series encoder to obtain the predicted concept labels based on a domain concept dictionary.

13 . The computer program product of claim 8 , wherein the adapter layers are inserted into all transformer layers of the PLM.

14 . The computer program product of claim 8 , wherein the PLM is fine-tuned with schedule sampling using a combination of teacher-forcing and free-running modes.

15 . A computer processing system for employing a time-series-to-text generation model to generate accurate description texts, comprising:

a memory device for storing program code; and

a processor device, operatively coupled to the memory device, for running the program code to:

pass time series data through a time series encoder and a multilayer perceptron (MLP) classifier to obtain predicted concept labels;

convert the predicted concept labels, by a serializer, to a text token sequence by concatenating an aspect term and an option term of every aspect;

input the text token sequence into a pretrained language model (PLM) including a bidirectional encoder and an autoregressive decoder;

iteratively train the PLM updated with adapter layers by updating parameters of the adapter layers by minimizing a sum of a cross-entropy loss of the predicted concept labels and a text token maximum likelihood estimation (MLE) loss of the text token sequence; and

use the adapter layers to fine-tune the PLM to generate description texts.

16 . The computer processing system of claim 15 , wherein a concept word parser is used to mine phrases related to the aspects from a training text of the times series data.

17 . The computer processing system of claim 15 , wherein concepts serve as a ground truth for training concept classifiers and are aggregated to form a domain concept dictionary.

18 . The computer processing system of claim 15 , wherein iteratively training the PLM further comprises computing the cross-entropy loss by aggregating the cross-entropy loss between the predicted concept label and a ground truth label for every aspect.

19 . The computer processing system of claim 15 , wherein the MLP classifier takes a last hidden vector of the time series encoder to obtain the predicted concept labels based on a domain concept dictionary.

20 . The computer processing system of claim 15 , wherein the adapter layers are inserted into all transformer layers of the pretrained language model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 18, 2025
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 072938/0913 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2023
From: CHEN, YUNCONG; LIU, YANCHI; YU, WENCHAO; CHEN, HAIFENG
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 064386/0773 →
Continuity (3)
Provisional Application 63402202 · Aug 30, 2022
Provisional Application 63399718 · Aug 21, 2022
Related Publication 20240061998A1 · Feb 22, 2024
References Cited (30)
US 12045568B1 · Shrivastava · 2024 [cited by examiner]
US 12271698B1 · Wang · 2025 [cited by examiner]
US 20090240729A1 · Zwol · 2009 [cited by examiner]
US 20180300400A1 · Paulus · 2018 [cited by examiner]
US 20200265060A1 · Crapo · 2020 [cited by examiner]
US 20210012179A1 · Kalia · 2021 [cited by examiner]
US 20220382975A1 · Gu · 2022 [cited by examiner]
US 20220391755A1 · Li · 2022 [cited by examiner]
US 20220414344A1 · Makki Niri · 2022 [cited by examiner]
US 20230135659A1 · Wu · 2023 [cited by examiner]
US 20230342559A1 · Bhardwaj · 2023 [cited by examiner]
US 20240045890A1 · Nguyen · 2024 [cited by examiner]
US 20240386015A1 · Crabtree · 2024 [cited by examiner]
US 20240412008A1 · Salim · 2024 [cited by examiner]
Van Aken, Betty, et al. “How does bert answer questions? a layer-wise analysis of transformer representations.” Proceedings of the 28th ACM international conference on information and knowledge management. 2019, pp. 182… [cited by examiner]
Cho, Jaemin, et al. “Unifying vision-and-language tasks via text generation.” International Conference on Machine Learning. PMLR, 2021, pp. 1-12. (Year: 2021). [cited by examiner]
Ding, Ning, et al. “Delta tuning: A comprehensive study of parameter efficient methods for pre-trained language models.” arXiv preprint arXiv:2203.06904. Mar. 2022, pp. 1-49. (Year: 2022). [cited by examiner]
Zhang, Zhengkun, et al. “Hyperpelt: Unified parameter-efficient language model tuning for both language and vision-and-language tasks.” arXiv preprint arXiv:2203.03878. Mar. 2022, pp. 1-14. (Year: 2022). [cited by examiner]
Jin, Xisen, et al. “Learn continually, generalize rapidly: Lifelong knowledge accumulation for few-shot learning.” arXiv preprint arXiv: 2104.08808. Aug. 2022, pp. 1-16. (Year: 2022). [cited by examiner]
Li, Jingye, et al. “Unified named entity recognition as word-word relation classification.” proceedings of the AAAI conference on artificial intelligence. vol. 36. No. 10. Jun. 2022, pp. 10965-10973. (Year: 2022). [cited by examiner]
Osei-Brefo, Emmanuel, et al. “UoR-NCL at SemEval-2022 Task 6: Using ensemble loss with BERT for intended sarcasm detection.” Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval—2022). Jul. 202… [cited by examiner]
Ustun, Ahmet, et al. “UDapter: Language adaptation for truly Universal Dependency parsing.” arXiv preprint arXiv:2004.14327. Oct. 2020, pp. 1-14. (Year: 2020). [cited by examiner]
Yan, Hang, et al. “A unified generative framework for various NER subtasks.” arXiv preprint arXiv:2106.01223. Jun. 2021, pp. 1-15. (Year: 2021). [cited by examiner]
Goyal, et al. “Professor forcing: A new algorithm for training recurrent networks.” Advances in neural information processing systems 29, 2016, pp. 1-9. (Year: 2016). [cited by examiner]
Wang, Xinyi, et al. “Efficient test time adapter ensembling for low-resource language varieties.” arXiv preprint arXiv:2109.04877. Sep. 2021, pp. 1-8. (Year: 2021). [cited by examiner]
Qin, Y., Song, D., Chen, H., Cheng, W., Jiang, G., & Cottrell, G. (Aug. 14, 2017). A dual-stage attention-based recurrent neural network for time series prediction. arXiv preprint arXiv:1704.02971. [cited by applicant]
Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (Apr. 22, 2019). The curious case of neural text degeneration. arXiv preprint arXiv:1904.09751. [cited by applicant]
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., . . . & Zettlemoyer, L. (Oct. 29, 2019). Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and compr… [cited by applicant]
He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T., & Neubig, G. (Oct. 8, 2021). Towards a unified view of parameter-efficient transfer learning. arXiv preprint arXiv:2110.04366. [cited by applicant]
Bengio, S., Vinyals, O., Jaitly, N., & Shazeer, N. (Dec. 11, 2015). Scheduled sampling for sequence prediction with recurrent neural networks. Advances in neural information processing systems, 28. [cited by applicant]