IP Library › Granted Patent US 12,554,939
Granted Patent B2
US 12,554,939 · App. 18/339,252 · Granted Feb 17, 2026

Forecasted diverse natural language generation models

Inventors: Aaron K. Baughman (Cary, NC); Chandankumar Johakhim Patel (Fairborn, OH); Mauro Marzorati (Lutz, FL); Jeremy R. Fox (Georgetown, TX)
Assignee: International Business Machines Corporation
G06F40/40G06F40/247
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,939
App. No.
18/339,252
Granted
Feb 17, 2026
Kind
B2
Abstract

A method, a structure, and a computer system for diverse natural language generation. The exemplary embodiments may include training a data-to-text neural network (D2T NN) and training a text-to-text neural network (T2T NN), wherein the D2T NN and the T2T NN have identical transformer architectures, and wherein the training of the D2T NN and the T2T NN are indexed by time. The exemplary embodiments may further include interleaving weights between the D2T NN and the T2T NN, as well as generating a sentence based on the interleaved D2T NN and T2T NN.

Claims (37)

1 . A method for diverse natural language generation, the method comprising:

training a data-to-text neural network (D2T NN) using first training data;

training a text-to-text neural network (T2T NN) using second training data, wherein the D2T NN and the T2T NN have identical transformer architectures, and wherein the training of the D2T NN and the T2T NN are indexed by time;

interleaving weights between the D2T NN and the T2T NN by transferring weights between encoding portions of the D2T NN and the T2T NN, wherein the interleaving minimizes an effect of the first training data and the second training data being collected at different times; and

generating a sentence based on the interleaved D2T NN and T2T NN.

2 . The method of claim 1 , wherein the D2T NN converts one or more statistics into one or more tuples and generates one or more base sentences based on the one or more tuples.

3 . The method of claim 2 , wherein the T2T NN paraphrases the one or more base sentences while adding context.

4 . The method of claim 1 , wherein interleaving the weights between the D2T NN and the T2T NN further comprises:

transferring weights between decoding portions of the D2T NN and the T2T NN.

5 . The method of claim 1 , further comprising:

smoothing the interleaved weights and original weights to a midpoint.

6 . The method of claim 5 , wherein the smoothing is performed via Gibs sampling.

7 . A computer program product for diverse natural language generation, the computer program product comprising:

one or more non-transitory computer-readable storage media and program instructions stored on the one or more non-transitory computer-readable storage media capable of performing a method, the method comprising:

training a data-to-text neural network (D2T NN) using first training data;

training a text-to-text neural network (T2T NN) using second training data, wherein the D2T NN and the T2T NN have identical transformer architectures, and wherein the training of the D2T NN and the T2T NN are indexed by time;

interleaving weights between the D2T NN and the T2T NN by transferring weights between encoding portions of the D2T NN and the T2T NN, wherein the interleaving minimizes an effect of the first training data and the second training data being collected at different times; and

generating a sentence based on the interleaved D2T NN and T2T NN.

8 . The computer program product of claim 7 , wherein the D2T NN converts one or more statistics into one or more tuples and generates one or more base sentences based on the one or more tuples.

9 . The computer program product of claim 8 , wherein the T2T NN paraphrases the one or more base sentences while adding context.

10 . The computer program product of claim 7 , wherein interleaving the weights between the D2T NN and the T2T NN further comprises:

transferring weights between a decoding portions of the D2T NN and the T2T NN.

11 . The computer program product of claim 7 , further comprising:

smoothing the interleaved weights and original weights to a midpoint.

12 . The computer program product of claim 11 , wherein the smoothing is performed via Gibs sampling.

13 . A computer system for diverse natural language generation, the system comprising:

one or more computer processors, one or more computer-readable storage media, and program instructions stored on the one or more of the computer-readable storage media for execution by at least one of the one or more processors capable of performing a method, the method comprising:

training a data-to-text neural network (D2T NN) using first training data;

training a text-to-text neural network (T2T NN) using second training data, wherein the D2T NN and the T2T NN have identical transformer architectures, and wherein the training of the D2T NN and the T2T NN are indexed by time;

interleaving weights between the D2T NN and the T2T NN by transferring weights between encoding portions of the D2T NN and the T2T NN, wherein the interleaving minimizes an effect of the first training data and the second training data being collected at different times; and

generating a sentence based on the interleaved D2T NN and T2T NN.

14 . The computer system of claim 13 , wherein the D2T NN converts one or more statistics into one or more tuples and generates one or more base sentences based on the one or more tuples.

15 . The computer system of claim 14 , wherein the T2T NN paraphrases the one or more base sentences while adding context.

16 . The computer system of claim 13 , wherein interleaving the weights between the D2T NN and the T2T NN further comprises:

transferring weights between decoding portions of the D2T NN and the T2T NN.

17 . The computer system of claim 13 , further comprising:

smoothing the interleaved weights and original weights to a midpoint.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 22, 2023
From: BAUGHMAN, AARON K.; PATEL, CHANDANKUMAR JOHAKHIM; MARZORATI, MAURO; FOX, JEREMY R.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064024/0315 →
Continuity (1)
Related Publication 20240428012A1 · Dec 26, 2024
References Cited (30)
US 7249117B2 · Estes · 2007 [cited by applicant]
US 11122341B1 · Decrop · 2021 [cited by applicant]
US 11392847B1 · Abdollahian · 2022 [cited by examiner]
US 20170235723A1 · Allen · 2017 [cited by applicant]
US 20220309245A1 · Baughman · 2022 [cited by applicant]
US 20230154188A1 · Li · 2023 [cited by examiner]
US 20240020337A1 · Maharana · 2024 [cited by examiner]
US 20240037335A1 · Akbari · 2024 [cited by examiner]
CN 106405352A · 2017 [cited by applicant]
CN 107578101B · 2018 [cited by applicant]
Kantharaj et al., Chart-to-text: A large-scale benchmark for chart summarization, arXiv:2203.06486, 2022 (Year: 2022). [cited by examiner]
Alayrac et al., Flamingo: a visual language model for few-shot learning, Advances in Neural Information Processing Systems, 35, 2022 (Year: 2022). [cited by examiner]
Gelfand, Gibbs sampling, Journal of the American statistical Association, vol. 95 No. 452, 2000 (Year: 2000). [cited by examiner]
Maekaku et al., Attention Weight Smoothing Using Prior Distributions for Transformer-Based End-to-End ASR, Interspeech, 2022 (Year: 2022). [cited by examiner]
Shan et al., ERNIE-ViL 2.0: Multi-view Contrastive Learning for Image-Text Pre-training, arXiv:2209.15270, 2022 (Year: 2022). [cited by examiner]
Xia et al., Tied transformers: Neural machine translation with shared encoder and decoder, Proceedings of the AAAI conference on artificial intelligence, vol. 33, No. 01, 2019 (Year: 2019). [cited by examiner]
Kale et al., Text-to-Text Pre-Training for Data-to-Text Tasks, arXiv:2005.10433, 2021 (Year: 2021). [cited by examiner]
Disclosed Anonymously, “Method and System for Providing Context-based, Customizable Virtual Commentator for Electronic Tournaments,” IP.com No. IPCOM000254001D, IP.com Publication Date: May 23, 2018, 3 pages. [cited by applicant]
Dong et al., “A Survey of Natural Language Generation”, ACM Computing Surveys, vol. 55, No. 8, Article 173. Publication date: Dec. 23, 2022., https://dl.acm.org/doi/10.1145/3554727, 38 pages. [cited by applicant]
Kim et al., “Automatic Baseball Commentary Generation Using Deep Learning,” The 35th ACM/SIGAPP Symposium on Applied Computing (SAC '20), Mar. 30-Apr. 3, 2020, ACM, pp. 1056-1065. [cited by applicant]
Kmet et al., “A 24h forecast of solar irradiance using echo state neural networks”, 16th EANN workshops, Sep. 25-28, 2015, ACM, https://dl.acm.org/doi/10.1145/2797143.2797166, 5 pages. [cited by applicant]
Monzon et al., “Incremental Local Model Networks for Time Series Prediction,” IJCNN'01. International Joint Conference on Neural Networks. Proceedings (Cat. No. 01CH37222), IEEE, Jul. 15-19, 2001, https://ieeexplore.iee… [cited by applicant]
Murphy, “The Shot Heard Round The World”, YouTube, Accessed: Apr. 12, 2023, https://www.youtube.com/watch?v=lrl7dVj90zs&t=19s, 9 pages. [cited by applicant]
Strauss et al., “Eric: A Generic Rule-based Framework for an Affective Embodied Commentary Agent,” Proc. of 7th Int. Conf. on Autonomous Agents and Multiagent Systems (AA-MAS 2008), May, 12-16, 2008, 8 pages. [cited by applicant]
Suglia et al., “Going for Goal: A Resource for Grounded Football Commentaries,” arXiv:2211.04534v1 [cs.CV] Nov. 8, 2022, https://arxiv.org/pdf/2211.04534.pdf, 16 pages. [cited by applicant]
Unterthiner et al., “Predicting Neural Network Accuracy from Weights,” arXiv:2002.11448v4 [stat.ML] Apr. 9, 2021, https://arxiv.org/pdf/2002.11448.pdf, 18 pages. [cited by applicant]
Wikipedia, “Mick Hubert”, Wikipedia—The Free Encyclopedia, Accessed Apr. 12, 2023, https://en.wikipedia.org/wiki/Mick_Hubert, 3 pages. [cited by applicant]
Wu et al., Connecting the Dots: Multivariate Time Series Forecasting with Graph Neural Networks, KDD '20: Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Aug. 2020, ACM,… [cited by applicant]
Xie et al., Integrating Image-Based and Knowledge-Based Representation Learning:, IEEE Transactions on Cognitive and Developmental Systems, vol. 12, Issue: 2, Jun. 2020, https://ieeexplore.ieee.org/document/8689107, pp.… [cited by applicant]
Yao et al., “Internet Traffic Forecasting using Temporal-Topological Graph Convolutional Networks,” 2021 International Joint Conference on Neural Networks (IJCNN), IEEE, Jul. 18-22, 2021, https://ieeexplore.ieee.org/doc… [cited by applicant]