IP Library › Granted Patent US 12,387,057
Granted Patent B2
US 12,387,057 · App. 18/208,083 · Granted Aug 12, 2025

System and method for utilizing weak learners on large language models

Inventors: Hariharan Manikandan (Pittsburgh, PA); Yiding Jiang (Pittsburgh, PA); Jeremy Kolter (Pittsburgh, PA); Chen Qiu (Sindelfingen, DE); Wan-Yi Lin (Wexford, PA); Filipe J. Cabrita Condessa (Pittsburgh, PA)
Assignees: Robert Bosch GmbH; Carnegie Mellon University
G06F40/40G06F40/157
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,057
App. No.
18/208,083
Granted
Aug 12, 2025
Kind
B2
Abstract

A computer-implemented method includes converting tabular data to a text representation, generating metadata associated with the text representation of the tabular data, outputting one or more natural language data descriptions indicative of the tabular data in response to utilizing a large language model (LLM) and zero-shot prompting of the metadata and text representation of the tabular data, outputting one or more summaries utilizing the LLM and appending a prompt on the one or more natural language data descriptions, selecting a single summary of the one or more summaries in response to the single summary having a smallest validation rate, receiving a query associated with the tabular data, outputting one or more predictions associated with the query, and in response to meeting a convergence threshold with the one or more predictions generated from the one or more iterations, output a final prediction associated with the query.

Claims (51)

1. A computer-implemented method for natural-language processing, comprising:

receiving tabular data associated with one or more records;

converting the tabular data to a text representation indicative of the tabular data;

generating metadata associated with the text representation of the tabular data, wherein the metadata is indicative of a description of the tabular data;

for one or more iterations, outputting one or more natural language data descriptions indicative of the tabular data in response to utilizing a large language model (LLM) and zero-shot prompting of the metadata and text representation of the tabular data, wherein the LLM includes a neural network with a plurality of parameters;

for one or more iterations, outputting one or more summaries utilizing the LLM, wherein the one or more summaries include less text than the one or more natural language data descriptions;

for one or more iterations, selecting a single summary of the one or more summaries in response to the single summary having a smallest validation rate;

receiving a query associated with the tabular data;

for one or more iterations, outputting one or more predictions associated with the query utilizing the LLM on the single summary and the query; and

in response to meeting a convergence threshold with the one or more predictions generated from the one or more iterations, outputting a final prediction associated with the query, wherein the final prediction is selected in response to a weighted-majority vote of all of the one or more predictions generated from the one or more iterations.

2. The method of claim 1 , wherein the LLM does not utilize fine-tuning or building of a new language model to output the final prediction.

3. The method of claim 1 , wherein the plurality of parameters is over one billion.

4. The method of claim 1 , wherein the LLM is GPT-4.

5. The method of claim 1 , wherein the text representation is a parsed JSON file.

6. The method of claim 1 , wherein any numerical features associated with the tabular data are encoded into percentiles.

7. The method of claim 1 , wherein outputting one or more summaries utilizing the LLM includes appending a prompt on the one or more natural language data descriptions.

8. The method of claim 1 , wherein the one or more predications are weak learners.

9. The method of claim 1 , wherein method includes utilizing weighted stratified sampling for selecting the single summary.

10. A system, comprising:

an input interface to the system, wherein the input interface is configured to receive data associated with the system; and

one or more processors in communication with the input interface, the one or more processors processor programmed to:

receive tabular data associated with one or more records;

convert the tabular data to a text representation indicative of the tabular data;

generate metadata associated with the text representation of the tabular data, wherein the metadata is indicative of a description of the tabular data;

output one or more natural language data descriptions indicative of the tabular data in response to utilizing a large language model (LLM) and zero-shot prompting of the metadata and text representation of the tabular data, wherein the LLM includes a neural network with a plurality of parameters;

output one or more summaries utilizing the LLM and the one or more natural language data descriptions, wherein the one or more summaries include less text than the one or more natural language data descriptions;

for one or more iterations, select a single summary of the one or more summaries in response to the single summary having a smallest validation rate;

receive a query associated with the tabular data;

for one or more iterations, output one or more predictions associated with the query utilizing the LLM on the single summary and the query; and

in response to meeting a convergence threshold with the one or more predictions generated from the one or more iterations of outputting one or more predictions, output a final prediction associated with the query.

11. The system of claim 10 , wherein the processor is further programmed to output the final prediction in response to a weighted-majority vote of all of the one or more predictions generated from the one or more iterations.

12. The system of claim 10 , wherein numerical features associated with the tabular data are encoded into percentiles.

13. The system of claim 10 , wherein the one or more processors are programmed to output one or more summaries utilizing the LLM and appending a prompt on the one or more natural language data descriptions.

14. The system of claim 10 , wherein the one or more processors are collectively programmed to execute steps.

15. The system of claim 10 , wherein the one or more processors includes a single processor programmed to execute steps.

16. The system of claim 10 , wherein the one or more processors are programmed to select a representative subset of the one or more summaries utilizing weighted stratified sampling.

17. The system of claim 10 , wherein the one or more processors are programmed to select a representative subset of the one or more summaries utilizing weighted stratified sampling.

18. The system of claim 10 , wherein the one or more processors are programmed to select a representative subset of the one or more summaries from a fixed number of iterations.

19. The system of claim 10 , wherein the processor is programmed to, for one or more iterations, output the one or more natural language data descriptions and output the one or more summaries.

20. A computer-implemented method for natural-language processing, comprising:

receiving tabular data associated with one or more records;

converting the tabular data to a text representation indicative of the tabular data;

generating metadata associated with the text representation of the tabular data, wherein the metadata is indicative of a description of the tabular data;

outputting one or more natural language data descriptions indicative of the tabular data in response to utilizing a large language model (LLM) and zero-shot prompting of the metadata and text representation of the tabular data;

outputting one or more summaries utilizing the LLM and appending a prompt on the one or more natural language data descriptions, wherein the one or more summaries include less text than the one or more natural language data descriptions;

for one or more iterations, selecting a single summary of the one or more summaries in response to the single summary having a smallest validation rate;

creating a subset of single summaries;

receiving a query associated with the tabular data;

for one or more iterations, output one or more predictions associated with the query utilizing the LLM on one of the single summaries of the subset of single summaries and the query;

creating a subset of one or more predictions; and

in response to meeting a convergence threshold with the subset of one or more predictions, output a final prediction associated with the query, wherein the final prediction is selected in response to a weighted-majority vote of the subset of one or more predictions.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2023
From: KOLTER, JEREMY; QIU, CHEN; LIN, WAN-YI; CABRITA CONDESSA, FILIPE J.
To: ROBERT BOSCH GMBH
Reel/Frame 064626/0419 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2023
From: MANIKANDAN, HARIHARAN; JIANG, YIDING
To: CARNEGIE MELLON UNIVERSITY
Reel/Frame 064626/0446 →
Continuity (1)
Related Publication 20240412004A1 · Dec 12, 2024
References Cited (54)
US 20200356891A1 · Saito · 2020 [cited by examiner]
US 20220005463A1 · Bender · 2022 [cited by examiner]
US 20240289371A1 · Martínez Galindo · 2024 [cited by examiner]
US 20240311619A1 · Licato · 2024 [cited by examiner]
US 20240330600A1 · Perlitz · 2024 [cited by examiner]
WO WO2015009586A2 · 2015 [cited by examiner]
Zhang, Shuo, et al. “Summarizing and exploring tabular data in conversational search.” Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval. 2020. (Year: 2020). [cited by examiner]
Vadim Borisova et al., “Deep Neural Networks and Tabular Data: A Survey,” arXiv:2110.01889v1 [cs.LG] Oct. 5, 2021, 19 Pages. [cited by applicant]
Tom B. Brown et al., “Language Models are Few-Shot Learners.” 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada, 25 Pages. [cited by applicant]
Tianqi Chen et al., “XGBoost: A Scalable Tree Boosting System.” KDD '16, Aug. 13-17, 2016, San Francisco, CA, USA, 10 Pages. [cited by applicant]
Ganqu Cui et al., “Prototypical Verbalizer for Prompt-based Few-shot Tuning.” arXiv:2203.09770v1 [cs.CL] Mar. 18, 2022, 11 Pages. [cited by applicant]
Jacob Devlin e al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.” arXiv:1810.04805v1 [cs.CL] Oct. 11, 2018, 14 Pages. [cited by applicant]
Shizhe Diao et al., “Active Prompting with Chain-of-Thought for Large Language Models.” arXiv:2302.12246v3 [cs.CL] May 23, 2023, 20 Pages. [cited by applicant]
Tuan Dinh et al., “LIFT: Language-Interfaced Fine-Tuning for Non-Language Machine Learning Tasks.” arXiv:2206.06565v4 [cs.LG] Oct. 31, 2022, 54 Pages. [cited by applicant]
Yoav Freund et al., “A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting.” journal of computer and system sciences 55, 119-139 (1997). [cited by applicant]
Jerome Friedman. “Greedy function approximation: a gradient boosting machine.” Annals of statistics 2001, 39 Pages. [cited by applicant]
Jerome Friedman. “Stochastic gradient boosting.” Computational statistics & data analysis 2002, vol. 38, No. 4, 10 Pages. [cited by applicant]
Yury Gorishniy et al., “Revisiting Deep Learning Models for Tabular Data.” 35th Conference on Neural Information Processing Systems (NeurIPS 2021), 12 Pages. [cited by applicant]
Patrick Haluptzok et al., “Language Models Can Teach Themselves to Program Better.” arXiv:2207.14502v1 [cs.LG] Jul. 29, 2022, 15 Pages. [cited by applicant]
Yaru Hao et al., “Structured Prompting: Scaling In-Context Learning to 1,000 Examples.” arXiv:2212.06713v1 [cs.CL] Dec. 13, 2022, 14 Pages. [cited by applicant]
Jonathan Herzig et al., “TAPAS: Weakly Supervised Table Parsing via Pre-training.” arXiv:2004.02349v2 [cs.IR] Apr. 21, 2020, 14 Pages. [cited by applicant]
Namgyu Ho et al., “Large Language Models Are Reasoning Teachers.” abiliarXiv:2212.10071v1 [cs.CL] Dec. 20, 2022, 22 Pages. [cited by applicant]
Noah Hollmann et al., “TABPFN: a Transformer That Solves Small Tabular Classification Problems in a Second.” arXiv:2207.01848v6 [cs.LG] Sep. 16, 2023, 37 Pages. [cited by applicant]
Bairu Hou et al., “Promptboosting: Black-Box Text Classification With Ten Forward Passes.” arXiv:2212.09257v1 [cs.CL] Dec. 19, 2022, 14 Pages. [cited by applicant]
Shengding Hu et al., “Knowledgeable Prompt-tuning: Incorporating Knowledge into Prompt Verbalizer for Text Classification.” arXiv:2108.02035v1 [cs.CL] Aug. 4, 2021, 12 Pages. [cited by applicant]
Jiaxin Huang et al., “Large Language Models Can Self-Improve.” arXiv:2210.11610v2 [cs.CL] Oct. 25, 2022, 19 Pages. [cited by applicant]
Takeshi Kojima et al., “Large Language Models are Zero-Shot Reasoners.” arXiv:2205.11916v1 [cs.CL] May 24, 2022, 36 Pages. [cited by applicant]
Brian Lester et al., “The Power of Scale for Parameter-Efficient Prompt Tuning.” bearXiv:2104.08691v2 [cs.CL] Sep. 2, 2021, 15 Pages. [cited by applicant]
Aitor Lewkowycz et al., “Solving Quantitative Reasoning Problems with Language Models.” arXiv:2206.14858v2 [csCL] Jul. 1, 2022, 54 Pages. [cited by applicant]
Pengfei Liu et al., “Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing.” ACM Computing Surveys, vol. 55, No. 9, Article 195. Publication date: Jan. 2023, 35 Pages. [cited by applicant]
Xiao Liu et al., “GPT Understands, Too.” arXiv:2103.10385v1 [cs.CL] Mar. 18, 2021, 10 Pages. [cited by applicant]
Yinhan Liu et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach.” arXiv:1907.11692v1 [cs.CL] Jul. 26, 2019, 13 Pages. [cited by applicant]
Avanika Narayan et al., “Can Foundation ModelsWrangle Your Data?” arXiv:2205.09911v2 [cs.LG] Dec. 24, 2022, 12 Pages. [cited by applicant]
Long Ouyang et al. “Training language models to follow instructions with human feedback.” 36th Conference on Neural Information Processing Systems (NeurIPS 2022), 15 Pages. [cited by applicant]
Guanghui Qin et al., “Learning How to Ask: Querying LMs with Mixtures of Soft Prompts.” comarXiv:2104.06599v1 [cs.CL] Apr. 14, 2021, 11 Pages. [cited by applicant]
Nathanael Carraz Rakotonirina et al., “Can Discrete Information Extraction Prompts Generalize Across Language Models?” arXiv:2302.09865v2 [cs.CL] Mar. 7, 2023, 17 Pages. [cited by applicant]
Laria Reynolds et al., “Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm.” arXiv:2102.07350v1 [cs.CL] Feb. 15, 2021, 10 Pages. [cited by applicant]
Swarnadeep Saha et al., “MURMUR: Modular Multi-Step Reasoning for Semi-Structured Data-to-Text Generation.” arXiv:2212.08607v1 [cs.CL] Dec. 16, 2022, 22 Pages. [cited by applicant]
Bernhard Schäfly et al., “Hopular: Modern Hopfield Networks for Tabular Data.” arXiv:2206.00664v1 [cs.LG] Jun. 1, 2022, 21 Pages. [cited by applicant]
Robert E. Schapire. “The strength of weak learnability.” Machine learning, 5:197-227, 1990. [cited by applicant]
Taylor Shin et al., “Autoprompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts.” arXiv:2010.15980v2 [cs.CL] Nov. 7, 2020, 15 Pages. [cited by applicant]
Ravid Shwartz-Ziv et al., “Tabular Data: Deep Learning is Not All You Need.” Information Fusion 2022, 11 Pages. [cited by applicant]
Joaquin Vanschoren et al., “OpenML: networked science in machine learning.” arXiv:1407.7722v1 [cs.LG] Jul. 29, 2014, 13 Pages. [cited by applicant]
Ashish Vaswani et al., “Attention Is All You Need.” 31stConferenceonNeuralInformationProcessingSystems (NIPS2017),LongBeach,CA,USA, 11 Pages. [cited by applicant]
Han Wang et al., “Automatic Multi-Label Prompting: Simple and Interpretable Few-Shot Classification.” arXiv:2204.06305v2 [cs.CL] Apr. 14, 2022, 10 Pages. [cited by applicant]
Xuezhi Wang et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models.” arXiv:2203.11171v1 [cs.CL] Mar. 21, 2022, 15 Pages. [cited by applicant]
Zifeng Wang et al., “Learning to Prompt for Continual Learning.” Proceedings of the IEEE/CVF Conference on 401 Computer Vision and Pattern Recognition, pp. 139-149, 2022. [cited by applicant]
Jason Wei et al., “Chain of Thought Prompting Elicits Reasoning in Large Language Models.” arXiv:2201.11903v1 [cs.CL] Jan. 28, 2022, 24 Pages. [cited by applicant]
Pengcheng Yin et al., “TABERT: Pretraining for Joint Understanding of Textual and Tabular Data.” modarXiv: 2005.08314v1 [cs.CL] May 17, 2020, 15 Pages. [cited by applicant]
Wenhao Yu et al., “Generate Rather Than Retrieve: Large Language Models Are Strong Context Generators.” arXiv:2209.10063v1 [cs.CL] Sep. 21, 2022, 24 Pages. [cited by applicant]
Zhuosheng Zhangy et al., “Automatic Chain of Thought Prompting in Large Language Models.” arXiv:2210.03493v1 [cs.CL] Oct. 7, 2022, 25 Pages. [cited by applicant]
Hattie Zhou et al., “Teaching Algorithmic Reasoning via In-context Learning.” arXiv:2211.09066v1 [cs.LG] Nov. 15, 2022, 37 Pages. [cited by applicant]
Yongchao Zhou et al., “Large Language Models Are Human-Level Prompt Engineers.” arXiv:2211.01910v1 [cs.LG] Nov. 3, 2022, 40 Pages. [cited by applicant]
Frank Nielsen et al., “Introduction to HPC with MPI for Data Science.” Springer International Publishing Switzerland 2016, 304 Pages. [cited by applicant]
Cited By (1)
US 12,675,831