IP Library › Granted Patent US 12,743,451
Granted Patent B2
US 12,743,451 · App. 19/271,920 · Granted Sep 22, 2026

Method and system for multi-level artificial intelligence supercomputer design featuring sequencing of large language models

Inventors: Vijay Madisetti (Alpharetta, GA); Arshdeep Bahga (Chandigarh, IN)
Assignee: Vijay Madisetti
G06F16/3329G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,451
App. No.
19/271,920
Granted
Sep 22, 2026
Kind
B2
Abstract

A system and method for creating a merged large language model (h-LLM) using a bagging approach including receiving input data at a computer system, creating a plurality of data subsets from the input data, training a plurality of h-LLMs, each h-LLM of the plurality of h-LLMs being trained on a respective data subset of the plurality of data subsets, creating a merged h-LLM by merging the plurality of h-LLMs, and outputting the merged h-LLM.

Claims (40)

1 . A method for creating a merged family of large language models (h-LLM) using a bagging approach comprising:

receiving input data at a computer system comprising a processor, non-transitory storage medium, and software stored on the non-transitory storage medium;

creating a plurality of data subsets from the input data;

training a plurality of h-LLMs, each h-LLM of the plurality of h-LLMs being trained on a respective data subset of the plurality of data subsets;

creating a merged h-LLM by merging the plurality of h-LLMs; and

outputting the merged h-LLM.

2 . The method of claim 1 wherein creating the plurality of data subsets comprises dividing the input data to create multiple data subsets.

3 . The method of claim 1 wherein merging the plurality of h-LLMs comprises combining model parameters from each h-LLM through a merging or fusing process.

4 . The method of claim 1 wherein the h-LLMs of the plurality of h-LLMs are trained on different computational resources within a distributed computing environment.

5 . The method of claim 1 wherein the h-LLMs of the plurality of h-LLMs are trained concurrently.

6 . The method of claim 1 wherein the merged h-LLM has at least one of a higher precision, a higher accuracy, and an improved stability than each h-LLM of the plurality of h-LLMs.

7 . A method for creating an enhanced family of large language models (h-LLM) using a boosting approach comprising:

training a first h-LLM using original input data;

testing the first h-LLM to generate first output results;

generating a first weighted data by assigning one or more weights to the first output results;

training a second h-LLM using the first weighted data;

generating a sequence of h-LLMs with increasing levels of at least one of precision and accuracy by iteratively testing the second h-LLM and subsequent h-LLMs, generating subsequent output results, generating subsequent weighted data from the subsequent output results, and training subsequent h-LLMs on the subsequent weighted data;

merging the sequence of h-LLMs to create a merged enhanced h-LLM; and

outputting the merged enhanced h-LLM for use in processing language tasks.

8 . The method of claim 7 wherein the sequence of h-LLMs comprises at least three h-LLMs, each subsequent h-LLM in the sequence having higher accuracy than an immediately preceding h-LLM in the sequence.

9 . A method for creating a specialized family of large language models (h-LLM) through extraction comprising:

receiving a general purpose h-LLM;

identifying a specialized task from a group of tasks;

extracting task-specific knowledge from the general purpose h-LLM corresponding to the specialized task;

creating the specialized h-LLM having reduced computational requirements compared to the general purpose h-LLM while maintaining performance for the specialized task; and

configuring the specialized h-LLM to process prompts related to the specialized task.

10 . The method of claim 9 wherein extracting task-specific knowledge comprises identifying and utilizing one or more neural network components relevant to the specialized task.

11 . The method of claim 9 wherein the specialized h-LLM comprises fewer parameters than the general purpose h-LLM while achieving comparable accuracy for the specialized task.

12 . The method of claim 9 wherein creating the specialized h-LLM comprises using the general purpose h-LLM to guide the creation process.

13 . The method of claim 9 wherein the general purpose h-LLM is trained on broad domain data.

14 . The method of claim 9 wherein the group of tasks consists of sentiment analysis, question answering, information extraction, image captioning, object recognition, instruction following, classification, inferencing, and sentence similarity.

15 . A system for creating specialized families of large language models (h-LLMs) comprising:

a processor;

a non-transitory computer-readable storage medium positioned in communication with the processor; and

software stored on the storage medium that, when executed by the processor, is operable to:

receive a general purpose h-LLM;

identify a specialized task;

extract task-specific knowledge from the general purpose h-LLM;

generate a specialized h-LLM with reduced computational requirements while maintaining task-specific performance related to the specialized task; and

deploy the specialized h-LLM for processing prompts related to the specialized task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2025
From: BAHGA, ARSHDEEP
To: MADISETTI, VIJAY
Reel/Frame 071753/0307 →
Continuity (7)
Continuation 19057610 · Feb 19, 2025
Continuation 18801421 · Aug 12, 2024
Continuation 18470487 · Sep 20, 2023
Continuation 18348692 · Jul 7, 2023
Provisional Application 63469571 · May 30, 2023
Provisional Application 63463913 · May 4, 2023
Related Publication 20250342186A1 · Nov 6, 2025
References Cited (17)
US 12210550B2 · Madisetti et al. · 2025 [cited by applicant]
US 12321371B1 · Madisetti et al. · 2025 [cited by applicant]
US 12386871B2 · Madisetti et al. · 2025 [cited by applicant]
US 12430370B2 · Madisetti et al. · 2025 [cited by applicant]
US 20080133434A1 · Asar · 2008 [cited by examiner]
US 20170124074A1 · Cama · 2017 [cited by examiner]
US 20170262770A1 · Purdy · 2017 [cited by examiner]
US 20180308003A1 · Singh · 2018 [cited by examiner]
US 20210182662A1 · Lai · 2021 [cited by examiner]
US 20220245343A1 · Brun · 2022 [cited by examiner]
US 20220269986A1 · Filikov · 2022 [cited by examiner]
US 20230351203A1 · Ozay · 2023 [cited by examiner]
US 20230418694A1 · Meen · 2023 [cited by examiner]
US 20240054338A1 · Korbak · 2024 [cited by examiner]
US 20240420491A1 · Park · 2024 [cited by examiner]
Raphael Tang, Yao Lu, Linqing Liu, Lili Mou, Olga Vechtomova, Jimmy Lin, “Distilling Task-Specific Knowledge from BERT into Simple Neural Networks”, 2019, https://arxiv.org/abs/1903.12136 (Year: 2019). [cited by examiner]
Non-Final Office Action received in U.S. Appl. No. 19/271,920 issued on Feb. 18, 2025; 33 pages. [cited by applicant]