IP Library › Granted Patent US 12,399,920
Granted Patent B2
US 12,399,920 · App. 19/057,610 · Granted Aug 26, 2025

Method and system for multi-level artificial intelligence supercomputer design featuring sequencing of large language models

Inventors: Vijay Madisetti (Alpharetta, GA); Arshdeep Bahga (Chandigarh, IN)
Assignee: Vijay Madisetti
G06F16/3329G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,399,920
App. No.
19/057,610
Filed
Feb 19, 2025
Granted
Aug 26, 2025
Kind
B2
Examiner
YEN, ERIC L
Art Unit
2658
USPC
704/9
Abstract

A method for assigning tasks to LLMs using h-including receiving documents relevant to embeddings, generating context-aware prompts responsive to at least one of a received prompt, derived prompts, and the received knowledge documents, each context-aware prompt corresponding to a specialized task of the one or more specialized tasks, transmitting each context-aware prompt to a respective h- that is configured to specialize in processing prompts having a specialty corresponding to the specialized task of the context-aware prompt, and receiving results from the h-LLMs.

Claims (51)

1. A method for assigning tasks to LLMs using one or more families of large language models (h-LLMs) by a computer comprising of a processor, non-transitory storage medium, and software on the storage medium, the method comprising:

receiving one or more received knowledge documents relevant to a plurality of prompt embeddings at an input broker;

generating a plurality of context-aware prompts by the input broker responsive to at least one of a received prompt, a plurality of derived prompts, and the one or more received knowledge documents, each context-aware prompt of the plurality of context-aware prompts corresponding to a specialized task of one or more specialized tasks;

transmitting each context-aware prompt of the plurality of context-aware prompts to a respective h-LLM of a plurality of h-LLMs that is configured to specialize in processing prompts having a specialty corresponding to the specialized task of the context-aware prompt; and

receiving a plurality of produced results from at least one of the plurality of h-LLMs by an output broker.

2. The method of claim 1 wherein the plurality of h-LLMs operate as a network of communicating h-LLMs.

3. The method of claim 2 wherein at least one of the received prompt or the plurality of produced results is processed and transmitted to another h-LLM of the plurality of h-LLMs in the network.

4. The method of claim 2 wherein the network of communicating h-LLMs is organized in one of a serial structure, a parallel structure, or a hybrid structure.

5. The method of claim 1 wherein at least one of the input broker and the output broker is configured to coordinate the plurality of context-aware prompts to be processed by a sequence of h-LLMs of the plurality of h-LLMs in response to at least one of the received prompt or the derived prompts.

6. The method of claim 5 wherein at least one of the input broker and the output broker are configured to coordinate the sequence of h-LLMs of the plurality of h-LLMs in response to each of workload, technical accuracy, and business value metrics of the plurality of h-LLMs.

7. The method of claim 1 wherein each h-LLM of the plurality of h-LLMs are configured to communicate at least one of the received prompt and the plurality of produced results to another h-LLM of the plurality of h-LLMs via one of the input broker and the output broker.

8. The method of claim 1 wherein each h-LLM of the plurality of h-LLMs communicates at least one of the received prompt and the plurality of produced results to another h-LLM of the plurality of h-LLMs sequentially responsive to receiving the at least one of the received prompt and the plurality of produced results from at least one of the input broker and the output broker.

9. The method of claim 1 wherein the output broker is further configured to:

generate reviewed produced results by assigning at least one of a priority to each produced result and a weight to each produced result;

generate filtered produced results by filtering the reviewed produced results;

rank each filtered produced result; and

cache each filtered produced result.

10. The method of claim 1 wherein:

at least one of the input broker and the output broker is configured to update an index of derived prompts responsive to the plurality of produced results; and

generating the plurality of derived prompts comprises identifying one or more derived prompts comprised by the index of derived prompts responsive to the received prompt.

11. The method of claim 1 wherein at least one of the input broker and the output broker comprises one or more LLMs.

12. A system for assigning tasks to LLMs using one or more families of large language models (h-LLMs) comprising:

a processor;

a network communication device positioned in communication with the processor and operable to communicate across a computer network; and

a non-transitory computer-readable storage medium positioned in communication with the processor and having stored thereon software that, when executed by the processor, is operable to:

operate an input broker to:

receive one or more received knowledge documents relevant to a plurality of prompt embeddings; and

generate a plurality of context-aware prompts responsive to at least one of a received prompt, a plurality of derived prompts, and the one or more received knowledge documents, each context-aware prompt of the plurality of context-aware prompts corresponding to a specialized task of one or more specialized tasks;

transmit each context-aware prompt of the plurality of context-aware prompts to a respective h-LLM of a plurality of h-LLMs that is configured to specialize in processing prompts having a specialty corresponding to the specialized task of the context-aware prompt; and

operate an output broker to receive a plurality of produced results from at least one of the plurality of h-LLMs.

13. The system of claim 12 wherein the plurality of h-LLMs operate as a network of communicating h-LLMs.

14. The system of claim 13 wherein at least one of the received prompt or the plurality of produced results is processed and transmitted to another h-LLM of the plurality of h-LLMs in the network.

15. The system of claim 13 wherein the network of communicating h-LLMs is organized in one of a serial structure, a parallel structure, or a hybrid structure.

16. The system of claim 12 wherein the software is operable to, when executed by the processor, operate at least one of the input broker and the output broker to coordinate the plurality of context-aware prompts to be processed by a sequence of h-LLMs of the plurality of h-LLMs in response to at least one of the received prompt or the derived prompts.

17. The system of claim 16 wherein the software is operable to, when executed by the processor, operate at least one of the input broker and the output broker to coordinate the sequence of h-LLMs of the plurality of h-LLMs in response to each of workload, technical accuracy, and business value metrics of the plurality of h-LLMs.

18. The system of claim 12 wherein each h-LLM of the plurality of h-LLMs are configured to communicate at least one of the received prompt and the plurality of produced results to another h-LLM of the plurality of h-LLMs via one of the input broker and the output broker.

19. The system of claim 12 wherein each h-LLM of the plurality of h-LLMs communicates at least one of the received prompt and the plurality of produced results to another h-LLM of the plurality of h-LLMs sequentially responsive to receiving the at least one of the received prompt and the plurality of produced results from at least one of the input broker and the output broker.

20. The system of claim 12 wherein the software is operable to, when executed by the processor, operate the output broker to:

generate reviewed produced results by assigning at least one of a priority to each produced result and a weight to each produced result;

generate filtered produced results by filtering the reviewed produced results;

rank each filtered produced result; and

cache each filtered produced result.

21. The system of claim 12 wherein the software is operable to, when executed by the processor:

operate at least one of the input broker and the output broker to update an index of derived prompts responsive to the plurality of produced results; and

generate the plurality of derived prompts by identifying one or more derived prompts comprised by the index of derived prompts responsive to the received prompt.

22. The system of claim 12 wherein the software is operable to, when executed by the processor, operate at least one of the input broker and the output broker to comprise one or more LLMs.

23. A system for assigning tasks to LLMs using one or more families of large language models (h-LLMs) comprising:

means for receiving one or more received knowledge documents relevant to a plurality of prompt embeddings at an input broker;

means for generating a plurality of context-aware prompts by the input broker responsive to at least one of a received prompt, a plurality of derived prompts, and the one or more received knowledge documents, each context-aware prompt of the plurality of context-aware prompts corresponding to a specialized task of one or more specialized tasks;

means for transmitting each context-aware prompt of the plurality of context-aware prompts to a respective h-LLM of a plurality of h-LLMs that is configured to specialize in processing prompts having a specialty corresponding to the specialized task of the context-aware prompt; and

means for receiving a plurality of produced results from at least one of the plurality of h-LLMs by an output broker.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2025
From: BAHGA, ARSHDEEP
To: MADISETTI, VIJAY
Reel/Frame 070348/0091 →
Continuity (6)
Continuation 18801421 · Aug 12, 2024
Continuation 18470487 · Sep 20, 2023
Continuation 18348692 · Jul 7, 2023
Provisional Application 63469571 · May 30, 2023
Provisional Application 63463913 · May 4, 2023
Related Publication 20250190462A1 · Jun 12, 2025
References Cited (27)
US 11663201B2 · Alakuijala · 2023 [cited by examiner]
US 11971914B1 · Watson · 2024 [cited by examiner]
US 20110153744A1 · Brown · 2011 [cited by examiner]
US 20140046978A1 · Krishnaprasad · 2014 [cited by examiner]
US 20210133264A1 · Tiwari · 2021 [cited by examiner]
US 20220292266A1 · Dey · 2022 [cited by examiner]
US 20230083512A1 · Newman · 2023 [cited by examiner]
US 20230315999A1 · Mohammed · 2023 [cited by examiner]
US 20230316104A1 · Ott · 2023 [cited by examiner]
US 20240126795A1 · Zhong · 2024 [cited by examiner]
US 20240202225A1 · Siebel · 2024 [cited by examiner]
US 20240281472A1 · LaRhette · 2024 [cited by examiner]
US 20240289560A1 · Kelly · 2024 [cited by examiner]
US 20240320444A1 · Maschmeyer · 2024 [cited by examiner]
US 20240330863A1 · Hajarnis · 2024 [cited by examiner]
US 20240346162A1 · Luitjens · 2024 [cited by examiner]
US 20240346256A1 · Qin · 2024 [cited by examiner]
US 20240354436A1 · Mukherjee · 2024 [cited by examiner]
US 20240404632A1 · Ho · 2024 [cited by examiner]
US 20240406125A1 · Hine · 2024 [cited by examiner]
US 20240412029A1 · Yang · 2024 [cited by examiner]
US 20240414108A1 · Sun · 2024 [cited by examiner]
US 20240414191A1 · Humphrey · 2024 [cited by examiner]
US 20250005293A1 · Nguyen · 2025 [cited by examiner]
US 20250077844A1 · Lin · 2025 [cited by examiner]
US 20250086394A1 · Reddy · 2025 [cited by examiner]
US 20250094703A1 · Malak · 2025 [cited by examiner]
Cited By (1)
US 12,639,346