IP Library › Granted Patent US 12,387,050
Granted Patent B1
US 12,387,050 · App. 19/051,193 · Granted Aug 12, 2025

Multi-stage LLM with unlimited context

Inventors: Brian Galvin (Silverdale, WA); Alan McCord (Forney, TX)
Assignee: ATOMBEAM TECHNOLOGIES INC.
G06F40/30G06F16/3325G06F16/3329
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,050
App. No.
19/051,193
Filed
Feb 12, 2025
Granted
Aug 12, 2025
Kind
B1
Art Unit
2657
USPC
704/9
Abstract

A system and method for efficient natural language processing combines large and small language models with a thought caching architecture. The system includes a router that directs prompts either to a large language model for thought generation or to a thought cache containing previously generated thoughts. When using the large model, generated thoughts are combined with the original prompt and routed through a smaller language model to produce responses. The thought cache stores reasoning patterns that can be retrieved and reused, eliminating the need to regenerate similar thoughts for related prompts. The system supports both local and cloud-based caching, enabling personal and enterprise-wide thought storage and retrieval. This architecture reduces computational overhead while maintaining reasoning capabilities, effectively extends context windows beyond traditional limits, and enables efficient scaling across different deployment scenarios. The system can operate with reduced resources by leveraging cached thoughts without requiring constant access to the large model.

Claims (52)

1. A computer system comprising a hardware memory, wherein the computer system is configured to execute software instructions stored on nontransitory machine-readable storage media that:

receive a prompt from a user;

process the prompt into a plurality of corresponding thoughts using a first large language model;

route both the prompt and the plurality of thoughts through a second large language model that has fewer parameters than the first large language model;

associate each corresponding thought in the plurality of corresponding thoughts to a portion of the prompt;

cache each associated corresponding thought, wherein a plurality of associated corresponding thoughts may be retrieved from a thought cache when the portion of the prompt they correspond to is present in a future prompt; and

generate a response to the prompt by processing the plurality of thoughts and the prompt through the second large language model.

2. The computer system of claim 1 , wherein the computer system is further configured to execute software instructions stored on nontransitory machine-readable storage media that:

analyze the prompt using a prompt analyzer to determine key concepts and requirements;

query the thought cache to determine if similar thoughts exist for the determined key concepts; and

synthesize new thoughts when similar thoughts exist but do not fully address the prompt requirements.

3. The computer system of claim 1 , wherein the thought cache comprises at least a local cache stored on an edge device and a global cache stored in a cloud environment, wherein the global cache is accessible by a plurality of edge devices.

4. The computer system of claim 3 , wherein the global cache is organized into specialized domains, and thoughts are categorized and stored according to their relevant domain.

5. The computer system of claim 1 , wherein caching each associated corresponding thought comprises:

evaluating relevance of the thought to the portion of the prompt;

assigning metadata tags based on the evaluation;

storing the thought with vector embeddings for similarity searching; and

indexing the thought for retrieval.

6. The computer system of claim 1 , wherein caching each associated corresponding thought comprises:

storing each associated corresponding thought in a short-term memory as explicit reasoning text;

progressively compressing a plurality of older associated corresponding thoughts into consolidated representations; and

storing a plurality of compressed older associated corresponding thoughts in a long-term memory.

7. The computer system of claim 1 , wherein the computer system is further configured to execute software instructions stored on nontransitory machine-readable storage media that:

maintain a shared thought cache accessible by multiple AI agents;

enable thought transfer between specialized reasoning modules; and

coordinate collaborative reasoning across multiple model instances.

8. A method for encrypted data compression with a hardware management layer, comprising the steps of:

receiving a prompt from a user;

processing the prompt into a plurality of corresponding thoughts using a first large language model;

routing both the prompt and the plurality of thoughts through a second large language model that has fewer parameters than the first large language model;

associating each corresponding thought in the plurality of corresponding thoughts to a portion of the prompt;

caching each associated corresponding thought, wherein a plurality of associated corresponding thoughts may be retrieved from a thought cache when the portion of the prompt they correspond to is present in a future prompt; and

generating a response to the prompt by processing the plurality of thoughts and the prompt through the second large language model.

9. The method of claim 8 , further comprising the steps of:

analyzing the prompt using a prompt analyzer to determine key concepts and requirements;

querying the thought cache to determine if similar thoughts exist for the determined key concepts; and

synthesizing new thoughts when similar thoughts exist but do not fully address the prompt requirements.

10. The method of claim 8 , wherein the thought cache comprises at least a local cache stored on an edge device and a global cache stored in a cloud environment, wherein the global cache is accessible by a plurality of edge devices.

11. The method of claim 10 , wherein the global cache is organized into specialized domains, and thoughts are categorized and stored according to their relevant domain.

12. The method of claim 8 , wherein caching each associated corresponding thought comprises:

evaluating relevance of the thought to the portion of the prompt;

assigning metadata tags based on the evaluation;

storing the thought with vector embeddings for similarity searching; and

indexing the thought for retrieval.

13. The method of claim 8 , wherein caching each associated corresponding thought comprises:

storing each associated corresponding thought in a short-term memory as explicit reasoning text;

progressively compressing a plurality of older associated corresponding thoughts into consolidated representations; and

storing a plurality of compressed older associated corresponding thoughts in a long-term memory.

14. The method of claim 8 , further comprising the steps of:

maintaining a shared thought cache accessible by multiple AI agents;

enabling thought transfer between specialized reasoning modules; and

coordinating collaborative reasoning across multiple model instances.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2025
From: GALVIN, BRIAN; MCCORD, ALAN
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 070916/0597 →
References Cited (29)
US 12223456B1 · Manohar · 2025 [cited by examiner]
US 20020091801A1 · Lewin · 2002 [cited by examiner]
US 20180174055A1 · Tirumale · 2018 [cited by examiner]
US 20190174514A1 · Ramesh · 2019 [cited by examiner]
US 20200034914A1 · Boss · 2020 [cited by examiner]
US 20200073983A1 · Sen · 2020 [cited by examiner]
US 20200285704A1 · Rajani · 2020 [cited by examiner]
US 20200336562A1 · Luft · 2020 [cited by examiner]
US 20200351344A1 · Das Gupta · 2020 [cited by examiner]
US 20210073808A1 · Gu · 2021 [cited by examiner]
US 20210406224A1 · Neufeld · 2021 [cited by examiner]
US 20220138156A1 · Wang · 2022 [cited by examiner]
US 20230316006A1 · Tunstall-Pedoe · 2023 [cited by examiner]
US 20240104391A1 · Higgins · 2024 [cited by examiner]
US 20240160955A1 · Zhao · 2024 [cited by examiner]
US 20240256965A1 · Chung · 2024 [cited by examiner]
US 20240354320A1 · Procter · 2024 [cited by examiner]
US 20240386015A1 · Crabtree · 2024 [cited by examiner]
US 20240411809A1 · Najafirad · 2024 [cited by examiner]
US 20240428008A1 · Abraham · 2024 [cited by examiner]
US 20250028882A1 · Ataei · 2025 [cited by examiner]
US 20250094455A1 · Bista · 2025 [cited by examiner]
US 20250148203A1 · Pan · 2025 [cited by examiner]
US 20250165718A1 · Seo · 2025 [cited by examiner]
US 20250191369A1 · Huang · 2025 [cited by examiner]
Ramirez et al., Cache & Distil: Optimising API Calls to Large Language Models, 2023, ournal reference: Findings of the Association for Computational Linguistics: ACL 2024, Subjects: Computation and Language (cs.CL); Mac… [cited by examiner]
Schroeder, title={VectorQ: Advanced Semantic Prompt Caching With Dynamic Thresholds and Performance-Based Clustering}, School OfComputation, Information and Technology—Informatics, pp. 1-63, 2024 (Year: 2024). [cited by examiner]
Gao et al., title={Memory sharing for large language model based agents}, journal={arXiv preprint arXiv:2404.09982}, pp. 1-14 (Year: 2024). [cited by examiner]
Gim, In et al., “Prompt Cache: Modular Attention Reuse for Low-Latency Inference”, Proceedings of the 5th MLSys Conference, Santa Clara, CA, USA, 2024. [cited by applicant]
Cited By (2)
US 12,572,830 US 12,626,167