IP Library › Granted Patent US 12,579,437
Granted Patent B2
US 12,579,437 · App. 19/294,125 · Granted Mar 17, 2026

Mobile-optimized multi-stage LLM with federated persistent cognitive architecture

Inventors: Brian Galvin (Silverdale, WA); Alan McCord (Forney, TX)
Assignee: ATOMBEAM TECHNOLOGIES INC.
G06N3/082
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,437
App. No.
19/294,125
Granted
Mar 17, 2026
Kind
B2
Abstract

A system and method for extending mobile-optimized multi-stage language model processing with federated persistent cognitive architecture. The system processes prompts through a first large language model to generate “thoughts,” which are cached and processed with the original prompt through a smaller language model. Building upon the three-tier thought caching, the system implements a federated multi-tier hierarchy with local device, domain-specific branch, and global collective caches. A federated cognitive orchestrator coordinates operations across multiple domain-specialized instances, managing thought routing, state synchronization, and cross-domain knowledge sharing while maintaining domain boundaries. During user inactivity, autonomous reasoning continues in cloud environments, generating insights from existing thoughts and interaction history. The system performs memory consolidation, thought cache optimization, and cross-domain pattern recognition without consuming mobile device resources, while maintaining privacy boundaries. This persistent cognitive architecture functions as an evolving reasoning partner rather than merely a responsive tool.

Claims (35)

1 . A computer system comprising a hardware memory, wherein the computer system is configured to execute software instructions stored on nontransitory machine-readable storage media that:

receive a prompt from a user;

process the prompt into a plurality of corresponding thoughts using a first reasoning large language model;

route both the prompt and the plurality of thoughts through a second large language model that has fewer parameters than the first reasoning large language model;

associate each corresponding thought in the plurality of corresponding thoughts to a portion of the prompt;

cache each associated corresponding thought in a federated multi-tier thought caching hierarchy comprising at least a local device cache, a domain-specific branch instance cache, and a global collective cache;

generate a response to the prompt by processing the plurality of thoughts and the prompt through the second large language model;

continue autonomous reasoning operations without user interaction in a cloud environment, generating new thoughts based on existing cached thoughts and user interaction history; and

coordinate federated cognitive operations across multiple domain-specialized instances through a federated cognitive orchestrator that manages thought routing, state synchronization, and cross-domain knowledge sharing while maintaining domain boundaries.

2 . The computer system of claim 1 , wherein the system implements a persistent reasoning state manager that ensures continuity of reasoning across sessions through comprehensive state capture, maintenance, and restoration.

3 . The computer system of claim 1 , wherein the federated cognitive orchestrator analyzes a domain relevance of received prompts and determines one or more appropriate processing paths across the multiple domain-specialized instances.

4 . The computer system of claim 1 , wherein the system implements an autonomous reasoning engine cluster comprising multiple domain-specialized instances, wherein each instance processes prompts using a model optimized for its respective knowledge domain.

5 . The computer system of claim 1 , wherein the system implements a cross-domain integration framework that establishes isolation boundaries between specialized domains while facilitating controlled knowledge transfers through graduated access controls.

6 . The computer system of claim 1 , wherein the system implements a hierarchical sleep management system that coordinates optimization during system idle periods across the multiple domain-specialized instances by performing memory consolidation, thought cache optimization, and insight generation through cross-domain pattern recognition.

7 . The computer system of claim 6 , wherein the system performs a memory consolidation that comprises creating multi-level abstractions that maintain one or more reasoning patterns while systematically reducing storage requirements through progressive information distillation.

8 . The computer system of claim 6 , wherein the system implements an insight generation process that comprises identifying structural similarities in reasoning patterns across domains and applying analogical reasoning to transfer solution patterns between different specialized knowledge areas.

9 . The computer system of claim 1 , wherein the system implements a collective thought generalization process that extracts common structural elements from similar reasoning patterns across domains while removing domain-specific details.

10 . The computer system of claim 1 , wherein the system implements a system auditor that monitors operational patterns across domains to identify potential efficiency improvements for the federated architecture.

11 . A method comprising:

receiving a prompt from a user; processing the prompt into a plurality of corresponding thoughts using a first reasoning large language model;

routing both the prompt and the plurality of thoughts through a second large language model that has fewer parameters than the first reasoning large language model;

associating each corresponding thought in the plurality of corresponding thoughts to a portion of the prompt;

caching each associated corresponding thought in a federated multi-tier thought caching hierarchy comprising at least a local device cache, a domain-specific branch instance cache, and a global collective cache;

generating a response to the prompt by processing the plurality of thoughts and the prompt through the second large language model;

continuing autonomous reasoning operations without user interaction in a cloud environment, generating new thoughts based on existing cached thoughts and user interaction history; and

coordinating federated cognitive operations across multiple domain-specialized instances through a federated cognitive orchestrator that manages thought routing, state synchronization, and cross-domain knowledge sharing while maintaining domain boundaries.

12 . The method of claim 11 , further comprising implementing a persistent reasoning state manager that ensures continuity of reasoning across sessions through comprehensive state capture, maintenance, and restoration.

13 . The method of claim 11 , wherein coordinating federated cognitive operations comprises analyzing a domain relevance of received prompts and determines one or more appropriate processing paths across the multiple domain-specialized instances.

14 . The method of claim 11 , further comprising implementing an autonomous reasoning engine cluster comprising multiple domain-specialized instances, wherein each instance processes prompts using a model optimized for its respective knowledge domain.

15 . The method of claim 11 , further comprising implementing a cross-domain integration framework that establishes isolation boundaries between specialized domains while facilitating controlled knowledge transfers through graduated access controls.

16 . The method of claim 11 , further comprising implementing a hierarchical sleep management system that coordinates optimization during system idle periods across the multiple domain-specialized instances by performing memory consolidation, thought cache optimization, and insight generation through cross-domain pattern recognition.

17 . The method of claim 16 , further comprising performing memory consolidation that comprises creating multi-level abstractions that maintain one or more reasoning patterns while systematically reducing storage requirements through progressive information distillation.

18 . The method of claim 16 , further comprising implementing an insight generation process that comprises identifying structural similarities in reasoning patterns across domains and applying analogical reasoning to transfer solution patterns between different specialized knowledge areas.

19 . The method of claim 11 , further comprising implementing a collective thought generalization process that extracts common structural elements from similar reasoning patterns across domains while removing domain-specific details.

20 . The method of claim 11 , further comprising implementing a system auditor that monitors operational patterns across domains to identify potential efficiency improvements for the federated architecture.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2026
From: GALVIN, BRIAN; MCCORD, ALAN
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 073557/0018 →
Continuity (15)
Continuation In Part 19203069 · Jun 3, 2025
Continuation 19205960 · May 12, 2025
Continuation In Part 19060794 · Feb 24, 2025
Continuation In Part 19044546 · Feb 3, 2025
Continuation In Part 19026276 · Jan 16, 2025
Continuation In Part 18928022 · Oct 26, 2024
Continuation In Part 18919417 · Oct 17, 2024
Continuation In Part 18918077 · Oct 17, 2024
Continuation In Part 18737906 · Jun 7, 2024
Continuation In Part 18736498 · Jun 6, 2024
Continuation In Part 19178873 · Apr 15, 2025
Continuation In Part 19177611 · Apr 13, 2025
Continuation In Part 19051193 · Feb 12, 2025
Provisional Application 63651359 · May 23, 2024
Related Publication 20250390750A1 · Dec 25, 2025
References Cited (53)
US 4780718A · Hudson et al. · 1988 [cited by applicant]
US 5708436A · Loiz et al. · 1998 [cited by applicant]
US 7411540B1 · Lopez et al. · 2008 [cited by applicant]
US 7629922B2 · Winstead et al. · 2009 [cited by applicant]
US 7876257B2 · Vetro et al. · 2011 [cited by applicant]
US 9524392B2 · Naehrig et al. · 2016 [cited by applicant]
US 11451242B2 · Choi et al. · 2022 [cited by applicant]
US 11656353B2 · Li et al. · 2023 [cited by applicant]
US 11972333B1 · Horesh · 2024 [cited by examiner]
US 11977854B2 · Tunstall-Pedoe · 2024 [cited by examiner]
US 20040017307A1 · Cirillo et al. · 2004 [cited by applicant]
US 20040160353A1 · Cirillo et al. · 2004 [cited by applicant]
US 20080231504A1 · Sartor et al. · 2008 [cited by applicant]
US 20110012778A1 · Nguyen et al. · 2011 [cited by applicant]
US 20150054678A1 · Wakayama · 2015 [cited by applicant]
US 20170048537A1 · Boufounos et al. · 2017 [cited by applicant]
US 20180196609A1 · Niesen · 2018 [cited by applicant]
US 20200258296A1 · Pennings et al. · 2020 [cited by applicant]
US 20220156631A1 · Kanso et al. · 2022 [cited by applicant]
US 20220404490A1 · Evans et al. · 2022 [cited by applicant]
US 20230131694A1 · Saber et al. · 2023 [cited by applicant]
US 20230169623A1 · Chen et al. · 2023 [cited by applicant]
US 20230184927A1 · Chen et al. · 2023 [cited by applicant]
US 20230316006A1 · Tunstall-Pedoe · 2023 [cited by examiner]
US 20240104391A1 · Higgins · 2024 [cited by examiner]
US 20240185037A1 · Park et al. · 2024 [cited by applicant]
US 20240195438A1 · Isik et al. · 2024 [cited by applicant]
US 20240354320A1 · Procter · 2024 [cited by examiner]
US 20240420491A1 · Park · 2024 [cited by examiner]
US 20250291825A1 · Park · 2025 [cited by examiner]
US 20250291866A1 · Park · 2025 [cited by examiner]
US 20250371317A1 · Thompson, III · 2025 [cited by examiner]
EP 3364212A1 · 2018 [cited by applicant]
GB 2620921A · 2024 [cited by applicant]
WO 2020104416A1 · 2020 [cited by applicant]
RamÃ-rez, Guillem, et al. “Cache & distil: Optimising api calls to large language models.” arXiv preprint arXiv:2310.13561 (2023). (Year: 2023). [cited by examiner]
Zhang, Dawen, et al. “A layered architecture for developing and enhancing capabilities in large language model-based software systems.” arXiv preprint arXiv:2411.12357 (2024). (Year: 2024). [cited by examiner]
Zhou, Hao, et al. “Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities.” arXiv preprint arXiv:2405.10825 v2 (2024). (Year: 2024). [cited by examiner]
Bai, Guangji, et al. “Beyond efficiency: A systematic survey of resource-efficient large language models.” arXiv preprint arXiv:2401.00625 (2024). (Year: 2024). [cited by examiner]
Yin, Wangsong, et al. “Llm as a system service on mobile devices.” arXiv preprint arXiv:2403.11805 (2024). (Year: 2024). [cited by examiner]
Chen, Daihang, et al. “LLM for Mobile: An Initial Roadmap.” arXiv preprint arXiv:2407.06573 (2024). (Year: 2024). [cited by examiner]
Fungwacharakorn, Wachara, et al. “Layer-of-Thoughts Prompting (LoT): Leveraging LLM-Based Retrieval with Constraint Hierarchies.” arXiv preprint arXiv:2410.12153 (2024). (Year: 2024). [cited by examiner]
Hu, Pengbo, and Xiang Ying. “Unified mind model: Reimagining autonomous agents in the llm era.” arXiv preprint arXiv:2503.03459 v2 (Mar. 2025). (Year: 2025). [cited by examiner]
Spivack, Nova, et al. “Cognition is All You Need—The Next Layer of AI Above Large Language Models.” arXiv preprint arXiv:2403.02164 (2024). (Year: 2024). [cited by examiner]
Panchal, Kunjal, et al. “Thinking forward: memory-efficient federated finetuning of language models.” Advances in Neural Information Processing Systems 37 (2024): 69069-69119. (Year: 2024). [cited by examiner]
Balaneshin-Kordan, Saeid et al; “Deep Neural Architecture for Multi-Modal Retrieval based on Joint Embedding Space for Text and Images”, Association for Computing Machinery, Feb. 5-9, 2018, pp. 1-9, Marina Del Rey, CA, … [cited by applicant]
Gim, In, et al; “Prompt Cache: Modular Attention Reuse for Low-Latency Inference”, arXiv:2311.04934v2, Apr. 2024. [cited by applicant]
Khan, Abdul Rafae, et al; “Coding Textual Inputs Boosts the Accuracy of Neural Networks”, 2020 Conference on Empirical Methods in Natural Language Processing, Nov. 16-20, 2020, pp. 1350-1360. [cited by applicant]
Messina, Nicola et al; “Towards Efficient Cross-Modal Visual Textual Retrieval using Transformer-Encoder Deep Features”, 2021 International Conference on Content-Based Multimedia Indexing, 2021, pp. 1-6, United States. [cited by applicant]
Seo, Beomsoek et al; “How Does A Transformer Learn Compression? An Attention Study on Huffman and LZ4”, Department of Electronic and Electrical Engineering, Dec. 12, 2023, vol. 11, Seoul, South Korea. [cited by applicant]
Vaswani, Ashish, et al; “Attention is All You Need”, arXiv:1706.03762v7, Aug. 2023. [cited by applicant]
Wang, Tianming, & Wan, Xiaojun; “T-CVAE: Transformer-based conditioned variational autoencoder for story completion”, Proceedings of the Twenty-Eighth Joint Conference on Artificial Intelligence, pp. 5233-5239, 2019. [cited by applicant]
Wieting, John et al; “A Bilingual Generative Transformer for Semantic Sentence Embedding”, arXiv:1911.03895v2, Nov. 2020. [cited by applicant]