IP Library › Granted Patent US 12,608,631
Granted Patent B1
US 12,608,631 · App. 19/178,873 · Granted Apr 21, 2026

Mobile-optimized multi-stage LLM with autonomous reasoning

Inventors: Brian Galvin (Silverdale, WA); Alan McCord (Forney, TX)
Assignee: ATOMBEAM TECHNOLOGIES INC.
G06N5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,631
App. No.
19/178,873
Filed
Apr 15, 2025
Granted
Apr 21, 2026
Kind
B1
Art Unit
2125
USPC
706/46
Abstract

A system and method for extending mobile-optimized multi-stage language model processing with autonomous reasoning capabilities. Building upon the three-tier thought caching architecture from the parent invention, the system implements a cognitive dyad framework that continues reasoning operations in cloud environments when mobile devices are inactive. The system enters a dream-state processing mode during periods of user inactivity, performing memory consolidation, thought cache optimization, and novel thought generation without consuming mobile device resources. Through persistent cognitive operation, the system maintains reasoning continuity across user interactions and devices while preserving mobile optimization benefits including battery-aware execution, offline functionality, and privacy protection. The cognitive dyad functions as a thinking partner rather than merely a responsive tool, generating novel insights through autonomous exploration while maintaining strict boundaries between private and shared thought spaces.

Claims (33)

1 . A computer system comprising a hardware memory, wherein the computer system is configured to execute software instructions stored on nontransitory machine-readable storage media that:

receive a prompt from a user of a mobile device;

process the prompt into a plurality of corresponding thoughts using a first reasoning large language model;

route both the prompt and the plurality of thoughts through a second large language model that has fewer parameters than the first reasoning large language model;

associate each corresponding thought in the plurality of corresponding thoughts to a portion of the prompt;

cache each associated corresponding thought in a multi-tier thought caching architecture;

generate a response to the prompt by processing the plurality of thoughts and the prompt through the second large language model; and

continue autonomous reasoning operations without user interaction in a cloud environment when the mobile device is inactive, generating new thoughts based on existing cached thoughts and user interaction history.

2 . The computer system of claim 1 , wherein the system enters a dream-state processing mode during periods of user inactivity to perform memory consolidation without consuming mobile device resources.

3 . The computer system of claim 1 , wherein the multi-tier thought caching architecture comprises a local device cache, a user-specific cloud cache, and a global generalized thought cache, and wherein the autonomous reasoning operations maintain strict separation between private user thoughts and generalized thought patterns.

4 . The computer system of claim 1 , wherein the system maintains cognitive continuity across multiple user devices by synchronizing cognitive state and continuing reasoning processes regardless of which device is currently active.

5 . The computer system of claim 2 , wherein memory consolidation comprises transferring information from short-term thought storage to long-term thought storage.

6 . The computer system of claim 2 , wherein novel thought generation comprises combinatorial processes that connect concepts from different domains.

7 . The computer system of claim 1 , wherein the system determines when to present autonomously generated thoughts to the user based on relevance scoring and alignment with identified user interests.

8 . The computer system of claim 1 , wherein the system implements self-directed reasoning by generating internal prompts based on user history and identified knowledge gaps.

9 . The computer system of claim 1 , wherein the system synchronizes autonomously generated thoughts between cloud and local caches.

10 . The computer system of claim 1 , wherein the system functions as a cognitive dyad that adapts to user preferences through continuous learning while maintaining operation during periods of user disengagement.

11 . A method comprising:

receiving a prompt from a user of a mobile device;

processing the prompt into a plurality of corresponding thoughts using a first large language model;

routing both the prompt and the plurality of corresponding thoughts through a second large language model that has fewer parameters than the first large language model;

associating each corresponding thought in the plurality of corresponding thoughts to a portion of the prompt;

caching each associated corresponding thought in a multi-tier thought caching architecture;

generating a response to the prompt by processing the plurality of corresponding thoughts and the prompt through the second large language model; and

continuing autonomous reasoning operations without user interaction in a cloud environment when the mobile device is inactive, generating new thoughts based on existing cached thoughts and user interaction history.

12 . The method of claim 11 , further comprising entering a dream-state processing mode during periods of user inactivity to perform memory consolidation without consuming mobile device resources.

13 . The method of claim 11 , wherein the multi-tier thought caching architecture comprises a local device cache, a user-specific cloud cache, and a global generalized thought cache, and wherein the autonomous reasoning operations maintain strict separation between private user thoughts and generalized thought patterns.

14 . The method of claim 11 , further comprising maintaining cognitive continuity across multiple user devices by synchronizing cognitive state and continuing reasoning processes regardless of which device is currently active.

15 . The method of claim 12 , wherein memory consolidation comprises transferring information from short-term thought storage to long-term thought storage.

16 . The method of claim 12 , wherein novel thought generation comprises combinatorial processes that connect concepts from different domains.

17 . The method of claim 11 , further comprising determining when to present autonomously generated thoughts to the user based on relevance scoring and alignment with identified user interests.

18 . The method of claim 11 , further comprising implementing self-directed reasoning by generating internal prompts based on user history and identified knowledge gaps.

19 . The method of claim 11 , further comprising synchronizing autonomously generated thoughts between cloud and local caches.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2025
From: GALVIN, BRIAN; MCCORD, ALAN
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 070952/0189 →
Continuity (2)
Continuation In Part 19177611 · Apr 13, 2025
Continuation In Part 19051193 · Feb 12, 2025
References Cited (31)
US 11556470B2 · Lin · 2023 [cited by examiner]
US 11972333B1 · Horesh · 2024 [cited by examiner]
US 12223456B1 · Manohar · 2025 [cited by examiner]
US 20220138156A1 · Wang · 2022 [cited by examiner]
US 20230316006A1 · Tunstall-Pedoe · 2023 [cited by examiner]
US 20240104391A1 · Higgins · 2024 [cited by examiner]
US 20240354320A1 · Procter · 2024 [cited by examiner]
US 20240420491A1 · Park · 2024 [cited by examiner]
US 20240428008A1 · Abraham · 2024 [cited by examiner]
US 20250028882A1 · Ataei · 2025 [cited by examiner]
US 20250165718A1 · Seo · 2025 [cited by examiner]
US 20250173554A1 · Lancioni · 2025 [cited by examiner]
US 20250291842A1 · Park · 2025 [cited by examiner]
US 20250291866A1 · Park · 2025 [cited by examiner]
US 20250342179A1 · Gangumalla · 2025 [cited by examiner]
US 20250342367A1 · Shama · 2025 [cited by examiner]
US 20250363311A1 · Chandarana · 2025 [cited by examiner]
US 20250371041A1 · Yang · 2025 [cited by examiner]
US 20250371433A1 · Bhat · 2025 [cited by examiner]
US 20250378396A1 · Ezrielev · 2025 [cited by examiner]
US 20250390352A1 · Crabtree · 2025 [cited by examiner]
US 20250390602A1 · Tran · 2025 [cited by examiner]
RamÃ-rez, Guillem, et al. “Cache & distil: Optimising api calls to large language models.” arXiv preprint arXiv:2310.13561 (2023). (Year: 2023). [cited by examiner]
Zhang, Dawen, et al. “A layered architecture for developing and enhancing capabilities in large language model-based software systems.” arXiv preprint arXiv:2411.12357 (2024). (Year: 2024). [cited by examiner]
Zhou, Hao, et al. “Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities.” arXiv preprint arXiv:2405.10825 v2 (2024). (Year: 2024). [cited by examiner]
Bai, Guangji, et al. “Beyond efficiency: A systematic survey of resource-efficient large language models.” arXiv preprint arXiv: 2401.00625 (2024). (Year: 2024). [cited by examiner]
Yin, Wangsong, et al. “Llm as a system service on mobile devices.” arXiv preprint arXiv:2403.11805 (2024). (Year: 2024). [cited by examiner]
Chen, Daihang, et al. “LLM for Mobile: An Initial Roadmap.” arXiv preprint arXiv:2407.06573 (2024). (Year: 2024). [cited by examiner]
Fungwacharakorn, Wachara, et al. “Layer-of-Thoughts Prompting (LoT): Leveraging LLM-Based Retrieval with Constraint Hierarchies.” arXiv preprint arXiv:2410.12153 (2024). (Year: 2024). [cited by examiner]
Gao, Hang, and Yongfeng Zhang. “Memory sharing for large language model based agents.” arXiv preprint arXiv:2404.09982 (2024). (Year: 2024). [cited by examiner]
Gim, In, et al; “Prompt Cache: Modular Attention Reuse for Low-Latency Inference”, arXiv:2311.04934v2, Apr. 2024. [cited by applicant]