IP Library › Granted Patent US 12,626,167
Granted Patent B2
US 12,626,167 · App. 19/339,302 · Granted May 12, 2026

System and method for large language model with integrated memory during inference using manifold traversal architecture

Inventors: Brian Galvin (Silverdale, WA); Alan McCord (Forney, TX)
Assignee: ATOMBEAM TECHNOLOGIES INC.
G06N5/045G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,167
App. No.
19/339,302
Filed
Sep 25, 2025
Granted
May 12, 2026
Kind
B2
Art Unit
2125
USPC
706/46
Abstract

A large language model system integrates persistent memory directly into inference operations through geometric manifold traversal rather than external retrieval. The system implements a memory-integrated inference engine that performs token generation with simultaneous memory access by navigating curved regions in a geometric memory manifold. Memories exist as navigable basins of increased curvature that are reinforced through usage rather than stored as discrete objects. An intent conditioning system formulates user queries as utility functions and generates vector fields that guide goal-directed memory traversal. A manifold geometry interface converts geometric memory coordinates into vectors compatible with language model attention mechanisms, augmenting standard key-value caches with memory-derived content. The system performs intentional remembering through path optimization that balances fidelity to prior cognitive trajectories with current intent guidance. Each memory access operation simultaneously retrieves information and strengthens accessed memory regions through bidirectional geometric shaping, enabling persistent cognitive evolution and cross-session memory continuity.

Claims (31)

1 . A memory-integrated large language model system comprising:

a memory-integrated inference engine that performs token generation with simultaneous memory access through geometric manifold traversal, wherein memory access occurs through goal-conditioned navigation of curved memory regions rather than discrete retrieval operations;

a geometric memory manifold representing memories as navigable regions of increased curvature in latent hyperspace, wherein memories exist as geometric basins that are repeatedly traversed and reinforced through usage;

an intent conditioning system that formulates user queries as utility functions and generates intent vector fields that provide goal-directed guidance for memory traversal through the geometric memory manifold; and

a manifold geometry interface that converts geometric memory coordinates into vector representations compatible with language model attention mechanisms and integrates memory-derived vectors with attention computations to produce memory-enhanced token generation;

wherein the memory-integrated inference engine performs intentional remembering by solving a path optimization that minimizes a functional balancing alignment with intent vector fields against manifold traversal costs, and wherein each memory access operation simultaneously retrieves information retrieve from accessed memory basins based on current manifold position and alignment with the intent vector field and reinforces accessed memory regions through bidirectional geometric shaping that strengthens frequently accessed memory pathways.

2 . The system of claim 1 , wherein the geometric memory manifold comprises a Riemannian manifold with time-evolving metric tensor that encodes memory strength through local curvature intensity.

3 . The system of claim 1 , wherein the geometric basins comprise episodic memory basins for high-resolution recent experiences, semantic memory basins for abstracted knowledge structures, and procedural memory basins for skill patterns.

4 . The system of claim 1 , wherein the manifold geometry interface augments standard attention key and value caches with memory-derived vectors to create memory-enhanced attention computations.

5 . The system of claim 1 , further comprising a dynamic compression engine that implements memory consolidation through geometric flow processes that preserve frequently accessed memory regions while compressing unused areas.

6 . The system of claim 5 , wherein the dynamic compression engine performs sleep-like consolidation cycles that replay memory trajectories and promote stable patterns across hierarchical memory substrates.

7 . The system of claim 1 , further comprising a persistent state manager that serializes manifold geometry and restores memory state across inference sessions to maintain cognitive continuity.

8 . The system of claim 1 , wherein the bidirectional geometric shaping increases curvature along successfully traversed memory paths and deepens memory basins based on access frequency patterns.

9 . The system of claim 1 , implemented as a federated architecture comprising multiple domain-specific instances that coordinate cross-domain memory access and synthesis.

10 . The system of claim 1 , wherein the memory-integrated inference engine evaluates memory access necessity for each token generation step and selectively performs manifold traversal based on context complexity requirements.

11 . A method for memory-integrated large language model inference comprising the steps of:

performing token generation with simultaneous memory access through geometric manifold traversal, wherein memory access occurs through goal-conditioned navigation of curved memory regions rather than discrete retrieval operations;

maintaining memories as navigable regions of increased curvature in a geometric memory manifold within latent hyperspace, wherein memories exist as geometric basins that are repeatedly traversed and reinforced through usage;

formulating user queries as utility functions and generating intent vector fields that provide goal-directed guidance for memory traversal through the geometric memory manifold;

converting geometric memory coordinates into vector representations compatible with language model attention mechanisms and integrating memory-derived vectors with attention computations to produce memory-enhanced token generation;

performing intentional remembering by solving a path optimization that minimizes a functional balancing alignment with intent vector fields against manifold traversal costs; and

simultaneously retrieving information from accessed memory basins based on current manifold position and alignment with the intent vector field and reinforcing accessed memory regions through bidirectional geometric shaping that strengthens frequently accessed memory pathways during each memory access operation.

12 . The method of claim 11 , wherein maintaining the geometric memory manifold comprises evolving a Riemannian manifold with time-evolving metric tensor that encodes memory strength through local curvature intensity.

13 . The method of claim 11 , wherein the geometric basins comprise episodic memory basins for high-resolution recent experiences, semantic memory basins for abstracted knowledge structures, and procedural memory basins for skill patterns.

14 . The method of claim 11 , wherein integrating memory-derived vectors comprises augmenting standard attention key and value caches with memory-derived vectors to create memory-enhanced attention computations.

15 . The method of claim 11 , further comprising the step of implementing memory consolidation through geometric flow processes that preserve frequently accessed memory regions while compressing unused areas.

16 . The method of claim 15 , further comprising the step of performing sleep-like consolidation cycles that replay memory trajectories and promote stable patterns across hierarchical memory substrates.

17 . The method of claim 11 , further comprising the step of serializing manifold geometry and restoring memory state across inference sessions to maintain cognitive continuity.

18 . The method of claim 11 , wherein the bidirectional geometric shaping comprises increasing curvature along successfully traversed memory paths and deepening memory basins based on access frequency patterns.

19 . The method of claim 11 , implemented across a federated architecture comprising coordinating cross-domain memory access and synthesis across multiple domain-specific instances.

20 . The method of claim 11 , further comprising the step of evaluating memory access necessity for each token generation step and selectively performing manifold traversal based on context complexity requirements.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2026
From: GALVIN, BRIAN; MCCORD, ALAN
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 073557/0142 →
Continuity (17)
Continuation In Part 19294125 · Aug 7, 2025
Continuation In Part 19203069 · Jun 3, 2025
Continuation In Part 19205960 · May 12, 2025
Continuation In Part 19178873 · Apr 15, 2025
Continuation In Part 19177611 · Apr 13, 2025
Continuation In Part 19060794 · Feb 24, 2025
Continuation In Part 19051193 · Feb 12, 2025
Continuation In Part 19044546 · Feb 3, 2025
Continuation In Part 19026276 · Jan 16, 2025
Continuation In Part 18928022 · Oct 26, 2024
Continuation In Part 18919417 · Oct 17, 2024
Continuation In Part 18918077 · Oct 17, 2024
Continuation In Part 18737906 · Jun 7, 2024
Continuation In Part 18736498 · Jun 6, 2024
Provisional Application 63847408 · Jul 20, 2025
Provisional Application 63651359 · May 23, 2024
Related Publication 20260023992A1 · Jan 22, 2026
References Cited (63)
US 4780718A · Hudson et al. · 1988 [cited by applicant]
US 5708436A · Loiz et al. · 1998 [cited by applicant]
US 7411540B1 · Lopez et al. · 2008 [cited by applicant]
US 7629922B2 · Winstead et al. · 2009 [cited by applicant]
US 7876257B2 · Vetro et al. · 2011 [cited by applicant]
US 9524392B2 · Naehrig et al. · 2016 [cited by applicant]
US 11451242B2 · Choi et al. · 2022 [cited by applicant]
US 11656353B2 · Li et al. · 2023 [cited by applicant]
US 11972333B1 · Horesh · 2024 [cited by examiner]
US 12387050B1 · Galvin · 2025 [cited by examiner]
US 12481688B1 · Galvin · 2025 [cited by examiner]
US 20040017307A1 · Cirillo et al. · 2004 [cited by applicant]
US 20040160353A1 · Cirillo et al. · 2004 [cited by applicant]
US 20080231504A1 · Sartor et al. · 2008 [cited by applicant]
US 20110012778A1 · Nguyen et al. · 2011 [cited by applicant]
US 20150054678A1 · Wakayama · 2015 [cited by applicant]
US 20170048537A1 · Boufounos et al. · 2017 [cited by applicant]
US 20180196609A1 · Niesen · 2018 [cited by applicant]
US 20190147349A1 · Ng · 2019 [cited by examiner]
US 20200258296A1 · Pennings et al. · 2020 [cited by applicant]
US 20220156631A1 · Kanso et al. · 2022 [cited by applicant]
US 20220404490A1 · Evans et al. · 2022 [cited by applicant]
US 20230131694A1 · Saber et al. · 2023 [cited by applicant]
US 20230169623A1 · Chen et al. · 2023 [cited by applicant]
US 20230184927A1 · Chen et al. · 2023 [cited by applicant]
US 20230316006A1 · Tunstall-Pedoe · 2023 [cited by examiner]
US 20230401262A1 · Orús Lacort · 2023 [cited by examiner]
US 20240104391A1 · Higgins · 2024 [cited by examiner]
US 20240135283A1 · Sahoo · 2024 [cited by examiner]
US 20240185037A1 · Park et al. · 2024 [cited by applicant]
US 20240195438A1 · Isik et al. · 2024 [cited by applicant]
US 20240354320A1 · Procter · 2024 [cited by examiner]
US 20240420491A1 · Park · 2024 [cited by examiner]
US 20250259082A1 · Crabtree · 2025 [cited by examiner]
US 20250291866A1 · Park · 2025 [cited by examiner]
EP 3364212A1 · 2018 [cited by applicant]
GB 2620921A · 2024 [cited by applicant]
WO 2020104416A1 · 2020 [cited by applicant]
RamÃ-rez, Guillem, et al. “Cache & distil: Optimising api calls to large language models.” arXiv preprint arXiv:2310.13561 (2023). (Year: 2023). [cited by examiner]
Lu, Meng. “A mathematical framework of intelligence and consciousness based on Riemannian Geometry.” arXiv preprint arXiv:2407.11024 (2024). (Year: 2024). [cited by examiner]
Zhang, Dawen, et al. “A layered architecture for developing and enhancing capabilities in large language model-based software systems.” arXiv preprint arXiv:2411.12357 (2024). (Year: 2024). [cited by examiner]
Zhou, Hao, et al. “Large Language Model (LLM) for Telecommunications: A Comprehensive Survey on Principles, Key Techniques, and Opportunities.” arXiv preprint arXiv:2405.10825 v2 (2024). (Year: 2024). [cited by examiner]
Bai, Guangji, et al. “Beyond efficiency: A systematic survey of resource-efficient large language models.” arXiv preprint arXiv:2401.00625 (2024). (Year: 2024). [cited by examiner]
Yin, Wangsong, et al. “Llm as a system service on mobile devices.” arXiv preprint arXiv:2403.11805 (2024). (Year: 2024). [cited by examiner]
Chen, Daihang, et al. “LLM for Mobile: An Initial Roadmap.” arXiv preprint arXiv:2407.06573 (2024). (Year: 2024). [cited by examiner]
Fungwacharakorn, Wachara, et al. “Layer-of-Thoughts Prompting (LoT): Leveraging LLM-Based Retrieval with Constraint Hierarchies.” arXiv preprint arXiv:2410.12153 (2024). (Year: 2024). [cited by examiner]
Spivack, Nova, et al. “Cognition is All You Need—The Next Layer of Al Above Large Language Models.” arXiv preprint arXiv:2403.02164 (2024). (Year: 2024). [cited by examiner]
Patil, Sarang, et al. “Hyperbolic Large Language Models.” arXiv preprint arXiv:2509.05757 (Sep. 6, 2025). (Year: 2025). [cited by examiner]
Vassilis, Koinis, et al. “Lexical manifold reconfiguration in large language models: A novel architectural approach for contextual modulation.” arXiv preprint arXiv:2502.08818 (Feb. 2025). (Year: 2025). [cited by examiner]
Kingswell, Jonathan, et al. “Sequential manifold regularization for large language model contextual stability.” (Aug. 2025). (Year: 2025). [cited by examiner]
Li, Xin. “Manifold Learning via Memory and Context.” arXiv preprint arXiv:2407.09488 v3 (Jul. 2025). (Year: 2025). [cited by examiner]
Morgoulis, Daria. “4D Semantic Coupling: A Mathematical Framework for Measuring Cognitive Complexity in AI Dialogue.” Authorea Preprints (Aug. 2025). (Year: 2025). [cited by examiner]
Hu, Pengbo, and Xiang Ying. “Unified mind model: Reimagining autonomous agents in the LLM era.” arXiv preprint arXiv:2503.03459 v2 (Mar. 2025). (Year: 2025). [cited by examiner]
Luo, Haoxiang, et al. “Toward Edge General Intelligence with Multiple-Large Language Model (Multi-LLM): Architecture, Trust, and Orchestration.” arXiv preprint arXiv:2507.00672 (Jul. 2025). (Year: 2025). [cited by examiner]
Xu, Wenchao, et al. “Deploying foundation model powered agent services: A survey.” IEEE Communications Surveys & Tutorials (Jun. 2025). (Year: 2025). [cited by examiner]
Balaneshin-Kordan, Saeid et al; “Deep Neural Architecture for Multi-Modal Retrieval based on Joint Embedding Space for Text and Images”, Association for Computing Machinery, Feb. 5-9, 2018, pp. 1-9, Marina Del Rey, CA, … [cited by applicant]
Gim, In, et al; “Prompt Cache: Modular Attention Reuse for Low-Latency Inference”, arXiv:2311.04934v2, Apr. 2024. [cited by applicant]
Khan, Abdul Rafae, et al; “Coding Textual Inputs Boosts the Accuracy of Neural Networks”, 2020 Conference on Empirical Methods in Natural Language Processing, Nov. 16-20, 2020, pp. 1350-1360. [cited by applicant]
Messina, Nicola et al; “Towards Efficient Cross-Modal Visual Textual Retrieval using Transformer-Encoder Deep Features”, 2021 International Conference on Content-Based Multimedia Indexing, 2021, pp. 1-6, United States. [cited by applicant]
Seo, Beomsoek et al; “How Does A Transformer Learn Compression? An Attention Study on Huffman and LZ4”, Department of Electronic and Electrical Engineering, Dec. 12, 2023, vol. 11, Seoul, South Korea. [cited by applicant]
Vaswani, Ashish, et al; “Attention is All You Need”, arXiv:1706.03762v7, Aug. 2023. [cited by applicant]
Wang, Tianming, & Wan, Xiaojun; “T-CVAE: Transformer-based conditioned variational autoencoder for story completion”, Proceedings of the Twenty-Eighth Joint Conference on Artificial Intelligence, pp. 5233-5239, 2019. [cited by applicant]
Wieting, John et al; “A Bilingual Generative Transformer for Semantic Sentence Embedding”, arXiv:1911.03895v2, Nov. 2020. [cited by applicant]