IP Library Granted Patent US 12,506,496
Granted Patent B2
US 12,506,496 · App. 19/242,953 · Granted Dec 23, 2025

Hierarchical smart caching for machine learning codeword responses

Inventor: Brian Galvin (Silverdale, WA)
Assignee: ATOMBEAM TECHNOLOGIES INC.
H03M7/3059G06N20/00H03M7/6005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,506,496
App. No.
19/242,953
Filed
Jun 18, 2025
Granted
Dec 23, 2025
Kind
B2
Art Unit
2845
USPC
707/693
Abstract

A system and method for deep learning using a large codeword model with hierarchical caching is disclosed. The system processes input prompts into tokens, maps them to codewords using a codebook, and processes these through a machine learning core to generate responses. A sophisticated caching architecture stores and retrieves responses across both local and global cache tiers. The local cache maintains frequently accessed responses on edge devices through short-term and persistent storage components, while the global cache enables knowledge sharing across multiple devices. A context aggregator identifies relationships between cached responses to form comprehensive contextual representations. This hierarchical caching system significantly reduces computational requirements by reusing previously generated responses for similar prompts, while continuously optimizing cache contents based on relevance scoring and usage patterns. The approach enables efficient scaling across distributed environments while maintaining response quality.

Claims (55)

1 . A computer system comprising:

a hardware memory, wherein the computer system is configured to execute software instructions stored on nontransitory machine-readable storage media that:

receive a prompt from a user;

process the prompt into a plurality of tokens;

map the plurality of tokens to a plurality of codewords using a codebook, where each token is assigned a specific codeword;

train a machine learning core on a plurality of codewords;

process the plurality of codewords through the trained machine learning core to generate a codeword response;

store the codeword response and associated processing outputs from the machine learning core in a hierarchical caching system comprising:

a local cache configured to store the codeword responses on an edge device; and

a global cache configured to store the codeword responses across multiple edge devices;

evaluate relevance of the stored codeword responses;

retrieve relevant codeword responses when processing similar prompts; and

translate the codeword response into a response which matches a modality of the prompt.

2 . The system of claim 1 , wherein the machine learning core is a conventional transformer-based architecture comprising:

an embedding layer;

a positional encoding layer; and

a series of transformer layers.

3 . The system of claim 2 , wherein computer system is further configured to:

split the codewords into smaller units before processing through the conventional transformer-based architecture.

4 . The system of claim 1 , wherein the machine learning core is a latent transformer core comprising:

a variational autoencoder with an encoder and a decoder; and

a transformer that processes latent space vectors, wherein the transformer does not include an embedding layer and a positional encoding layer.

5 . The system of claim 4 , wherein processing the plurality of codewords through the machine learning core comprises:

generating a plurality of latent space vectors by processing the plurality of codewords through the variational autoencoder's encoder;

learning relationships between the plurality of latent space vectors by processing them through the transformer;

using the learned relationships to generate a plurality of output latent space vectors; and

generating the codeword response by passing the output latent space vectors through the variational autoencoder's decoder.

6 . The system of claim 1 , wherein computer system is further configured to:

split the latent space vectors into smaller units before processing through the transformer.

7 . A method for hierarchical smart caching for machine learning codeword responses, comprising the steps of:

receiving a prompt from a user;

processing the prompt into a plurality of tokens;

mapping the plurality of tokens to a plurality of codewords using a codebook, where each token is assigned a specific codeword;

training a machine learning core on a plurality of codewords;

processing the plurality of codewords through the trained machine learning core to generate a codeword response;

storing the codeword response and associated processing outputs from the machine learning core in a hierarchical caching system comprising:

a local cache configured to store the codeword responses on an edge device; and

a global cache configured to store the codeword responses across multiple edge devices;

evaluating relevance of the stored codeword responses;

retrieving relevant codeword responses when processing similar prompts; and

translating the codeword response into a response which matches a modality of the prompt.

8 . The method of claim 7 , wherein the machine learning core is a conventional transformer-based architecture comprising:

an embedding layer;

a positional encoding layer; and

a series of transformer layers.

9 . The method of claim 8 , further comprising splitting the codewords into smaller units before processing through the conventional transformer-based architecture.

10 . The method of claim 7 , wherein the machine learning core is a latent transformer core comprising:

a variational autoencoder with an encoder and a decoder; and

a transformer that processes latent space vectors, wherein the transformer does not include an embedding layer and a positional encoding layer.

11 . The method of claim 10 , wherein processing the plurality of codewords through the machine learning core comprises:

generating a plurality of latent space vectors by processing the plurality of codewords through the variational autoencoder's encoder;

learning relationships between the plurality of latent space vectors by processing them through the transformer;

using the learned relationships to generate a plurality of output latent space vectors; and

generating the codeword response by passing the output latent space vectors through the variational autoencoder's decoder.

12 . The method of claim 11 , further comprising splitting the latent space vectors into smaller units before processing through the transformer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 27, 2025
From: GALVIN, BRIAN
To: ATOMBEAM TECHNOLOGIES INC.
Reel/Frame 072557/0164 →
Continuity (6)
Continuation In Part 19088978 · Mar 24, 2025
Continuation 18909976 · Oct 9, 2024
Continuation In Part 18737906 · Jun 7, 2024
Continuation In Part 18736498 · Jun 6, 2024
Provisional Application 63651359 · May 23, 2024
Related Publication 20250365007A1 · Nov 27, 2025
References Cited (42)
US 4780718A · Hudson et al. · 1988 [cited by applicant]
US 5708436A · Loiz et al. · 1998 [cited by applicant]
US 7411540B1 · Lopez et al. · 2008 [cited by applicant]
US 7629922B2 · Winstead et al. · 2009 [cited by applicant]
US 7804428B2 · Oslick · 2010 [cited by applicant]
US 7876257B2 · Vetro et al. · 2011 [cited by applicant]
US 9524392B2 · Naehrig et al. · 2016 [cited by applicant]
US 9846574B2 · Raman · 2017 [cited by examiner]
US 11086843B2 · Swaminathan et al. · 2021 [cited by applicant]
US 11100420B2 · Dirac et al. · 2021 [cited by applicant]
US 11216742B2 · Bhattacharyya · 2022 [cited by applicant]
US 11451242B2 · Choi et al. · 2022 [cited by applicant]
US 11557323B1 · Stewart · 2023 [cited by examiner]
US 11656353B2 · Li et al. · 2023 [cited by applicant]
US 20040017307A1 · Cirillo et al. · 2004 [cited by applicant]
US 20040160353A1 · Cirillo et al. · 2004 [cited by applicant]
US 20080231504A1 · Sartor et al. · 2008 [cited by applicant]
US 20110012778A1 · Nguyen et al. · 2011 [cited by applicant]
US 20150054678A1 · Wakayama · 2015 [cited by applicant]
US 20160179799A1 · Raman · 2016 [cited by examiner]
US 20170048537A1 · Boufounos et al. · 2017 [cited by applicant]
US 20180196609A1 · Niesen · 2018 [cited by applicant]
US 20180300606A1 · Corkery · 2018 [cited by examiner]
US 20200258296A1 · Pennings et al. · 2020 [cited by applicant]
US 20200395955A1 · Choi et al. · 2020 [cited by applicant]
US 20220156631A1 · Kanso et al. · 2022 [cited by applicant]
US 20220404490A1 · Evans et al. · 2022 [cited by applicant]
US 20230131694A1 · Saber et al. · 2023 [cited by applicant]
US 20230169623A1 · Chen et al. · 2023 [cited by applicant]
US 20230184927A1 · Chen et al. · 2023 [cited by applicant]
US 20240185037A1 · Park et al. · 2024 [cited by applicant]
US 20240195438A1 · Isik et al. · 2024 [cited by applicant]
EP 3364212A1 · 2018 [cited by applicant]
GB 2620921A · 2024 [cited by applicant]
WO 2020104416A1 · 2020 [cited by applicant]
Balaneshin-Kordan, Saeid et al; “Deep Neural Architecture for Multi-Modal Retrieval based on Joint Embedding Space for Text and Images”, Association for Computing Machinery, Feb. 5-9, 2018, pp. 1-9, Marina Del Rey, CA, … [cited by applicant]
Khan, Abdul Rafae, et al; “Coding Textual Inputs Boosts the Accuracy of Neural Networks”, 2020 Conference on Empirical Methods in Natural Language Processing, Nov. 16-20, 2020, pp. 1350-1360. [cited by applicant]
Messina, Nicola et al; “Towards Efficient Cross-Modal Visual Textual Retrieval using Transformer-Encoder Deep Features”, 2021 International Conference on Content-Based Multimedia Indexing, 2021, pp. 1-6, United States. [cited by applicant]
Seo, Beomsoek et al; “How Does a Transformer Learn Compression? An Attention Study on Huffman and LZ4”, Department of Electronic and Electrical Engineering, Dec. 12, 2023, vol. 11, Seoul, South Korea. [cited by applicant]
Vaswani, Ashish, et al; “Attention is All You Need”, arXiv:1706.03762v7, Aug. 2023. [cited by applicant]
Wang, Tianming, & Wan, Xiaojun; “T-CVAE: Transformer-based conditioned variational autoencoder for story completion”, Proceedings of the Twenty-Eighth Joint Conference on Artificial Intelligence, pp. 5233-5239, 2019. [cited by applicant]
Wieting, John et al; “A Bilingual Generative Transformer for Semantic Sentence Embedding”, arXiv:1911.03895v2, Nov. 2020. [cited by applicant]