IP Library › Granted Patent US 12,450,168
Granted Patent B2
US 12,450,168 · App. 18/428,422 · Granted Oct 21, 2025

Dynamic updating of content addressable associative memories for large language models

Inventors: Georgios Kollias (White Plains, NY); Elliot Nelson (Malvern, PA); Payel Das (Yorktown Heights, NY); Subhajit Chaudhury (White Plains, NY); Aurelie Chloe Lozano (Scarsdale, NY); Pin-Yu Chen (White Plains, NY)
Assignee: International Business Machines Corporation
G06F12/12G06F12/1466
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,168
App. No.
18/428,422
Granted
Oct 21, 2025
Kind
B2
Abstract

A computerized system includes a generative large language model (LLM). The LLM includes an encoder configured to provide a vector input to a content addressable memory system and a decoder configured to receive a vector output from the content addressable memory system and generate a data output. The content addressable memory system is configured to receive the vector input and generate a vector output based on data contents of the content addressable memory system. The content addressable memory system is configured to respond to receiving at least one memory update command by updating the data contents of the content addressable memory system.

Claims (24)

1. A computerized system comprising:

a generative large language model (LLM) including an encoder configured to provide a vector input to a content addressable memory system and a decoder configured to receive a vector output from the content addressable memory system and generate a data output; and

the content addressable memory system being configured to receive the vector input and generate a vector output based on data contents of the content addressable memory system, and the content addressable memory system being configured to respond to receiving at least one memory update command by updating the data contents of the content addressable memory system.

2. The computerized system of claim 1 , wherein the content addressable memory system is further configured to dynamically modify an associative memory matrix on which the generative LLM operates without modifying the generative LLM.

3. The computerized system of claim 1 , wherein the content addressable memory system includes a static addressing mode.

4. The computerized system of claim 3 , wherein the data contents of the content addressable memory system includes distinct data elements, and wherein an address of each data element in the data contents is statically dependent on the data element such that the address of each data element is dependent on the data element in a fixed, constant way, across a sequence of operations to the content addressable memory system.

5. The computerized system of claim 3 , wherein the static addressing mode defines an injective mapping of data elements such that each data element is mapped to a corresponding memory address vector.

6. The computerized system of claim 1 , wherein the content addressable memory system includes a dynamic addressing mode.

7. The computerized system of claim 6 , the data contents of the content addressable memory system includes distinct data elements, and wherein addresses of data elements in the data contents are dynamically dependent on the data element and a current status of the content addressable memory system.

8. The computerized system of claim 1 , wherein the content addressable memory system includes a key generation subsystem and a memory update system configured to cooperatively update the data contents of the content addressable memory system based on the at least one memory update command.

9. A method comprising:

receiving a memory update command at a memory system of a generative large language model (LLM), wherein the generative LLM includes an encoder configured to provide a vector input to a content addressable memory system and a decoder configured to receive a vector output from the content addressable memory system and generate a data output; and

responding to the memory update command by editing at least one data element of the content addressable memory system in accordance with the update command.

10. The method of claim 9 , further comprising identifying the at least one data element to be edited using an addressing system configured to identify the at least one data element based on one of a static memory address and a dynamic memory address.

11. The method of claim 10 , wherein updating the at least one data element comprises dynamically modifying an associative memory on which the generative LLM operates without modifying the generative LLM.

12. The method of claim 10 , wherein the content addressable memory system includes static memory addresses, and wherein data contents of the content addressable memory system includes distinct data elements, and the address of any given data element is statically dependent on the data element, such that the address of each data element is dependent on the data element in a fixed, constant way, across a sequence of operations to the memory system.

13. The method of claim 12 , wherein the static memory addresses defines an injective mapping of data elements such that each data element is mapped to a corresponding memory address vector.

14. The method of claim 10 , the content addressable memory system includes dynamic memory addresses, and wherein data contents of the content addressable memory system includes distinct data elements, and wherein an address of each data element in the data contents is dynamically dependent on the data element and a current status of the content addressable memory system.

15. The method of claim 10 , wherein the update command includes at least one of an add operation, a blurry add operation, a forget operation, and a weak forget operation.

16. The method of claim 15 , wherein the add operation adds a data element to a memory and the forget operation removes a data element from the memory.

17. The method of claim 15 , wherein the update command includes a sparsification component.

18. The method of claim 17 , further comprising responding to the sparsification component by sparsifying a memory address of the data element.

19. A computer program product comprising a non-transitory memory storing a content addressable memory system configured to interface with a generative LLM, wherein the content addressable memory system is configured to receive a vector input from an encoder of the generative LLM and is configured to provide a vector output to a decoder of the generative LLM.

20. The computer program product of claim 19 , wherein the content addressable memory system includes a key generation subsystem and a memory update system configured to cooperatively update data contents of the content addressable memory system based on at least one memory update command.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2024
From: KOLLIAS, GEORGIOS; NELSON, ELLIOT; DAS, PAYEL; CHAUDHURY, SUBHAJIT; LOZANO, AURELIE CHLOE; CHEN, PIN-YU
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 066322/0726 →
Continuity (1)
Related Publication 20250245170A1 · Jul 31, 2025
References Cited (95)
US 8589163B2 · Ljolje et al. · 2013 [cited by applicant]
US 11182028B2 · Lee · 2021 [cited by examiner]
US 12223269B2 · He · 2025 [cited by examiner]
US 20230316006A1 · Tunstall-Pedoe et al. · 2023 [cited by applicant]
US 20230343324A1 · Baeuml et al. · 2023 [cited by applicant]
CN 116127020A · 2023 [cited by applicant]
CN 116935288A · 2023 [cited by applicant]
CN 116975192A · 2023 [cited by applicant]
CN 117093679A · 2023 [cited by applicant]
Longuet-Higgins et al., “Theories of associative recall.” Quarterly reviews of biophysics 3.2 (1970): 223-244. [cited by applicant]
Meng et al. “Mass-editing memory in a transformer.” arXiv preprint arXiv:2210.07229 (2023): 21 pages. [cited by applicant]
Meng et al., Locating and editing factual associations in GPT, Advances in Neural Information Processing Systems. (2022): 35 pages. [cited by applicant]
Meng, Yuanliang et al., “Context-aware neural model for temporal information extraction.” Proceedings of the 56th annual meeting of the association for computational linguistics (2018): 10 pages. [cited by applicant]
Mitchell et al., “Fast Model Editing at Scale.” International Conference on Learning Representations. (2021): 21 pages. [cited by applicant]
Mitchell et al., “Memory-based model editing at scale.” International Conference on Machine Learning. PMLR, (2022): 15 pages. [cited by applicant]
Murdock, Bennet B. “Developing TODAM: Three models for serial-order information.” Memory & Cognition 23.5 (1995): 631-645. [cited by applicant]
Personnaz et al., “Collective computational properties of neural networks: New learning mechanisms.” Physical Review A 34.5 (1986): 4217-4228. [cited by applicant]
Petroni et al., “Language models as knowledge bases?. ” arXiv preprint arXiv:1909.01066 (2019): 11 pages. [cited by applicant]
Pham et. al. “Generative Pseudo-Inverse Memory” ICLR (2022): 18 pages. [cited by applicant]
Plate, Tony A. “Holographic reduced representations.” IEEE Transactions on Neural networks 6.3 (1995): 21 pages. [cited by applicant]
Rae, Jack, et al. “Scaling memory-augmented neural networks with sparse reads and writes.” Advances in Neural Information Processing Systems 29 (2016): 9 pages. [cited by applicant]
Ramsauer et al., “Hopfield networks is all you need.” arXiv preprint arXiv:2008.02217 (2020): 94 pages. [cited by applicant]
Raunak et al., “Rank-One Editing of Encoder-Decoder Models.” arXiv preprint arXiv:2211.13317 (2022): 6 pages. [cited by applicant]
Redhakrishnan et al., “Overparameterized neural networks implement associative memory.” Proceedings of the National Academy of Sciences 117.44 (2020): 27162-27170. [cited by applicant]
Saha et al., “Gradient projection memory for continual learning.” arXiv preprint arXiv:2103.09762 (2021): 18 pages. [cited by applicant]
Salvatori et al., “Associative memories via predictive coding.” Advances in Neural Information Processing Systems 34 (2021): 13 pages. [cited by applicant]
Schlag et al., “Linear transformers are secretly fast weight programmers.” International Conference on Machine Learning. PMLR, (2021) 12 pages. [cited by applicant]
Shun-Ichi Amari. “Learning patterns and pattern sequences by self-organizing nets of threshold elements.” IEEE Transactions on computers 100.11 (1972): 1197-1206. [cited by applicant]
Smolensky, Paul. “Tensor product variable binding and the representation of symbolic structures in connectionist systems.” Artificial intelligence 46.1-2 (1990): 159-216. [cited by applicant]
Stiles et al., “On the effect of noise on the Moore-Penrose generalized inverse associative memory.” IEEE transactions on pattern analysis and machine intelligence 3 (1985): 358-360. [cited by applicant]
Sukhbaatar et al. “End-to-end memory networks.” Advances in neural information processing systems 28 (2015): 9 pages. [cited by applicant]
Sukhbaatar et al., “Augmenting self-attention with persistent memory.” arXiv preprint arXiv:1907.01470 (2019): 11 pages. [cited by applicant]
Sukhbaatar et al., “Not all memories are created equal: Learning to forget by expiring.” International Conference on Machine Learning. PMLR, (2021): 11 pages. [cited by applicant]
Valle-Lisboa et al., “Multiplicative processing in the modeling of cognitive activities in large neural networks.” Biophysical Reviews (2023): 1-19. [cited by applicant]
Willshaw et al. “Non-holographic associative memory.” Nature 222.5197 (1969): 960-962. [cited by applicant]
Whittington et al., “Relating transformers to models and neural representations of the hippocampal formation.” arXiv preprint arXiv:2112.04035 (2021): 20 pages. [cited by applicant]
Wu et al. “The Kanerva machine: A generative distributed memory.” arXiv preprint arXiv:1804.01756 (2018): 16 pages. [cited by applicant]
Wu et al., “Learning attractor dynamics for generative memory.” Advances in Neural Information Processing Systems 31 (2018): 10 pages. [cited by applicant]
Wu, Yuhuai et al. “Memorizing Transformers.” International Conference on Learning Representations. (2021): 19 pages. [cited by applicant]
Yen et al., “A learning and forgetting algorithm in associative memories. The eigenstructure method.” [1991] Proceedings of the 30th IEEE Conference on Decision and Control. IEEE, (1991): pp. 847-852. [cited by applicant]
Zhang et al., “Hippocampal spatial representations exhibit a hyperbolic geometry that expands with experience.” Nature Neuroscience 26.1 (2023): 131-139. [cited by applicant]
Yen et al., “A learning and forgetting algorithm in associative memories: results involving pseudo inverses.” 1991., IEEE International Symposium on Circuits and Systems. IEEE, (1991): pp. 778-781. [cited by applicant]
Steinbuch, K “Die Lernmatric” Kybernetik, 1(1): Jan. 1961. pp. 36-45. [cited by applicant]
Anderson, J. A., “A simple neural network generating an interactive memory,” Mathematical Biosciences, vol. 14, (1972): pp. 197-220. [cited by applicant]
Bau et al. “Rewriting a Deep Generative Model.” arXiv preprint arXiv:2007.15646 (2020): 31 pages. [cited by applicant]
Weston et al., “Memory Networks.” in 3rd ICLR (2015): 15 pages. [cited by applicant]
Armand Joulin et al., “Inferring algorithmic patterns with stack-augmented recurrent nets.” Advances in neural information processing systems 28 (2015): 9 pages. [cited by applicant]
Bietti et al., “Birth of a Transformer: A Memory Viewpoint.” arXiv preprint arXiv:2306.00802 (2023): 28 pages. [cited by applicant]
Bricken et al., “Attention approximates sparse distributed memory.” Advances in Neural Information Processing Systems 35 (2021): 15 pages. [cited by applicant]
Bricken et al., “Sparse Distributed Memory is a Continual Learner.” arXiv preprint arXiv:2303.11934 (2023): 57 pages. [cited by applicant]
Burtsev et al., “Memory transformer.” arXiv preprint arXiv:2006.11527 (2020): 17 pages. [cited by applicant]
Cabannes et al, “Scaling laws for associative memories.” arXiv preprint arXiv:2310.02984 (2023): 32 pages. [cited by applicant]
Caplan et al., “Associative recognition without hippocampal associations.” Psychological Review 129.6 (2022): 54 pages. [cited by applicant]
Cheng et al., “Language model with Plug-in Knowldge Memory.” ICLR (2022): 22 pages. [cited by applicant]
Dai et al., “Knowledge neurons in pretrained transformers.” arXiv preprint arXiv:2104.08696 (2021): 10 pages. [cited by applicant]
Dai et al., “Neural knowledge bank for pretrained transformers.” CCF International Conference on Natural Language Processing and Chinese Computing. Cham: Springer Nature Switzerland, (2023): 11 pages. [cited by applicant]
De Cao, Nicola et al., “Editing Factual Knowledge in Language Models.” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. (2021): pp. 6491-6506. [cited by applicant]
Dong et al., “Calibrating factual knowledge in pretrained language models.” arXiv preprint arXiv:2210.03329 (2022): 11 pages. [cited by applicant]
Eldan et al., “Who's Harry Potter? Approximate unlearning in LLMs,” arXiv:2310.02238v2, (2023): 21 pages. [cited by applicant]
Elhage et al., “A mathematical framework for transformer circuits.” Transformer Circuits Thread 1 https://transformer-circuits.pub/2021/framework/index.html (retrieved Jan. 29, 2024), 47 pages. [cited by applicant]
Fan et al., “Augmenting transformers with KNN-based composite memory for dialog.” Transactions of the Association for Computational Linguistics 9 (2021): 82-99. [cited by applicant]
Farajtabar et al., “Orthogonal gradient descent for continual learning.” International Conference on Artificial Intelligence and Statistics. PMLR, (2020): 11 pages. [cited by applicant]
Feldman et al., “What neural networks memorize and why: Discovering the long tail via influence estimation.” Advances in Neural Information Processing Systems 34 (2020): 11 pages. [cited by applicant]
Feldman, Vitaly. “Does learning require memorization? a short tale about a long tail.” Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing. (2020): pp. 954-959. [cited by applicant]
Geva et al., “Transformer Feed-Forward Layers are Key-Value Memories.” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. (2021): pp. 5484-5495. [cited by applicant]
Geva et al., “Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.” arXiv preprint arXiv:2203.14680 (2022): 16 pages. [cited by applicant]
Gorban et al., “Blessing of dimensionality: mathematical foundations of the statistical physics of data.” Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 376.2118 (2018… [cited by applicant]
Grave, Edouard et al “Improving neural language models with a continuous cache.” arXiv preprint arXiv:1612.04426 (2016): 9 pages. [cited by applicant]
Graves et al., “Unbounded cache model for online language modeling with open vocabulary.” Advances in neural information processing systems 30 (2017): 11 pages. [cited by applicant]
Graves, Alex et al.,“Neural turing machines.” arXiv preprint arXiv:1410.5401 (2014): 26 pages. [cited by applicant]
Graves, Alex, et al. “Hybrid computing using a neural network with dynamic external memory.” Nature 538.7626 (2016): 21 pages. [cited by applicant]
Gulcehre et al. “Dynamic neural turing machine with continuous and discrete addressing schemes.” Neural computation 30.4 (2018): 24 pages. [cited by applicant]
Hopefield, John J. “Neural networks and physical systems with emergent collective computational abilities.” Proceedings of the national academy of sciences 79.8 (1982): 2554-2558. [cited by applicant]
Howard, Marc W. “Formal models of memory based on temporally-varying representations.” The new handbook of mathematical psychology 3 (2022): 40 pages. [cited by applicant]
Huang et al, “Transformers-patcher: One mistake with one neuron,” ICLR (2023): 16 pages. [cited by applicant]
Iscen et al., “Improving image recognition by retrieving from web-scale image-text data.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. (2023): pp. 19295-19304. [cited by applicant]
Kanerva, Pentti, “Sparse distributed memory,” MIT Press chapter 3, (1988): 53 pages. [cited by applicant]
Kanerva, Pentti. “Sparse distributed memory and related models,” No. NASA-CR-190553. (1992): 58 pages. [cited by applicant]
Kohonen et al., “Representation of associated data by matrix operators.” IEEE Transactions on Computers 100.7 (1973): 701-702. [cited by applicant]
Kohonen, Teuvo. “Correlation matrix memories.” IEEE transactions on computers 100.4 (1972): 353-359. [cited by applicant]
Kohonen, Teuvo , “Self-organization and associative memory,” vol. 8, chapter 1, Springer Science & Business Media, (2012) pp. 1-11, 14-18, 21-25. & 28. [cited by applicant]
Krotov, Dmitry et al., “Large associative memory problem in neurobiology and machine learning.” arXiv preprint arXiv:2008.06996 (2020): 12 pages. [cited by applicant]
Krotov, Dmitry. “Hierarchical Associative Memory.” arXiv preprint arXiv:2107.06446 (2021): 13 pages. [cited by applicant]
Kuh, Anthony. “Performance measures for associative memories that learn and forget.” Neural Information Processing Systems. (1987): pp. 432-441. [cited by applicant]
Lample et al. “Large memory layers with product keys.” Advances in Neural Information Processing Systems 32 (2019): 12 pages. [cited by applicant]
Le et al. “Variational memory encoder-decoder.” Advances in neural information processing systems 31 (2018): 11 pages. [cited by applicant]
Le et al., “Learning to remember more with less memorization.” arXiv preprint arXiv:1901.01347 (2019): 20 pages. [cited by applicant]
Le, Hung et al., “Self-attentive associative memory.” International Conference on Machine Learning. PMLR, (2020): 10 pages. [cited by applicant]
Liang et al., “Associative Learning for Network Embedding.” arXiv preprint arXiv:2208.14376 (2022): 5 pages. [cited by applicant]
Little, William A. “The existence of persistent states in the brain.” Mathematical biosciences 19.1-2 (1974): 101-120. [cited by applicant]
Liu et al., “Think-in-Memory: Recalling and Post-thinking Enable LLMs with Long-Term Memory.” arXiv preprint arXiv:2311.08719 (2023): 9 pages. [cited by applicant]
Mablestone, Adam et al., “Product kanerva machines: Factorized bayesian memory.” arXiv preprint arXiv:2002.02385 (2020): 20 pages. [cited by applicant]
Maekawa et al., “Generative Replay Inspired by Hippocampal Memory Indexing for Continual Language Learning.” Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. (… [cited by applicant]
McClelland et al., “Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory.” Psychological review 102.3 (19… [cited by applicant]
Li et al., “Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space” arXiv, Apr. 5, 2020, 22 pages, doi: https://arxiv.org/abs/2004.04092v4. [cited by applicant]