US 8589163B2
· Ljolje et al.
· 2013
[cited by applicant]
CN 116127020A
· 2023
[cited by applicant]
CN 116935288A
· 2023
[cited by applicant]
CN 116975192A
· 2023
[cited by applicant]
CN 117093679A
· 2023
[cited by applicant]
Longuet-Higgins et al., “Theories of associative recall.” Quarterly reviews of biophysics 3.2 (1970): 223-244.
[cited by applicant]
Meng et al. “Mass-editing memory in a transformer.” arXiv preprint arXiv:2210.07229 (2023): 21 pages.
[cited by applicant]
Meng et al., Locating and editing factual associations in GPT, Advances in Neural Information Processing Systems. (2022): 35 pages.
[cited by applicant]
Meng, Yuanliang et al., “Context-aware neural model for temporal information extraction.” Proceedings of the 56th annual meeting of the association for computational linguistics (2018): 10 pages.
[cited by applicant]
Mitchell et al., “Fast Model Editing at Scale.” International Conference on Learning Representations. (2021): 21 pages.
[cited by applicant]
Mitchell et al., “Memory-based model editing at scale.” International Conference on Machine Learning. PMLR, (2022): 15 pages.
[cited by applicant]
Murdock, Bennet B. “Developing TODAM: Three models for serial-order information.” Memory & Cognition 23.5 (1995): 631-645.
[cited by applicant]
Personnaz et al., “Collective computational properties of neural networks: New learning mechanisms.” Physical Review A 34.5 (1986): 4217-4228.
[cited by applicant]
Petroni et al., “Language models as knowledge bases?. ” arXiv preprint arXiv:1909.01066 (2019): 11 pages.
[cited by applicant]
Pham et. al. “Generative Pseudo-Inverse Memory” ICLR (2022): 18 pages.
[cited by applicant]
Plate, Tony A. “Holographic reduced representations.” IEEE Transactions on Neural networks 6.3 (1995): 21 pages.
[cited by applicant]
Rae, Jack, et al. “Scaling memory-augmented neural networks with sparse reads and writes.” Advances in Neural Information Processing Systems 29 (2016): 9 pages.
[cited by applicant]
Ramsauer et al., “Hopfield networks is all you need.” arXiv preprint arXiv:2008.02217 (2020): 94 pages.
[cited by applicant]
Raunak et al., “Rank-One Editing of Encoder-Decoder Models.” arXiv preprint arXiv:2211.13317 (2022): 6 pages.
[cited by applicant]
Redhakrishnan et al., “Overparameterized neural networks implement associative memory.” Proceedings of the National Academy of Sciences 117.44 (2020): 27162-27170.
[cited by applicant]
Saha et al., “Gradient projection memory for continual learning.” arXiv preprint arXiv:2103.09762 (2021): 18 pages.
[cited by applicant]
Salvatori et al., “Associative memories via predictive coding.” Advances in Neural Information Processing Systems 34 (2021): 13 pages.
[cited by applicant]
Schlag et al., “Linear transformers are secretly fast weight programmers.” International Conference on Machine Learning. PMLR, (2021) 12 pages.
[cited by applicant]
Shun-Ichi Amari. “Learning patterns and pattern sequences by self-organizing nets of threshold elements.” IEEE Transactions on computers 100.11 (1972): 1197-1206.
[cited by applicant]
Smolensky, Paul. “Tensor product variable binding and the representation of symbolic structures in connectionist systems.” Artificial intelligence 46.1-2 (1990): 159-216.
[cited by applicant]
Stiles et al., “On the effect of noise on the Moore-Penrose generalized inverse associative memory.” IEEE transactions on pattern analysis and machine intelligence 3 (1985): 358-360.
[cited by applicant]
Sukhbaatar et al. “End-to-end memory networks.” Advances in neural information processing systems 28 (2015): 9 pages.
[cited by applicant]
Sukhbaatar et al., “Augmenting self-attention with persistent memory.” arXiv preprint arXiv:1907.01470 (2019): 11 pages.
[cited by applicant]
Sukhbaatar et al., “Not all memories are created equal: Learning to forget by expiring.” International Conference on Machine Learning. PMLR, (2021): 11 pages.
[cited by applicant]
Valle-Lisboa et al., “Multiplicative processing in the modeling of cognitive activities in large neural networks.” Biophysical Reviews (2023): 1-19.
[cited by applicant]
Willshaw et al. “Non-holographic associative memory.” Nature 222.5197 (1969): 960-962.
[cited by applicant]
Whittington et al., “Relating transformers to models and neural representations of the hippocampal formation.” arXiv preprint arXiv:2112.04035 (2021): 20 pages.
[cited by applicant]
Wu et al. “The Kanerva machine: A generative distributed memory.” arXiv preprint arXiv:1804.01756 (2018): 16 pages.
[cited by applicant]
Wu et al., “Learning attractor dynamics for generative memory.” Advances in Neural Information Processing Systems 31 (2018): 10 pages.
[cited by applicant]
Wu, Yuhuai et al. “Memorizing Transformers.” International Conference on Learning Representations. (2021): 19 pages.
[cited by applicant]
Yen et al., “A learning and forgetting algorithm in associative memories. The eigenstructure method.” [1991] Proceedings of the 30th IEEE Conference on Decision and Control. IEEE, (1991): pp. 847-852.
[cited by applicant]
Zhang et al., “Hippocampal spatial representations exhibit a hyperbolic geometry that expands with experience.” Nature Neuroscience 26.1 (2023): 131-139.
[cited by applicant]
Yen et al., “A learning and forgetting algorithm in associative memories: results involving pseudo inverses.” 1991., IEEE International Symposium on Circuits and Systems. IEEE, (1991): pp. 778-781.
[cited by applicant]
Steinbuch, K “Die Lernmatric” Kybernetik, 1(1): Jan. 1961. pp. 36-45.
[cited by applicant]
Anderson, J. A., “A simple neural network generating an interactive memory,” Mathematical Biosciences, vol. 14, (1972): pp. 197-220.
[cited by applicant]
Bau et al. “Rewriting a Deep Generative Model.” arXiv preprint arXiv:2007.15646 (2020): 31 pages.
[cited by applicant]
Weston et al., “Memory Networks.” in 3rd ICLR (2015): 15 pages.
[cited by applicant]
Armand Joulin et al., “Inferring algorithmic patterns with stack-augmented recurrent nets.” Advances in neural information processing systems 28 (2015): 9 pages.
[cited by applicant]
Bietti et al., “Birth of a Transformer: A Memory Viewpoint.” arXiv preprint arXiv:2306.00802 (2023): 28 pages.
[cited by applicant]
Bricken et al., “Attention approximates sparse distributed memory.” Advances in Neural Information Processing Systems 35 (2021): 15 pages.
[cited by applicant]
Bricken et al., “Sparse Distributed Memory is a Continual Learner.” arXiv preprint arXiv:2303.11934 (2023): 57 pages.
[cited by applicant]
Burtsev et al., “Memory transformer.” arXiv preprint arXiv:2006.11527 (2020): 17 pages.
[cited by applicant]
Cabannes et al, “Scaling laws for associative memories.” arXiv preprint arXiv:2310.02984 (2023): 32 pages.
[cited by applicant]
Caplan et al., “Associative recognition without hippocampal associations.” Psychological Review 129.6 (2022): 54 pages.
[cited by applicant]
Cheng et al., “Language model with Plug-in Knowldge Memory.” ICLR (2022): 22 pages.
[cited by applicant]
Dai et al., “Knowledge neurons in pretrained transformers.” arXiv preprint arXiv:2104.08696 (2021): 10 pages.
[cited by applicant]
Dai et al., “Neural knowledge bank for pretrained transformers.” CCF International Conference on Natural Language Processing and Chinese Computing. Cham: Springer Nature Switzerland, (2023): 11 pages.
[cited by applicant]
De Cao, Nicola et al., “Editing Factual Knowledge in Language Models.” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. (2021): pp. 6491-6506.
[cited by applicant]
Dong et al., “Calibrating factual knowledge in pretrained language models.” arXiv preprint arXiv:2210.03329 (2022): 11 pages.
[cited by applicant]
Eldan et al., “Who's Harry Potter? Approximate unlearning in LLMs,” arXiv:2310.02238v2, (2023): 21 pages.
[cited by applicant]
Elhage et al., “A mathematical framework for transformer circuits.” Transformer Circuits Thread 1 https://transformer-circuits.pub/2021/framework/index.html (retrieved Jan. 29, 2024), 47 pages.
[cited by applicant]
Fan et al., “Augmenting transformers with KNN-based composite memory for dialog.” Transactions of the Association for Computational Linguistics 9 (2021): 82-99.
[cited by applicant]
Farajtabar et al., “Orthogonal gradient descent for continual learning.” International Conference on Artificial Intelligence and Statistics. PMLR, (2020): 11 pages.
[cited by applicant]
Feldman et al., “What neural networks memorize and why: Discovering the long tail via influence estimation.” Advances in Neural Information Processing Systems 34 (2020): 11 pages.
[cited by applicant]
Feldman, Vitaly. “Does learning require memorization? a short tale about a long tail.” Proceedings of the 52nd Annual ACM SIGACT Symposium on Theory of Computing. (2020): pp. 954-959.
[cited by applicant]
Geva et al., “Transformer Feed-Forward Layers are Key-Value Memories.” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. (2021): pp. 5484-5495.
[cited by applicant]
Geva et al., “Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space.” arXiv preprint arXiv:2203.14680 (2022): 16 pages.
[cited by applicant]
Gorban et al., “Blessing of dimensionality: mathematical foundations of the statistical physics of data.” Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 376.2118 (2018…
[cited by applicant]
Grave, Edouard et al “Improving neural language models with a continuous cache.” arXiv preprint arXiv:1612.04426 (2016): 9 pages.
[cited by applicant]
Graves et al., “Unbounded cache model for online language modeling with open vocabulary.” Advances in neural information processing systems 30 (2017): 11 pages.
[cited by applicant]
Graves, Alex et al.,“Neural turing machines.” arXiv preprint arXiv:1410.5401 (2014): 26 pages.
[cited by applicant]
Graves, Alex, et al. “Hybrid computing using a neural network with dynamic external memory.” Nature 538.7626 (2016): 21 pages.
[cited by applicant]
Gulcehre et al. “Dynamic neural turing machine with continuous and discrete addressing schemes.” Neural computation 30.4 (2018): 24 pages.
[cited by applicant]
Hopefield, John J. “Neural networks and physical systems with emergent collective computational abilities.” Proceedings of the national academy of sciences 79.8 (1982): 2554-2558.
[cited by applicant]
Howard, Marc W. “Formal models of memory based on temporally-varying representations.” The new handbook of mathematical psychology 3 (2022): 40 pages.
[cited by applicant]
Huang et al, “Transformers-patcher: One mistake with one neuron,” ICLR (2023): 16 pages.
[cited by applicant]
Iscen et al., “Improving image recognition by retrieving from web-scale image-text data.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. (2023): pp. 19295-19304.
[cited by applicant]
Kanerva, Pentti, “Sparse distributed memory,” MIT Press chapter 3, (1988): 53 pages.
[cited by applicant]
Kanerva, Pentti. “Sparse distributed memory and related models,” No. NASA-CR-190553. (1992): 58 pages.
[cited by applicant]
Kohonen et al., “Representation of associated data by matrix operators.” IEEE Transactions on Computers 100.7 (1973): 701-702.
[cited by applicant]
Kohonen, Teuvo. “Correlation matrix memories.” IEEE transactions on computers 100.4 (1972): 353-359.
[cited by applicant]
Kohonen, Teuvo , “Self-organization and associative memory,” vol. 8, chapter 1, Springer Science & Business Media, (2012) pp. 1-11, 14-18, 21-25. & 28.
[cited by applicant]
Krotov, Dmitry et al., “Large associative memory problem in neurobiology and machine learning.” arXiv preprint arXiv:2008.06996 (2020): 12 pages.
[cited by applicant]
Krotov, Dmitry. “Hierarchical Associative Memory.” arXiv preprint arXiv:2107.06446 (2021): 13 pages.
[cited by applicant]
Kuh, Anthony. “Performance measures for associative memories that learn and forget.” Neural Information Processing Systems. (1987): pp. 432-441.
[cited by applicant]
Lample et al. “Large memory layers with product keys.” Advances in Neural Information Processing Systems 32 (2019): 12 pages.
[cited by applicant]
Le et al. “Variational memory encoder-decoder.” Advances in neural information processing systems 31 (2018): 11 pages.
[cited by applicant]
Le et al., “Learning to remember more with less memorization.” arXiv preprint arXiv:1901.01347 (2019): 20 pages.
[cited by applicant]
Le, Hung et al., “Self-attentive associative memory.” International Conference on Machine Learning. PMLR, (2020): 10 pages.
[cited by applicant]
Liang et al., “Associative Learning for Network Embedding.” arXiv preprint arXiv:2208.14376 (2022): 5 pages.
[cited by applicant]
Little, William A. “The existence of persistent states in the brain.” Mathematical biosciences 19.1-2 (1974): 101-120.
[cited by applicant]
Liu et al., “Think-in-Memory: Recalling and Post-thinking Enable LLMs with Long-Term Memory.” arXiv preprint arXiv:2311.08719 (2023): 9 pages.
[cited by applicant]
Mablestone, Adam et al., “Product kanerva machines: Factorized bayesian memory.” arXiv preprint arXiv:2002.02385 (2020): 20 pages.
[cited by applicant]
Maekawa et al., “Generative Replay Inspired by Hippocampal Memory Indexing for Continual Language Learning.” Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. (…
[cited by applicant]
McClelland et al., “Why there are complementary learning systems in the hippocampus and neocortex: insights from the successes and failures of connectionist models of learning and memory.” Psychological review 102.3 (19…
[cited by applicant]
Li et al., “Optimus: Organizing Sentences via Pre-trained Modeling of a Latent Space” arXiv, Apr. 5, 2020, 22 pages, doi: https://arxiv.org/abs/2004.04092v4.
[cited by applicant]