We augment recurrent neural networks with an external memory mechanism that builds upon recent progress in metalearning. We conceptualize this memory as a rapidly adaptable function that we parameterize as a deep neural network. Reading from the neural memory function amounts to pushing an input (the key vector) throug…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Recurrent neural networks can learn complex transduction problems that require maintaining and actively exploiting a memory of their inputs. Such models traditionally consider memory and input-output functionalities indissolubly entangled. We introduce a novel recurrent architecture based on the conceptual separation b…
A new memory replay mechanism improves reinforcement learning stability and speed.
Enhanced Hopfield model boosts memory retrieval capacity.
A general market model with memory is considered in terms of stochastic functional differential equations. We aim at representation formulae for the sensitivity analysis of the dependence of option prices on the memory. This implies a generalization of the concept of delta.
Memory-constrained algorithms need superlinear memory for efficient convex optimization.
Novel method learns memory kernels in Langevin equations.
Study on memory effects in RNNs learning temporal data.
A model of associative memory is studied, which stores and reliably retrieves many more patterns than the number of neurons in the network. We propose a simple duality between this dense associative memory and neural networks commonly used in deep learning. On the associative memory side of this duality, a family of mo…
We note that known methods achieving the optimal oracle complexity for first order convex optimization require quadratic memory, and ask whether this is necessary, and more broadly seek to characterize the minimax number of first order queries required to optimize a convex Lipschitz function subject to a memory constra…
Typical neural networks with external memory do not effectively separate capacity for episodic and working memory as is required for reasoning in humans. Applying knowledge gained from psychological studies, we designed a new model called Differentiable Working Memory (DWM) in order to specifically emulate human workin…
Memory networks are neural networks with an explicit memory component that can be both read and written to by the network. The memory is often addressed in a soft way using a softmax function, making end-to-end training with backpropagation possible. However, this is not computationally scalable for applications which …
New memory models improve LSSVM's generalization and reduce time cost.
Generalizes memory and forecasting capacities for nonlinear recurrent networks with dependent inputs.
This paper introduces a hierarchical associative memory model with multiple layers.
This work analyzes and optimizes memory and compute costs of learned optimizers.
New memory-query tradeoffs for convex optimization algorithms.
A new model for sequential memory using temporal predictive coding.
Study identifies three quantization regimes for ReLU networks.
Generative diffusion models mimic biological memory networks, encoding associative dynamics in deep neural weights.
A technique identifies memoryless algorithms approximating memory-dependent optimization methods.
Sparse Hopfield model improves memory retrieval with fewer connections.
Improves energy efficiency of neuromorphic hardware by optimizing memory organization and encoding schemes.
We propose a stochastic process driven by memory effect with novel distributions including both exponential and leptokurtic heavy-tailed distributions. A class of distribution is analytically derived from the continuum limit of the discrete binary process with the renormalized auto-correlation and the closed form momen…
Energy Transformer integrates attention, energy models, and associative memory.
A new theory explains large associative memory with biological plausibility.
A central challenge faced by memory systems is the robust retrieval of a stored pattern in the presence of interference due to other stored patterns and noise. A theoretically well-founded solution to robust retrieval is given by attractor dynamics, which iteratively clean up patterns during recall. However, incorporat…
Despite their attractiveness, popular perception is that techniques for nonparametric function approximation do not scale to streaming data due to an intractable growth in the amount of storage they require. To solve this problem in a memory-affordable way, we propose an online technique based on functional stochastic …
Financial market dynamics is rigorously studied via the exact generalized Langevin equation. Assuming market Brownian self-similarity, the market return rate memory and autocorrelation functions are derived, which exhibit an oscillatory-decaying behavior with a long-time tail, similar to empirical observations. Individ…
New neural operators model turbulence with memory and randomness.
Deep learning typically requires training a very capable architecture using large datasets. However, many important learning problems demand an ability to draw valid inferences from small size datasets, and such problems pose a particular challenge for deep learning. In this regard, various researches on "meta-learning…
We obtain option pricing formulas for stock price models in which the drift and volatility terms are functionals of a continuous history of the stock prices. That is, the stock dynamics follows a nonlinear stochastic functional differential equation. A model with full memory is obtained via approximation through a stoc…
Tensor models decode human perception and memory using SPO triples.
Develops hyperparameter transfer methods for Dense Associative Memories.
HiPPO framework optimizes memory compression for sequential data.
For the London Stock Exchange we demonstrate that the signs of orders obey a long-memory process. The autocorrelation function decays roughly as with , corresponding to a Hurst exponent . This implies that the signs of future orders are quite predictable from the signs of past orde…
V-HMN integrates memory mechanisms for improved image recognition.
Quadratic memory is essential for optimal convex optimization queries.
Survey on statistical inference under memory constraints.
New matrix approximation method using RBF components for better memory efficiency.
A method for learning complex functions from data with reduced memory usage.
MEM learns set functions from permutation-invariant data.
The origin of the long-range memory in the non-equilibrium systems is still an open problem as the phenomenon can be reproduced using models based on Markov processes. In these cases a notion of spurious memory is introduced. A good example of Markov processes with spurious memory is stochastic process driven by a non-…
We focus on emergence of the power-law cross-correlations from processes with both short and long term memory properties. In the case of correlated error-terms, the power-law decay of the cross-correlation function comes automatically with the characteristics of separate processes. Bivariate Hurst exponent is then equa…
Memory capacity of DAM scales exponentially with feature separation, unaffected by correlations.
We study the problem of learning associative memory -- a system which is able to retrieve a remembered pattern based on its distorted or incomplete version. Attractor networks provide a sound model of associative memory: patterns are stored as attractors of the network dynamics and associative retrieval is performed by…
We consider the problem of performing linear regression over a stream of -dimensional examples, and show that any algorithm that uses a subquadratic amount of memory exhibits a slower rate of convergence than can be achieved without memory constraints. Specifically, consider a sequence of labeled examples $(a_1,b_1)…
New algorithms for constrained online optimization with memory and predictions.