New perspective on SGD reveals short-range memory effects in deep learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we present the results of Monte Carlo simulations for two popular techniques of long-range correlations detection - classical and modified rescaled range analyses. A focus is put on an effect of different distributional properties on an ability of the methods to efficiently distinguish between short and …
Embed-KCPD segments text without labels, outperforming baselines.
Study finds phase transition in context-sensitive language model with short-range interactions.
The problem of recovering the asymptotics of a short range perturbation of the Euclidean metric on R^n from fixed energy scattering data is studied. It is shown that if two such metrics, g1, g2, have scattering data at some fixed energy which are equal up to smoothing, then there exists a diffeomorphism ψ`fixing infini…
Proves energy expression on Poincaré-Einstein spaces.
We introduce a general framework of the Mixed-correlated ARFIMA (MC-ARFIMA) processes which allows for various specifications of univariate and bivariate long-term memory. Apart from a standard case when , MC-ARFIMA also allows for processes with but also for long-range …
Split conformal prediction works well for time series despite temporal dependence.
Study evaluates neural networks based on random graph structures and finds key performance indicators.
News sentiment in U.S. economic newspapers has become more persistent over 45 years.
We present a simple microstructure model of financial returns that combines (i) the well-known ARFIMA process applied to tick-by-tick returns, (ii) the bid-ask bounce effect, (iii) the fat tail structure of the distribution of returns and (iv) the non-Poissonian statistics of inter-trade intervals. This model allows us…
We consider Markov models of stochastic processes where the next-step conditional distribution is defined by a kernel density estimator (KDE), similar to Markov forecast densities and certain time-series bootstrap schemes. The KDE Markov models (KDE-MMs) we discuss are nonlinear, nonparametric, fully probabilistic repr…
We present a modification of the so-called Parrondo's paradox where one is allowed to choose in each turn the game that a large number of individuals play. It turns out that, by choosing the game which gives the highest average earnings at each step, one ends up with systematic loses, whereas a periodic or random seque…
V1 cortex reconstructs images as Poisson equation solutions with varying weights.
We study consumption behaviour in systems with heterogeneous interacting agents. Two different models are introduced, respectively with long and short range interactions among agents. At any time step an agent decides whether or not to consume a good, doing so if this provides positive utility. Utility is affected by i…
In this paper we consider certain asymptotically Euclidean spaces, namely compact manifolds with boundary X equipped with a scattering metric g, as defined by Melrose. We then consider Hamiltonians H which are `short-range' self-adjoint perturbations of the Laplacian of g. Melrose and Zworski have given a detailed desc…
We review the spectral analysis and the time-dependent approach of scattering theory for manifolds with asymptotically cylindrical ends. For the spectral analysis, higher order resolvent estimates are obtained via Mourre theory for both short-range and long-range behaviors of the metric and the perturbation at infinity…
Convolutional neural networks are commonly used to control the steering angle for autonomous cars. Most of the time, multiple long range cameras are used to generate lateral failure cases. In this paper we present a novel model to generate this data and label augmentation using only one short range fisheye camera. We p…
We propose a Standing Wave Decomposition (SWD) approximation to Gaussian Process regression (GP). GP involves a costly matrix inversion operation, which limits applicability to large data analysis. For an input space that can be approximated by a grid and when correlations among data are short-ranged, the kernel matrix…
Clinical outcome prediction based on the Electronic Health Record (EHR) plays a crucial role in improving the quality of healthcare. Conventional deep sequential models fail to capture the rich temporal patterns encoded in the longand irregular clinical event sequences. We make the observation that clinical events at a…
Heretofore, neural networks with external memory are restricted to single memory with lossy representations of memory interactions. A rich representation of relationships between memory pieces urges a high-order and segregated relational memory. In this paper, we propose to separate the storage of individual experience…
In this paper, we propose a dual memory structure for reinforcement learning algorithms with replay memory. The dual memory consists of a main memory that stores various data and a cache memory that manages the data and trains the reinforcement learning agent efficiently. Experimental results show that the dual memory …
The paper analyzes time-dependent streaming data with biased gradient estimates and proposes improved stochastic optimization methods.
Autoregressive networks can achieve promising performance in many sequence modeling tasks with short-range dependence. However, when handling high-dimensional inputs and outputs, the huge amount of parameters in the network lead to expensive computational cost and low learning efficiency. The problem can be alleviated …
Stable Hadamard Memory improves reinforcement learning by efficiently managing memory.
We analyze the resolvent of Schrödinger operators with short range potential on asymptotically conic manifolds (this setting includes asymptotically Euclidean manifolds) near . We make the assumption that the dimension is greater or equal to 3 and that has no null …
Image partitioning, or segmentation without semantics, is the task of decomposing an image into distinct segments, or equivalently to detect closed contours. Most prior work either requires seeds, one per segment; or a threshold; or formulates the task as multicut / correlation clustering, an NP-hard problem. Here, we …
Fractal analysis is carried out on the stock market indices of seven European countries and the US. We find evidence of long range dependence in the log return series of the Mibtel (Italy) and the PX Glob (Czech Republic). Long range dependence implies that predictable patterns in the log returns do not dissipate quick…
New memory allocation scheme improves image generation performance.
DAM with MRL improves relational reasoning in MANNs.
We discuss memory models which are based on tensor decompositions using latent representations of entities and events. We show how episodic memory and semantic memory can be realized and discuss how new memory traces can be generated from sensory input: Existing memories are the basis for perception and new memories ar…
mGRN improves multivariate time series prediction by managing marginal and joint memories.
Paper tackles non-stationary bandits with various examples.
Model combines long-term and short-term memory using conceptors.
Learning to remember long sequences remains a challenging task for recurrent neural networks. Register memory and attention mechanisms were both proposed to resolve the issue with either high computational cost to retain memory differentiability, or by discounting the RNN representation learning towards encoding shorte…
In recent years, memory-augmented neural networks(MANNs) have shown promising power to enhance the memory ability of neural networks for sequential processing tasks. However, previous MANNs suffer from complex memory addressing mechanism, making them relatively hard to train and causing computational overheads. Moreove…
Enhanced Hopfield model boosts memory retrieval capacity.
Current generation of memory-augmented neural networks has limited scalability as they cannot efficiently process data that are too large to fit in the external memory storage. One example of this is lifelong learning scenario where the model receives unlimited length of data stream as an input which contains vast majo…
Memory networks are neural networks with an explicit memory component that can be both read and written to by the network. The memory is often addressed in a soft way using a softmax function, making end-to-end training with backpropagation possible. However, this is not computationally scalable for applications which …
Memory-constrained algorithms need superlinear memory for efficient convex optimization.
We present an end-to-end trained memory system that quickly adapts to new data and generates samples like them. Inspired by Kanerva's sparse distributed memory, it has a robust distributed reading and writing mechanism. The memory is analytically tractable, which enables optimal on-line compression via a Bayesian updat…
Generative diffusion models mimic biological memory networks, encoding associative dynamics in deep neural weights.
A new model for sequential memory using temporal predictive coding.
We augment recurrent neural networks with an external memory mechanism that builds upon recent progress in metalearning. We conceptualize this memory as a rapidly adaptable function that we parameterize as a deep neural network. Reading from the neural memory function amounts to pushing an input (the key vector) throug…
Accelerators with power-law memory are proposed in the framework of the discrete time approach. To describe discrete accelerators we use the capital stock adjustment principle, which has been suggested by Matthews.The suggested discrete accelerators with memory describe the economic processes with the power-law memory …
A new memory system handles non-stationary environments by self-sizing and retaining memories.
Paper questions RNN and LSTM's long-term memory and introduces a new definition.
Deep learning typically requires training a very capable architecture using large datasets. However, many important learning problems demand an ability to draw valid inferences from small size datasets, and such problems pose a particular challenge for deep learning. In this regard, various researches on "meta-learning…