RNNs are suboptimal at compressing past sensory inputs for future prediction.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Modeling a temporal process as if it is Markovian assumes the present encodes all of the process's history. When this occurs, the present captures all of the dependency between past and future. We recently showed that if one randomly samples in the space of structured processes, this is almost never the case. So, how d…
The study examines how investor protection and past information affect stock returns and interest rates.
EvoRate metric assesses learnability of sequential data by measuring predictive information.
We consider the problem of building a state representation model in a continual fashion. As the environment changes, the aim is to efficiently compress the sensory state's information without losing past knowledge. The learned features are then fed to a Reinforcement Learning algorithm to learn a policy. We propose to …
Past lightcones of certain points in globally hyperbolic spacetimes determine the entire spacetime.
One of the most fundamental questions one can ask about a pair of random variables X and Y is the value of their mutual information. Unfortunately, this task is often stymied by the extremely large dimension of the variables. We might hope to replace each variable by a lower-dimensional representation that preserves th…
PI-SAC agents learn predictive information to improve RL efficiency.
Paper analyzes iterative learning for concept classes and learns half-spaces.
We consider the problem of building a state representation model for control, in a continual learning setting. As the environment changes, the aim is to efficiently compress the sensory state's information without losing past knowledge, and then use Reinforcement Learning on the resulting features for efficient policy …
Neural model uses deductive database to predict events from past patterns.
We propose a new method to study the internal memory used by reinforcement learning policies. We estimate the amount of relevant past information by estimating mutual information between behavior histories and the current action of an agent. We perform this estimation in the passive setting, that is, we do not interven…
Dynamical-VAE learns causal dynamics from POMDPs using future information.
We study the flow of information and the evolution of internal representations during deep neural network (DNN) training, aiming to demystify the compression aspect of the information bottleneck theory. The theory suggests that DNN training comprises a rapid fitting phase followed by a slower compression phase, in whic…
Stochastic volatility models describe stock returns as driven by an unobserved process capturing the random dynamics of volatility . The present paper quantifies how much information about volatility and future stock returns can be inferred from past returns in stochastic volatility models in terms of …
Meta Optimal Transport learns from past problems to solve similar OT problems faster.
Algorithm learns actions from past states in complex tasks.
Study on RNNs' ability to approximate past-dependent Hölder functions and their application to regression.
Bayesian method predicts future network configurations from past snapshots.
Bayesian mixture models are widely applied for unsupervised learning and exploratory data analysis. Markov chain Monte Carlo based on Gibbs sampling and split-merge moves are widely used for inference in these models. However, both methods are restricted to limited types of transitions and suffer from torpid mixing and…
We present two sampled quasi-Newton methods (sampled LBFGS and sampled LSR1) for solving empirical risk minimization problems that arise in machine learning. Contrary to the classical variants of these methods that sequentially build Hessian or inverse Hessian approximations as the optimization progresses, our proposed…
We discovered that past changes in the market correlation structure are significantly related with future changes in the market volatility. By using correlation-based information filtering networks we device a new tool for forecasting the market volatility changes. In particular, we introduce a new measure, the "correl…
AdaX improves Adam by exponentially accumulating past gradients, leading to better performance in machine learning tasks.
New algorithms assign credit to past decisions based on hindsight.
Learning long-term dependencies in extended temporal sequences requires credit assignment to events far back in the past. The most common method for training recurrent neural networks, back-propagation through time (BPTT), requires credit information to be propagated backwards through every single step of the forward c…
Transformers learn to integrate information from past positions incrementally, specializing heads in distinct patterns.
Financial markets, with their vast range of different investment opportunities, can be seen as a system of many different simultaneous games with diverse and often unknown levels of risk and reward. We introduce generalizations to the classic Kelly investment game [Kelly (1956)] that incorporates these features, and us…
A major drawback of backpropagation through time (BPTT) is the difficulty of learning long-term dependencies, coming from having to propagate credit information backwards through every single step of the forward computation. This makes BPTT both computationally impractical and biologically implausible. For this reason,…
Surface mount technology (SMT) is a process for producing printed circuit boards. Solder paste printer (SPP), package mounter, and solder reflow oven are used for SMT. The board on which the solder paste is deposited from the SPP is monitored by solder paste inspector (SPI). If SPP malfunctions due to the printer defec…
Develops new Markov processes with switching rates and past dependence.
The paper introduces neural INGARCH models for time series of counts.
DBULL learns new clusters without forgetting past knowledge in streaming unlabelled data.
IAM improves deep RL by selectively storing influential past observations.
We study an intrinsic distribution, called polar, on the space of -dimensional integral elements of the higher order contact structure on jet spaces. The main result establishes that this exterior differential system is the prolongation of a natural system of PDEs, named pasting conditions, on sections of the bundle…
The proliferation of automated inference algorithms in Bayesian statistics has provided practitioners newfound access to fast, reproducible data analysis and powerful statistical models. Designing automated methods that are also both computationally scalable and theoretically sound, however, remains a significant chall…
We consider the problem of predicting the next observation given a sequence of past observations, and consider the extent to which accurate prediction requires complex algorithms that explicitly leverage long-range dependencies. Perhaps surprisingly, our positive results show that for a broad class of sequences, there …
We note a simple mechanism that may at least partially resolve several outstanding economic puzzles, including why the cyclically adjusted price to earnings ratio of the S&P 500 index has been oddly high for the past two decades, why gains to capital have outpaced gains to wages, and the persistence of the equity premi…
Modeling financial crises and cryptocurrency shocks using copulae clustering.
Improved ARMA-GARCH model for illiquid assets like cryptocurrencies.
A technique identifies memoryless algorithms approximating memory-dependent optimization methods.
How technology affects growth or employment has long been debated. With a hiatus, the debate revived once again in the form of how Information and Communications Technology, as a form of new technology, exerts on productivity and employment. Information and Communications Technology perceived as General Purpose Technol…
Market impact is a key concept in the study of financial markets and several models have been proposed in the literature so far. The Transient Impact Model (TIM) posits that the price at high frequency time scales is a linear combination of the signs of the past executed market orders, weighted by a so-called propagato…
RPO uses past and future state-action info for better policy optimization.
The study extracts market direction from transaction data.
New method achieves superlinear convergence rate with limited memory.
This paper studies a composite problem involving the decision making of the optimal entry time and dynamic consumption afterwards. In stage-1, the investor has access to full market information subjecting to some information costs and needs to choose an optimal stopping time to initiate stage-2; in stage-2, the investo…
Predictive rate-distortion analysis suffers from the curse of dimensionality: clustering arbitrarily long pasts to retain information about arbitrarily long futures requires resources that typically grow exponentially with length. The challenge is compounded for infinite-order Markov processes, since conditioning on fi…
Approach collects missing outcomes to improve fairness in classification.