Unified framework for adaptive learning systems using consolidation and expansion operations.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
EVCL combines VCL and EWC to prevent forgetting new tasks.
EWC helps prevent forgetting in neural networks by adjusting weights dynamically.
The brain optimizes memory by forgetting what's predictable, improving generalization.
Technological progress is leading to proliferation and diversification of trading venues, thus increasing the relevance of the long-standing question of market fragmentation versus consolidation. To address this issue quantitatively, we analyse systems of adaptive traders that choose where to trade based on their previ…
Algorithm improves model performance on shifted concepts without retraining.
Collecting the large datasets needed to train deep neural networks can be very difficult, particularly for the many applications for which sharing and pooling data is complicated by practical, ethical, or legal concerns. However, it may be the case that derivative datasets or predictive models developed within individu…
Artificial neural networks face the well-known problem of catastrophic forgetting. What's worse, the degradation of previously learned skills becomes more severe as the task sequence increases, known as the long-term catastrophic forgetting. It is due to two facts: first, as the model learns more tasks, the intersectio…
Dual representations for robust risk measures and uncertainty sets.
Elastic weight consolidation (EWC, Kirkpatrick et al, 2017) is a novel algorithm designed to safeguard against catastrophic forgetting in neural networks. EWC can be seen as an approximation to Laplace propagation (Eskin et al, 2004), and this view is consistent with the motivation given by Kirkpatrick et al (2017). In…
FedFMC improves federated learning on non-iid data without sharing data or increasing communication costs.
Paper uses AI to predict stock market volatility with neural networks and genetic algorithms.
Class incremental learning refers to a special multi-class classification task, in which the number of classes is not fixed but is increasing with the continual arrival of new data. Existing researches mainly focused on solving catastrophic forgetting problem in class incremental learning. To this end, however, these m…
Blog post discusses various implementations of Fisher Information for EWC in continual learning.
This paper consists of two parts. The first part is devoted to empirical analysis of consolidated order book (COB) for the index RTS futures. In the second part we consider Poissonian multi--agent model of the COB. By varying parameters of different groups of agents submitting orders to the book we are able to model va…
We propose a method for tackling catastrophic forgetting in deep reinforcement learning that is \textit{agnostic} to the timescale of changes in the distribution of experiences, does not require knowledge of task boundaries, and can adapt in \textit{continuously} changing environments. In our \textit{policy consolidati…
Proposes a framework for semi-supervised continual learning from sequentially arriving data.
Proposes a CL technique to improve accuracy and reduce forgetting.
Humans do not acquire perceptual abilities in the way we train machines. While machine learning algorithms typically operate on large collections of randomly-chosen, explicitly-labeled examples, human acquisition relies more heavily on multimodal unsupervised learning (as infants) and active learning (as children). Wit…
The 2008 financial crisis revealed banking consolidation paradoxically increased systemic fragility and global financial contagion with negligible spatial decay.
Both the scientific community and the popular press have paid much attention to the speed of the Securities Information Processor, the data feed consolidating all trades and quotes across the US stock market. Rather than the speed of the Securities Information Processor, or SIP, we focus here on its accuracy. Relying o…
This paper investigates the relevance of the No-Ponzi game condition for public debt (i.e. the public debt growth rate has to be lower than the real interest rate, a necessary assumption for Ricardian equivalence) and of the transversality condition for the GDP growth rate (i.e. the GDP growth rate has to be lower than…
Under Solvency II the computation of capital requirements is based on value at risk (V@R). V@R is a quantile-based risk measure and neglects extreme risks in the tail. V@R belongs to the family of distortion risk measures. A serious deficiency of V@R is that firms can hide their total downside risk in corporate network…
Paper benchmarks CF mitigation in federated time series forecasting.
This research proposes a CL model for RNNs to handle sequential data without forgetting.
New method improves ABI for sequential data, reducing forgetting and improving accuracy.
Analyzes self-attention in recurrent networks, proving it mitigates vanishing gradients.
The concept of soliton, in its most general version, allows us to find canonical or distinguished elements on any set provided with an equivalence relation and an `optimal' tangent direction at each point. We study in this paper solitons on homogeneous spaces, which have consolidated its role as a quite useful tool to …
We discuss memory models which are based on tensor decompositions using latent representations of entities and events. We show how episodic memory and semantic memory can be realized and discuss how new memory traces can be generated from sensory input: Existing memories are the basis for perception and new memories ar…
A model retains learned knowledge for longer by adding a plastic component to neural networks.
The paper surveys the topic of tensor decompositions in modern machine learning applications. It focuses on three active research topics of significant relevance for the community. After a brief review of consolidated works on multi-way data analysis, we consider the use of tensor decompositions in compressing the para…
Paper proves uniqueness of a complex construction.
Modified PCA algorithm with continual learning preserves features of previous modes for multimode process monitoring.
Sequential learning of multiple tasks in artificial neural networks using gradient descent leads to catastrophic forgetting, whereby previously learned knowledge is erased during learning of new, disjoint knowledge. Here, we propose a new approach to sequential learning which leverages the recent discovery of adversari…
We present the "Annotation and Benchmarking on Understanding and Transparency of Machine Learning Lifecycles" (ABOUT ML) project as an initiative to operationalize ML transparency and work towards a standard ML documentation practice. We make the case for the project's relevance and effectiveness in consolidating dispa…
We address the problem of classifying discrete differential-geometric Poisson brackets (dDGPBs) of any fixed order on target space of dimension 1. It is proved that these Poisson brackets (PBs) are in one-to-one correspondence with the intersection points of certain projective hypersurfaces. In addition, they can be re…
Our research is focused on understanding and applying biological memory transfers to new AI systems that can fundamentally improve their performance, throughout their fielded lifetime experience. We leverage current understanding of biological memory transfer to arrive at AI algorithms for memory consolidation and repl…
Unsupervised learning on imbalanced data is challenging because, when given imbalanced data, current model is often dominated by the major category and ignores the categories with small amount of data. We develop a latent variable model that can cope with imbalanced data by dividing the latent space into a shared space…
A machine learning approach to record fusion with high accuracy.
This article provides a new representation for pricing adjustments in derivatives.
Unified framework for Gaussian process methods in differential equations.
We define the (total) center of mass for suitably asymptotically hyperbolic time-slices of asymptotically anti-de Sitter spacetimes in general relativity. We do so in analogy to the picture that has been consolidated for the (total) center of mass of suitably asymptotically Euclidean time-slices of asymptotically Minko…
TMNs model brain memory systems for continual learning.
We study positive scalar curvature on the regular part of Riemannian manifolds with singular, uniformly Euclidean () metrics that consolidate Gromov's scalar curvature polyhedral comparison theory and edge metrics that appear in the study of Einstein manifolds. We show that, in all dimensions, edge singularit…
The two key issues of modern Bayesian statistics are: (i) establishing principled approach for distilling statistical prior that is consistent with the given data from an initial believable scientific prior; and (ii) development of a Bayes-frequentist consolidated data analysis workflow that is more effective than eith…
A streaming GNN model tackles continual learning for updating node representations in real-time.
Learning-to-learn or meta-learning leverages data-driven inductive bias to increase the efficiency of learning on a novel task. This approach encounters difficulty when transfer is not advantageous, for instance, when tasks are considerably dissimilar or change over time. We use the connection between gradient-based me…
Survey of methods to recover CI graphs from feature relationships.