A model improves continual learning in reinforcement learning without forgetting.
problem Catastrophic forgetting in deep reinforcement learning.
method Policy consolidation model interacting with a cascade of hidden networks.
result Improves continual learning on various continuous control tasks.
The 2008 financial crisis revealed banking consolidation paradoxically increased systemic fragility and global financial contagion with negligible spatial decay.
problem Fundamental vulnerabilities in interconnected banking systems during the 2008 financial crisis were inadequately addressed by existing frameworks.
method Developed a unified spatial-network framework using spectral analysis of network Laplacian operators combined with spatial difference-in-differences identification.
result Banking consolidation paradoxically increased systemic fragility and global financial contagion with negligible spatial decay.
Develops a method to define admissible rewards for robust policy evaluation in RL.
problem Defining a reward function for robust off-policy evaluation in RL with limited data.
method Identifies an admissible set of reward functions ensuring policies are close to past behavior and can be evaluated with high confidence.
result Demonstrates the approach on synthetic and real-world domains, including a critical care application.
RPG improves sample efficiency in RL by learning optimal action ranks.
problem Sample inefficiency in reinforcement learning.
method Ranking Policy Gradient (RPG) method that learns optimal rank of discrete actions.
result RPG reduces sample complexity and improves sample efficiency for large-scale problems.
Study shows how adaptive traders decide between fragmented or consolidated markets based on venue demand.
problem Understanding market fragmentation and consolidation in adaptive trading systems.
method Analysis of adaptive traders choosing trading venues based on past experience, considering aggregate parameters like demand to supply ratio.
result Conditions for market fragmentation and stability of steady states are identified, showing fragmented states are metastable.
This paper investigates the relevance of the No-Ponzi game condition for public debt (i.e. the public debt growth rate has to be lower than the real interest rate, a necessary assumption for Ricardian equivalence) and of the transversality condition for the GDP growth rate (i.e. the GDP growth rate has to be lower than…
Unified framework for adaptive learning systems using consolidation and expansion operations.
problem Managing the balance between consolidating known knowledge and expanding into new evidence in adaptive learning systems.
method Introduces Consolidation-Expansion Operator Mechanics (OpMech) with the order-gap metric to control the balance.
result The order-gap signal provides real-time control and termination guarantees for adaptive learning systems.
This paper merges deterministic policy gradient estimations to improve deep reinforcement learning performance.
problem The bias-variance tradeoff in estimating and using policy gradients for deep reinforcement learning.
method Introduces elite policy gradients and a two-step merging method to balance bias-variance tradeoffs.
result Two-step merging outperforms interpolation merging and state-of-the-art algorithms on benchmark control tasks.
EVCL combines VCL and EWC to prevent forgetting new tasks.
problem Preventing catastrophic forgetting in continual learning.
method Hybrid model integrating VCL and EWC.
result Consistently outperforms baselines in learning new tasks.
EWC helps prevent forgetting in neural networks by adjusting weights dynamically.
problem Preventing forgetting in neural networks during training.
method EWC adjusts weights dynamically to prevent forgetting.
result EWC effectively prevents catastrophic forgetting in neural networks.
The brain optimizes memory by forgetting what's predictable, improving generalization.
problem Memory consolidation struggles with representational drift, semanticisation, and offline replay.
method Proposes predictive forgetting as a mechanism to optimize generalization by reducing complexity.
result Predictive forgetting improves information-theoretic generalization bounds on stored representations.
ANPyC combats forgetting by pruning and consolidating neural parameters.
problem Catastrophic forgetting in neural networks, especially with long-term tasks.
method Adversarial Neural Pruning and Synaptic Consolidation (ANPyC) to balance task-relevant and irrelevant parameters.
result ANPyC prevents forgetting while enabling efficient learning of multiple tasks.
Algorithm improves model performance on shifted concepts without retraining.
problem Improving model performance on shifted concepts with limited source data.
method Model consolidation of intermediate internal distributions after adaptation.
result Effective improvement in model performance on shifted concepts.
Collecting the large datasets needed to train deep neural networks can be very difficult, particularly for the many applications for which sharing and pooling data is complicated by practical, ethical, or legal concerns. However, it may be the case that derivative datasets or predictive models developed within individu…
A new approach to meta-RL reduces sample inefficiency.
problem Meta-RL requires impractical amounts of on-policy experience.
method Federated learning approach where off-policy learners solve individual tasks and consolidate solutions into a meta-learner.
result Significant improvements in meta-RL sample efficiency.
Dual representations for robust risk measures and uncertainty sets.
problem Characterizing continuity of robust risk measures and their uncertainty sets.
method Develop dual representations for robust risk measures and uncertainty sets based on distinct geometric assumptions.
result Two dual frameworks for consolidated uncertainty sets are complementary, not interchangeable.
Elastic weight consolidation (EWC, Kirkpatrick et al, 2017) is a novel algorithm designed to safeguard against catastrophic forgetting in neural networks. EWC can be seen as an approximation to Laplace propagation (Eskin et al, 2004), and this view is consistent with the motivation given by Kirkpatrick et al (2017). In…
This review explores causal decision-making to improve decision quality.
problem Effective decision-making requires understanding causal relationships.
method Causal structure learning, causal effect learning, and causal policy learning.
result Challenges in causal decision-making are identified and recent advances are discussed.
FedFMC improves federated learning on non-iid data without sharing data or increasing communication costs.
problem Efficiently updating a global model on non-iid data in federated learning.
method FedFMC dynamically forks devices into different global models, merges and consolidates them.
result FedFMC substantially improves upon earlier approaches to non-iid data in federated learning.
Paper uses AI to predict stock market volatility with neural networks and genetic algorithms.
problem Traditional methods for predicting stock market volatility have high errors.
method Back-propagation neural network and genetic algorithm integrated model.
result The model predicts future volatility with low errors and high accuracy.
Blog post discusses various implementations of Fisher Information for EWC in continual learning.
problem Improving Elastic Weight Consolidation (EWC) results by optimizing Fisher Information computation.
method Empirically compares different implementations of Fisher Information for EWC.
result Many reported EWC results can be improved by changing Fisher Information computation methods.
This paper consists of two parts. The first part is devoted to empirical analysis of consolidated order book (COB) for the index RTS futures. In the second part we consider Poissonian multi--agent model of the COB. By varying parameters of different groups of agents submitting orders to the book we are able to model va…
TREK uses distillation to help students solve hard problems.
problem Stalled progress on hard prompts when current policy lacks useful reasoning trajectories.
method TREK combines distillation and reinforcement learning to expand student support.
result TREK significantly improves student performance on mathematical reasoning and agentic tasks.
New model LMRC tackles class incremental learning without needing old classes.
problem Softmax suppression problem in class incremental learning.
method Label Mapping with Response Consolidation (LMRC) and multi-head neural network.
result LMRC achieves better performance than related methods in different scenarios.
Proposes a framework for semi-supervised continual learning from sequentially arriving data.
problem Learning from data with changing task distribution over time, especially in domains with a mix of labeled and unlabeled data.
method Meta-Consolidation for Continual Semi-Supervised Learning (MCSSL) framework with a hypernetwork and semi-supervised auxiliary classifier.
result Significant improvements in continual semi-supervised learning setting.
Proposes a CL technique to improve accuracy and reduce forgetting.
problem Sequential task learners struggle with forgetting information from previous tasks.
method Extracts modular parts of neural networks and estimates task relatedness.
result Remarkable performance gain in robustness to forgetting for EWC and GEM methods.
Develops a generic two-layer framework for adaptive ABMs.
problem Bi-level adaptation problem in ABMs: agents adapt to environment, and environment adapts to agents.
method Formalizes bi-level problem as a Stackelberg game with conditional policies, solving coupled non-linear equations.
result Unified framework for adaptive ABMs, addressing traditional ABM limitations.
Both the scientific community and the popular press have paid much attention to the speed of the Securities Information Processor, the data feed consolidating all trades and quotes across the US stock market. Rather than the speed of the Securities Information Processor, or SIP, we focus here on its accuracy. Relying o…
Under Solvency II the computation of capital requirements is based on value at risk (V@R). V@R is a quantile-based risk measure and neglects extreme risks in the tail. V@R belongs to the family of distortion risk measures. A serious deficiency of V@R is that firms can hide their total downside risk in corporate network…
New unsupervised learning framework for sound recognition.
problem Learning sound recognition without explicit labels.
method Combines self-supervised and clustering objectives with active learning.
result Achieves state-of-the-art unsupervised audio representation with reduced labels.
Paper benchmarks CF mitigation in federated time series forecasting.
problem Catastrophic forgetting in federated learning for time series forecasting.
method Comprehensive evaluation of CF mitigation strategies in federated time series forecasting.
result Introduction of a new benchmark for CF in time series federated learning.
This research proposes a CL model for RNNs to handle sequential data without forgetting.
problem Learning in dynamic environments without forgetting previous knowledge for sequential data.
method A Recurrent Neural Network (RNN) model with Elastic Weight Consolidation (EWC) for CL.
result The proposed model outperforms EWC and RNNs on CL benchmarks for sequential data.
New method improves ABI for sequential data, reducing forgetting and improving accuracy.
problem Performance degradation of ABI under model misspecification and distribution shifts.
method Decouples simulation-based pre-training from unsupervised SC fine-tuning, using memory buffer and elastic weight consolidation.
result Significant mitigation of forgetting and improved posterior estimates compared to standard simulation-based training.
Analyzes self-attention in recurrent networks, proving it mitigates vanishing gradients.
problem Vanishing gradients in recurrent networks when capturing long-term dependencies.
method Formal analysis of self-attention's effect on gradient propagation, proposing a relevancy screening mechanism.
result Self-attention mitigates vanishing gradients in recurrent networks, providing guarantees.
We discuss memory models which are based on tensor decompositions using latent representations of entities and events. We show how episodic memory and semantic memory can be realized and discuss how new memory traces can be generated from sensory input: Existing memories are the basis for perception and new memories ar…
Three scenarios for continual learning tasks are described and compared.
problem Difficulty in comparing continual learning methods due to different evaluation protocols.
method Three continual learning scenarios based on task identity and inference requirements.
result Different scenarios have varying difficulty and require different approaches.
This paper optimizes multi-currency AMMs to reduce forex trading costs.
problem Lack of direct liquid markets for currency pairs.
method Constant-mean AMM architecture, hierarchical agglomerative clustering algorithm.
result Optimized multi-currency pools reduce trading costs by ~13%.
EXPO framework eliminates need for reward model, achieving better optimization.
problem Optimizing large language model responses without a separate reward model.
method Introduces EXPO framework that avoids reparameterization, directly optimizing preferences.
result Demonstrates better regularization and intuitive interpolation behaviors.
A model retains learned knowledge for longer by adding a plastic component to neural networks.
problem Catastrophic forgetting in neural networks when learning new tasks.
method Differentiable Hebbian Consolidation model with a DHP Softmax layer.
result Reduces forgetting in benchmarks like Permuted MNIST and Vision Datasets Mixture.
A new memory replay mechanism improves reinforcement learning stability and speed.
problem Forgetting in reinforcement learning with continuous control.
method Augmented Memory Replay (AMR) that optimizes the replay of past experiences.
result AMR enhances stability and convergence speed of learning algorithms.
Researchers study solitons on homogeneous spaces, finding useful geometric structures.
problem Finding solitons on homogeneous spaces.
method Investigating solitons with a focus on equivalence relations and optimal tangent directions.
result Solitons on homogeneous spaces have been identified as a useful tool in various geometric contexts.
Paper proves uniqueness of a complex construction.
problem Proving uniqueness of a complex construction.
method Defined a total complex and proved uniqueness of terms.
result Total complex CTot(L) is unique. Optimal Control Theory optimizes neural networks, improving robustness and efficiency.
problem Optimizing deep neural networks (DNNs) for better performance and efficiency.
method Integrating Optimal Control Theory with Backpropagation to develop a new optimizer.
result Optimal Control Theoretic Neural Optimizer (OCNOpt) improves upon existing methods in robustness and efficiency.
Modified PCA algorithm with continual learning preserves features of previous modes for multimode process monitoring.
problem Catastrophic forgetting of previous modes in monitoring models for successive modes.
method Modified PCA algorithm with elastic weight consolidation (EWC) to preserve features of previous modes.
result PCA-EWC algorithm effectively monitors multimode processes without performance decrease.
Sequential learning of multiple tasks in artificial neural networks using gradient descent leads to catastrophic forgetting, whereby previously learned knowledge is erased during learning of new, disjoint knowledge. Here, we propose a new approach to sequential learning which leverages the recent discovery of adversari…
Deep RL multi-task learning outperforms single-task learning on new tasks.
problem Improving performance on new tasks in multi-task reinforcement learning.
method Investigation of multi-task reinforcement learning algorithms with and without Elastic Weight Consolidation (EWC).
result Multi-task reinforcement learning algorithms outperform single-task learning on new tasks.
We address the problem of classifying discrete differential-geometric Poisson brackets (dDGPBs) of any fixed order on target space of dimension 1. It is proved that these Poisson brackets (PBs) are in one-to-one correspondence with the intersection points of certain projective hypersurfaces. In addition, they can be re…
Unsupervised learning on imbalanced data is challenging because, when given imbalanced data, current model is often dominated by the major category and ignores the categories with small amount of data. We develop a latent variable model that can cope with imbalanced data by dividing the latent space into a shared space…