Secure aggregation for buffered asynchronous federated learning without TEEs.
problem Privacy and convergence in buffered asynchronous federated learning.
method Developed a new protocol (BASecAgg) that ensures privacy without TEEs by carefully designing masks.
result BASecAgg achieves similar convergence guarantees as FedBuff without TEEs.
A new buffer system improves continual learning in RL agents by adapting to changing environments.
problem Improving RL agents' ability to learn from changing environments over time.
method Multi-timescale replay buffer combined with invariant risk minimization.
result The method shows improvement over baselines in continual learning settings.
MER algorithm speeds up VI solving with Markovian data.
problem Solving stochastic variational inequalities with Markovian data.
method MER algorithm using multi-scale sampling from a Markovian buffer.
result Achieves faster convergence without knowing Markov chain mixing time.
Neural Episodic Control learns faster than other reinforcement learning agents.
problem Inefficient reinforcement learning methods requiring vast amounts of data.
method Uses a semi-tabular value function representation with a buffer of past experiences.
result Significantly faster learning across various environments.
LiDER refreshes past experiences in RL by dreaming about them.
problem Improving data efficiency in off-policy RL algorithms.
method Refreshing past experiences in a replay buffer using the current policy.
result LiDER consistently improves performance in Atari games.
A new model for predicting market order book dynamics using a buffer Hawkes process.
problem Predicting the evolution of limit order books in financial markets.
method Introducing a Markovian single point process with a buffer mechanism and self-exciting effect.
result The model accurately predicts market order book dynamics and converges to Brownian motion.
Experience replay helps neural networks learn new tasks without forgetting old knowledge.
problem Catastrophic forgetting in neural networks trained on non-stationary data.
method Experience replay buffers with a mixture of on- and off-policy learning.
result Experience replay can learn new tasks quickly and reduce catastrophic forgetting.
Paper introduces a dynamic reference frame strategy to predict events with a buffer time.
problem Lack of time buffer for predictions to enable timely action.
method Introduces a new concept of dynamic reference frame creation.
result Enables organizations to act on predictions with a buffer time.
Regulator allocates buffers to prevent financial contagion in networks with common assets.
problem Containment of default contagion in financial networks with common asset exposures.
method Allocates nonnegative buffer vectors under linear budget constraints to maximize default or insolvency resilience margins or minimize worst-case systemic losses.
result Exact synthesis results for buffer allocation under ℓ∞ and ℓ1 uncertainty sets, showing significant gains over uniform and exposure-proportional allocations. This work improves reinforcement learning with sparse rewards by following diverse past trajectories.
problem Challenges in reinforcement learning with sparse rewards and myopic behavior.
method Proposes a trajectory-conditioned policy to learn from a memory buffer of diverse past trajectories.
result Significantly outperforms existing methods on complex tasks with local optima.
Proposes a new policy gradient algorithm to improve reinforcement learning efficiency and stability.
problem Inefficiency and instability of DDPG in practical applications, and difficulty in controlling Q estimation bias and variance.
method Introduces a Regularly Updated Deterministic (RUD) policy gradient algorithm.
result The RUD algorithm makes better use of new data and has lower Q value variance, leading to improved performance.
Hierarchical GANs reduce anomaly detection costs.
problem Balancing anomaly detection accuracy and sampling costs.
method Hierarchical GANs for nonuniform sampling and buffer zones.
result Proposed GAN-based detector outperforms baseline in detection delay and average cost of error.
A distributed system identification method for LTI systems using reverse experience replay.
problem Online system identification of LTI systems over multi-agent networks.
method DSGD-RER, a distributed variant of SGD-RER with backward updates.
result The estimation error decreases as the network size grows.
ARCHER counters bias in HER to improve sample efficiency in RL.
problem Sample inefficiency in deep RL due to biased replay buffer experiences.
method ARCHER extends HER with aggressive hindsight rewards to counter bias.
result ARCHER increases sample efficiency in RL applications with limited computing budget.
MAPO uses a memory buffer to improve policy optimization in structured prediction tasks.
problem Improving sample efficiency and robustness in policy optimization for structured prediction tasks.
method Memory Augmented Policy Optimization (MAPO) uses a memory buffer to reduce policy gradient variance.
result MAPO achieves state-of-the-art results in program synthesis and semantic parsing tasks.
ReaPER improves learning efficiency by prioritizing reliable experiences.
problem Inefficient sampling of past experiences in reinforcement learning.
method Introducing a novel measure of reliability to prioritize experiences in PER.
result ReaPER outperforms PER in various environments, including Atari-10.
Proposes a new sampling method for deep Q-learning to improve efficiency and convergence.
problem Challenges in learning state-action value function from replay buffer.
method State distribution-aware sampling method to balance replay times for transitions.
result Reduces unnecessary TD updates and increases updates for uncertain state-action values.
This work improves continual learning by selecting diverse samples for replay buffers.
problem Overcoming catastrophic forgetting in online continual learning.
method Formulates sample selection as a constraint reduction problem and uses gradient-based diversity maximization.
result Demonstrates improved performance compared to existing methods that rely on task boundaries.
This study examines how risky investments affect insurance capital valuation.
problem Standard cost-of-capital assumptions do not account for risky investments.
method Analyzed effects of allowing buffer capital investments in risky assets.
result Decomposition of buffer capital contributions varies with riskiness.
BUZz defends images from adversarial attacks using simple transformations.
problem Adversarial attacks on deep neural networks for image classification.
method Combination of deep neural networks and simple image transformations.
result Achieves significant improvement over state-of-the-art defenses with a modest drop in clean accuracy.
New algorithm optimizes linear system estimation from single trajectory.
problem Estimating LTI systems from a single trajectory.
method SGD with Reverse Experience Replay (SGD−RER) result Optimal guarantees for parameter and prediction errors.
New method scales Bayesian inference for nonlinear SSMs using buffered stochastic gradient.
problem Inference for nonlinear, non-Gaussian SSMs is computationally challenging and particle degeneracy increases with longer series.
method Extends stochastic gradient MCMC to nonlinear SSMs using particle methods and error bounds.
result Demonstrates the importance of particle buffered stochastic gradient for long sequential data.
ETGL-DDPG improves DDPG for sparse reward control with new exploration and replay techniques.
problem Sparse reward continuous control in reinforcement learning.
method Introduces εt-greedy search and GDRB framework for efficient exploration and reward use. result ETGL-DDPG outperforms DDPG and other methods on sparse-reward continuous benchmarks.
Efficiently combines autoregressive and set-based models for joint distributions.
problem Joint distributions over multiple predictions from set-based models.
method Causal autoregressive buffer that caches context and captures dependencies.
result Up to 20x faster joint sampling and density evaluation, up to 7x lower memory usage.
We propose a streaming submodular maximization algorithm "stream clipper" that performs as well as the offline greedy algorithm on document/video summarization in practice. It adds elements from a stream either to a solution set S or to an extra buffer B based on two adaptive thresholds, and improves S by a final…
The paper proposes a method to learn from both simulation and real-world data.
problem Training autonomous systems in simulation and applying them to real-world environments.
method Balancing samples from simulation and real-world data using a replay buffer.
result The method achieves better performance in real-world tasks compared to training only in simulation.
We present a new algorithm that significantly improves the efficiency of exploration for deep Q-learning agents in dialogue systems. Our agents explore via Thompson sampling, drawing Monte Carlo samples from a Bayes-by-Backprop neural network. Our algorithm learns much faster than common exploration strategies such as …
A novel RL objective and prioritization framework improve performance and sample-efficiency in multi-goal tasks.
problem Learning diverse goals in multi-goal reinforcement learning.
method Maximum entropy regularization for objective and prioritization framework.
result Promising improvements in performance and sample-efficiency on multi-goal robotic tasks.
The paper investigates the effectiveness of reusing experience in Deep Q-Learning for FPS environments.
problem The high number of interactions required for reinforcement learning limits its practicality.
method The authors test the effectiveness of applying learning update steps multiple times per environmental step in the VizDoom environment.
result Updating learning steps less frequently than every 4th environmental step does not improve performance and can degrade performance.
A mean-reverting financial instrument is optimally traded by buying it when it is sufficiently below the estimated `mean level' and selling it when it is above. In the presence of linear transaction costs, a large amount of value is paid away crossing bid-offers unless one devises a `buffer' through which the price mus…
Study compares AI and static analysis for detecting buffer overflows.
problem Detecting buffer overflows in code.
method Developed s-bAbI to generate code samples, compared AI with static analysis tools.
result AI system requires extensive training data to match static analysis precision and recall.
Combines policy gradient and Q-learning for improved data efficiency and stability.
problem Improving data efficiency and stability in reinforcement learning.
method Combines policy gradient with off-policy Q-learning using a replay buffer and fixed points of policy gradient.
result Achieved performance exceeding A3C and Q-learning on Atari games.
Investigates multi-period portfolio optimization for DC plans using buffered Probability of Exceedance.
problem Optimizing long-term Defined Contribution plans with realistic constraints and dynamic dynamics.
method Formulates and solves bilevel optimization problems for pre-commitment and time-consistent Mean-bPoE and Mean-CVaR portfolio optimization.
result Time-consistent Mean-bPoE strategies maintain investor preferences for minimum terminal wealth, unlike Mean-CVaR.
The paper improves reinforcement learning stability and efficiency with a new theoretical framework.
problem Stability and efficiency in reinforcement learning, especially in data-scarce scenarios.
method Theoretical framework using resampled U- and V-statistics to model experience replay, applied to policy evaluation and kernel ridge regression. result Significant improvements in stability and efficiency, particularly in data-scarce scenarios.
Improves on-policy RL by reusing data from multiple policies.
problem Lack of reuse of data from previous policies in on-policy RL.
method Adapts replay buffer concept to combine on- and off-policy learning.
result Method outperforms state-of-the-art on-policy RL algorithms.
DAC enhances exploration in reinforcement learning with entropy regularization.
problem Improving exploration efficiency in reinforcement learning.
method Sample-aware entropy regularization using replay buffer action distributions.
result DAC significantly outperforms existing algorithms in reinforcement learning tasks.
The paper studies how memory replay affects reinforcement learning performance.
problem Understanding the effects of memory size and prioritization in experience replay.
method Formulated a dynamical systems ODE model of Q-learning with experience replay, derived analytic solutions, and proposed an adaptive memory buffer size algorithm.
result The amount of memory kept and prioritization significantly impact learning dynamics; too much or too little memory slows down learning.
CDP improves RL performance and sample-efficiency by prioritizing rare goal states.
problem Learning from imbalanced data in RL.
method Curiosity-Driven Prioritization (CDP) framework.
result CDP improves both performance and sample-efficiency of RL agents.
A simple technique improves continual learning by 50% on image datasets.
problem Challenges in training neural networks on a stream of shifting data.
method Experience Replay (ER) with five tricks to mitigate its shortcomings.
result ER, enhanced with tricks, achieves significant accuracy gains.
Flexible VHDL design for multiple neural networks on FPGAs.
problem Inflexible neural network designs for FPGAs.
method Proposes a flexible VHDL structure with multiple processor groups.
result Allows training and testing of multiple neural networks on multiple FPGAs.
New method improves continual learning by anchoring past knowledge.
problem Catastrophic forgetting in continual learning.
method Bilevel optimization to update current task knowledge while keeping past task predictions.
result Improves accuracy and forgetting metrics compared to experience replay.
SOCP uses SOM to find groups and local calibration buffers for better regional coverage.
problem Heterogeneous regional coverage gaps in conformal prediction.
method Self-Organizing Map (SOM) for group discovery; local calibration buffers at BMU or fixed grid.
result Reduces regional coverage gaps on 7/8 benchmarks by 7.1%.
Investigates optimal pension policies in PAYG systems with forward utility and ageing population.
problem Optimal investment and pension policies in PAYG systems with sustainability and adequacy constraints.
method Non-zero volatility forward CRRA utilities, closed-form optimal policies, detailed numerical analysis.
result Characterization of optimal policies and detailed impact analysis under various scenarios.
Proportional transaction costs present difficult theoretical problems in trading algorithm design, on account of their lack of analytical tractability. The author derives a solution of DT-NT-DT form for an arbitrary model in which the the traded asset has diffusive dynamics described by one or more stochastic risk fact…
Novel asynchronous SGD method resists Byzantine attacks without server storage.
problem Asynchronous distributed learning with Byzantine attacks and failures.
method Buffered Asynchronous SGD (BASGD) and its momentum variant (BASGDm).
result BASGD and BASGDm resist non-omniscient and omniscient attacks without server storage.
Tiled Squeeze-and-Excite improves channel attention with local spatial context.
problem Improving channel attention mechanisms in neural networks.
method Proposes tiled squeeze-and-excite (TSE) framework for channel attention.
result Local context of 7 rows or columns is sufficient for matching global context performance.
Study finds AI-generated financial advice influences life cycle investing patterns.
problem Understanding how AI-generated financial advice impacts life cycle investing.
method Sentiment analysis of prompts from AI-generated financial advice and simulation of lifetime effects.
result AI-generated financial advice leads to life cycle investing patterns, influenced by gender and AI experience.
Improved diffusion models for sampling from given distributions.
problem Training diffusion models to sample from a given distribution.
method Benchmarked and improved off-policy methods for diffusion sampling.
result A novel exploration strategy improves sample quality.