This paper tackles safe global optimization of noisy functions with a Lipschitz condition.
problem Safe global maximization of expensive, noisy, Lipschitz functions.
method Develops a δ-Lipschitz framework and two algorithms to ensure safety constraints are met.
result The proposed methods ensure safety constraints are met before evaluating noisy functions.
Proposes a method to accelerate safe sequential learning using offline data.
problem Limited exploration due to disconnected safe regions and slow task learning.
method Safe transfer sequential learning using Gaussian processes and offline data.
result Enhances global exploration across multiple disjoint safe regions with lower data consumption.
This paper introduces dynamic safe interruptibility for multi-agent reinforcement learning.
problem Preventing dangerous situations in decentralized multi-agent reinforcement learning.
method Introduces dynamic safe interruptibility, studies it in two learning frameworks, and gives sufficient conditions for its implementation.
result Dynamic safe interruptibility can be enabled for joint action learners but not for independent learners.
Safe screening improves generalized CGM's feature selection stability.
problem Improving feature selection stability in generalized CGM.
method Coupling safe screening with generalized CGM.
result Safe screening matches solution support at rate O ( 1 / ( t δ 2 ) ) O(1/(tδ^2)) O ( 1/ ( t δ 2 )) . SAMBA improves safe reinforcement learning with active exploration metrics.
problem Safe reinforcement learning in dynamic systems.
method Combines probabilistic modelling, information theory, and statistics. Uses novel metrics for out-of-sample Gaussian process evaluation.
result Orders of magnitude reduction in samples and violations compared to state-of-the-art methods.
Study markets without safe assets, deriving option pricing equations.
problem No riskless asset in financial markets.
method Derive Black-Scholes-Merton equations for various risky asset dynamics.
result Option pricing equations for different risky asset dynamics.
Optimal and safe semi-supervised learning estimator for high-dimensional data.
problem Improving regression parameter estimation with unlabeled data in high-dimensional settings.
method Established minimax lower bound, proposed optimal and safe semi-supervised estimators.
result Optimal semi-supervised estimator achieves the minimax lower bound.
A novel approach for safe offline RL using latent safety constraints.
problem Balancing safety constraints and reward maximization in offline RL.
method Conditional Variational Autoencoders for latent safety modeling, Constrained Reward-Return Maximization.
result Our approach maintains safety compliance while optimizing rewards, outperforming existing methods.
Safe exploration method for RL under disturbance ensures safety with probabilistic guarantees.
problem Safe reinforcement learning in real environments with disturbance.
method Uses partial prior knowledge and conservative inputs to ensure state constraint satisfaction.
result Guaranteed safety with pre-specified probability in the presence of stochastic disturbance.
A new screening test for Lasso improves solution speed.
problem Improving the efficiency of Lasso solution methods.
method Safe region with dome geometry based on dual cutting half-spaces.
result The new screening test leads to significant acceleration.
Safe learning for optimal control with known and unknown dynamics.
problem Safe learning of control strategies for systems with unknown dynamics.
method Reachability analysis and Gaussian Process regression for updating disturbances.
result Algorithm learns optimal control policies without violating safety constraints.
Sparse learning techniques have been routinely used for feature selection as the resulting model usually has a small number of non-zero entries. Safe screening, which eliminates the features that are guaranteed to have zero coefficients for a certain value of the regularization parameter, is a technique for improving t…
In classical reinforcement learning, when exploring an environment, agents accept arbitrary short term loss for long term gain. This is infeasible for safety critical applications, such as robotics, where even a single unsafe action may cause system failure. In this paper, we address the problem of safely exploring fin…
The paper tackles safe reinforcement learning with convex regularization.
problem Safe reinforcement learning in complex, high-dimensional settings with safety constraints.
method Doubly-regularized RL framework combining reward and parameter regularization, formulated as a convex regularized objective with parametrized policies on an infinite-dimensional statistical manifold.
result Exponential convergence guarantees under sufficient regularization, robust theoretical insights and guarantees for safe RL.
Safe linear stochastic bandits ensure safe exploration with optimal regret.
problem Ensuring safe exploration in stochastic bandits while minimizing regret.
method Combining known safe arms with exploratory arms to safely expand the set of safe arms over time.
result The algorithm achieves an expected regret of O ( T log ( T ) ) O(\sqrt{T}\log (T)) O ( T log ( T )) . Safe RL for autonomous vehicles using PCPO with trust regions and parallel learners.
problem Unexplainable behaviours and lack of safety guarantees in RL for real vehicles.
method PCPO framework with trust regions and parallel learners.
result Safe learning confirmed for autonomous vehicles with fast convergence.
New approach to reinforcement learning that balances safety and performance against adversaries.
problem Balancing safety and performance in reinforcement learning against potential adversaries.
method Developed a new reinforcement learning framework that integrates interruptibility, resilience, and safe exploration.
result Achieved both interruptibility and resilience to adversaries without sacrificing optimal policy probability.
CoCoRL learns safe constraints from demonstrations with unknown rewards.
problem Learning safe constraints from demonstrations with different unknown rewards.
method Convex Constraint Learning for Reinforcement Learning (CoCoRL) constructs a convex safe set based on demonstrations.
result CoCoRL learns constraints that lead to safe driving behavior and can safely transfer to different tasks and environments.
We consider rules for discarding predictors in lasso regression and related problems, for computational efficiency. El Ghaoui et al (2010) propose "SAFE" rules that guarantee that a coefficient will be zero in the solution, based on the inner products of each predictor with the outcome. In this paper we propose strong …
Improves policies with high certainty, even in small samples.
problem Ensuring new policies are better than the baseline with high probability.
method Leverages powerful safety tests and multiple testing for threshold policies.
result Controls the rate of adopting a worse policy to pre-specified error level.
SafePILCO is a Python tool for safe reinforcement learning.
problem Safe and efficient policy synthesis in reinforcement learning.
method Extends PILCO algorithm with safety features, implemented in Python.
result Safe and data-efficient policy synthesis achieved.
Robust regression model for safe exploration in control problems.
problem Learning and exploring safely in sequential control problems.
method Deep robust regression model trained to predict uncertainty bounds.
result Empirically outperforms conventional GP-based safe exploration.
Safe screening reduces the number of triplets in metric learning.
problem Optimizing a metric over many triplets is computationally expensive and impractical.
method Safe triplet screening identifies and removes redundant triplets.
result Safe triplet screening maintains optimality without increasing computational cost.
Safe learning of stochastic dynamics with safety constraints.
problem Learning controlled stochastic dynamics with safety constraints.
method Iterative expansion of a safe control set using kernel-based confidence bounds.
result The method ensures safe exploration and efficient estimation of system dynamics.
High dimensional regression benefits from sparsity promoting regularizations. Screening rules leverage the known sparsity of the solution by ignoring some variables in the optimization, hence speeding up solvers. When the procedure is proven not to discard features wrongly the rules are said to be \emph{safe}. In this …
Study on neural networks to identify redundancy issues in safe machine learning.
problem Identifying redundancy in neural network architectures for safe machine learning.
method Experiments with MNIST database using neural network classifiers.
result Underlines difficulties in using neural network classifiers for safe systems.
Proposes a new Q-learning method to improve sample complexity.
problem Improving sample complexity in Q-learning with limited data.
method Integrates Q-function from a source task into a target task under safe conditions.
result The method converges faster than standard Q-learning under certain conditions.
Paper derives uniform error bounds for Gaussian process regression for safer control applications.
problem Quantifying model error in Gaussian process regression for safety-critical applications.
method Employing Gaussian process distribution and continuity arguments, derive uniform error bounds under weaker assumptions.
result Derives novel uniform error bounds for Gaussian process regression under weaker assumptions.
Bitcoin fails to prove safe haven status during pandemic.
problem Determining if Bitcoin is a reliable safe haven asset during crises.
method Quantile correlations of Bitcoin with S&P500, VIX, and gold.
result Gold is a better safe haven during crises, not Bitcoin.
Meta-learning priors improves safe Bayesian optimization.
problem Optimizing robot controllers under safety constraints.
method Meta-learning priors from offline data using F-PACOH.
result Meta-learned priors accelerate safe BO convergence.
Revel tackles safe exploration in RL with verified symbolic policies.
problem Computational infeasibility of verifying neural networks in RL learning loops.
method Two policy classes: neurosymbolic with approximate gradients and symbolic policies for efficient verification. Mirror descent over policies to safely update and project policies.
result Revel discovers policies that outperform prior approaches to verified exploration.
Safe exploration framework for IML algorithms.
problem Safe decision-making in IML without unsafe outcomes.
method Exploits Gaussian process prior to efficiently learn safe decisions.
result Outperforms other algorithms empirically.
Bitcoin became a strong safe haven after Trump's win, but its status varies over time.
problem Determining Bitcoin's role as a safe haven during market turmoil.
method Ensemble Empirical Mode Decomposition-based approach to analyze correlations.
result Bitcoin's safe-haven property is time-varying, being a weak safe haven in the short term and long term.
The study reveals gold's effectiveness as a hedge and safe haven varies with uncertainty levels.
problem Gold's role as a hedge and safe haven is not constant and depends on uncertainty levels.
method Quantile-on-quantile regression and dynamic factor model to analyze gold returns and uncertainty.
result Gold returns positively and strongly with high uncertainty, suggesting it can be a protective asset.
ASE safely explores unknown MDPs with unknown dynamics, improving sample efficiency.
problem Balancing exploration and safety in unknown MDPs with stochastic dynamics.
method Exploits analogies between state-action pairs to safely learn near-optimal policies.
result Empirically improves sample efficiency compared to existing methods.
Framework for safe reinforcement learning using expert demonstrations.
problem Ensuring safe behavior in reinforcement learning with unknown reward functions.
method Combines expert demonstrations with optimization methods to find safe reward functions.
result Trained agent safely avoids harmful states while mimicking expert behavior.
Safe reinforcement learning for autonomous vehicles using prediction constraints.
problem Safe reinforcement learning for safety-critical applications like autonomous vehicles.
method Use prediction to constrain exploration in reinforcement learning models.
result Successfully learned intersection handling behaviors on an autonomous vehicle.
Safe-House secures DeFi by limiting losses and enhancing security.
problem Ongoing hacks and security concerns in DeFi.
method Safe-House uses blockchain principles to secure asset movements.
result Safe-House limits maximum one-time loss to specified limits.
Safe policy optimization using Gaussian process models.
problem Optimizing safe policies for task performance.
method Training a Gaussian process model to capture dynamics, ensuring safe policies only.
result Closed-form computation of error gradients and constraint violation probability.
StageOpt efficiently optimizes safe decisions by separating safety and utility stages.
problem Optimizing unknown utility with safety constraints in sequential decisions.
method Develops StageOpt, a two-stage safe Bayesian optimization algorithm.
result StageOpt is more efficient and applicable to broader problems than existing methods.
OSIL learns safe policies from unsafe demonstrations.
problem Offline safe imitation learning with implicit safety.
method Formulates CMDP, infers safety from non-preferred trajectories, learns cost model.
result OSIL learns safer policies without degrading reward performance.
The problem of learning a sparse model is conceptually interpreted as the process of identifying active features/samples and then optimizing the model over them. Recently introduced safe screening allows us to identify a part of non-active features/samples. So far, safe screening has been individually studied either fo…
A new screening rule 'dynamic Sasvi' improves sparse optimization speed.
problem Sparse optimization problem identification.
method Flexible framework based on Fenchel-Rockafellar duality for norm-regularized least squares.
result Dynamic Sasvi can eliminate more features and increase solver speed.
Safe screening rules improve variable selection speed in high-dimensional regression.
problem Efficiently selecting important variables in high-dimensional regression problems.
method Developing Gap Safe screening rules for generalized linear models with sparsity enforcing penalties.
result Significant speed-ups in variable selection compared to previous methods on various learning tasks.
New algorithm for safe bandits with unknown parameters and constraints.
problem Safety-critical systems with unknown parameters and constraints.
method Safe-LUCB algorithm with two phases: pure exploration and safe exploration-exploitation.
result General and problem-dependent regret bounds for the Safe-LUCB algorithm.
SRF learns sparse rule models by screening out features efficiently.
problem Learning optimal sparse rule models is computationally intractable due to the large number of possible rules.
method SRF uses meta safe screening (mSS) to efficiently screen out multiple features, improving the learning of sparse rule models.
result SRF provides a general framework for fitting sparse rule models and can handle group regularization.
Safe autonomous decisions made with machine learning predictions using Conformal Decision Theory.
problem Safe decisions from imperfect machine learning predictions.
method Conformal Decision Theory framework for producing safe decisions.
result Safe decisions with provable statistical guarantees of low risk.
Crypto-assets perform better than gold as safe-havens during market crashes.
problem Evaluating safe-haven properties of crypto-assets and gold during the 2020 market crash.
method Comparative analysis of Crypto-assets (Tether, Cardano, Dogecoin, Bitcoin, Ethereum, Litecoin, Ripple) and gold for European indices.
result Tether, Cardano, and Dogecoin exhibited hedging properties similar to gold, while gold was not more efficient as a safe-haven.