Deep RL for safe, multi-agent driving policies.
problem Safe negotiation with other road users in autonomous driving.
method Policy gradient iterations, decomposed into desires and constraints, hierarchical temporal abstraction.
result Significant reduction in gradient variance for safer driving policies.
Deep learning agents negotiate contracts with prosocial or selfish behaviors.
problem Training agents to negotiate contracts with varying behaviors.
method Multi-Agent Reinforcement Learning, modeling prosocial and selfish behaviors, training a meta agent.
result Trained agents hold their own against human players and emulate human behavior.
MOANOFS tackles online feature selection for big data classification.
problem Online supervised feature selection for binary classification in big data.
method Hybrid of online learning and automated negotiation.
result MOANOFS achieves high accuracy with real-world applications.
Paper automates car negotiation in intersections using Q-learning.
problem Automated vehicles negotiate with human-driven cars in intersections.
method Deep Q-learning applied to simulated traffic with various driver behaviors.
result 98% success rate in avoiding collisions with other vehicles.
A framework for fair derivative contract pricing and risk-sharing between parties with funding differences.
problem Price asymmetry due to funding differences in bilateral contracts.
method Defines a negotiation problem that maximizes the sum of utilities for two parties, deriving optimal prices and collateral.
result Optimal negotiation price and collateral can be used to interpret margin requirements.
We consider two risk-averse financial agents who negotiate the price of an illiquid indivisible contingent claim in an incomplete semimartingale market environment. Under the assumption that the agents are exponential utility maximizers with non-traded random endowments, we provide necessary and sufficient conditions f…
The paper analyzes a game where players must balance short-term and long-term interests, leading to cooperative or competitive outcomes.
problem Analyzing time inconsistency in inter-personal decision-making under non-exponential discounting.
method Iterative procedures and Zorn's lemma to find Nash equilibria between players' intra-personal equilibria.
result Inter-personal equilibria exist and depend on the impatience levels of the players.
IDAS approach for autonomous vehicles to make decisions under merging scenarios.
problem Decision making for autonomous vehicles in merging scenarios with varying driver cooperativeness.
method IDAS approach using multi-agent reinforcement learning (MARL) with curriculum learning and masking mechanism.
result IDAS approach can handle uncertainties in real-world scenarios and make strategic decisions.
We describe an agent-based simulation of a fictional (but feasible) information trading business. The Gas Price Information Trader (GPIT) buys information about real-time gas prices in a metropolitan area from drivers and resells the information to drivers who need to refuel their vehicles. Our simulation uses real wor…
Safe linear stochastic bandits ensure safe exploration with optimal regret.
problem Ensuring safe exploration in stochastic bandits while minimizing regret.
method Combining known safe arms with exploratory arms to safely expand the set of safe arms over time.
result The algorithm achieves an expected regret of O ( T log ( T ) ) O(\sqrt{T}\log (T)) O ( T log ( T )) . A rapid pattern-recognition approach to characterize driver's curve-negotiating behavior is proposed. To shorten the recognition time and improve the recognition of driving styles, a k-means clustering-based support vector machine ( kMC-SVM) method is developed and used for classifying drivers into two types: aggressiv…
New approach to reinforcement learning that balances safety and performance against adversaries.
problem Balancing safety and performance in reinforcement learning against potential adversaries.
method Developed a new reinforcement learning framework that integrates interruptibility, resilience, and safe exploration.
result Achieved both interruptibility and resilience to adversaries without sacrificing optimal policy probability.
CoCoRL learns safe constraints from demonstrations with unknown rewards.
problem Learning safe constraints from demonstrations with different unknown rewards.
method Convex Constraint Learning for Reinforcement Learning (CoCoRL) constructs a convex safe set based on demonstrations.
result CoCoRL learns constraints that lead to safe driving behavior and can safely transfer to different tasks and environments.
SafePILCO is a Python tool for safe reinforcement learning.
problem Safe and efficient policy synthesis in reinforcement learning.
method Extends PILCO algorithm with safety features, implemented in Python.
result Safe and data-efficient policy synthesis achieved.
Robust regression model for safe exploration in control problems.
problem Learning and exploring safely in sequential control problems.
method Deep robust regression model trained to predict uncertainty bounds.
result Empirically outperforms conventional GP-based safe exploration.
Safe screening reduces the number of triplets in metric learning.
problem Optimizing a metric over many triplets is computationally expensive and impractical.
method Safe triplet screening identifies and removes redundant triplets.
result Safe triplet screening maintains optimality without increasing computational cost.
Safe learning of stochastic dynamics with safety constraints.
problem Learning controlled stochastic dynamics with safety constraints.
method Iterative expansion of a safe control set using kernel-based confidence bounds.
result The method ensures safe exploration and efficient estimation of system dynamics.
High dimensional regression benefits from sparsity promoting regularizations. Screening rules leverage the known sparsity of the solution by ignoring some variables in the optimization, hence speeding up solvers. When the procedure is proven not to discard features wrongly the rules are said to be \emph{safe}. In this …
Study on neural networks to identify redundancy issues in safe machine learning.
problem Identifying redundancy in neural network architectures for safe machine learning.
method Experiments with MNIST database using neural network classifiers.
result Underlines difficulties in using neural network classifiers for safe systems.
This research improves debt collection strategies using advanced machine learning.
problem Accurate estimation of propensity to pay and cashflow for optimal debt collection.
method Developed a machine learning framework with pre-processing and model selection.
result The proposed model outperforms current industry strategies.
Bitcoin fails to prove safe haven status during pandemic.
problem Determining if Bitcoin is a reliable safe haven asset during crises.
method Quantile correlations of Bitcoin with S&P500, VIX, and gold.
result Gold is a better safe haven during crises, not Bitcoin.
Proposes a method to accelerate safe sequential learning using offline data.
problem Limited exploration due to disconnected safe regions and slow task learning.
method Safe transfer sequential learning using Gaussian processes and offline data.
result Enhances global exploration across multiple disjoint safe regions with lower data consumption.
Meta-learning priors improves safe Bayesian optimization.
problem Optimizing robot controllers under safety constraints.
method Meta-learning priors from offline data using F-PACOH.
result Meta-learned priors accelerate safe BO convergence.
Revel tackles safe exploration in RL with verified symbolic policies.
problem Computational infeasibility of verifying neural networks in RL learning loops.
method Two policy classes: neurosymbolic with approximate gradients and symbolic policies for efficient verification. Mirror descent over policies to safely update and project policies.
result Revel discovers policies that outperform prior approaches to verified exploration.
Safe exploration framework for IML algorithms.
problem Safe decision-making in IML without unsafe outcomes.
method Exploits Gaussian process prior to efficiently learn safe decisions.
result Outperforms other algorithms empirically.
Bitcoin became a strong safe haven after Trump's win, but its status varies over time.
problem Determining Bitcoin's role as a safe haven during market turmoil.
method Ensemble Empirical Mode Decomposition-based approach to analyze correlations.
result Bitcoin's safe-haven property is time-varying, being a weak safe haven in the short term and long term.
ASE safely explores unknown MDPs with unknown dynamics, improving sample efficiency.
problem Balancing exploration and safety in unknown MDPs with stochastic dynamics.
method Exploits analogies between state-action pairs to safely learn near-optimal policies.
result Empirically improves sample efficiency compared to existing methods.
Framework for safe reinforcement learning using expert demonstrations.
problem Ensuring safe behavior in reinforcement learning with unknown reward functions.
method Combines expert demonstrations with optimization methods to find safe reward functions.
result Trained agent safely avoids harmful states while mimicking expert behavior.
Safe reinforcement learning for autonomous vehicles using prediction constraints.
problem Safe reinforcement learning for safety-critical applications like autonomous vehicles.
method Use prediction to constrain exploration in reinforcement learning models.
result Successfully learned intersection handling behaviors on an autonomous vehicle.
Safe-House secures DeFi by limiting losses and enhancing security.
problem Ongoing hacks and security concerns in DeFi.
method Safe-House uses blockchain principles to secure asset movements.
result Safe-House limits maximum one-time loss to specified limits.
Safe policy optimization using Gaussian process models.
problem Optimizing safe policies for task performance.
method Training a Gaussian process model to capture dynamics, ensuring safe policies only.
result Closed-form computation of error gradients and constraint violation probability.
StageOpt efficiently optimizes safe decisions by separating safety and utility stages.
problem Optimizing unknown utility with safety constraints in sequential decisions.
method Develops StageOpt, a two-stage safe Bayesian optimization algorithm.
result StageOpt is more efficient and applicable to broader problems than existing methods.
Optimal margin loan agreements for sophisticated gamblers and brokers.
problem Finding fair interest rates and loan sizes between gamblers and brokers.
method Derives formulas for optimal arrangements based on gamblers' risk preferences and market conditions.
result Gambler gains higher capital growth with lower interest rates, broker gains intermediary profit.
OSIL learns safe policies from unsafe demonstrations.
problem Offline safe imitation learning with implicit safety.
method Formulates CMDP, infers safety from non-preferred trajectories, learns cost model.
result OSIL learns safer policies without degrading reward performance.
The problem of learning a sparse model is conceptually interpreted as the process of identifying active features/samples and then optimizing the model over them. Recently introduced safe screening allows us to identify a part of non-active features/samples. So far, safe screening has been individually studied either fo…
A new screening rule 'dynamic Sasvi' improves sparse optimization speed.
problem Sparse optimization problem identification.
method Flexible framework based on Fenchel-Rockafellar duality for norm-regularized least squares.
result Dynamic Sasvi can eliminate more features and increase solver speed.
In an incomplete semimartingale model of a financial market, we consider several risk-averse financial agents who negotiate the price of a bundle of contingent claims. Assuming that the agents' risk preferences are modelled by convex capital requirements, we define and analyze their demand functions and propose a notio…
Safe screening rules improve variable selection speed in high-dimensional regression.
problem Efficiently selecting important variables in high-dimensional regression problems.
method Developing Gap Safe screening rules for generalized linear models with sparsity enforcing penalties.
result Significant speed-ups in variable selection compared to previous methods on various learning tasks.
This paper tackles safe global optimization of noisy functions with a Lipschitz condition.
problem Safe global maximization of expensive, noisy, Lipschitz functions.
method Develops a δ-Lipschitz framework and two algorithms to ensure safety constraints are met.
result The proposed methods ensure safety constraints are met before evaluating noisy functions.
New algorithm for safe bandits with unknown parameters and constraints.
problem Safety-critical systems with unknown parameters and constraints.
method Safe-LUCB algorithm with two phases: pure exploration and safe exploration-exploitation.
result General and problem-dependent regret bounds for the Safe-LUCB algorithm.
SRF learns sparse rule models by screening out features efficiently.
problem Learning optimal sparse rule models is computationally intractable due to the large number of possible rules.
method SRF uses meta safe screening (mSS) to efficiently screen out multiple features, improving the learning of sparse rule models.
result SRF provides a general framework for fitting sparse rule models and can handle group regularization.
Safe autonomous decisions made with machine learning predictions using Conformal Decision Theory.
problem Safe decisions from imperfect machine learning predictions.
method Conformal Decision Theory framework for producing safe decisions.
result Safe decisions with provable statistical guarantees of low risk.
Crypto-assets perform better than gold as safe-havens during market crashes.
problem Evaluating safe-haven properties of crypto-assets and gold during the 2020 market crash.
method Comparative analysis of Crypto-assets (Tether, Cardano, Dogecoin, Bitcoin, Ethereum, Litecoin, Ripple) and gold for European indices.
result Tether, Cardano, and Dogecoin exhibited hedging properties similar to gold, while gold was not more efficient as a safe-haven.
Safe Bayesian optimization method using information theory.
problem Optimizing unknown functions while respecting safety constraints.
method Information-theoretic exploration criterion for continuous domains.
result The method learns the value of the safe optimum up to arbitrary precision.
Safe sample screening improves RSVM performance without sacrificing accuracy.
problem Improving RSVM performance under noisy conditions.
method Proposed two safe sample screening rules based on CCCP framework for RSVM.
result Significant reduction in computational time for RSVM.
SafeRNet uses IoT and cloud computing to provide real-time safe routes.
problem High traffic fatality rates despite advanced technology.
method Bayesian network for safe route modeling.
result Demonstrated effectiveness with real traffic data.
Safe screening rules reduce computation time in logistic regression with ℓ 0 − ℓ 2 \ell_0-\ell_2 ℓ 0 − ℓ 2 regularization.
problem Efficiently solving logistic regression with many features and regularization.
method Screening rules based on Fenchel dual lower bounds of strong conic relaxations.
result A high percentage of features can be safely removed before solving, leading to substantial speed-up.
Safe-M 3 ^3 3 -UCRL learns safe policies for multi-agent systems with global constraints.
problem Global constraints in mean-field reinforcement learning for multi-agent systems.
method Safe-M 3 ^3 3 -UCRL uses epistemic uncertainty and log-barrier approach to ensure constraints satisfaction. result Safe-M 3 ^3 3 -UCRL learns safe policies for multi-agent systems with global constraints.