The notion of developing statistical methods in machine learning which are robust to adversarial perturbations in the underlying data has been the subject of increasing interest in recent years. A common feature of this work is that the adversarial robustification often corresponds exactly to regularization methods whi…
The paper connects ABC to GBI, suggesting ABC as a robustification strategy.
problem Approximate Bayesian Computation struggles with tractability in complex simulators.
method Reinterpreting ABC as an implicitly defined error model and suggesting GBI.
result ABC can be seen as a robustification strategy for approximating Bayesian posteriors.
In this paper, we establish a robustification of an on-line algorithm for modelling asset prices within a hidden Markov model (HMM). In this HMM framework, parameters of the model are guided by a Markov chain in discrete time, parameters of the asset returns are therefore able to switch between different regimes. The p…
This work explores robust multi-objective optimisation with scalarisation and robustification.
problem Optimizing functions with uncertainty and varying objectives.
method Identifies the importance of robustification and scalarisation in multi-objective optimisation.
result Different orders of scalarisation and robustification lead to different solutions.
Robust biclustering method tackles heavy-tailed data issues.
problem Discovering local correlation in heavy-tailed data.
method Convex biclustering with Huber loss and tuning-free parameter selection.
result Outperforms traditional biclustering methods in heavy-tailed noise.
Unified framework for robust risk measures beyond convexity.
problem Developing risk measures for uncertainty beyond classical convexity.
method Constructing robust quasi-convex measures through uncertainty sets.
result Unified framework for robust quasi-convex risk measures.
Adapting robust statistics to neural networks, researchers found neural networks can be more robust with certain loss functions.
problem The robustness of neural networks in complex learning tasks.
method Adapting the regression breakdown point from robust statistics to neural networks and comparing different configurations and contamination settings.
result Neural networks can benefit from robust loss functions, as demonstrated in extensive simulations.
New method designs experiments robustly for nonlinear estimation, improving parameter knowledge.
problem Designing robust experiments for nonlinear estimation under parametric uncertainty.
method Multi-stage robust optimization framework for sequential experiments.
result Identifies experiments better conducted early for improved parameter knowledge.
RLGP model improves robustness and accuracy for discontinuous response surfaces.
problem Challenges in modeling abrupt jumps and discontinuities in response surfaces.
method Integrates adaptive nearest-neighbor selection with robustification mechanism.
result Consistently delivers high predictive accuracy and robustness in higher dimensions.
Label smoothing improves model robustness against misspecification.
problem Improving model robustness against model misspecification.
method Introducing modified label smoothing (MLSLR) that maintains consistent probability estimation while modifying the loss function.
result MLSLR exhibits higher robustness against model misspecification than conventional label smoothing.
Mechanism design for large item sets using topic models.
problem Designing optimal mechanisms for large item sets when exact priors are unknown or intractable.
method Proposes a framework that disentangles statistical estimation and mechanism design, leveraging topic models to reduce complexity.
result Reduces the complexity of mechanism design for large item sets, making it feasible with topic models.
It is well known that Principal Component Analysis (PCA) is strongly affected by outliers and a lot of effort has been put into robustification of PCA. In this paper we present a new algorithm for robust PCA minimizing the trimmed reconstruction error. By directly minimizing over the Stiefel manifold, we avoid deflatio…
Paper presents a GAN-based method to automate robust hedging.
problem Addressing uncertainty in market data generating process.
method Adversarial approach inspired by GANs for automated robustification.
result Automated robustification of hedging objective through three modular components.
We propose a general framework for increasing local stability of Artificial Neural Nets (ANNs) using Robust Optimization (RO). We achieve this through an alternating minimization-maximization procedure, in which the loss of the network is minimized over perturbed examples that are generated at each parameter update. We…
Data-driven Distributionally Robust Optimization (DD-DRO) via optimal transport has been shown to encompass a wide range of popular machine learning algorithms. The distributional uncertainty size is often shown to correspond to the regularization parameter. The type of regularization (e.g. the norm used to regularize)…
We study statistical inference and distributionally robust solution methods for stochastic optimization problems, focusing on confidence intervals for optimal values and solutions that achieve exact coverage asymptotically. We develop a generalized empirical likelihood framework---based on distributional uncertainty se…
The paper characterizes law-invariant star-shaped risk measures.
problem Understanding and characterizing law-invariant star-shaped risk measures.
method Developed characterizations for positively homogeneous and star-shaped functionals, derived Kusuoka-type representations, and offered representations of general law-invariant star-shaped functionals.
result Characterizations of law-invariant star-shaped functionals, including their connections to Value-at-Risk and Expected Shortfall.
To improve the off-sample generalization of classical procedures minimizing the empirical risk under potentially heavy-tailed data, new robust learning algorithms have been proposed in recent years, with generalized median-of-means strategies being particularly salient. These procedures enjoy performance guarantees in …
In this paper, we present two new communication-efficient methods for distributed minimization of an average of functions. The first algorithm is an inexact variant of the DANE algorithm that allows any local algorithm to return an approximate solution to a local subproblem. We show that such a strategy does not affect…
A Bayesian framework models adversarial uncertainty for robust machine learning.
problem Vulnerability of machine learning models to adversarial attacks.
method Formal Bayesian framework that models adversarial uncertainty through a stochastic channel, articulating probabilistic assumptions.
result Explicitly modeling adversarial uncertainty leads to improved robustification strategies.
In this paper, we address a problem of machine learning system vulnerability to adversarial attacks. We propose and investigate a Key based Diversified Aggregation (KDA) mechanism as a defense strategy. The KDA assumes that the attacker (i) knows the architecture of classifier and the used defense strategy, (ii) has an…
Robust quickest change detection method for unknown score functions.
problem Detecting changes in data streams with unknown pre- and post-change distributions.
method Selects least-favorable distributions and robustifies score-based detection algorithm.
result Demonstrates improved performance in simulations.
Robustness to outliers is a central issue in real-world machine learning applications. While replacing a model to a heavy-tailed one (e.g., from Gaussian to Student-t) is a standard approach for robustification, it can only be applied to simple models. In this paper, based on Zellner's optimization and variational form…
Regularized policies are robust to adversarial rewards.
problem Understanding the effects of regularization on policy exploration and robustness.
method Using Fenchel duality to derive the dual problem of the regularized RL objective, showing the optimal policy is robust to adversarial rewards.
result Regularized policies are optimal for a reinforcement learning problem under adversarial reward conditions.
Improved scalable machine learning under heavy-tailed data.
problem Machine learning scalability under heavy-tailed data without strong convexity.
method Simple robust validation sub-routine to boost confidence in gradient-based sub-processes.
result Substantial improvement in dimension dependence without strong convexity.
Study optimal transport for robust optimization, showing how adversary's strategy relates to regularization.
problem Optimizing under uncertain parameters with a fictitious adversary reshaping a reference distribution.
method Introduces optimal transport and regularization to relate robustification to variation and Lipschitz norms.
result Conditions for existence and computability of Nash equilibrium between decision-maker and adversary.
Robustifies Markowitz portfolios to reduce transaction costs and improve performance.
problem Markowitz portfolios are unreliable due to estimation errors and extreme weights.
method Projected gradient descent and robust statistics for stable weights and costs.
result Robustified Markowitz portfolios have lower turnover and maintain or improve performance.
Robust Trimmed k-means improves clustering with outliers and mixed data.
problem Real-world data often contains outliers and mixed membership clusters, complicating traditional clustering methods.
method Proposes Robust Trimmed k-means (RTKM) that robustifies k-means for both single- and multi-membership data.
result RTKM outperforms other methods on multi-membership data with outliers and single membership data with outliers.
In many situations where the interest lies in identifying clusters one might expect that not all available variables carry information about these groups. Furthermore, data quality (e.g. outliers or missing entries) might present a serious and sometimes hard-to-assess problem for large and complex datasets. In this pap…
SRO optimizes decisions against worst-case sampler induced by generative models.
problem Operational uncertainty shifts from explicit probability law to sampler induced by learned generators.
method SRO optimizes decisions against the worst-case sampler induced by perturbing the learned generator.
result Empirical worst-case objective provides high-probability upper certificate for true population objective.
The paper uses EVT to improve tail risk measures under ambiguity sets.
problem Misspecification of tail risk measures leads to inflated risk estimates.
method Applies Extreme Value Theory to derive worst-case tail risk under ambiguity sets.
result Proposes a tail-calibrated ambiguity design that preserves nominal tail asymptotic scaling.
Nonparametric methods are widely applicable to statistical inference problems, since they rely on a few modeling assumptions. In this context, the fresh look advocated here permeates benefits from variable selection and compressive sampling, to robustify nonparametric regression against outliers - that is, data markedl…
New robustness test for kernel goodness-of-fit tests.
problem Lack of robustness in existing kernel goodness-of-fit tests.
method Proposes a new robust kernel goodness-of-fit test using kernel Stein discrepancy (KSD) balls.
result First robust kernel goodness-of-fit test addressing both qualitative and quantitative robustness.
New algorithm poLinUCB improves online learning in content recommendation platforms.
problem Improving efficiency in content recommendation platforms with post-serving context.
method Novel contextual bandit problem with post-serving contexts and a new algorithm, poLinUCB.
result Achieves tight regret under standard assumptions and significant benefit of utilizing post-serving contexts.
Super learner with Huber loss improves cost prediction and causal effect estimation in healthcare expenditure data.
problem Challenges in modeling healthcare expenditure distributions with standard super learning methods.
method Proposes a super learner using Huber loss, a robust loss function that down-weights outliers.
result Demonstrates appreciable finite-sample gains in cost prediction and causal effect estimation.
Principal component analysis (PCA) is widely used for dimensionality reduction, with well-documented merits in various applications involving high-dimensional data, including computer vision, preference measurement, and bioinformatics. In this context, the fresh look advocated here permeates benefits from variable sele…
We develop the idea of using Monte Carlo sampling of random portfolios to solve portfolio investment problems. In this first paper we explore the need for more general optimization tools, and consider the means by which constrained random portfolios may be generated. A practical scheme for the long-only fully-invested …
Study dynamic risk measures with distributional uncertainty using optimal transport.
problem Risk robustification under distributional uncertainty in Markovian models.
method Characterize risk measures via convex monotone semigroups and optimal transport costs.
result Identify generator and correction terms for dynamic risk measures under different scaling regimes.
Paper tackles efficient HMM learning with conditional samples.
problem Cryptographic hardness in learning HMMs from i.i.d. samples.
method Interactive access model, polynomial-time algorithms for conditional probabilities and latent low rank structures.
result Efficient algorithms for HMM learning in both exact and approximate conditional settings.
GE finds failures in autonomous systems without domain heuristics.
problem Finding failures in autonomous systems without domain-specific heuristics.
method Adaptive stress testing using go-explore (GE) algorithm.
result GE finds failures in scenarios other RL techniques cannot solve.
Study on sequential prediction with log-loss, focusing on well-specified and misspecified cases.
problem Sequential prediction with log-loss under different specification conditions.
method Analysis of cumulative regret in well-specified and misspecified cases for a Gaussian location hypothesis class.
result Cumulative regrets in well-specified and misspecified cases asymptotically coincide for the d-dimensional Gaussian location hypothesis class. Machine learning has begun to play a central role in many applications. A multitude of these applications typically also involve datasets that are distributed across multiple computing devices/machines due to either design constraints (e.g., multiagent systems) or computational/privacy reasons (e.g., learning on smartp…
Study extends DRO with IPMs, linking robustness to regularization and GANs.
problem Addressing robustness of deep neural networks to adversarial attacks.
method Distributionally Robust Optimization (DRO) with Integral Probability Metrics (IPMs).
result DRO under any IPM corresponds to a family of regularization penalties.