New research extends optimal transport map breakdown properties to general costs.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We formalize notions of robustness for composite estimators via the notion of a breakdown point. A composite estimator successively applies two (or more) estimators: on data decomposed into disjoint parts, it applies the first estimator on each part, then the second estimator on the outputs of the first estimator. And …
We analyze the performance of the Tukey median estimator under total variation (TV) distance corruptions. Previous results show that under Huber's additive corruption model, the breakdown point is 1/3 for high-dimensional halfspace-symmetric distributions. We show that under TV corruptions, the breakdown point reduces …
The support vector machine (SVM) is one of the most successful learning methods for solving classification problems. Despite its popularity, SVM has a serious drawback, that is sensitivity to outliers in training samples. The penalty on misclassification is defined by a convex loss called the hinge loss, and the unboun…
Paper analyzes robustness of MDPDE under INH setups.
New algorithm estimates edge density of random graphs robustly, achieving optimal breakdown point.
Adapting robust statistics to neural networks, researchers found neural networks can be more robust with certain loss functions.
Paper solves outlier robust mean estimation near breakdown point.
Unified framework for Byzantine robust gossip algorithms with guaranteed performance.
New method prevents neural network breakdown by combining trimmed loss and variation regularization.
We consider the dimensionality-reduction problem (finding a subspace approximation of observed data) for contaminated data in the high dimensional regime, where the number of observations is of the same magnitude as the number of variables of each observation, and the data set contains some (arbitrarily) corrupted obse…
WPCA improves subspace recovery robustness to outliers.
Paper shows how to use geometric median for robust SGD in high dimensions.
Efficient SVD algorithm robust to outliers.
Proposes MPCA for robust PCA using mode estimation.
We propose a framework for distributed robust statistical learning on {\em big contaminated data}. The Distributed Robust Learning (DRL) framework can reduce the computational time of traditional robust learning methods by several orders of magnitude. We analyze the robustness property of DRL, showing that DRL not only…
Proposes a robust factor analysis for matrix data.
This work uses neural density estimation to analyze laser-induced breakdown spectroscopy data, enabling accurate predictions and uncertainty quantification.
We consider principal component analysis for contaminated data-set in the high dimensional regime, where the dimensionality of each observation is comparable or even more than the number of observations. We propose a deterministic high-dimensional robust PCA algorithm which inherits all theoretical properties of its ra…
Robust estimation methods find global minima efficiently via quasi-gradients.
The recent "correlation breakdown" in the modeling of credit default swaps, in which model correlations had to exceed 100% in order to reproduce market prices of supersenior tranches, is analyzed and argued to be a fundamental market inconsistency rather than an inadequacy of the specific model. As a consequence, marke…
In this paper we propose a new method to assist in labeling data arriving from fast running processes using anomaly detection. A result is the possibility to manually classify data arriving at a high rates to train machine learning models. To circumvent the problem of not having a real ground truth we propose specific …
Study improves robustness of Bayesian inference for cognitive models.
Paper introduces robust methods for consensus ranking in AI systems.
Residential homes constitute roughly one-fourth of the total energy usage worldwide. Providing appliance-level energy breakdown has been shown to induce positive behavioral changes that can reduce energy consumption by 15%. Existing approaches for energy breakdown either require hardware installation in every target ho…
The paper is concerned with regularity properties of boundaries of causal pasts of points in a 3+1-dimensional Einstein-vacuum spacetime. In a Lorentzian manifold such boundaries play crucial role in propagation of linear and nonlinear waves. We prove a uniform lower bound on the radius of injectivity of these null bou…
Complex models are commonly used in predictive modeling. In this paper we present R packages that can be used to explain predictions from complex black box models and attribute parts of these predictions to input features. We introduce two new approaches and corresponding packages for such attribution, namely live and …
Unified AI detection framework for various artifacts.
Let $\M_*=\cup_{t\in [t_0, t_*)} Σ_t$ be a part of vacuum globally hyperbolic space-time $(\bM, \bg)$, foliated by constant mean curvature hypersurfaces with . We show that the foliation can be extended beyond if the second fundamental form and the lapse function satisfy $$ \int_{t_0}^{t_…
New robust learning framework for regression NNs using β-divergences.
The paper reports the construction of artificial stock market that emerges the similar statistical facts with real data in Indonesian stock market. We use the individual but dominant data, i.e.: PT TELKOM in hourly interval. The artificial stock market shows standard statistical facts, e.g.: volatility clustering, the …
Optimizes search times by resetting agents when a threshold is reached.
Superstatistics is a widely employed tool of non-equilibrium statistical physics which plays an important role in analysis of hierarchical complex dynamical systems. Yet, its "canonical" formulation in terms of a single nuisance parameter is often too restrictive when applied to complex empirical data. Here we show tha…
We study the robustness properties of norm minimization for the classical linear regression problem with a given design matrix and contamination restricted to the dependent variable. We perform a fine error analysis of the estimator for measurements errors consisting of outliers coupled with noise. We…
Predicting unscheduled breakdowns of plasma etching equipment can reduce maintenance costs and production losses in the semiconductor industry. However, plasma etching is a complex procedure and it is hard to capture all relevant equipment properties and behaviors in a single physical model. Machine learning offers an …
Instead of investigating the Willmore flow for two-dimensional, closed immersed surfaces directly we turn to its inversion. We give a lower bound on the lifespan of this inverse Willmore flow, depending on the concentration of curvature in space and the extension of the initial surface, as well as a characterization of…
We consider the issue of the slice invariance of refined topological string amplitudes, which means that they are independent of the choice of the preferred direction of the refined topological vertex. We work out two examples. The first example is a geometric engineering of five-dimensional U(1) gauge theory with a ma…
Study on dropout in neural networks using percolation theory.
The main objective of this paper is to control the geometry of null cones with time foliation in Einstein vacuum spacetime under the assumptions of small curvature flux and a weaker condition on the deformation tensor for $\bT$. We establish a series of estimates on Ricci coefficients, which plays a crucial role to pro…
A new robust PCA estimator combining M-estimators and minimum divergence estimators.
I derive practical formulas for optimal arrangements between sophisticated stock market investors (namely, continuous-time Kelly gamblers or, more generally, CRRA investors) and the brokers who lend them cash for leveraged bets on a high Sharpe asset (i.e. the market portfolio). Rather than, say, the broker posting a m…
Graph auto-encoders predict stock market instability by measuring graph structure changes.
Theoretical framework for M-posteriors connects Bayesian and frequentist statistics.
The study uncovers the breakdown of Gaussian universality in high-dimensional empirical risk minimization.
The importance weighted autoencoder (IWAE) (Burda et al., 2016) is a popular variational-inference method which achieves a tighter evidence bound (and hence a lower bias) than standard variational autoencoders by optimising a multi-sample objective, i.e. an objective that is expressible as an integral over Mont…
This paper relates parameter distance to gradient breakdown for a broad class of nonlinear compositional functions. The analysis leads to a new distance function called deep relative trust and a descent lemma for neural networks. Since the resulting learning rule seems to require little to no learning rate tuning, it m…
New findings show Gaussian universality breaks down in high-dimensional linear factor mixtures.
End-to-end learning refers to training a possibly complex learning system by applying gradient-based learning to the system as a whole. End-to-end learning system is specifically designed so that all modules are differentiable. In effect, not only a central learning machine, but also all "peripheral" modules like repre…