We assess cluster stability by trimming extreme points and tracking data range reduction.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Improved Frank-Wolfe algorithm for constrained convex optimization with nearest extreme point oracle.
We show that any n-dimensional nonnegatively curved Alexandrov space with the maximal possible number of extremal points is isometric to a quotient space of Euclidean n -space by an action of a crystallographic group. We describe all such actions.
Trimming helps in conformal prediction when it separates anomaly scores.
We consider the problem of robustifying high-dimensional structured estimation. Robust techniques are key in real-world applications which often involve outliers and data corruption. We focus on trimmed versions of structurally regularized M-estimators in the high-dimensional setting, including the popular Least Trimme…
New method trims network data to resist adversarial contamination.
Nonconvex penalty methods for sparse modeling in linear regression have been a topic of fervent interest in recent years. Herein, we study a family of nonconvex penalty functions that we call the trimmed Lasso and that offers exact control over the desired level of sparsity of estimators. We analyze its structural prop…
We describe a general framework for measuring risks, where the risk measure takes values in an abstract cone. It is shown that this approach naturally includes the classical risk measures and set-valued risk measures and yields a natural definition of vector-valued risk measures. Several main constructions of risk meas…
We relate trimmed sums of twists in cylinders along a typical Teichmuller geodesic to the area Siegel-Veech constant.
New method solves sparse approximation problem using trimmed lasso and generalized soft-min penalties.
New core inflation measure predicts future headline inflation.
Gaussian Graphical Models (GGMs) are popular tools for studying network structures. However, many modern applications such as gene network discovery and social interactions analysis often involve high-dimensional noisy data with outliers or heavier tails than the Gaussian distribution. In this paper, we propose the Tri…
A method for estimating parameters from entangled single-sample distributions, robust to high-noise data.
New conditions ensure Dantzig-Wolfe relaxation matches rank-constrained optimization problems.
Alpha-trimming prunes trees in random forests to improve predictive performance.
New method prevents neural network breakdown by combining trimmed loss and variation regularization.
Paper proposes robust gossip algorithms for mean and trimmed mean estimation.
We consider the signed density of the extremal points of (two-dimensional) scalar fields with a Gaussian distribution. We assign a positive unit charge to the maxima and minima of the function and a negative one to its saddles. At first, we compute the average density for a field in half-space with Dirichlet boundary c…
TrIM improves gradient-based dimension reduction and regression.
New algorithm robustly estimates sparse models in high dimensions with corrupted data.
Estimates GLMs robustly against label corruptions.
We propose a robust elastic net (REN) model for high-dimensional sparse regression and give its performance guarantees (both the statistical error bound and the optimization bound). A simple idea of trimming the inner product is applied to the elastic net model. Specifically, we robustify the covariance matrix by trimm…
This paper considers the problem of removing costly features from a Bayesian network classifier. We want the classifier to be robust to these changes, and maintain its classification behavior. To this end, we propose a closeness metric between Bayesian classifiers, called the expected classification agreement (ECA). Ou…
A new robust GP regression algorithm that trims outliers improves model accuracy.
Given an iterated function system of affine dilations with fixed points the vertices of a regular polygon, we characterize which points in the limit set lie on the boundary of its convex hull.
TRIM improves interpretability of deep neural networks in cosmology.
Analyzes convex structures in Teichmüller space unit tangent spheres.
Let be the class of complete simply connected dimensional manifolds without conjugate points. The hyperbolic space as well as Euclidean space are good examples of such manifolds. Let and let be a subset of . This article aims at characterization and bu…
We develop a fast, tractable technique called Net-Trim for simplifying a trained neural network. The method is a convex post-processing module, which prunes (sparsifies) a trained network layer by layer, while preserving the internal responses. We present a comprehensive analysis of Net-Trim from both the algorithmic a…
New method measures model variability from stochastic optimization.
Density ratio estimation is a vital tool in both machine learning and statistical community. However, due to the unbounded nature of density ratio, the estimation procedure can be vulnerable to corrupted data points, which often pushes the estimated ratio toward infinity. In this paper, we present a robust estimator wh…
New algorithm reduces ERM problem size while maintaining accuracy.
Improved robust regression for heavy-tailed and contaminated data.
Combines VaR and ES forecasts from a large pool of methods.
Robust Trimmed k-means improves clustering with outliers and mixed data.
Paper develops robust OPF method using contextual information.
Using a trimming approach, we investigate a k-means type method based on Bregman divergences for clustering data possibly corrupted with clutter noise. The main interest of Bregman divergences is that the standard Lloyd algorithm adapts to these distortion measures, and they are well-suited for clustering data sampled …
Clustering, or unsupervised classification, is a task often plagued by outliers. Yet there is a paucity of work on handling outliers in clustering. Outlier identification algorithms tend to fall into three broad categories: outlier inclusion, outlier trimming, and post hoc outlier identification methods, with the forme…
A new robust regression method handles outliers in high-dimensional data.
Proposes a sensitivity framework to handle limited overlap in causal inference.
Identifying features that leak information about sensitive attributes is a key challenge in the design of information obfuscation mechanisms. In this paper, we propose a framework to identify information-leaking features via information density estimation. Here, features whose information densities exceed a pre-defined…
Efficiently clusters data with weak assumptions, robust to contamination.
We introduce and analyze a new technique for model reduction for deep neural networks. While large networks are theoretically capable of learning arbitrarily complex models, overfitting and model redundancy negatively affects the prediction accuracy and model variance. Our Net-Trim algorithm prunes (sparsifies) a train…
Convex geometry explains optimal neural network parameters.
We study polar orbitopes, i.e. convex hulls of orbits of a polar representation of a compact Lie group. The face structure is studied by means of the gradient momentum map and it is shown that every face is exposed and is again a polar orbitope. Up to conjugation the faces are completely determined by the momentum poly…
Recent studies on automatic neural architectures search have demonstrated significant performance, competitive to or even better than hand-crafted neural architectures. However, most of the existing network architecture tend to use residual, parallel structures and concatenation block between shallow and deep features …
Robust CG methods avoid data corruption and solve structured statistical estimation problems.
We establish a new uniqueness theorem for the three dimensional Schwarzschild-de Sitter metrics. For this some new or improved tools are developed. These include a reverse Lojasiewicz inequality, which holds in a neighborhood of the extremal points of any smooth function. We further prove smoothness of the set of maxim…