Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

126253379505 · Jun 202019922001200920172026
48 results for regression breakdown point

Adapting robust statistics to neural networks, researchers found neural networks can be more robust with certain loss functions.

problem The robustness of neural networks in complex learning tasks.
method Adapting the regression breakdown point from robust statistics to neural networks and comparing different configurations and contamination settings.
result Neural networks can benefit from robust loss functions, as demonstrated in extensive simulations.

We formalize notions of robustness for composite estimators via the notion of a breakdown point. A composite estimator successively applies two (or more) estimators: on data decomposed into disjoint parts, it applies the first estimator on each part, then the second estimator on the outputs of the first estimator. And …

2016-09-05abs ↗pdf ↗

New robust learning framework for regression NNs using β-divergences.

problem Outliers and data contamination in regression NNs training.
method Proposes rRNet based on β-divergence for robust learning of regression NNs.
result rRNet achieves optimal 50% asymptotic breakdown point for all β ∈ (0, 1].

We analyze the performance of the Tukey median estimator under total variation (TV) distance corruptions. Previous results show that under Huber's additive corruption model, the breakdown point is 1/3 for high-dimensional halfspace-symmetric distributions. We show that under TV corruptions, the breakdown point reduces …

2020-01-21abs ↗pdf ↗

The support vector machine (SVM) is one of the most successful learning methods for solving classification problems. Despite its popularity, SVM has a serious drawback, that is sensitivity to outliers in training samples. The penalty on misclassification is defined by a convex loss called the hinge loss, and the unboun…

2014-09-03abs ↗pdf ↗

We propose a framework for distributed robust statistical learning on {\em big contaminated data}. The Distributed Robust Learning (DRL) framework can reduce the computational time of traditional robust learning methods by several orders of magnitude. We analyze the robustness property of DRL, showing that DRL not only…

2014-09-21abs ↗pdf ↗

New algorithm estimates edge density of random graphs robustly, achieving optimal breakdown point.

problem Estimating edge density of Erdős-Rényi graphs under adversarial edge manipulation.
method Sum-of-Squares (SoS) hierarchy, constructing constant-degree certificates for concentration.
result First polynomial-time algorithm with optimal breakdown point and matching error guarantees.

Unified framework for Byzantine robust gossip algorithms with guaranteed performance.

problem Vulnerability of decentralized machine learning to misbehaving devices.
method Introduces F-RG framework and CS+ robust aggregation rule for Byzantine resilience.
result CS+-RG has near-optimal breakdown tolerance and outperforms existing methods.

New method prevents neural network breakdown by combining trimmed loss and variation regularization.

problem Outlier contamination in neural network training.
method Integrates transformed trimmed loss and higher-order variation regularization.
result Ensures robustness to outlier contamination with a high functional breakdown point.

Paper shows how to use geometric median for robust SGD in high dimensions.

problem Robustifying SGD for high-dimensional optimization problems with gross corruption.
method Applying geometric median to only chosen blocks of coordinates at a time.
result Retains optimal breakdown point of 0.5 for smooth non-convex problems.

Study enhances robustness of In-CVaR based regression models under perturbation and contamination.

problem Enhancing robustness of nonlinear regression models under perturbation and contamination.
method Introduces interval conditional value-at-risk (In-CVaR) and rigorously analyzes its robustness properties under both perturbation and contamination.
result The In-CVaR based estimator is qualitatively robust in terms of the Prokhorov metric if and only if the largest portion of losses is trimmed.

New algorithms robust to adversarial data achieve optimal performance.

problem Adversarial robustness in high-dimensional online learning problems.
method Alternating minimization scheme combining least-squares and convex reweighting.
result Achieves optimal robustness guarantees without distributional assumptions.

This work uses neural density estimation to analyze laser-induced breakdown spectroscopy data, enabling accurate predictions and uncertainty quantification.

problem Inference of probability densities in high-dimensional spectral data is often intractable.
method Normalizing flows on structured spectral latent spaces for density estimation and uncertainty quantification.
result The approach enables generation of realistic spectral samples and accurate prediction of state vectors with well-calibrated uncertainties.

We consider principal component analysis for contaminated data-set in the high dimensional regime, where the dimensionality of each observation is comparable or even more than the number of observations. We propose a deterministic high-dimensional robust PCA algorithm which inherits all theoretical properties of its ra…

2012-06-18abs ↗pdf ↗

The paper develops AMP theory for sparse and robust regression with polynomial iterations.

problem Challenges in high-dimensional statistical estimation due to asymptotic theory breakdown.
method Non-asymptotic distributional theory of AMP for sparse and robust regression.
result First finite-sample non-asymptotic distributional theory of AMP for polynomial iterations.

The recent "correlation breakdown" in the modeling of credit default swaps, in which model correlations had to exceed 100% in order to reproduce market prices of supersenior tranches, is analyzed and argued to be a fundamental market inconsistency rather than an inadequacy of the specific model. As a consequence, marke…

2009-08-31abs ↗pdf ↗

Study improves robustness of Bayesian inference for cognitive models.

problem Outliers and contaminants affect parameter estimation in cognitive models.
method Robustness of parameter estimation using amortized Bayesian inference (ABI) with neural networks.
result Introducing contaminants from a Cauchy distribution increases robustness.

Residential homes constitute roughly one-fourth of the total energy usage worldwide. Providing appliance-level energy breakdown has been shown to induce positive behavioral changes that can reduce energy consumption by 15%. Existing approaches for energy breakdown either require hardware installation in every target ho…

2019-09-02abs ↗pdf ↗

The paper is concerned with regularity properties of boundaries of causal pasts of points in a 3+1-dimensional Einstein-vacuum spacetime. In a Lorentzian manifold such boundaries play crucial role in propagation of linear and nonlinear waves. We prove a uniform lower bound on the radius of injectivity of these null bou…

2006-03-01abs ↗pdf ↗

Complex models are commonly used in predictive modeling. In this paper we present R packages that can be used to explain predictions from complex black box models and attribute parts of these predictions to input features. We introduce two new approaches and corresponding packages for such attribution, namely live and …

2018-04-05abs ↗pdf ↗

Let $\M_*=\cup_{t\in [t_0, t_*)} Σ_t$ be a part of vacuum globally hyperbolic space-time $(\bM, \bg)$, foliated by constant mean curvature hypersurfaces ΣtΣ_t with t0<t<0t_0<t_*<0. We show that the foliation can be extended beyond tt_* if the second fundamental form kk and the lapse function nn satisfy $$ \int_{t_0}^{t_…

2010-04-17abs ↗pdf ↗

The paper reports the construction of artificial stock market that emerges the similar statistical facts with real data in Indonesian stock market. We use the individual but dominant data, i.e.: PT TELKOM in hourly interval. The artificial stock market shows standard statistical facts, e.g.: volatility clustering, the …

2004-08-16abs ↗pdf ↗

A new method estimates parameters in heavy-tailed corrupted regression with unknown covariance and heterogeneous noise.

problem Estimating parameters in regression with heavy-tailed errors and unknown covariance.
method Near-optimal computationally tractable estimator based on power method and Multiplicative Weight Update algorithm.
result The estimator achieves the optimal statistical rate and breakdown-point under near-optimal sample size.

Optimizes search times by resetting agents when a threshold is reached.

problem Improving search efficiency in systems with thresholds.
method Develops a framework for correlated stochastic processes with threshold resetting.
result Optimal resetting can prevent larger losses and is applicable to various stochastic systems.

Study improves robust nonparametric regression in heavy-tailed noise.

problem Robust nonparametric regression with heavy-tailed noise and unbounded functions.
method Huber regression in reproducing kernel Hilbert spaces (RKHS), probabilistic effective hypothesis space, new comparison theorems.
result Explicit finite-sample error bounds and convergence rates for Huber regression in RKHS under heavy-tailed noise.

The paper presents anomaly detection in time series data using InfluxDB and Python.

problem Anomalous data points in time series data affect decision making in water and environmental systems.
method Data cleaning, cost-sensitive machine learning (Logistic Regression, Random Forest, SVM), feature selection, and InfluxDB integration.
result Random Forest outperformed other models in detecting anomalies.

Instead of investigating the Willmore flow for two-dimensional, closed immersed surfaces directly we turn to its inversion. We give a lower bound on the lifespan of this inverse Willmore flow, depending on the concentration of curvature in space and the extension of the initial surface, as well as a characterization of…

2015-08-31abs ↗pdf ↗

We consider the issue of the slice invariance of refined topological string amplitudes, which means that they are independent of the choice of the preferred direction of the refined topological vertex. We work out two examples. The first example is a geometric engineering of five-dimensional U(1) gauge theory with a ma…

2009-03-31abs ↗pdf ↗

A new robust PCA estimator combining M-estimators and minimum divergence estimators.

problem Adverse effect of outlying observations in PCA for high-dimensional data.
method Minimum density power divergence estimator combined with a computationally efficient algorithm.
result High breakdown guarantee regardless of data dimension with theoretical support and practical applications.