This work proposes robustness curves to analyze model robustness.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Recently theoretical guarantees have been obtained for matrix completion in the non-uniform sampling regime. In particular, if the sampling distribution aligns with the underlying matrix's leverage scores, then with high probability nuclear norm minimization will exactly recover the low rank matrix. In this article, we…
This work analyzes how to choose regularization norms for adversarial training in high dimensions.
SGD achieves a bound for minimizing gradient norm of smooth functions.
We show that minimum-norm interpolation in the Reproducing Kernel Hilbert Space corresponding to the Laplace kernel is not consistent if input dimension is constant. The lower bound holds for any choice of kernel bandwidth, even if selected based on data. The result supports the empirical observation that minimum-norm …
Recently, matrix norm has been widely applied to many areas such as computer vision, pattern recognition, biological study and etc. As an extension of vector norm, the mixed matrix norm is often used to find jointly sparse solutions. Moreover, an efficient iterative algorithm has been designed…
Overconfidence and underconfidence in machine learning classifiers is measured by calibration: the degree to which the probabilities predicted for each class match the accuracy of the classifier on that prediction. How one measures calibration remains a challenge: expected calibration error, the most popular metric, ha…
In binary classification and regression problems, it is well understood that Lipschitz continuity and smoothness of the loss function play key roles in governing generalization error bounds for empirical risk minimization algorithms. In this paper, we show how these two properties affect generalization error bounds in …
Learning linear combinations of multiple kernels is an appealing strategy when the right choice of features is unknown. Previous approaches to multiple kernel learning (MKL) promote sparse kernel combinations to support interpretability and scalability. Unfortunately, this 1-norm MKL is rarely observed to outperform tr…
Optimal financial strategies minimize risk under uncertain models.
We study the implicit bias of generic optimization methods, such as mirror descent, natural gradient descent, and steepest descent with respect to different potentials and norms, when optimizing underdetermined linear regression or separable linear classification problems. We explore the question of whether the specifi…
Proposes a new criterion for selecting Nash equilibria considering both utility and inequality.
The paper ranks items based on top choices in multiway comparisons.
We address the problem of {\it adaptivity} in the framework of reproducing kernel Hilbert space (RKHS) regression. More precisely, we analyze estimators arising from a linear regularization scheme $g_\lam$. In practical applications, an important task is to choose the regularization parameter $\lam$ appropriately, i.e.…
Muon optimizer improves deep learning with spectral norm constraints.
The paper tackles system identification via Hankel nuclear norm regularization, improving estimation rates and singular value gaps.
Equivalent tests for SGD batch size selection found.
Study shows momentum-based optimizers like Muon and MomentumGD bias towards KKT points in smooth homogeneous models.
In this paper, we study the effects of different prior and likelihood choices for Bayesian matrix factorisation, focusing on small datasets. These choices can greatly influence the predictive performance of the methods. We identify four groups of approaches: Gaussian-likelihood with real-valued priors, nonnegative prio…
We consider the problem of learning the preferences of a heterogeneous population by observing choices from an assortment of products, ads, or other offerings. Our observation model takes a form common in assortment planning applications: each arriving customer is offered an assortment consisting of a subset of all pos…
We apply the OSCAR (octagonal selection and clustering algorithms for regression) in recovering group-sparse matrices (two-dimensional---2D---arrays) from compressive measurements. We propose a 2D version of OSCAR (2OSCAR) consisting of the norm and the pair-wise norm, which is convex but non-d…
This study shows how optimizer choice affects adversarial robustness in neural networks.
We study a norm for structured sparsity which leads to sparse linear predictors whose supports are unions of prede ned overlapping groups of variables. We call the obtained formulation latent group Lasso, since it is based on applying the usual group Lasso penalty on a set of latent variables. A detailed analysis of th…
Unified formula for higher traces of linear maps on finite-dimensional normed spaces.
Kolmogorov-Arnold Networks offer improved interpretability and parsimony in science tasks.
Mirror descent algorithm recovers low-rank matrices in matrix sensing.
Proposes an efficient method for sparse index tracking with -norm constraints.
The recent proposed Tensor Nuclear Norm (TNN) [Lu et al., 2016; 2018a] is an interesting convex penalty induced by the tensor SVD [Kilmer and Martin, 2011]. It plays a similar role as the matrix nuclear norm which is the convex surrogate of the matrix rank. Considering that the TNN based Tensor Robust PCA [Lu et al., 2…
New algorithms for linear bandits avoid norm knowledge, reducing regret.
Recently, there has been focus on penalized log-likelihood covariance estimation for sparse inverse covariance (precision) matrices. The penalty is responsible for inducing sparsity, and a very common choice is the convex norm. However, the best estimator performance is not always achieved with this penalty. The …
Paper addresses concentration of distances for fractional quasi p-norms, identifying conditions for concentration and anti-concentration.
Proposes a new adaptive gradient method based on gradient differences.
New method stabilizes machine learning for physics-informed inverse problems.
Consider the recovery of an unknown signal from quantized linear measurements. In the one-bit compressive sensing setting, one typically assumes that is sparse, and that the measurements are of the form . Since such measurements give no informati…
Adaptive gradient methods such as AdaGrad and its variants update the stepsize in stochastic gradient descent on the fly according to the gradients received along the way; such methods have gained widespread use in large-scale optimization for their ability to converge robustly, without the need to fine-tune the stepsi…
Significant attention has been given to minimizing a penalized least squares criterion for estimating sparse solutions to large linear systems of equations. The penalty is responsible for inducing sparsity and the natural choice is the so-called norm. In this paper we develop a Momentumized Iterative Shrinkage Th…
We revisit the choice of SGD for training deep neural networks by reconsidering the appropriate geometry in which to optimize the weights. We argue for a geometry invariant to rescaling of weights that does not affect the output of the network, and suggest Path-SGD, which is an approximate steepest descent method with …
The extraction of clusters from a dataset which includes multiple clusters and a significant background component is a non-trivial task of practical importance. In image analysis this manifests for example in anomaly detection and target detection. The traditional spectral clustering algorithm, which relies on the lead…
The paper proposes using function approximations to reduce the computational burden in measuring counterparty credit exposure.
New Muon and Momo variants improve neural network optimization robustness.
We investigate the effect of explicitly enforcing the Lipschitz continuity of neural networks with respect to their inputs. To this end, we provide a simple technique for computing an upper bound to the Lipschitz constant---for multiple -norms---of a feed forward neural network composed of commonly used layer types.…
In applications such as recommendation systems and revenue management, it is important to predict preferences on items that have not been seen by a user or predict outcomes of comparisons among those that have never been compared. A popular discrete choice model of multinomial logit model captures the structure of the …
Adversarial training improves linear regression solutions, offering robustness against small perturbations.
The paper studies how regularization parameters affect sparsity in deep neural networks.
Gradient descent constructs tight fusion frames.
EF21-Muon optimizes deep learning with error feedback, improving efficiency and accuracy.
New measure shows various training techniques control model complexity.
In this work we propose to fit a sparse logistic regression model by a weakly convex regularized nonconvex optimization problem. The idea is based on the finding that a weakly convex function as an approximation of the pseudo norm is able to better induce sparsity than the commonly used norm. For a cl…