Abstract compares two norms in holomorphic quadratic differentials.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New algorithms adapt to both gradient norms and comparator norms in online learning.
New algorithms reduce regret for convex bandits with small comparator norms.
We propose a set of convex low rank inducing norms for a coupled matrices and tensors (hereafter coupled tensors), which shares information between matrices and tensors through common modes. More specifically, we propose a mixture of the overlapped trace norm and the latent norms with the matrix trace norm, and then, w…
We suggest using the max-norm as a convex surrogate constraint for clustering. We show how this yields a better exact cluster recovery guarantee than previously suggested nuclear-norm relaxation, and study the effectiveness of our method, and other related convex relaxations, compared to other clustering approaches.
Minimum-norm solutions generalize well in over-parametrized neural networks.
Differential calculus on metric spaces is contained in the algebraic study of normed groupoids with -structures. Algebraic study of normed groups endowed with dilatation structures is contained in the differential calculus on metric spaces. Thus all algebraic properties of the small world of normed groups with dilat…
Exact spectral norm regularization improves neural network generalization.
In a recent paper, McMullen showed an inequality between the Thurston norm and the Alexander norm of a 3-manifold. This generalizes the well-known fact that twice the genus of a knot is bounded from below by the degree of the Alexander polynomial. We extend the Bennequin inequality for links to an inequality for all po…
CNN layers with large norms are still robust to adversarial attacks.
A toolkit for path-norms enhances neural network generalization bounds.
We consider a variant of online convex optimization in which both the instances (input vectors) and the comparator (weight vector) are unconstrained. We exploit a natural scale invariance symmetry in our unconstrained setting: the predictions of the optimal comparator are invariant under any linear transformation of th…
The Schatten quasi-norm was introduced to bridge the gap between the trace norm and rank function. However, existing algorithms are too slow or even impractical for large-scale problems. Motivated by the equivalence relation between the trace norm and its bilinear spectral penalty, we define two tractable Schatten norm…
A new method compares synthetic power networks to actual ones using multiscale flat norm.
HBR improves normative modeling of neuroimaging data across multiple sites.
Simple regional perturbations maintain model transferability while reducing adversarial example distortion.
New method for efficient proximal mapping of 1-path-norm in shallow networks.
We provide a necessary and sufficient condition that -norms, , of eigenfunctions of the square root of minus the Laplacian on 2-dimensional compact boundaryless Riemannian manifolds are small compared to a natural power of the eigenvalue . The condition that ensures this is that their norms ove…
This paper compares two invariants of foliated manifolds which seem to measure the non-Hausdorffness of the leaf space: the transversal length on the fundamental group and the foliated Gromov norm on the homology. We consider foliations with the property that the set of singular simplices transverse to the foliation sa…
Learning linear combinations of multiple kernels is an appealing strategy when the right choice of features is unknown. Previous approaches to multiple kernel learning (MKL) promote sparse kernel combinations to support interpretability and scalability. Unfortunately, this 1-norm MKL is rarely observed to outperform tr…
Proposes a new criterion for selecting Nash equilibria considering both utility and inequality.
Paper develops algorithms for sparse linear regression with generalized elastic net penalty.
Sparse alpha-norm regularization has many data-rich applications in Marketing and Economics. Alpha-norm, in contrast to lasso and ridge regularization, jumps to a sparse solution. This feature is attractive for ultra high-dimensional problems that occur in demand estimation and forecasting. The alpha-norm objective is …
Principal component analysis (PCA) is often used to reduce the dimension of data by selecting a few orthonormal vectors that explain most of the variance structure of the data. L1 PCA uses the L1 norm to measure error, whereas the conventional PCA uses the L2 norm. For the L1 PCA problem minimizing the fitting error of…
Extended Gauss-Markov theorem for linear estimation with bounded bias.
The paper analyzes the risk of a least squares estimator under a spike covariance model.
Paper addresses fault-tolerance in distributed machine learning with stochastic gradient descent.
This paper develops a new class of nonconvex regularizers for low-rank matrix recovery. Many regularizers are motivated as convex relaxations of the matrix rank function. Our new factor group-sparse regularizers are motivated as a relaxation of the number of nonzero columns in a factorization of the matrix. These nonco…
The Group-Lasso is a well-known tool for joint regularization in machine learning methods. While the l_{1,2} and the l_{1,\infty} version have been studied in detail and efficient algorithms exist, there are still open questions regarding other l_{1,p} variants. We characterize conditions for solutions of the l_{1,p} G…
This thesis presents the Conditional Value-at-Risk concept and combines an analysis that covers its application as a risk measure and as a vector norm. For both areas of application the theory is revised in detail and examples are given to show how to apply the concept in practice. In the first part, CVaR as a risk mea…
Analysis of non-asymptotic estimation error and structured statistical recovery based on norm regularized regression, such as Lasso, needs to consider four aspects: the norm, the loss function, the design matrix, and the noise model. This paper presents generalizations of such estimation error analysis on all four aspe…
Optimal a priori estimates are derived for the population risk, also known as the generalization error, of a regularized residual network model. An important part of the regularized model is the usage of a new path norm, called the weighted path norm, as the regularization term. The weighted path norm treats the skip c…
We simplify Thurston norm computation for 2-bridge link complements.
In this paper we investigate panel regression models with interactive fixed effects. We propose two new estimation methods that are based on minimizing convex objective functions. The first method minimizes the sum of squared residuals with a nuclear (trace) norm regularization. The second method minimizes the nuclear …
Noise injection before gradient steps helps in regularization for neural networks.
We analyze low rank tensor completion (TC) using noisy measurements of a subset of the tensor. Assuming a rank-, order-, tensor where , the best sampling complexity that was achieved is , which is obtained by solving a tensor nuclear-norm minimizatio…
This paper investigates the problem of sparse signal recovery in the presence of additive impulsive noise. The heavytailed impulsive noise is well modelled with stable distributions. Since there is no explicit formulation for the probability density function of distribution, alternative approximations like Genera…
Max-convolution is an important problem closely resembling standard convolution; as such, max-convolution occurs frequently across many fields. Here we extend the method with fastest known worst-case runtime, which can be applied to nonnegative vectors by numerically approximating the Chebyshev norm $\| \cdot \|_\infty…
Theoretical framework for neural network compression using sparsity norms.
QAPCA uses quantum annealing for robust PCA.
IRKSN algorithm achieves sparse recovery with wider applicability conditions.
This study uses neural networks to solve interpolation problems with sparse, infinitely wide layers.
Robustness is an increasingly important property of machine learning models as they become more and more prevalent. We propose a defense against adversarial examples based on a k-nearest neighbor (kNN) on the intermediate activation of neural networks. Our scheme surpasses state-of-the-art defenses on MNIST and CIFAR-1…
The paper tackles multi-armed bandits with vector losses, focusing on minimizing the -norm of relative losses.
The recent proposed Tensor Nuclear Norm (TNN) [Lu et al., 2016; 2018a] is an interesting convex penalty induced by the tensor SVD [Kilmer and Martin, 2011]. It plays a similar role as the matrix nuclear norm which is the convex surrogate of the matrix rank. Considering that the TNN based Tensor Robust PCA [Lu et al., 2…
GAS-Norm improves deep learning time series forecasting in non-stationary settings.
In this paper, we propose -norm regularized models to seek near-optimal sparse portfolios. These sparse solutions reduce the complexity of portfolio implementation and management. Theoretical results are established to guarantee the sparsity of the second-order KKT points of the -norm regularized models…
In this work, we present a method to compute the Kantorovich-Wasserstein distance of order one between a pair of two-dimensional histograms. Recent works in Computer Vision and Machine Learning have shown the benefits of measuring Wasserstein distances of order one between histograms with bins, by solving a classic…