Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

55111166221 · Jun 202019922001200920172026
48 results for support norm

The kk-support norm is a regularizer which has been successfully applied to sparse vector prediction problems. We show that it belongs to a general class of norms which can be formulated as a parameterized infimum over quadratics. We further extend the kk-support norm to matrices, and we observe that it is a special …

2014-03-06abs ↗pdf ↗

We study a regularizer which is defined as a parameterized infimum of quadratics, and which we call the box-norm. We show that the k-support norm, a regularizer proposed by [Argyriou et al, 2012] for sparse vector prediction problems, belongs to this family, and the box-norm can be generated as a perturbation of the fo…

2015-12-27abs ↗pdf ↗

The spectral kk-support norm enjoys good estimation properties in low rank matrix learning problems, empirically outperforming the trace norm. Its unit ball is the convex hull of rank kk matrices with unit Frobenius norm. In this paper we generalize the norm to the spectral (k,p)(k,p)-support norm, whose additional para…

2016-01-04abs ↗pdf ↗

Support Vector Machine (SVM) is an efficient classification approach, which finds a hyperplane to separate data from different classes. This hyperplane is determined by support vectors. In existing SVM formulations, the objective function uses L2 norm or L1 norm on slack variables. The number of support vectors is a me…

2018-04-06abs ↗pdf ↗

We derive a novel norm that corresponds to the tightest convex relaxation of sparsity combined with an 2\ell_2 penalty. We show that this new {\em kk-support norm} provides a tighter relaxation than the elastic net and is thus a good replacement for the Lasso or the elastic net in sparse prediction problems. Through …

2012-04-23abs ↗pdf ↗

We propose a Generalized Dantzig Selector (GDS) for linear models, in which any norm encoding the parameter structure can be leveraged for estimation. We investigate both computational and statistical aspects of the GDS. Based on conjugate proximal operator, a flexible inexact ADMM framework is designed for solving GDS…

2014-06-20abs ↗pdf ↗

Proposes a new SVM model for binary classification with theoretical and practical advantages.

problem Binary classification in supervised learning.
method Quadratic surface support vector machine with L1 norm regularization.
result The model can detect true sparsity patterns and is efficient for both synthetic and real data.

IRKSN algorithm achieves sparse recovery with wider applicability conditions.

problem Sparse recovery challenges due to NP-hard nature and restrictive conditions.
method IRKSN algorithm based on kk-support norm regularizer.
result Achieves sparse recovery with explicit constants and standard linear rate.

In this paper, we propose an unifying view of several recently proposed structured sparsity-inducing norms. We consider the situation of a model simultaneously (a) penalized by a set- function de ned on the support of the unknown parameter vector which represents prior knowledge on supports, and (b) regularized in Lp-n…

2012-05-06abs ↗pdf ↗

We describe Milnor open books and Legendrian surgery diagrams for canonical contact structures of links of some rational surface singularities. We also describe an infinite family of Milnor fillable contact 3-manifolds so that the Milnor genus (resp. Milnor norm) is strictly greater than the support genus (resp. suppor…

2009-12-21abs ↗pdf ↗

Proves singular support of sheaves is γ-coisotropic, with implications for symplectic homeomorphisms.

problem Understanding the singular support of sheaves and its properties.
method Proves γ-coisotropic property and relates it to spectral norm.
result Singular support of sheaves is γ-coisotropic, with invariance under symplectic homeomorphisms.

Paper introduces new risk norms based on ES with flexible distortion functions.

problem Risk quantification and anomaly detection in financial data.
method Developed generalized Expected-Shortfall (ES) norms using distortion risk measures and duality theory.
result Unified analytical framework for risk quantification and practical applications.

Sparse methods for supervised learning aim at finding good linear predictors from as few variables as possible, i.e., with small cardinality of their supports. This combinatorial selection problem is often turned into a convex optimization problem by replacing the cardinality function by its convex envelope (tightest c…

2010-08-25abs ↗pdf ↗

In this note we define three invariants of contact structures in terms of open books supporting the contact structures. These invariants are the support genus (which is the minimal genus of a page of a supporting open book for the contact structure), the binding number (which is the minimal number of binding components…

2006-05-16abs ↗pdf ↗

Adversarial training linked to operator norm regularization, proving network sensitivity to attacks.

problem Robustifying neural networks against adversarial attacks.
method Theoretical link established between adversarial training and operator norm regularization.
result Adversarial training is equivalent to data-dependent operator norm regularization.

Improved greedy 2-coordinate updates for optimization problems with constraints.

problem Minimizing smooth functions subject to constraints.
method Exploiting a connection to steepest descent in the 1-norm, we give faster convergence rates and efficient computation.
result Greedy selection converges faster than random selection and can be computed in O(nlogn)O(n \log n) time.

Gradient flow on softmax attention minimizes nuclear norm of weight matrices.

problem Classification with separate key and query weight matrices.
method Gradient flow on exponential loss, separability assumption, reparameterization, approximate KKT conditions.
result Gradient flow implicitly minimizes nuclear norm of weight matrices, contrasting with Frobenius norm minimization.

Geodesics of contactomorphisms on a specific manifold are characterized by Hamiltonian functions.

problem Characterizing geodesics of contactomorphisms on a manifold with a standard contact structure.
method Analyzing geodesics defined by different norms on the identity component of the group of contactomorphisms.
result The norm of a geodesic contactomorphism can be expressed in terms of the maximum of the Hamiltonian function.

The aim of this paper is to discuss some applications of the relation between Seiberg-Witten theory and two natural norms defined on the first cohomology group of a closed 3-manifold N - the Alexander and Thurston norms. We start by giving a "new" proof of McMullen's inequality between these norms, and then use these n…

2002-04-16abs ↗pdf ↗

Deep neural networks with adversarial training achieve sup-norm convergence for nonparametric regression.

problem Achieving sup-norm convergence for deep neural network estimators in nonparametric regression.
method Developed an adversarial training scheme to address the sup-norm convergence issue.
result Deep neural network estimators achieve optimal sup-norm convergence with the proposed adversarial training.

Learning linear combinations of multiple kernels is an appealing strategy when the right choice of features is unknown. Previous approaches to multiple kernel learning (MKL) promote sparse kernel combinations to support interpretability and scalability. Unfortunately, this 1-norm MKL is rarely observed to outperform tr…

2010-02-27abs ↗pdf ↗

The higher order singular value decomposition (HOSVD) of tensors is a generalization of matrix SVD. The perturbation analysis of HOSVD under random noise is more delicate than its matrix counterpart. Recently, polynomial time algorithms have been proposed where statistically optimal estimates of the singular subspaces …

2017-07-05abs ↗pdf ↗

It is shown that the compactly supported identity component of the diffeomorphism group of the 2-dimensional punctured torus Tp2\mathbb T^2_p is an unbounded group. It follows that the fragmentation norm of Tp2\mathbb T^2_p is unbounded.

2011-03-18abs ↗pdf ↗

We study the relationship between geometry and capacity measures for deep neural networks from an invariance viewpoint. We introduce a new notion of capacity --- the Fisher-Rao norm --- that possesses desirable invariance properties and is motivated by Information Geometry. We discover an analytical characterization of…

2017-11-05abs ↗pdf ↗

New evidence supports the Euler class one conjecture for tight contact structures.

problem Euler class one conjecture for taut foliations and tight contact structures.
method Analysis of tight contact structures and counterexamples to the conjecture.
result Counterexamples to the Euler class one conjecture for taut foliations are also Euler classes of tight contact structures.

Neyshabur and Srebro proposed Simple-LSH, which is the state-of-the-art hashing method for maximum inner product search (MIPS) with performance guarantee. We found that the performance of Simple-LSH, in both theory and practice, suffers from long tails in the 2-norm distribution of real datasets. We propose Norm-rangin…

2018-09-24abs ↗pdf ↗

Given a normed plane P\mathcal{P}, we call P\mathcal{P}-cycloids the planar curves which are homothetic to their double P\mathcal{P}-evolutes. It turns out that the radius of curvature and the support function of a P\mathcal{P}-cycloid satisfy a differential equation of Sturm-Liouville type. By studying this equati…

2016-08-04abs ↗pdf ↗

Optimal scaling found to depend on operator norm across large models and datasets.

problem Lack of unifying principle for optimal hyperparameter scaling across models and datasets.
method Discovered that optimal scaling is conditioned on the operator norm of the output layer.
result The optimal learning rate/batch size pair (η,B)(η^{\ast}, B^{\ast}) consistently has the same operator norm value.

Study shows unusual non-monotonic risk behavior in minimum-norm interpolants for various data scaling.

problem Understanding the risk behavior of minimum-norm interpolants in RKHS for different data scaling.
method Analysis of spectral properties of the random kernel matrix restricted to eigen-spaces of the population covariance operator.
result Minimum-norm interpolants in RKHS exhibit multiple descent in risk for d=nαd = n^α with α(0,1)α\in(0,1).

AdaGrad-Norm achieves optimal convergence rates for non-convex objectives without tuning.

problem Optimal convergence rates for non-convex, smooth objectives with adaptive step sizes.
method Adaptive SGD (AdaGrad-Norm) with self-tuning step sizes, analyzing under unbounded gradients and affine variance scaling.
result AdaGrad-Norm achieves order optimal convergence rate of $\mathcal{O}\left(\frac{\mathrm{poly}\log(T)}{\sqrt{T}} ight)$ under optimal assumptions.