Deep networks learn sparse hierarchical features without CoD.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Unified algorithm for minimizing composite functions with flexible design.
The paper accelerates ISTA and FISTA algorithms for composite optimization problems.
Low-rank inducing unitarily invariant norms have been introduced to convexify problems with low-rank/sparsity constraint. They are the convex envelope of a unitary invariant norm and the indicator function of an upper bounding rank constraint. The most well-known member of this family is the so-called nuclear norm. To …
New method selects variables in groups with few nonzeros, improving support recovery.
DNNs can learn complex functions efficiently by breaking the curse of dimensionality.
We consider multi-level composite optimization problems where each mapping in the composition is the expectation over a family of random smooth mappings or the sum of some finite number of smooth mappings. We present a normalized proximal approximate gradient (NPAG) method where the approximate gradients are obtained v…
Training neural networks under a strict Lipschitz constraint is useful for provable adversarial robustness, generalization bounds, interpretable gradients, and Wasserstein distance estimation. By the composition property of Lipschitz functions, it suffices to ensure that each individual affine transformation or nonline…
Efficiently regularizes deep learning models using Jacobian nuclear norm.
Study excess capacity in neural networks using Rademacher complexity.
We introduce and study deformation of Minkowski norms in , determined by a set of linearly independent 1-forms and a smooth positive function of variables. In particular, the -image of a Euclidean norm is a Minkowski norm, whose indicat…
We study the Berezin-Toeplitz quantization on Kaehler manifolds. We explain first how to compute various associated asymptotic expansions, then we compute explicitly the first terms of the expansion of the kernel of the Berezin-Toeplitz operators, and of the composition of two Berezin-Toeplitz operators. As application…
In many applications one may acquire a composition of several signals that may be corrupted by noise, and it is a challenging problem to reliably separate the components from one another without sacrificing significant details. Adding to the challenge, in a compressive sensing framework, one is given only an undersampl…
We consider a group of mean-variance investors with mimicking desire such that each investor is willing to penalize deviations of his portfolio composition from compositions of other group members. Penalizing norm constraints are already applied for statistical improvement of Markowitz portfolio procedure in order to c…
This work addresses the convergence of SGD's final iterate without restrictive assumptions.
This paper advances FL algorithms for composite optimization and statistical recovery.
Paper analyzes convergence of PAM method for low-rank factorization models.
In this paper we develop proximal methods for statistical learning. Proximal point algorithms are useful in statistics and machine learning for obtaining optimization solutions for composite functions. Our approach exploits closed-form solutions of proximal operators and envelope representations based on the Moreau, Fo…
We study a hybrid conditional gradient - smoothing algorithm (HCGS) for solving composite convex optimization problems which contain several terms over a bounded set. Examples of these include regularization problems with several norms as penalties and a norm constraint. HCGS extends conditional gradient methods to cas…
The adaptive gradient online learning method known as AdaGrad has seen widespread use in the machine learning community in stochastic and adversarial online learning problems and more recently in deep learning methods. The method's full-matrix incarnation offers much better theoretical guarantees and potentially better…
We show that deep networks are better than shallow networks at approximating functions that can be expressed as a composition of functions described by a directed acyclic graph, because the deep networks can be designed to have the same compositional structure, while a shallow network cannot exploit this knowledge. Thu…
We consider the stochastic nested composition optimization problem where the objective is a composition of two expected-value functions. We proposed the stochastic ADMM to solve this complicated objective. In order to find an stationary point where the expected norm of the subgradient of corresponding augmented Lag…
Unified analysis of multi-task functional linear regression with manifold and composite penalties.
Using sparse-inducing norms to learn robust models has received increasing attention from many fields for its attractive properties. Projection-based methods have been widely applied to learning tasks constrained by such norms. As a key building block of these methods, an efficient operator for Euclidean projection ont…
New framework explains deep neural networks using variational spline theory.
Develops minibatch stochastic proximal gradient for large-scale learning models.
Paper shows linear convergence of ISTA and FISTA for ill-conditioned images.
We study the problem of estimating multiple predictive functions from a dictionary of basis functions in the nonparametric regression setting. Our estimation scheme assumes that each predictive function can be estimated in the form of a linear combination of the basis functions. By assuming that the coefficient matrix …
Neural networks outperform NTK on compositional tasks, revealing a complexity gap.
Existing Rademacher complexity bounds for neural networks rely only on norm control of the weight matrices and depend exponentially on depth via a product of the matrix norms. Lower bounds show that this exponential dependence on depth is unavoidable when no additional properties of the training data are considered. We…
The paper shows how Hamiltonian diffeomorphisms and homeomorphisms can be broken down into smaller, manageable pieces.
This paper is concerned with the factorization form of the rank regularized loss minimization problem. To cater for the scenario in which only a coarse estimation is available for the rank of the true matrix, an -norm regularized term is added to the factored loss function to reduce the rank adaptively; and…
Paper studies signal detection in noisy environments with limited communication.
The paper studies Minkowski norms and Hessian isometries induced by isoparametric foliations on spheres.
Lipschitz constraints under L2 norm on deep neural networks are useful for provable adversarial robustness bounds, stable training, and Wasserstein distance estimation. While heuristic approaches such as the gradient penalty have seen much practical success, it is challenging to achieve similar practical performance wh…
DoRA improves adaptation efficiency for large models by factoring norms and fusing kernels.
Sparse coding consists in representing signals as sparse linear combinations of atoms selected from a dictionary. We consider an extension of this framework where the atoms are further assumed to be embedded in a tree. This is achieved using a recently introduced tree-structured sparse regularization norm, which has pr…
Study shows LLMs can extrapolate rules from out-of-distribution prompts.
In this paper we consider the problem of minimizing composite objective functions consisting of a convex differentiable loss function plus a non-smooth regularization term, such as norm or nuclear norm, under Rényi differential privacy (RDP). To solve the problem, we propose two stochastic alternating direction m…
New variance-reduction methods solve stochastic composite inclusions.
The paper analyzes and improves a deep learning optimization technique using matrix gradient orthogonality.
Data augmentation improves robustness in adversarial training.
We present a method of rank-optimal weighting which can be used to explore the best possible position of a subject in a ranking based on a composite indicator by means of a mathematical optimization problem. As an example, we explore the dataset of the OECD Better Life Index and compute for each country a weight vector…
This work is an analytical and numerical study of the composition of several fractals into one and of the relation between the composite dimension and the dimensions of the component fractals. In the case of composition of standard IFS with segments of equal size, the composite dimension can be expressed as a function …
Kähler information manifolds for signal filters in weighted Hardy spaces are explored.
New geometric approach for analyzing compositional data like gut microbiomes.
In this paper we develop a randomized block-coordinate descent method for minimizing the sum of a smooth and a simple nonsmooth block-separable convex function and prove that it obtains an -accurate solution with probability at least in at most iterations, where is the numbe…
Study on deep neural networks using branching processes and Mehler's formula.