New algorithms optimize non-smooth, non-convex objectives with improved complexity.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
AsylADMM improves gossip-based learning for non-smooth objectives.
We consider the problem of finding local minimizers in non-convex and non-smooth optimization. Under the assumption of strict saddle points, positive results have been derived for first-order methods. We present the first known results for the non-smooth case, which requires different analysis and a different algorithm…
New method tackles non-smooth tensor data for better recovery.
Adaptive data fusion boosts efficiency in multi-task optimization.
New algorithms optimize spectral risk measures, improving interpolation between average and worst-case performance.
Stochastic Gradient Descent (SGD) is one of the simplest and most popular stochastic optimization methods. While it has already been theoretically studied for decades, the classical analysis usually required non-trivial smoothness assumptions, which do not apply to many modern applications of SGD with non-smooth object…
In this paper, we develop a novel {\bf ho}moto{\bf p}y {\bf s}moothing (HOPS) algorithm for solving a family of non-smooth problems that is composed of a non-smooth term with an explicit max-structure and a smooth term or a simple non-smooth term whose proximal mapping is easy to compute. The best known iteration compl…
Stochastic gradient descent's long-term fluctuations are described by a diffusion limit.
Safe-EF improves federated learning for non-smooth, constrained optimization.
Improves robustness of high-dimensional regression with rank objective and group lasso regularization.
This paper proposes a novel proximal-gradient algorithm for a decentralized optimization problem with a composite objective containing smooth and non-smooth terms. Specifically, the smooth and nonsmooth terms are dealt with by gradient and proximal updates, respectively. The proposed algorithm is closely related to a p…
Expanding FCCO to non-smooth weakly-convex problems, improving deep learning performance.
Novel method for shape optimization of non-smooth PDEs.
We propose a new proximal, path-following framework for a class of constrained convex problems. We consider settings where the nonlinear---and possibly non-smooth---objective part is endowed with a proximity operator, and the constraint set is equipped with a self-concordant barrier. Our approach relies on the followin…
Variable projection solves structured optimization problems by completely minimizing over a subset of the variables while iterating over the remaining variables. Over the last 30 years, the technique has been widely used, with empirical and theoretical results demonstrating both greater efficacy and greater stability c…
We study the dynamics of a particle in a space that is non-differentiable. Non-smooth geometrical objects have an inherently probabilistic nature and, consequently, introduce stochasticity in the motion of a body that lives in their realm. We use the mathematical concept of fiber bundle to characterize the multivalued …
New method smooths optimization for sparse regularization.
In this paper, we revisit the problem of private stochastic convex optimization. We propose an algorithm based on noisy mirror descent, which achieves optimal rates both in terms of statistical complexity and number of queries to a first-order stochastic oracle in the regime when the privacy parameter is inversely prop…
Generative Adversarial Networks (GANs) are one of the most practical methods for learning data distributions. A popular GAN formulation is based on the use of Wasserstein distance as a metric between probability distributions. Unfortunately, minimizing the Wasserstein distance between the data distribution and the gene…
We investigate the theoretical limits of pipeline parallel learning of deep learning architectures, a distributed setup in which the computation is distributed per layer instead of per example. For smooth convex and non-convex objective functions, we provide matching lower and upper complexity bounds and show that a na…
The three operator splitting scheme was recently proposed by [Davis and Yin, 2015] as a method to optimize composite objective functions with one convex smooth term and two convex (possibly non-smooth) terms for which we have access to their proximity operator. In this short note we provide an alternative proof for the…
MARINA-P improves non-smooth federated optimization with adaptive stepsizes.
Extends curve theory to non-smooth data with finite curvature and torsion.
Survey on preserving curvature bounds for non-smooth Ricci flow.
Paper analyzes convergence of stochastic methods under heavy-tailed noise.
We introduce non-smooth symplectic forms on manifolds and describe corresponding Poisson structures on the algebra of Colombeau generalized functions. This is achieved by establishing an extension of the classical map of smooth functions to Hamiltonian vector fields to the setting of non-smooth geometry. For mildly sin…
Cross-validation (CV) is a popular approach for assessing and selecting predictive models. However, when the number of folds is large, CV suffers from a need to repeatedly refit a learning procedure on a large number of training datasets. Recent work in empirical risk minimization (ERM) approximates the expensive refit…
SAPPHIRE tackles ill-conditioned rERM problems with faster convergence.
The paper explores various stationarity concepts in non-smooth optimization.
Smoothness analysis of adversarial training reveals constraints cause more non-smoothness.
Boosting is a popular way to derive powerful learners from simpler hypothesis classes. Following previous work (Mason et al., 1999; Friedman, 2000) on general boosting frameworks, we analyze gradient-based descent algorithms for boosting with respect to any convex objective and introduce a new measure of weak learner p…
In this paper we consider the problem of minimizing composite objective functions consisting of a convex differentiable loss function plus a non-smooth regularization term, such as norm or nuclear norm, under Rényi differential privacy (RDP). To solve the problem, we propose two stochastic alternating direction m…
Conditional stochastic optimization covers a variety of applications ranging from invariant learning and causal inference to meta-learning. However, constructing unbiased gradient estimators for such problems is challenging due to the composition structure. As an alternative, we propose a biased stochastic gradient des…
Stochastic algorithm achieves sublinear convergence for bi-objective optimization.
In this paper we develop proximal methods for statistical learning. Proximal point algorithms are useful in statistics and machine learning for obtaining optimization solutions for composite functions. Our approach exploits closed-form solutions of proximal operators and envelope representations based on the Moreau, Fo…
FeDualEx tackles saddle point optimization in federated learning with composite objectives.
In the framework of Lorentzian warped products, we study the Friedmann-Robertson-Walker cosmological model to investigate non-smooth curvatures associated with multiple discontinuities involved in the evolution of the universe. In particular we analyze non-smooth features of the spatially flat Friedmann-Robertson-Walke…
This work speeds up hyperparameter selection for non-smooth convex models using implicit differentiation.
In recent literature, a general two step procedure has been formulated for solving the problem of phase retrieval. First, a spectral technique is used to obtain a constant-error initial estimate, following which, the estimate is refined to arbitrary precision by first-order optimization of a non-convex loss function. N…
In spite of several notable efforts, explaining the generalization of deterministic non-smooth deep nets, e.g., ReLU-nets, has remained challenging. Existing approaches for deterministic non-smooth deep nets typically need to bound the Lipschitz constant of such deep nets but such bounds are quite large, may even incre…
New algorithm solves non-convex, non-differentiable min-max games.
Positive mass theorem for non-smooth metrics on flat manifolds with corners.
A new method for fast and robust sparsity learning over networks.
Synthetic framework for null hypersurfaces in non-smooth spacetimes.
We investigate a generalization of the so-called metric splitting of globally hyperbolic space-times to non-smooth Lorentzian manifolds and show the existence of this metric splitting for a class of wave-type space-times. Our approach is based on smooth approximations of non-smooth space-times by families (or sequences…
New SGD covering technique yields dimension-independent generalization bounds.
Advances smooth over-parameterization for solving non-smooth optimization problems.