Conditional gradients constitute a class of projection-free first-order algorithms for smooth convex optimization. As such, they are frequently used in solving smooth convex optimization problems over polytopes, for which the computational cost of orthogonal projections would be prohibitive. However, they do not enjoy …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Develops accelerated methods for optimization using low-dimensional projected-gradient information.
Paper solves robust multi-dimensional scaling with accelerated projections.
Sinh-acceleration speeds up B-spline option pricing.
PF-LaCG removes the need for knowing smoothness and strong convexity parameters for locally accelerated CG.
New projection techniques reduce the frequency of projections in solving LCPs.
Accelerates Birkhoff projection for manifold-constrained hyper-connections with high accuracy and speed.
Accelerated optimization methods improve robustness and privacy in estimation.
Recently, deep neural networks (DNNs) have shown advantages in accelerating optimization algorithms. One approach is to unfold finite number of iterations of conventional optimization algorithms and to learn parameters in the algorithms. However, these are forward methods and are indeed neither iterative nor convergent…
ADASAP accelerates GP inference for large datasets.
In this work we introduce a conditional accelerated lazy stochastic gradient descent algorithm with optimal number of calls to a stochastic first-order oracle and convergence rate improving over the projection-free, Online Frank-Wolfe based stochastic gradient descent of Hazan an…
GPU-accelerates multiuser detection for 5G URLLC systems.
New algorithms accelerate model-based optimization for stochastic problems.
Paper develops MMOT framework for financial applications with neural acceleration.
Kaczmarz++ accelerates convergence for ill-conditioned systems.
LightOn OPUs accelerate randomized numerical linear algebra, reducing computational costs.
Improved neural network training in low-dimensional random bases.
We develop a projected Nesterov's proximal-gradient (PNPG) approach for sparse signal reconstruction that combines adaptive step size with Nesterov's momentum acceleration. The objective function that we wish to minimize is the sum of a convex differentiable data-fidelity (negative log-likelihood (NLL)) term and a conv…
Sketching techniques have become popular for scaling up machine learning algorithms by reducing the sample size or dimensionality of massive data sets, while still maintaining the statistical power of big data. In this paper, we study sketching from an optimization point of view: we first show that the iterative Hessia…
Two Kähler metrics on a complex manifold are called c-projectively equivalent if their -planar curves coincide. These curves are defined by the property that the acceleration is complex proportional to the velocity. We give an explicit local description of all pairs of c-projectively equivalent Kähler metrics of arb…
We study --both in theory and practice-- the use of momentum motions in classic iterative hard thresholding (IHT) methods. By simply modifying plain IHT, we investigate its convergence behavior on convex optimization criteria with non-convex constraints, under standard assumptions. In diverse scenaria, we observe that …
UMAP speeds up significantly with GPU acceleration.
Two Kaehler metrics on one complex manifold are said to be c-projectively equivalent if their J-planar curves, i.e., curves defined by the property that their acceleration is complex proportional to their velocity, coincide. The degree of mobility of a Kaehler metric is the dimension of the space of metrics that are c-…
Skeinformer accelerates self-attention for long sequences with linear complexity.
New algorithms optimize constrained problems faster, avoiding full set optimization.
Momentum methods such as Polyak's heavy ball (HB) method, Nesterov's accelerated gradient (AG) as well as accelerated projected gradient (APG) method have been commonly used in machine learning practice, but their performance is quite sensitive to noise in the gradients. We study these methods under a first-order stoch…
GPU acceleration speeds up financial machine learning training time.
GRU models with Adam optimizer outperform other combinations in stock market forecasting.
We formalize the problem of trading-off DNN training time and memory requirements as the tensor rematerialization optimization problem, a generalization of prior checkpointing strategies. We introduce Checkmate, a system that solves for optimal rematerialization schedules in reasonable times (under an hour) using off-t…
Paper proposes an efficient online Newton method with Nesterov's acceleration for streaming data.
Efficient kernel methods for large datasets using GPU acceleration.
With the rise of self-driving vehicles comes the risk of accidents and the need for higher safety, and protection for pedestrian detection in the following scenarios: imminent crashes, thus the car should crash into an object and avoid the pedestrian, and in the case of road intersections, where it is important for the…
This work accelerates constrained sampling using large deviation principles.
Random Fourier features improve tabular deep learning convergence.
New method accelerates large margin metric learning for nearest neighbor classification.
Optical co-processor speeds up neural network training.
Random projections are able to perform dimension reduction efficiently for datasets with nonlinear low-dimensional structures. One well-known example is that random matrices embed sparse vectors into a low-dimensional subspace nearly isometrically, known as the restricted isometric property in compressed sensing. In th…
The study characterizes spacetime and modified gravity models using projective curvature tensor.
Network embedding, which learns low-dimensional vector representation for nodes in the network, has attracted considerable research attention recently. However, the existing methods are incapable of handling billion-scale networks, because they are computationally expensive and, at the same time, difficult to be accele…
We propose a generic framework based on a new stochastic variance-reduced gradient descent algorithm for accelerating nonconvex low-rank matrix recovery. Starting from an appropriate initial estimator, our proposed algorithm performs projected gradient descent based on a novel semi-stochastic gradient specifically desi…
On-device inference of machine learning models for mobile phones is desirable due to its lower latency and increased privacy. Running such a compute-intensive task solely on the mobile CPU, however, can be difficult due to limited computing power, thermal constraints, and energy consumption. App developers and research…
New analysis shows SNG's effectiveness in small samples.
The problem of minimizing a continuously differentiable convex function over an intersection of closed convex sets is ubiquitous in applied mathematics. It is particularly interesting when it is easy to project onto each separate set, but nontrivial to project onto their intersection. Algorithms based on Newton's metho…
Learning big data by matrix decomposition always suffers from expensive computation, mixing of complicated structures and noise. In this paper, we study more adaptive models and efficient algorithms that decompose a data matrix as the sum of semantic components with incoherent structures. We firstly introduce "GO decom…
A new method reduces the complexity of tensor products from cubic to quadratic, improving both speed and accuracy.
The problem of joint feature selection across a group of related tasks has applications in many areas including biomedical informatics and computer vision. We consider the l2,1-norm regularized regression model for joint feature selection from multiple tasks, which can be derived in the probabilistic framework by assum…
Accelerates Riemannian gradient methods with extrapolation.
This paper advances extragradient methods for solving inclusions under co-hypomonotonicity.