A new method automatically and dynamically sets learning rates in deep learning.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New method for CMS derivatives pricing using Watanabe's expansions.
The paper debiases mini-batch approximations in deep learning for more accurate optimization and uncertainty quantification.
The paper tackles learning smooth distance functions using query-based methods.
Gradients help find global optima in complex functions.
We propose a DC proximal Newton algorithm for solving nonconvex regularized sparse learning problems in high dimensions. Our proposed algorithm integrates the proximal Newton algorithm with multi-stage convex relaxation based on the difference of convex (DC) programming, and enjoys both strong computational and statist…
HA-SME models SGD dynamics with Hessian info for better escaping behaviors.
The DANE algorithm is an approximate Newton method popularly used for communication-efficient distributed machine learning. Reasons for the interest in DANE include scalability and versatility. Convergence of DANE, however, can be tricky; its appealing convergence rate is only rigorous for quadratic objective, and for …
AQFC method estimates mesh curvatures using quadratic surfaces.
Most of machine learning approaches have stemmed from the application of minimizing the mean squared distance principle, based on the computationally efficient quadratic optimization methods. However, when faced with high-dimensional and noisy data, the quadratic error functionals demonstrated many weaknesses including…
We derive caplet volatilities for quadratic models, providing an asymptotic approximation.
Given a negatively curved geodesic metric space , we study the statistical asymptotic penetration behavior of (locally) geodesic lines of in small neighborhoods of points, of closed geodesics, and of other compact (locally) convex subsets of . We prove Khintchine-type and logarithme law-type results for the s…
Quadratic points of a surface in the projective 3-space are the points which can be exceptionally well approximated by a quadric. They are also singularities of a 3-web in the elliptic part and of a line field in the hyperbolic part of the surface. We show that generically the index of the 3-web at a quadratic point is…
A quadratic point on a surface in is a point at which the surface can be approximated by a quadric abnormally well (up to order 3). We conjecture that the least number of quadratic points on a generic compact non-degenerate hyperbolic surface is 8; the relation between this and the classic Carathéodory conjectur…
We propose HAMSI (Hessian Approximated Multiple Subsets Iteration), which is a provably convergent, second order incremental algorithm for solving large-scale partially separable optimization problems. The algorithm is based on a local quadratic approximation, and hence, allows incorporating curvature information to sp…
Recently, deep learning has achieved huge successes in many important applications. In our previous studies, we proposed quadratic/second-order neurons and deep quadratic neural networks. In a quadratic neuron, the inner product of a vector of data and the corresponding weights in a conventional neuron is replaced with…
Complex-valued neural networks avoid spurious local minima.
This paper concerns a method of selecting a subset of features for a sequential logit model. Tanaka and Nakagawa (2014) proposed a mixed integer quadratic optimization formulation for solving the problem based on a quadratic approximation of the logistic loss function. However, since there is a significant gap between …
The paper extends a variance gamma model to quadratic functions, reducing arbitrage and computational costs.
A novel optimisation framework through quadratic nonlinear projection is introduced for credit portfolio when the portfolio risk is measured by Conditional Value-at-Risk (CVaR). The whole optimisation procedure to search toward the optimal portfolio state is conducted by a series of single-step optimisations under the …
Recently several methods were proposed for sparse optimization which make careful use of second-order information [10, 28, 16, 3] to improve local convergence rates. These methods construct a composite quadratic approximation using Hessian information, optimize this approximation using a first-order method, such as coo…
This paper shows how to create quadratic differentials with any given singularities.
Detects adversarial directions to make reinforcement learning policies more robust.
We propose a strategy for approximating Pareto optimal sets based on the global analysis framework proposed by Smale (Dynamical systems, New York, 1973, pp. 531-544). The method highlights and exploits the underlying manifold structure of the Pareto sets, approximating Pareto optima by means of simplicial complexes. Th…
Closed-form polynomial approximations replace MLPs in transformers, enabling new interpretability methods.
Semidefinite programs (SDP) are important in learning and combinatorial optimization with numerous applications. In pursuit of low-rank solutions and low complexity algorithms, we consider the Burer--Monteiro factorization approach for solving SDPs. We show that all approximate local optima are global optima for the pe…
Developed a theory of local convexity for second order differential equations on Lie algebroids.
We consider the problem of solving a large-scale Quadratically Constrained Quadratic Program. Such problems occur naturally in many scientific and web applications. Although there are efficient methods which tackle this problem, they are mostly not scalable. In this paper, we develop a method that transforms the quadra…
Derivative-free method solves stochastic optimization problems with noisy objectives and constraints.
Given a loss function that can be written as the sum of losses over a large set of inputs , it is often desirable to approximate by subsampling the input points. Strong theoretical guarantees require taking into account the importance of each point, measured by how …
Novel approximation hierarchy for sparse quadratic programs.
Four decades after their invention, quasi-Newton methods are still state of the art in unconstrained numerical optimization. Although not usually interpreted thus, these are learning algorithms that fit a local quadratic approximation to the objective function. We show that many, including the most popular, quasi-Newto…
The paper proves signatures of non-geometric rough paths can approximate functionals uniformly.
We present local discriminative Gaussian (LDG) dimensionality reduction, a supervised dimensionality reduction technique for classification. The LDG objective function is an approximation to the leave-one-out training error of a local quadratic discriminant analysis classifier, and thus acts locally to each training po…
Two log-linear approximations speed up optimal transport for deep learning applications.
Proves curvature of conference graphs and finds local matchings.
A new method for exponentially weighted moving models using approximations.
QLA improves Bayesian uncertainty estimation for DNNs without increasing computational cost.
In this paper we consider regularized convex cone programming problems. In particular, we first propose an iterative hard thresholding (IHT) method and its variant for solving regularized box constrained convex programming. We show that the sequence generated by these methods converges to a local minimizer.…
The runtime for Kernel Partial Least Squares (KPLS) to compute the fit is quadratic in the number of examples. However, the necessity of obtaining sensitivity measures as degrees of freedom for model selection or confidence intervals for more detailed analysis requires cubic runtime, and thus constitutes a computationa…
Proposes a new framework for invariant quadratic P&L predictions in option books.
Local search algorithms applied to optimization problems often suffer from getting trapped in a local optimum. The common solution for this deficiency is to restart the algorithm when no progress is observed. Alternatively, one can start multiple instances of a local search algorithm, and allocate computational resourc…
Method identifies shifts leading to large model performance differences.
GEORCE computes geodesics quickly and accurately.
In this paper, we propose a simple, fast and easy to implement algorithm LOSSGRAD (locally optimal step-size in gradient descent), which automatically modifies the step-size in gradient descent during neural networks training. Given a function , a point , and the gradient of , we aim to find the s…
QMME balances cost and speed in convex optimization.
Market maker optimizes SPX and VIX spread using quadratic rough Heston model.
Let be a pinched negatively curved Riemannian manifold, whose unit tangent bundle is endowed with a Gibbs measure associated to a potential . We compute the Hausdorff dimension of the conditional measures of . We study the -almost sure asymptotic penetration behaviour of locally geodesic lines of…