The CSA-ES is an Evolution Strategy with Cumulative Step size Adaptation, where the step size is adapted measuring the length of a so-called cumulative path. The cumulative path is a combination of the previous steps realized by the algorithm, where the importance of each step decreases with time. This article studies …
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper provides a block coordinate descent algorithm to solve unconstrained optimization problems. In our algorithm, computation of function values or gradients is not required. Instead, pairwise comparison of function values is used. Our algorithm consists of two steps; one is the direction estimate step and the o…
Researchers describe Casimir functions for 3- and 4-step nilpotent Lie groups.
Random Function Descent improves optimization in high dimensions.
Paper develops an online learning algorithm for functional data models.
Polyak step size GD reaches final radius of convergence after log iterations.
FHBI enhances generalization in Bayesian inference with iterative steps in functional spaces.
Exchangeable graphs arise via a sampling procedure from measurable functions known as graphons. A natural estimation problem is how well we can recover a graphon given a single graph sampled from it. One general framework for estimating a graphon uses step-functions obtained by partitioning the nodes of the graph accor…
The Douglas Rachford algorithm is an algorithm that converges to a minimizer of a sum of two convex functions. The algorithm consists in fixed point iterations involving computations of the proximity operators of the two functions separately. The paper investigates a stochastic version of the algorithm where both funct…
As a popular meta-learning approach, the model-agnostic meta-learning (MAML) algorithm has been widely used due to its simplicity and effectiveness. However, the convergence of the general multi-step MAML still remains unexplored. In this paper, we develop a new theoretical framework to provide such convergence guarant…
Efficiently optimizes expensive functions with multi-step lookahead using one-shot optimization.
A new machine learning method for Bayesian inverse problems in function spaces.
Proposes an exponentially increasing step-size for faster parameter estimation in statistical models.
New RL method learns K-step lookahead Q-functions for fixed-horizon MDPs.
Stochastic algorithm achieves sublinear convergence for bi-objective optimization.
A new Bayesian method optimizes time-dependent expensive functions with lookahead.
We propose a mini-batching scheme for improving the theoretical complexity and practical performance of semi-stochastic gradient descent applied to the problem of minimizing a strongly convex composite function represented as the sum of an average of a large number of smooth convex functions, and simple nonsmooth conve…
The main goal of this work is equipping convex and nonconvex problems with Barzilai-Borwein (BB) step size. With the adaptivity of BB step sizes granted, they can fail when the objective function is not strongly convex. To overcome this challenge, the key idea here is to bridge (non)convex problems and strongly convex …
We construct a deep portfolio theory. By building on Markowitz's classic risk-return trade-off, we develop a self-contained four-step routine of encode, calibrate, validate and verify to formulate an automated and general portfolio selection process. At the heart of our algorithm are deep hierarchical compositions of p…
The paper introduces a multi-step loss function to improve model-based reinforcement learning.
The purpose of this paper is to present the first continuous families of Riemannian manifolds isospectral on functions but not on 1-forms, and simultaneously, the first continuous families of Riemannian manifolds with the same marked length spectrum but not the same 1-form spectrum. The examples presented here are Riem…
The paper tackles noisy combinations of continuous and step functions, providing conditions for their identification.
Stochastic Gradient Descent (SGD) is a popular tool in training large-scale machine learning models. Its performance, however, is highly variable, depending crucially on the choice of the step sizes. Accordingly, a variety of strategies for tuning the step sizes have been proposed, ranging from coordinate-wise approach…
Stochastic (sub)gradient methods require step size schedule tuning to perform well in practice. Classical tuning strategies decay the step size polynomially and lead to optimal sublinear rates on (strongly) convex problems. An alternative schedule, popular in nonconvex optimization, is called \emph{geometric step decay…
Proposes a neural network for learning step-size policies for L-BFGS optimization.
Method learns graphons from graphs via Gromov-Wasserstein barycenters.
In order to decode the human brain, Multivariate Pattern (MVP) classification generates cognitive models by using functional Magnetic Resonance Imaging (fMRI) datasets. As a standard pipeline in the MVP analysis, brain patterns in multi-subject fMRI dataset must be mapped to a shared space and then a classification mod…
W-Flow generates images in one step, faster and better than multi-step methods.
We propose mS2GD: a method incorporating a mini-batching scheme for improving the theoretical complexity and practical performance of semi-stochastic gradient descent (S2GD). We consider the problem of minimizing a strongly convex function represented as the sum of an average of a large number of smooth convex function…
Paper uses RL to optimize daily step distribution for better health biomarkers.
New tests for VaR and ES forecast encompassing using flexible link functions.
A new method automatically and dynamically sets learning rates in deep learning.
Proposes SD-KDE for density estimation using debiased kernel density with score-based adjustments.
The paper extends risk measures to two-step approximations and studies log-concave distributions.
Community detection in hypergraphs is explored. Under a generative hypergraph model called "d-wise hypergraph stochastic block model" (d-hSBM) which naturally extends the Stochastic Block Model from graphs to d-uniform hypergraphs, the asymptotic minimax mismatch ratio is characterized. For proving the achievability, w…
Method learns neural network to overestimate reference function with guarantees.
The signature function of a knot is an integer-valued step function on the unit circle in the complex plane. Necessary and sufficient conditions for a function to be the signature function of a knot are presented.
We propose a novel Bayesian approach to solve stochastic optimization problems that involve finding extrema of noisy, nonlinear functions. Previous work has focused on representing possible functions explicitly, which leads to a two-step procedure of first, doing inference over the function space and second, finding th…
The practical performance of online stochastic gradient descent algorithms is highly dependent on the chosen step size, which must be tediously hand-tuned in many applications. The same is true for more advanced variants of stochastic gradients, such as SAGA, SVRG, or AdaGrad. Here we propose to adapt the step size by …
Model-based reinforcement learning is an appealing framework for creating agents that learn, plan, and act in sequential environments. Model-based algorithms typically involve learning a transition model that takes a state and an action and outputs the next state---a one-step model. This model can be composed with itse…
Introduces a new length functional for Ricci flow to detect steady solitons.
New -step policy gradient method avoids local optima in restricted policy classes.
The paper develops methods for causal function estimation and inference with multiway clustered data.
Three-hidden-layer neural networks can approximate Hölder continuous functions uniformly with exponential rate.
Estimates variance function using aggregation methods in regression models.
New embedding method in function spaces improves expressiveness.
New adaptive step-size method for convex optimization without tuning.
We consider the dynamics of a linear stochastic approximation algorithm driven by Markovian noise, and derive finite-time bounds on the moments of the error, i.e., deviation of the output of the algorithm from the equilibrium point of an associated ordinary differential equation (ODE). We obtain finite-time bounds on t…