New protocols show 1-bit mean estimation can be order-optimal without interaction.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Stochastic Gradient Langevin Dynamics infuses isotropic gradient noise to SGD to help navigate pathological curvature in the loss landscape for deep networks. Isotropic nature of the noise leads to poor scaling, and adaptive methods based on higher order curvature information such as Fisher Scoring have been proposed t…
Various factorization-based methods have been proposed to leverage second-order, or higher-order cross features for boosting the performance of predictive models. They generally enumerate all the cross features under a predefined maximum order, and then identify useful feature interactions through model training, which…
Adaptive market-making strategy improves profit by adjusting to order flow.
Study improves BN TTA under distribution shift using higher-order asymptotics.
In this paper, we presented a novel convolutional neural network framework for graph modeling, with the introduction of two new modules specially designed for graph-structured data: the -th order convolution operator and the adaptive filtering module. Importantly, our framework of High-order and Adaptive Graph Convo…
Improves adaptivity in sequence models by over-parameterizing.
New method for pricing options in stochastic volatility models.
Let S be a compact Riemann surfaces of genus g >= 2 and G a conformal automoprhism group of order n acting on S. In this paper we give the definition of an adapted generating set and an adapted basis for the first homology group of such a compact Riemann surface. This generating set and basis reflect the action of G in…
Gradient-based meta-learning has proven to be highly effective at learning model initializations, representations, and update rules that allow fast adaptation from a few samples. The core idea behind these approaches is to use fast adaptation and generalization -- two second-order metrics -- as training signals on a me…
Improved text-to-image and multimodal understanding through adaptive generation order optimization.
K-priors enable quick adaptation with minimal retraining.
In this paper we find a unique normal form for the symplectic matrix representation of the conjugacy class of a prime order element of the mapping-class group. We find a set of generators for the fundamental group of a surface with a conformal automorphism of prime order which reflects the action the automorphism in an…
We establish adaptive results for trend filtering: least squares estimation with a penalty on the total variation of order differences. Our approach is based on combining a general oracle inequality for the -penalized least squares estimator with "interpolating vectors" to upper-bound the "effe…
We investigate the use of regularized Newton methods with adaptive norms for optimizing neural networks. This approach can be seen as a second-order counterpart of adaptive gradient methods, which we here show to be interpretable as first-order trust region methods with ellipsoidal constraints. In particular, we prove …
The paper proposes a new order slicing strategy to reduce market impact in large-volume trading.
RL optimizes meta-order execution by adapting to market conditions.
Optimistic method adapted for faster convex-concave min-max problems.
A new method for faster optimization of noisy functions.
New methods using natural gradient for structured optimization.
We study adaptive (or online) nonlinear regression with Long-Short-Term-Memory (LSTM) based networks, i.e., LSTM-based adaptive learning. In this context, we introduce an efficient Extended Kalman filter (EKF) based second-order training algorithm. Our algorithm is truly online, i.e., it does not assume any underlying …
New federated learning methods improve model performance on non-IID data.
Enhances CEV model pricing with high-order scheme and adaptive time stepping.
Adaptive learning model forecasts financial prices using order book data.
New method improves zeroth-order stochastic optimization with adaptive sampling.
SpeqNets improve graph neural networks by scaling and adapting to graph sparsity.
New adaptive methods solve weakly convex stochastic optimization problems.
Paper develops an efficient mean estimator for 1-bit communication constraints.
Perfect adaptation in systems is identified and tested using graphical tools.
New algorithms avoid a dominant lower-order term in heavy-tailed loss settings.
The present paper proposes generalized Gaussian kernel adaptive filtering, where the kernel parameters are adaptive and data-driven. The Gaussian kernel is parametrized by a center vector and a symmetric positive definite (SPD) precision matrix, which is regarded as a generalization of the scalar width parameter. These…
New clustering method adapts to data structure.
Adaptive gradient methods are workhorses in deep learning. However, the convergence guarantees of adaptive gradient methods for nonconvex optimization have not been thoroughly studied. In this paper, we provide a fine-grained convergence analysis for a general class of adaptive gradient methods including AMSGrad, RMSPr…
Networks are a natural representation of complex systems across the sciences, and higher-order dependencies are central to the understanding and modeling of these systems. However, in many practical applications such as online social networks, networks are massive, dynamic, and naturally streaming, where pairwise inter…
Unified framework for adaptive learning systems using consolidation and expansion operations.
ADAHESSIAN optimizes machine learning models with adaptive second-order methods.
We consider first order gradient methods for effectively optimizing a composite objective in the form of a sum of smooth and, potentially, non-smooth functions. We present accelerated and adaptive gradient methods, called FLAG and FLARE, which can offer the best of both worlds. They can achieve the optimal convergence …
The geometrical theory of partial differential equations in the absolute sense, without any additional structures, is developed. In particular the symmetries need not preserve the hierarchy of independent and dependent variables. The order of derivatives can be changed and the article is devoted to the higher--order in…
Recently there has been renewed interest in the mapping-class group of a compact surface of genus and also in its finite order elements. A finite order element of the mapping-class group will be a conformal automorphisms on some Riemann surface of genus . Here we give the details of the proof that there is…
The adaptive momentum method (AdaMM), which uses past gradients to update descent directions and learning rates simultaneously, has become one of the most popular first-order optimization methods for solving machine learning problems. However, AdaMM is not suited for solving black-box optimization problems, where expli…
Adaptive algorithm AMSGrad converges for weakly convex constrained optimization problems.
This paper bridges MTL and meta-learning, showing their shared structure and efficiency.
Adaptive methods such as Adam and RMSProp are widely used in deep learning but are not well understood. In this paper, we seek a crisp, clean and precise characterization of their behavior in nonconvex settings. To this end, we first provide a novel view of adaptive methods as preconditioned SGD, where the precondition…
We aim to design strategies for sequential decision making that adjust to the difficulty of the learning problem. We study this question both in the setting of prediction with expert advice, and for more general combinatorial decision tasks. We are not satisfied with just guaranteeing minimax regret rates, but we want …
New adaptive first-order methods improve on quasi-Newton variants.
Proposes a new method for estimating non-pathwise differentiable functional parameters.
AdAdaGrad optimizes batch sizes for deep learning models, reducing the generalization gap.
This paper improves volatility forecasting using dynamic subset selection in genetic programming.