A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
We prove sharp bounds for the growth rate of eigenfunctions of the Ornstein-Uhlenbeck operator and its natural generalizations. The bounds are sharp even up to lower order terms and have important applications to geometric flows.
In this paper, we give a new sharp generalization bound of lp-MKL which is a generalized framework of multiple kernel learning (MKL) and imposes lp-mixed-norm regularization instead of l1-mixed-norm regularization. We utilize localization techniques to obtain the sharp learning rate. The bound is characterized by the d…
New adaptive scheduler improves SAM for better model training.
problem Training machine learning models requires selecting a learning rate, which is often difficult and time-consuming.
method Derive Polyak schedulers tailored to SAM-style updates, proving linear convergence for strongly convex objectives and an O(1/T) rate for convex objectives.
result Polyak schedulers achieve comparable or better performance than tuned SAM baselines, reducing the need for learning-rate tuning.
EM algorithm converges linearly and achieves sharp rate in estimating mixtures of pairwise differences.
problem Estimating mixtures of pairwise differences from noisy data.
method Sharp analysis of the EM algorithm locally around the ground truth.
result The EM sequence converges linearly with an ℓ∞-norm guarantee on the estimation error and achieves the sharp rate of estimation in the ℓ2-norm.
Forward regression is a statistical model selection and estimation procedure which inductively selects covariates that add predictive power into a working statistical regression model. Once a model is selected, unknown regression parameters are estimated by least squares. This paper analyzes forward regression in high-…
Perpetual futures offer leverage without maturity, with prices influenced by funding rates.
problem Understanding and pricing perpetual futures with funding rates.
method Derive no-arbitrage prices and bounds in markets with trading costs. Empirically analyze deviations and Sharpe ratios of implied arbitrage strategies.
result Implied arbitrage strategies in crypto markets yield high Sharpe ratios, indicating significant pricing inefficiencies.
Stochastic (sub)gradient methods require step size schedule tuning to perform well in practice. Classical tuning strategies decay the step size polynomially and lead to optimal sublinear rates on (strongly) convex problems. An alternative schedule, popular in nonconvex optimization, is called \emph{geometric step decay…
Classifiers built with neural networks handle large-scale high dimensional data, such as facial images from computer vision, extremely well while traditional statistical methods often fail miserably. In this paper, we attempt to understand this empirical success in high dimensional classification by deriving the conver…
Maximal initial learning rate for deep ReLU networks identified.
problem Finding the optimal initial learning rate for deep neural networks.
method Simple approach to estimate maximal initial learning rate η∗, analyzing its behavior in constant-width fully-connected ReLU networks.
result Maximal initial learning rate η∗ is well predicted as a power of depth × width, with specific conditions for network width and input layer training.
We describe a post hoc test for the Sharpe ratio, analogous to Tukey's test for pairwise equality of means. The test can be applied after rejection of the hypothesis that all population Signal-Noise ratios are equal. The test is applicable under a simple correlation structure among asset returns. Simulations indicate t…
We prove that the boundary of a (not necessarily connected) bounded smooth set with constant nonlocal mean curvature is a sphere. More generally, and in contrast with what happens in the classical case, we show that the Lipschitz constant of the nonlocal mean curvature of such a boundary controls its C2-distance fro…
Stochastic Gradient Descent (SGD) and its variants are mainstream methods for training deep networks in practice. SGD is known to find a flat minimum that often generalizes well. However, it is mathematically unclear how deep learning can select a flat minimum among so many minima. To answer the question quantitatively…
We construct continuous-time equilibrium models based on a finite number of exponential utility investors. The investors' income rates as well as the stock's dividend rate are governed by discontinuous Levy processes. Our main result provides the equilibrium (i.e., bond and stock price dynamics) in closed-form. As an a…
We show that if F is a convex class of functions that is L-subgaussian, the error rate of learning problems generated by independent noise is equivalent to a fixed point determined by `local' covering estimates of the class, rather than by the gaussian averages. To that end, we establish new sharp upper and lower e…
We study statistical risk minimization problems under a privacy model in which the data is kept confidential even from the learner. In this local privacy framework, we establish sharp upper and lower bounds on the convergence rates of statistical estimation procedures. As a consequence, we exhibit a precise tradeoff be…
We apply the procedure of Lee et al. to the problem of performing inference on the signal-noise ratio of the asset which displays maximum sample Sharpe ratio over a set of possibly correlated assets. We find a multivariate analogue of the commonly used approximate standard error of the Sharpe ratio to use in this condi…