Improved analysis and new algorithm for gradient-free optimization of smooth functions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Proposes a new method for estimating non-pathwise differentiable functional parameters.
PCHAL and PCHAR use principal components to speed up HAL and HAR methods.
Deep, wide ConvResNets can approximate functions and their smoothness.
Sharp representation theorems show depth benefits for ReLU networks.
New algorithms for interpreting complex multivariate functions.
Deep learning has exhibited superior performance for various tasks, especially for high-dimensional datasets, such as images. To understand this property, we investigate the approximation and estimation ability of deep learning on anisotropic Besov spaces. The anisotropic Besov space is characterized by direction-depen…
Continuous control imitation learning fails if expert actions are smooth.
We develop and analyze an asynchronous algorithm for distributed convex optimization when the objective writes a sum of smooth functions, local to each worker, and a non-smooth function. Unlike many existing methods, our distributed algorithm is adjustable to various levels of communication cost, delays, machines compu…
Paper improves stochastic bilevel optimization methods for highly-smooth problems.
New method for faster convergence in non-convex optimization with unbounded smoothness.
We consider first order gradient methods for effectively optimizing a composite objective in the form of a sum of smooth and, potentially, non-smooth functions. We present accelerated and adaptive gradient methods, called FLAG and FLARE, which can offer the best of both worlds. They can achieve the optimal convergence …
Survival trees exhibit end-cut preference, leading to biased splits.
Paper introduces methods to handle missing data in probabilistic regression trees.
It is well-known that overparametrized neural networks trained using gradient-based methods quickly achieve small training error with appropriate hyperparameter settings. Recent papers have proved this statement theoretically for highly overparametrized networks under reasonable assumptions. These results either assume…
Improved MLMC method for robust and efficient probability and density estimation.
We investigate the theoretical limits of pipeline parallel learning of deep learning architectures, a distributed setup in which the computation is distributed per layer instead of per example. For smooth convex and non-convex objective functions, we provide matching lower and upper complexity bounds and show that a na…
New method smooths integrands for efficient option pricing.
In statistical learning theory, convex surrogates of the 0-1 loss are highly preferred because of the computational and theoretical virtues that convexity brings in. This is of more importance if we consider smooth surrogates as witnessed by the fact that the smoothness is further beneficial both computationally- by at…
The Normalizing Flow (NF) models a general probability density by estimating an invertible transformation applied on samples drawn from a known distribution. We introduce a new type of NF, called Deep Diffeomorphic Normalizing Flow (DDNF). A diffeomorphic flow is an invertible function where both the function and its i…
The paper finds manifold structures on complex spaces.
Despite achieving impressive performance, state-of-the-art classifiers remain highly vulnerable to small, imperceptible, adversarial perturbations. This vulnerability has proven empirically to be very intricate to address. In this paper, we study the phenomenon of adversarial perturbations under the assumption that the…
Let P be a closed smooth (4j-2)-connected 8j-manifold. We complete Wilkens' classification of the manifolds P for j = 1,2 and give an alternative proof to Wall's classification of the manifolds for j > 2. The Hopf-invariant-one dimensions (j=1,2) are characteristed by the fact that the quadratic linking functions which…
We study smooth bundles over surfaces with highly connected almost parallelizable fiber of even dimension, providing necessary conditions for a manifold to be bordant to the total space of such a bundle and showing that, in most cases, these conditions are also sufficient. Using this, we determine the characteristi…
Here we develop an option pricing method based on Legendre series expansion of the density function. The key insight, relying on the close relation of the characteristic function with the series coefficients, allows to recover the density function rapidly and accurately. Based on this representation for the density fun…
We equip many non compact non simply connected surfaces with smooth Riemannian metrics whose isoperimetric profile is smooth, a highly non generic property. The computation of the profile is based on a calibration argument, a rearrangement argument, the Bol-Fiala curvature dependent inequality, together with new result…
We consider the combinatorial multi-armed bandit (CMAB) problem, where the reward function is nonlinear. In this setting, the agent chooses a batch of arms on each round and receives feedback from each arm of the batch. The reward that the agent aims to maximize is a function of the selected arms and their expectations…
Paper introduces arctan pinball loss for XGBoost quantile regression.
In this paper, we consider a generalized multivariate regression problem where the responses are monotonic functions of linear transformations of predictors. We propose a semi-parametric algorithm based on the ordering of the responses which is invariant to the functional form of the transformation function. We prove t…
Proposes a deep neural network approach for image response regression.
Adapting functional gradients improves FGD's practicality and theoretical guarantees.
Rough path theory is focused on capturing and making precise the interactions between highly oscillatory and non-linear systems. It draws on the analysis of LC Young and the geometric algebra of KT Chen. The concepts and the uniform estimates, have widespread application and have simplified proofs of basic questions fr…
In this paper we study smooth orientation-preserving free actions of the cyclic group on a class of -connected -manifolds, , where is a homotopy -sphere. When we obtain a classification up to topological conjugation. When we obtain a classi…
BOOOM optimizes orthonormal matrices without needing gradients.
Adjusting the learning rate schedule in stochastic gradient methods is an important unresolved problem which requires tuning in practice. If certain parameters of the loss function such as smoothness or strong convexity constants are known, theoretical learning rate schedules can be applied. However, in practice, such …
New activations improve deep network reproducibility without sacrificing accuracy.
Machine Learning models incorporating multiple layered learning networks have been seen to provide effective models for various classification problems. The resulting optimization problem to solve for the optimal vector minimizing the empirical risk is, however, highly nonconvex. This alone presents a challenge to appl…
We propose a Fourier-based learning algorithm for highly nonlinear multiclass classification. The algorithm is based on a smoothing technique to calculate the probability distribution of all classes. To obtain the probability distribution, the density distribution of each class is smoothed by a low-pass filter separate…
We prove that the space of complete, finite volume, pinched negatively curved Riemannian metrics on a smooth high-dimensional manifold is either empty or it is highly non-connected, provided their behavior at infinity is similar.
Advances M-polyfolds for complex geometry applications.
Polyhedral surfaces are fundamental objects in architectural geometry and industrial design. Whereas closeness of a given mesh to a smooth reference surface and its suitability for numerical simulations were already studied extensively, the aim of our work is to find and to discuss suitable assessments of smoothness of…
New framework establishes positivity of DNTK for PINNs.
We consider an optimal investment and consumption problem for a Black-Scholes financial market with stochastic coefficients driven by a diffusion process. We assume that an agent makes consumption and investment decisions based on CRRA utility functions. The dynamical programming approach leads to an investigation of t…
We propose a generic spatiotemporal event forecasting method, which we developed for the National Institute of Justice's (NIJ) Real-Time Crime Forecasting Challenge. Our method is a spatiotemporal forecasting model combining scalable randomized Reproducing Kernel Hilbert Space (RKHS) methods for approximating Gaussian …
The paper develops efficient estimators for semi-parametric binary models in distributed computing.
Paper studies Adam's convergence under relaxed assumptions, proving a rate of O(poly(log T)/sqrt(T)).
New algorithms achieve optimal robustness in stochastic convex optimization under contamination.
Alesker has introduced the space of {\it smooth valuations} on a smooth manifold , and shown that it admits a natural commutative multiplication. Although Alesker's original construction is highly technical, from a moral perspective this product is simply an artifact of the operation of inters…