New algorithm optimizes smooth functions with Hölder exponent > 1.
problem Optimizing smooth functions with unknown Hölder exponent > 1.
method Two-layer algorithms using misspecified linear/polynomial bandit algorithms in bins.
result Regret bound of O ~ ( T d + α d + 2 α ) \tilde{O}(T^{\frac{d+\alpha}{d+2\alpha}}) O ~ ( T d + 2 α d + α ) for α > 1 \alpha > 1 α > 1 . Deep neural networks with piecewise-polynomial activations can approximate smooth functions and their derivatives.
problem Approximating smooth functions and their derivatives with neural networks.
method Derives the depth, width, and sparsity required for approximation in Hölder norms.
result Deep neural networks with bounded weights can approximate Hölder smooth functions and their derivatives.
Improved learning rates with new smoothness measure.
problem Learning with noisy data and unknown function class.
method Generalized Hölder smoothness to average smoothness, proving upper and lower bounds.
result Achieved nearly optimal learning rates in realizable and agnostic settings.
New algorithm optimizes Hölder smooth functions in RKHS with tighter regret bounds.
problem Optimizing Hölder smooth functions in RKHS with bounded norm.
method Proposes a new algorithm ( exttt{LP-GP-UCB}) using Local Polynomial (LP) estimators and multi-scale UCB.
result Derives high probability bounds on simple and cumulative regret, matching optimal performance for SE kernel and uniformly tighter bounds for Matérn kernels.
Study on RNNs' ability to approximate past-dependent Hölder functions and their application to regression.
problem Understanding and optimizing the approximation capacity of RNNs for regression tasks.
method Derivation of upper bounds on RNN approximation error for Hölder smooth functions and application to regression.
result Achievement of minimax optimal prediction error bounds for RNNs under various data assumptions.
Adapts Hölder smoothness with normalized gradients.
problem Improving smoothness adaptation methods.
method Black-box adaptation of Levy's method using normalized gradients.
result Bound depends on local Hölder smoothness.
Gradient-free optimization for additive models achieves optimal error.
problem Optimizing noisy functions with zero-order information.
method Proposed a randomized gradient estimator for gradient-free optimization.
result Achieves minimax optimal error of order d T − ( β − 1 ) / β dT^{-(β-1)/β} d T − ( β − 1 ) / β . The paper bounds the expectation of empirical processes indexed by Hölder classes.
problem Estimating the expectation of the supremum of empirical processes for distributions on bounded sets.
method Providing upper bounds on the expectation of the supremum of empirical processes indexed by Hölder classes.
result Deriving non-asymptotic risk bounds for estimating distributions using empirical processes and IPM.
Paper establishes tight lower bounds for minimizing certain smooth and convex functions.
problem Minimizing high-order Hölder smooth and uniformly convex functions.
method Analyzes two asymmetric cases of q > p + ν q > p + ν q > p + ν and q < p + ν q < p + ν q < p + ν using worst-case oracle complexities. result Establishes worst-case oracle complexities for reaching an ε-approximate solution.
Optimal nonparametric regression estimator adapts to unknown smoothness.
problem Nonparametric regression with unknown smoothness.
method Constructs an interpolating estimator that adapts to unknown smoothness.
result Minimax optimal rates achieved on Hölder classes.
Derivatives of sub-Riemannian geodesics are always L p L_p L p -Hölder continuous.
problem Smoothness of sub-Riemannian geodesics
method Proving L p L_p L p -Hölder continuity of derivatives result Derivatives of sub-Riemannian geodesics are L p L_p L p -Hölder continuous Paper relaxes SGD privacy and generalization guarantees for non-smooth convex losses.
problem Privacy and generalization in SGD for non-smooth convex losses.
method Relaxes Lipschitz and strong smoothness assumptions to Hölder smoothness, proving ( ε , δ ) (ε,δ) ( ε , δ ) -DP and optimal excess risk. result Noisy SGD with α α α -Hölder smooth losses achieves optimal excess risk with linear gradient complexity for α ≥ 1 / 2 α \geq 1/2 α ≥ 1/2 . Smooth functions preserve Zygmund class on curves.
problem Characterizing functions based on their behavior on smooth curves.
method Analyzing functions through their behavior on smooth curves and mappings between Banach spaces.
result Functions in Zygmund class preserve Zygmund class on smooth curves.
The paper links set cuspidality to function regularity and flatness.
problem Linking set cuspidality to function regularity and flatness.
method Analyzes arc-smooth functions and their properties on various sets.
result Establishes a precise link between set cuspidality and function regularity.
New algorithm adapts to unknown demand smoothness for dynamic pricing.
problem Dynamic pricing with unknown Hölder smoothness of demand function.
method Self-similarity condition and adaptive algorithm.
result Adaptive algorithm achieves minimax optimal regret without prior knowledge of smoothness.
Smoothness of graphs evolving by fractional mean curvature is proven.
problem Evolution of graphs by fractional mean curvature.
method Analytic semigroup approach to nonlocal quasilinear evolution equation.
result Short time existence, uniqueness, and optimal Hölder regularity of classical solutions.
Whereas subriemannian geometry usually deals with smooth horizontal distributions, partially hyperbolic dynamical systems provide many examples of subriemannian geometries defined by non-smooth (namely, Hölder continuous) distributions. These distributions are of great significance for the behavior of the parent dynami…
This paper recovers smooth functions from noisy modulo samples using a three-stage strategy.
problem Recovering Hölder smooth functions from noisy modulo samples.
method Three-stage strategy: denoising with local polynomial estimators, unwrapping, and spline-based quasi-interpolant.
result Uniform error rates for Hölder class functions with high probability.
AdaGrad fails to adapt to Hölder-smoothness in composite optimization problems.
problem AdaGrad's convergence rate is suboptimal for composite objectives.
method Exhibited a simple one-dimensional convex problem to highlight AdaGrad's limitations.
result AdaGrad does not achieve the classical convergence rate for Hölder-smooth objectives.
The paper proves optimal smoothness for certain Lagrangian graphs with specific Hölder continuity.
problem Optimal regularity for Hölder continuous Hamiltonian stationary Lagrangian graphs.
method Establishing smoothness conditions based on Hölder exponent and Lagrangian phase properties.
result Smoothness of graphs is achieved when Hölder exponent is strictly greater than 1/3 and Lagrangian phase is supercritical.
Study asymptotic behavior of second Chern forms on degenerating Kähler-Einstein surfaces.
problem Asymptotic behavior of second Chern forms on degenerating Kähler-Einstein surfaces with ADE singularities.
method Investigates a function on the unit disc defined by fiber integrals of the forms with a smooth test function, showing a lower bound of Hölder exponent at the origin for both cscK-metrics and Ricci-flat metrics.
result Shows bounds of Hölder exponent for both cscK-metrics and Ricci-flat metrics.
The paper bounds the excess risk of deep neural networks for weakly dependent processes.
problem Learning with weakly dependent data using deep neural networks.
method Approximation of smooth functions by deep neural networks and a bound on excess risk.
result The excess risk bound for deep learning under weak dependence is close to O ( n − 1 / 2 ) \mathcal{O}(n^{-1/2}) O ( n − 1/2 ) for sufficiently smooth functions. Develops a mathematical framework for causal fermion systems in infinite dimensions.
problem Analysis of causal fermion systems in infinite-dimensional settings.
method Introduces Banach manifold structure and expedient differential calculus.
result Establishes Hölder continuity of causal Lagrangian and integrated causal Lagrangian.
Two definitions quantify C 2 , α C^{2,α} C 2 , α regularity of Riemannian surfaces.
problem Quantify the regularity of Riemannian surfaces.
method Intrinsic and extrinsic definitions using Hölder norms and smooth local representations.
result Intrinsic and extrinsic definitions are equivalent up to a constant.
Develops a deep learning framework for various data types.
problem Handling nonparametric regression and classification across different data types.
method Introduces a general framework with two estimators: NPDNN and SPDNN, based on data satisfying generalized Bernstein-type inequalities.
result Both NPDNN and SPDNN estimators are minimax optimal in many classical settings.
New method extends low-rank MDPs to continuous action spaces.
problem Limited applicability of current low-rank MDP methods to continuous action spaces.
method Extending FLAMBE algorithm to continuous action spaces with Hölder smoothness conditions.
result Similar PAC bound achieved for continuous actions with polynomial dependence on smoothness order.
Improved estimators for causal inference using cross-fitting and undersmoothing.
problem Estimating expected conditional covariance in causal inference.
method Double cross-fit doubly robust (DCDR) estimators with undersmoothing for non-smooth nuisance functions.
result DCDR estimators achieve n \sqrt{n} n -consistency and asymptotic normality under minimal conditions. New inequality for eigenfunctions on curved spaces.
problem Eigenfunctions on non-smooth spaces with Ricci curvature.
method Sharp reverse-Hölder inequality for Dirichlet Laplacian eigenfunctions.
result Generalizes classical comparison theorem to curved spaces.
Deep neural networks can learn smooth functions without parameters.
problem Learning smooth functions from shallow ReLU neural networks.
method Using over-parameterized shallow ReLU neural networks with norm constraints.
result Least squares estimators based on shallow neural networks are minimax optimal.
Convex solutions to a specific equation are smooth when the phase is smooth enough.
problem Regularity of solutions to the Lagrangian mean curvature equation.
method Showed regularity for convex solutions under Hölder continuity conditions on the phase.
result Convex viscosity solutions are regular if the Lagrangian phase is Hölder continuous.
Transformer networks approximate Hölder and Sobolev functions with fixed-depth networks.
problem Nonparametric regression with dependent observations.
method Established novel upper bounds for Transformer networks approximating Hölder and Sobolev functions under various β β β -mixing data assumptions. result Explicit convergence rates for nonparametric regression problems under β β β -mixing data assumptions. By an influential theorem of Boman, a function f f f on an open set U U U in R d \mathbb R^d R d is smooth ( C ∞ \mathcal C^\infty C ∞ ) if and only if it is arc-smooth, i.e., f ∘ c f\circ c f ∘ c is smooth for every smooth curve c : R → U c : \mathbb R \to U c : R → U . In this paper we investigate the validity of this result on closed sets. Our main focus is on s…
We consider the non-parametric regression problem under Huber's ε ε ε -contamination model, in which an ε ε ε fraction of observations are subject to arbitrary adversarial noise. We first show that a simple local binning median step can effectively remove the adversary noise and this median estimator is minimax optimal up t…
Deep neural networks without regularization can achieve consistent estimates with good convergence rates.
problem The necessity of regularization in deep neural networks for consistent estimates.
method Gradient descent on an over-parametrized neural network without regularization, with specific initialization, step size, and number of steps.
result An estimate without regularization is universally consistent and achieves good convergence rates.
Maps in Heisenberg groups can be extended to Hölder continuous functions.
problem Maximizing Hölder continuity of maps in Heisenberg groups.
method Analyzing one-form constraints and using distributional extensions.
result Smooth horizontal maps can be extended to Hölder continuous functions.
Adaptive smooth non-stationary bandits achieve optimal regret rates without knowing parameters.
problem Smooth non-stationary bandits with Hölder class rewards.
method Established optimal dynamic regret rate and adaptive algorithm.
result Optimal dynamic regret can be attained adaptively without knowing Hölder exponent and coefficient.
Unique solutions found for diffusive martingale problems.
problem Finding unique solutions to Cauchy problems for diffusive real-valued strict local martingales.
method Provided sets of smooth functions under local Hölder and Engelbert-Schmidt conditions for unique classical and weak solutions.
result Unique solutions found for specific martingale models.
Study shows attention-style models learn pairwise interactions efficiently.
problem Learning pairwise interactions in attention-style models.
method Proved minimax rate of convergence for learning pairwise interactions.
result Minimax rate is M − 2 β 2 β + 1 M^{-\frac{2β}{2β+1}} M − 2 β + 1 2 β independent of embedding dimension and token number. Adaptive NN method improves matrix completion for non-smooth data.
problem Matrix completion with non-smooth non-linear functions under high missingness.
method Two-sided nearest neighbors with \Holder function class non-linearity.
result NN error rate matches oracle's for latent factors, non-trivial for wide range of missingness.
New algorithm improves matrix estimation with one-sided covariates.
problem Estimating matrix means with unobserved row covariates.
method Proposes an algorithm for nonparametric matrix estimation with observed column covariates.
result Achieves minimax optimal nonparametric rate in moderately proportioned matrices.
Proves Hölder continuity of complex Monge-Ampère solutions.
problem Global Hölder continuity of solutions to complex Monge-Ampère equation.
method Analyzes Dirichlet problem on strictly pseudoconvex domains or Hermitian manifolds.
result Proves global Hölder continuity of solutions under given conditions.
New method efficiently interpolates nonparametric density estimators.
problem Efficient evaluation of nonparametric density estimators.
method Piecewise multivariate polynomial interpolation scheme.
result New estimator with low space requirements and efficient querying.
Study explores relationship between Hölder and FDPD divergences.
problem Understanding the relationship between Hölder and FDPD divergences.
method Intersection and generalization of divergence families, proving nonnegativity, deriving inequalities.
result Established ξ ξ ξ -Hölder divergence and derived inequalities. We study perturbations of a partially hyperbolic toral automorphism L which is diagonalizable over C and has a dense center foliation. For a small perturbation of L with a smooth center foliation we establish existence of a smooth leaf conjugacy to L. We also show that if a small perturbation of an ergodic irreducible …
Novel oracle-type inequality for logistic loss in DNNs achieves sharp convergence rates.
problem Generalization analysis for binary classification with DNNs and logistic loss.
method Established an oracle-type inequality to handle the boundedness of the target function.
result Optimal convergence rates for fully connected ReLU DNN classifiers trained with logistic loss.
The paper extends properties of smooth functions to closed sets and maps.
problem Properties of smooth functions on closed sets and maps.
method Extending properties of smooth functions to closed sets and maps, proving isomorphisms with natural topologies.
result Bornological isomorphisms of function spaces are established.
Improved computational efficiency for estimating Wasserstein distance.
problem Inefficient computation of Wasserstein distance for large samples.
method Developed Sample-Sketch-Solve paradigm using grid sketches.
result Approximates Wasserstein distance within ε error in ε^(-max(2, (d+1+o(1))/(1+α))) time.
New algorithm optimizes Hölder continuous functions efficiently.
problem Optimizing Hölder continuous multivariate functions.
method Uses a query creation rule for global optimization, avoiding proxy functions.
result Achieves an average regret bound of $O(T^{-racα{n}})$ for Hölder exponent α α α .