SAM improves generalization in overparameterized models, but its behavior in tensorized models is less understood.
problem Understanding the implicit regularization of SAM in tensorized models.
method Scale-invariance analysis and gradient flow analysis to derive Norm Deviation as a measure of core norm imbalance, and propose Deviation-Aware Scaling (DAS).
result DAS achieves competitive or improved performance over SAM, while offering reduced computational overhead.
Proves inequality linking function deviation to gradient norm on compact manifolds.
problem Analyzing coupled elliptic systems on compact manifolds.
method Develops a new Poincaré-Sobolev inequality with a density-free reference average.
result Poincaré constant depends on the density's gradient norm.
Unified normative modeling for neuroimaging phenotypes using denoising diffusion models.
problem Discarding multivariate dependence in neuroimaging pipelines.
method Denoising diffusion probabilistic models (DDPMs) with FiLM and SAINT backbones.
result Unified multivariate normative modeling with better calibration and dependence preservation.
The paper provides a new uniform tail bound for empirical processes.
problem Developing a uniform tail bound for empirical processes indexed by a class of functions.
method Introducing a deflation step to the standard generic chaining argument, and using a natural seminorm based on Cramér functions.
result Established a new uniform tail bound for empirical processes.
SVR analyzed within RQ framework for risk management.
problem Risk management in stochastic optimization.
method Risk Quadrangle (RQ) theory applied to SVR.
result SVR formulations as minimization of Vapnik error and CVaR norm.
The higher order singular value decomposition (HOSVD) of tensors is a generalization of matrix SVD. The perturbation analysis of HOSVD under random noise is more delicate than its matrix counterpart. Recently, polynomial time algorithms have been proposed where statistically optimal estimates of the singular subspaces …
Vertex distortion measures how far lattice knots deviate from straight lines.
problem Measuring how much lattice knots deviate from straight paths.
method Analogous to smooth knots, study vertex distortion in lattice knots.
result Vertex distortion is 1 only for the unknot and can be arbitrarily high.
Gradient descent dynamics in wide neural networks are analyzed using a dynamical CLT.
problem Understanding the fluctuations in wide shallow neural networks trained via gradient descent.
method Dynamical Central Limit Theorem (CLT) applied to neural network dynamics.
result Asymptotic fluctuations remain bounded in mean square throughout training.
The paper provides bounds for high-dimensional U-statistics with novel order-explicit inequalities.
problem Bounding the deviation of high-dimensional U-statistics from their Hájek projections.
method Develops novel order-explicit moment inequalities for higher-order Hoeffding components.
result The maximum deviation of a high-dimensional U-statistic from its Hájek projection is of order Op(φbn−1log2(dn)). So-called sparse estimators arise in the context of model fitting, when one a priori assumes that only a few (unknown) model parameters deviate from zero. Sparsity constraints can be useful when the estimation problem is under-determined, i.e. when number of model parameters is much higher than the number of data point…
Recently theoretical guarantees have been obtained for matrix completion in the non-uniform sampling regime. In particular, if the sampling distribution aligns with the underlying matrix's leverage scores, then with high probability nuclear norm minimization will exactly recover the low rank matrix. In this article, we…
Improves risk and variability measures continuity and consistency.
problem Improving the continuity and consistency of risk and variability measures.
method Analyzes convex and order bounded above functionals on Frechet lattices and Orlicz spaces.
result Order-continuous, law-invariant functionals on Orlicz spaces are strongly consistent everywhere.
Paper explores SVGD for Bayesian inference, linking deterministic and stochastic dynamics.
problem Bayesian inference and Markov chain Monte Carlo methods.
method Stein variational gradient descent (SVGD) with deterministic and stochastic dynamics.
result Identifies Stein-Fisher information as the leading order contribution in the long-time and many-particle regime.
New method for estimating covariance with robustness to outliers.
problem Estimating covariance from noisy data with outliers.
method Cross-fitted clipped covariance estimator with computable Bernstein certificates.
result The method balances certified stochastic error and robust hold-out proxy for clipping bias.
Level-set optimization formulations with data-driven constraints minimize a regularization functional subject to matching observations to a given error level. These formulations are widely used, particularly for matrix completion and sparsity promotion in data interpolation and denoising. The misfit level is typically …
Study shows limits of volume-constrained sets are finite unions of Wulff shapes.
problem Analyzing the behavior of sets with degenerating ellipticity.
method Proving rigidity of L1-accumulation points of volume-constrained almost-critical sets. result Limits of volume-constrained sets are finite unions of φ-Wulff shapes. MCE reduces embedding instability in nonlinear dimensionality reduction.
problem Embedding instability caused by random initialization.
method Median of multiple embeddings (MCE) based on large deviation theory.
result MCE achieves consistency at an exponential rate and effectively mitigates instability.
Study on discrepancy principle for learning algorithms in nonparametric regression.
problem Determining optimal iteration number in nonparametric regression with unknown optimal iteration.
method Investigates discrepancy principle and modified principles for kernelized spectral filters, using deviation inequalities and change-of-norm arguments.
result Classical discrepancy principle is adaptive for slow rates, while modified principles are adaptive for faster rates.
SWRLDA improves LDA for multi-class classification with edge classes.
problem LDA's vulnerability to edge classes causing biased mean and large distances.
method Self-weighted robust LDA with l21-norm distance criterion.
result SWRLDA outperforms other methods on synthetic and real-world datasets.
Novel metrics improve machine learning models for ICU patient care.
problem Predicting vital sign trajectories for early detection of adverse events.
method Developed novel performance metrics aligned with clinical contexts, validated on simulated and real datasets, and optimized neural networks using these metrics.
result Neural networks trained with these metrics excel in predicting clinically significant events.
Study connects covariance cleaning theory to information theory for heavy-tailed distributions.
problem Optimizing covariance matrices for heavy-tailed distributions using information theory.
method Minimizing Frobenius norm and information loss between true and estimated covariance matrices.
result Asymptotic regime of large matrices minimizes information loss for Student's t distributions.
New bounds link Schwarzian derivative to hyperbolic geometry.
problem Quantify the relationship between Schwarzian derivative and hyperbolic geometry.
method Established explicit quantitative bounds between Schwarzian norm and bending norm.
result Explicit bounds on bending norm for univalent maps with small Schwarzian norm.
This paper introduces a new bound to explain generalization in over-parameterized models.
problem Understanding why some over-parameterized models generalize well while others do not.
method PAC-Chernoff bounds and smoothness measures based on large deviation theory.
result Interpolators with smoother structures generalize better, according to the new theoretical framework.
Studies report that firms do not invest in cost-effective green technologies. While economic barriers can explain parts of the gap, behavioural aspects cause further under-valuation. This could be partly due to systematic deviations of decision-making agents' perceptions from normative benchmarks, and partly due to the…
New robust estimators achieve subgaussian bounds using VC-dimension.
problem Robust estimation of sparse and corrupted data.
method Use of VC-dimension to measure statistical complexity.
result First robust estimators for sparse estimation with subgaussian rate.
Tensor completion and robust principal component analysis have been widely used in machine learning while the key problem relies on the minimization of a tensor rank that is very challenging. A common way to tackle this difficulty is to approximate the tensor rank with the ℓ1−norm of singular values based on its …
We consider a group of mean-variance investors with mimicking desire such that each investor is willing to penalize deviations of his portfolio composition from compositions of other group members. Penalizing norm constraints are already applied for statistical improvement of Markowitz portfolio procedure in order to c…
Introduces Star-Shaped deviation measures for risk analysis.
problem Risk measurement and analysis in finance.
method Characterizes Star-Shaped deviation measures through acceptance sets and convex deviation measures.
result Exposes the relationship between Star-Shaped risk measures and deviation measures.
New model tackles real-world distribution mismatches in machine learning.
problem Real-world applications often have training and test distributions that differ.
method Developed a learning model based on information theory using importance sampling.
result The model performs better under large distribution deviations.
Paper characterizes monotonic mean-deviation risk measures.
problem Developing consistent risk measures from mean-deviation models.
method Applying a risk-weighting function to the deviation part of a mean-deviation model.
result Characterizes monotonic mean-deviation measures as consistent risk measures.
We extend previous large deviations results for the randomised Heston model to the case of moderate deviations. The proofs involve the Gärtner-Ellis theorem and sharp large deviations tools.
Paper proves large deviation principle for stochastic approximations.
problem Asymptotic estimates of learning algorithm deviations.
method Weak convergence approach to large deviations.
result Identifies appropriate scaling sequence and new representation for rate function.
We propose a general framework for reduced-rank modeling of matrix-valued data. By applying a generalized nuclear norm penalty we can directly model low-dimensional latent variables associated with rows and columns. Our framework flexibly incorporates row and column features, smoothing kernels, and other sources of sid…
We propose a version of least-mean-square (LMS) algorithm for sparse system identification. Our algorithm called online linearized Bregman iteration (OLBI) is derived from minimizing the cumulative prediction error squared along with an l1-l2 norm regularizer. By systematically treating the non-differentiable regulariz…
In this paper we propose the notion of dynamic deviation measure, as a dynamic time-consistent extension of the (static) notion of deviation measure. To achieve time-consistency we require that a dynamic deviation measures satisfies a generalised conditional variance formula. We show that, under a domination condition,…
The paper introduces a diagnostic method to detect grokking transitions in models before test accuracy improves.
problem Detecting the transition from training to generalization in machine learning models.
method Summarize task-dependent observables as empirical distributions, map them to Wasserstein/quantile coordinates, and analyze using Hankel dynamic mode decomposition.
result The diagnostic method achieves AUROC \(\approx\) 0.93 for grokking-vs-non-grokking discrimination at the run level.
Study large deviations in life insurance portfolios without identical distributions.
problem Large deviations in life insurance portfolios with bounded losses and variances.
method Upper bound from standard large deviations, counterexample for full large deviation principle.
result Exponential bound for average loss exceeding a threshold.
New stability theorems for H-type Carnot groups established.
problem Characterize H-type Carnot groups using stability theorems.
method Introduced H-type deviation, computed for families, established new characterizations.
result H-type Carnot groups are uniquely characterized by specific conditions.
New clustering method improves climate data analysis in Lesser Antilles.
problem Inducing undesirable effects in clustering algorithms using Euclidean distance.
method Replacing Euclidean distance with Expert Deviation (ED) based on symmetrized Kullback-Leibler divergence.
result KMS-ED produces more interpretable clusters with better discrimination of daily situations.
Study large deviations for hypoelliptic diffusion on sub-Riemannian manifolds.
problem Large deviations for hypoelliptic diffusion measures on sub-Riemannian manifolds.
method Rough path theory and manifold-valued Malliavin calculus.
result Proved a large deviation principle for pinned hypoelliptic diffusion measures.
Proposes new deviation measures using Minkowski gauges.
problem Lack of suitable acceptance sets for deviation measures.
method Derives deviation measures through Minkowski gauges of acceptable sets.
result Any positive homogeneous deviation measure can be accommodated in the framework.
In this paper we analyze a dynamic recursive extension of the (static) notion of a deviation measure and its properties. We study distribution invariant deviation measures and show that the only dynamic deviation measure which is law invariant and recursive is the variance. We also solve the problem of optimal risk-sha…
This work challenges the Neural Tangent Kernel's role in overparameterized neural networks, especially with large width and depth.
problem The Neural Tangent Kernel's behavior in overparameterized neural networks with large width and depth is unclear.
method Experimental and theoretical analysis of ReLU networks with large width and depth.
result The aggregate norm of hidden neuron deviations does not vanish in infinitely-wide ReLU networks, indicating non-trivial behavior.
We provide a unifying treatment of pathwise moderate deviations for models commonly used in financial applications, and for related integrated functionals. Suitable scaling allows us to transfer these results into small-time, large-time and tail asymptotics for diffusions, as well as for option prices and realised vari…
Importance sampling has become an important tool for the computation of tail-based risk measures. Since such quantities are often determined mainly by rare events standard Monte Carlo can be inefficient and importance sampling provides a way to speed up computations. This paper considers moderate deviations for the wei…
We present a novel human-aware navigation approach, where the robot learns to mimic humans to navigate safely in crowds. The presented model, referred to as DeepMoTIon, is trained with pedestrian surveillance data to predict human velocity in the environment. The robot processes LiDAR scans via the trained network to n…
Connections between Lie derivatives and the deviation equation has been investigated in spaces with affine connection. The deviation equations of the geodesics as well as deviation equations of non-geodesics trajectories have been obtained on this base. This is done via imposing certain conditions on the Lie derivative…
Unified approach to stochastic Volterra systems' deviations.
problem Large and moderate deviations for stochastic Volterra systems.
method Weak convergence approach by Budhijara, Dupuis and Ellis.
result Unified treatment of deviations for a broad class of stochastic Volterra equations.