The paper provides bounds for the empirical angular measure and applies them to improve statistical learning in extreme regions.
problem Estimating the angular measure in high-dimensional data with different distributions.
method Established bounds for the maximal deviations of the empirical angular measure from the true measure, using rank transformation and analyzing the most extreme observations.
result The bounds provide performance guarantees for statistical learning procedures in extreme regions, such as binary classification and anomaly detection.
This paper reformulates systemic risk measures and finds new properties and estimators.
problem Understanding and measuring systemic risk in financial networks.
method Representation of systemic risk measures in terms of univariate risk measures and quantiles determined by copulas. Empirical properties and estimators derived.
result MES is not suitable for measuring extreme risks. ES-based measures are more sensitive to power-law tails and large losses.
A new framework tightens risk measure confidence bounds.
problem Improving confidence bounds for various risk measures.
method Distribution optimization framework with two estimation schemes based on concentration bounds.
result Consistently tighter confidence bounds compared to previous methods.
The paper has 2 main goals: 1. We propose a variant of the CAPM based on coherent risk. 2. In addition to the real-world measure and the risk-neutral measure, we propose the third one: the extreme measure. The introduction of this measure provides a powerful tool for investigating the relation between the first two mea…
Efficient classifier error estimation without re-training.
problem Estimating classifier error without re-training.
method Generalized resubstitution based on empirical measures.
result Consistent and asymptotically unbiased error estimation.
New PAC-Bayes bounds derived using Legendre transform and f-divergences.
problem Deriving PAC-Bayes bounds under various assumptions.
method Combining Legendre transform and Fenchel--Young inequality to derive change-of-measure inequalities.
result Extended PAC-Bayesian guarantees under tailored assumptions.
Importance sampling has become an important tool for the computation of tail-based risk measures. Since such quantities are often determined mainly by rare events standard Monte Carlo can be inefficient and importance sampling provides a way to speed up computations. This paper considers moderate deviations for the wei…
We study compressing empirical measures in finite RKHSs using convex optimization.
problem Efficiently approximating empirical measures in high-dimensional spaces.
method Convex optimization and lower bounds on ball size.
result High probability lower bounds on ball size under various conditions.
We propose three measures of mutual dependence between multiple random vectors. All the measures are zero if and only if the random vectors are mutually independent. The first measure generalizes distance covariance from pairwise dependence to mutual dependence, while the other two measures are sums of squared distance…
Sharp bounds for max-sliced Wasserstein distances derived for empirical distributions.
problem Estimating the expected max-sliced Wasserstein distance between a probability measure and its empirical distribution.
method Banach space version and operator norm approach for upper bounds.
result Upper bounds for max-sliced Wasserstein distances are essentially matching and sharp up to a log factor.
This paper presents a unified approach based on Wasserstein distance to derive concentration bounds for empirical estimates for two broad classes of risk measures defined in the paper. The classes of risk measures introduced include as special cases well known risk measures from the finance literature such as condition…
Upper bound for max-sliced 2-Wasserstein distance between measures.
problem Estimating distance between probability measures and their empirical counterparts.
method Same technique as previous work, upper bound approach.
result Upper bound for expected max-sliced 2-Wasserstein distance.
Study sharp convergence rates of empirical UOT for spatio-temporal point processes.
problem Statistical analysis of UOT for spatio-temporal point processes.
method Empirical plug-in estimators for Kantorovich-Rubinstein distance between intensity measures.
result Sharp convergence rates of empirical UOT in terms of intrinsic dimensions of measures.
Generalization of deep networks has been of great interest in recent years, resulting in a number of theoretically and empirically motivated complexity measures. However, most papers proposing such measures study only a small set of models, leaving open the question of whether the conclusion drawn from those experiment…
New bounds show empirical EOT adapts to simpler measure.
problem Statistical performance of empirical EOT estimators.
method Novel statistical bounds, empirical process theory, dual formulation.
result Empirical EOT and its unregularized version follow lower complexity adaptation.
Transformers can interpolate between arbitrary measures.
problem Understanding the expressive power of Transformers as measure-to-measure maps.
method Provided an explicit choice of parameters for a single Transformer to match N arbitrary input measures to N arbitrary target measures.
result A single Transformer can interpolate between arbitrary measures.
New measure quantifies function similarity for optimization.
problem Measuring similarity between functions for optimization.
method Quantifies sub-optimality gaps and operation rules.
result Unified measure for various functional similarities.
Entropy asymmetry affects regularization in ERM, leading to biased solutions.
problem Analyzing the impact of relative entropy asymmetry in ERM regularization.
method Examined Type-I and Type-II ERM-RER, comparing their solutions and properties.
result Type-II ERM-RER regularization introduces a strong bias against training data.
We information-theoretically reformulate two measures of capacity from statistical learning theory: empirical VC-entropy and empirical Rademacher complexity. We show these capacity measures count the number of hypotheses about a dataset that a learning algorithm falsifies when it finds the classifier in its repertoire …
This research evaluates generalization measures in deep learning.
problem Understanding why deep learning models generalize well despite small training error.
method Empirical evaluation of generalization bounds and measures.
result Generalization measures should be evaluated using distributional robustness.
We provide upper bounds of the expected Wasserstein distance between a probability measure and its empirical version, generalizing recent results for finite dimensional Euclidean spaces and bounded functional spaces. Such a generalization can cover Euclidean spaces with large dimensionality, with the optimal dependence…
New metrics avoid high-dimensional analysis challenges, proving convergence without 'curse of dimensionality'.
problem High-dimensional analysis challenges in empirical measure convergence.
method Proposed a new class of probability metrics free of the curse of dimensionality.
result Convergence of empirical measures is free of the curse of dimensionality.
Optimal quantization of measures on Carnot groups
problem Quantization of probability measures on Carnot groups
method Zador-type asymptotic formula and weak convergence of empirical measures
result Convergence of quantization error and density of absolutely continuous part
Study provides convergence rates for risk measure estimation.
problem Estimating risk measures from limited data.
method Plug-in estimation using empirical measures.
result Non-asymptotic convergence rates for risk measure estimation.
New method improves scalability of SGD for large datasets.
problem High variance in stochastic gradient descent.
method Adaptive measure reduction with Carathéodory's theorem.
result Improved scalability to high-dimensional spaces.
Deep convolutional neural networks (CNNs) have been shown to be able to fit a random labeling over data while still being able to generalize well for normal labels. Describing CNN capacity through a posteriori measures of complexity has been recently proposed to tackle this apparent paradox. These complexity measures a…
Optimal algorithm identifies best arm for risk measures in heavy-tailed distributions.
problem Identifying the arm with smallest CVaR, VaR, or weighted sum of CVaR and mean from heavy-tailed distributions.
method Multi-armed bandit best-arm identification framework, solving non-convex optimization problem.
result Optimal δ-correct algorithm with matching lower bound on expected samples.
New regularization method reduces support of empirical risk minimization solutions.
problem Regularization in empirical risk minimization with relative entropy.
method Introduces Type-II regularization, characterizes solutions, analyzes properties of relative entropy.
result Type-II regularization collapses solution support into reference measure's support.
Economics does not need a scientific revolution. Economics needs accurate measurements according to high standards of natural sciences and meticulous work on revealing empirical relationships between measured variables.
Majorizing measures control sequential complexities for online learning.
problem Extending classical empirical processes theory to sequential cases.
method Generic chaining, majorizing measures, fractional covering numbers.
result Sharp control of worst-case sequential Rademacher complexity.
Many recent works have shown that adversarial examples that fool classifiers can be found by minimally perturbing a normal input. Recent theoretical results, starting with Gilmer et al. (2018b), show that if the inputs are drawn from a concentrated metric probability space, then adversarial examples with small perturba…
The study measures systemic risk using common and tail dependence factors.
problem Measuring systemic risk accurately during economic downturns.
method Modeling systemic risk with a common factor for market-wide shocks and a tail dependence factor for extreme events.
result Measures including a tail dependence factor offer better forecasting of financial stress than measures based solely on a common factor.
The paper provides theoretical guarantees for optimized sampling in compressed sensing, showing error vanishes with more measurements.
problem Theoretical and practical improvements in compressed sensing with optimized sampling schemes.
method Theoretical analysis and empirical experiments with optimized sampling schemes for subsampled unitary matrices.
result The error caused by measurement noise vanishes with an increasing number of measurements for optimized sampling schemes, assuming Gaussian noise.
When eliciting judgements from humans for an unknown quantity, one often has the choice of making direct-scoring (cardinal) or comparative (ordinal) measurements. In this paper we study the relative merits of either choice, providing empirical and theoretical guidelines for the selection of a measurement scheme. We pro…
The paper studies PCA of probability measures with varying sample sizes and finds optimal convergence rates.
problem PCA of multiple probability measures with varying sample sizes.
method Double asymptotic regime analysis with convergence rates n−1/2+m−α for empirical covariance and PCA risk. result Optimal convergence rates for empirical covariance and PCA risk in the dense regime are proven.
The coefficient of determination, known as R2, is commonly used as a goodness-of-fit criterion for fitting linear models. R2 is somewhat controversial when fitting nonlinear models, although it may be generalised on a case-by-case basis to deal with specific models such as the logistic model. Assume we are fittin…
Robust portfolio optimization considers uncertainty in market probabilities.
problem Uncertainty in market probabilities in multiperiod portfolio selection.
method Robust mean-variance optimization using Wasserstein ball centered at empirical data.
result Numerical simulations show improved performance compared to other strategies.
LEEP measures transferability of learned representations efficiently.
problem Evaluating the transferability of learned representations in machine learning.
method LEEP: Log Expected Empirical Prediction, a simple measure requiring one pass through the target data set.
result LEEP predicts transfer and meta-transfer learning performance and convergence speed, outperforming existing measures.
Study inverse problems with measure samples, improving estimator calibration and recovery.
problem Inverse problems with unknown potentials observed through measure samples.
method Introduced convex empirical objectives and sharpened Fenchel--Young losses for finite-dimensional potential classes.
result High-probability parameter recovery bounds for inverse entropic unbalanced optimal transport and inverse JKO learning.
New findings on null measurability in symmetrization interface of VC learning.
problem Null measurability issues in symmetrization interface of VC learning.
method Formalized in Lean 4, using Choquet capacitability and patching properties.
result Null-measurable bad event not Borel measurable, separating regularity levels.
An important task in computational statistics and machine learning is to approximate a posterior distribution p(x) with an empirical measure supported on a set of representative points {xi}i=1n. This paper focuses on methods where the selection of points is essentially deterministic, with an emphasis on achi…
Paper tackles measure estimation in barycentric coding model.
problem Estimating an unknown measure in the barycentric coding model.
method Geometric, statistical, and computational insights; quadratic optimization problem; empirical i.i.d. samples algorithm.
result Proves precise rates of convergence for algorithm, ensuring statistical consistency.
We use multiple measures of graph complexity to evaluate the realism of synthetically-generated networks of human activity, in comparison with several stylized network models as well as a collection of empirical networks from the literature. The synthetic networks are generated by integrating data about human populatio…
Recent works investigated the generalization properties in deep neural networks (DNNs) by studying the Information Bottleneck in DNNs. However, the mea- surement of the mutual information (MI) is often inaccurate due to the density estimation. To address this issue, we propose to measure the dependency instead of MI be…
We consider statistical estimation of superhedging prices using historical stock returns in a frictionless market with d traded assets. We introduce a plugin estimator based on empirical measures and show it is consistent but lacks suitable robustness. To address this we propose novel estimators which use a larger set …
New measure of robustness for estimators, with tight bounds for Gaussian mean estimation.
problem Developing robust statistical estimators for datasets with noise or outliers.
method Introducing empirical sensitivity as a new robustness measure and proving lower bounds for Gaussian mean estimation.
result Empirical sensitivity bounds for optimal estimators are tight, showing obstructions on mean and variance.
We propose a robust risk measurement approach that minimizes the expectation of overestimation plus underestimation costs. We consider uncertainty by taking the supremum over a collection of probability measures, relating our approach to dual sets in the representation of coherent risk measures. We provide results that…
We solve robust optimization problems using Wasserstein balls and apply it to mean-CVaR optimization.
problem Distributionally robust optimization with Wasserstein ambiguity sets.
method Transformed robust optimization into non-robust with penalty term, selecting ambiguity set size.
result Impressive results in robust mean-CVaR optimization compared to other strategies.