Unified approach to private statistics from empirical to population data.
problem Divided focus on empirical vs population statistics in private statistics.
method Unified methods for both types of statistics.
result Methods for empirical statistics can be applied to population statistics.
Raising statistical hurdles may not be justified due to data bias.
problem Data bias leads to unobserved results that weaken identification of revised hurdles.
method Theoretical and empirical analysis of statistical hurdles and data bias.
result Statistics targeting only published findings can be strongly identified.
This article introduces a framework to estimate the value of evidence-based decision making.
problem Lack of empirical tools to assess the value of evidence-based decision making and optimize statistical precision.
method Empirical framework using parametric and nonparametric empirical Bayes methods.
result The value of statistical evidence depends on how organizations translate it into policy decisions.
Develops an empirical likelihood framework for random forests and ensembles.
problem Quantifying the statistical uncertainty of random forests and ensembles.
method Empirical likelihood framework exploiting the incomplete U-statistic structure of ensemble predictions. result Modified empirical likelihood statistic achieves accurate coverage and practical reliability.
We provide sharp empirical estimates of expectation, variance and normal approximation for a class of statistics whose variation in any argument does not change too much when another argument is modified. Examples of such weak interactions are furnished by U- and V-statistics, Lipschitz L-statistics and various error f…
New method for efficient inference in large datasets.
problem Statistical inference in massive datasets.
method Combines divide-and-conquer method and empirical likelihood.
result Reduces computation burden and demonstrates effectiveness.
Active inference framework improves U-statistic estimation efficiency.
problem Costly acquisition of labels for U-statistics. method Active inference framework with optimal sampling rule.
result Substantial gains in estimation efficiency over baseline methods.
The study provides theoretical guarantees for the statistical performance of optimal decision trees.
problem Theoretical limits on the statistical performance of globally optimal decision trees.
method Sharp oracle inequalities and uniform concentration framework based on Rademacher complexity.
result Derivation of minimax optimal rates for piecewise sparse heterogeneous anisotropic Besov space.
The paper provides bounds for the empirical angular measure and applies them to improve statistical learning in extreme regions.
problem Estimating the angular measure in high-dimensional data with different distributions.
method Established bounds for the maximal deviations of the empirical angular measure from the true measure, using rank transformation and analyzing the most extreme observations.
result The bounds provide performance guarantees for statistical learning procedures in extreme regions, such as binary classification and anomaly detection.
A new trading strategy using reinforcement learning for statistical arbitrage.
problem Traditional statistical arbitrage models rely on model assumptions and price deviations from a long-term mean.
method Empirical reversion time metric, reinforcement learning framework, and state space optimization.
result Optimal mean reversion strategy identified through reinforcement learning.
Neural networks estimate statistical divergences with performance guarantees.
problem Estimating statistical divergences with theoretical performance guarantees.
method Parametrizing empirical variational form by a neural network and optimizing over parameter space.
result Established non-asymptotic absolute error bounds for neural estimators of four f-divergences. The paper develops statistical inference for gradient flows in optimization.
problem Uncertainty quantification along the entire optimization path.
method Uniform central limit theorem and algorithm-aware covariance estimator.
result Asymptotically valid confidence intervals for target parameter.
This paper tightens the law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.
problem Developing nonasymptotic concentration bounds for empirical KL_inf with optimal constants and rates.
method Presenting a tight law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.
result A tight law of the iterated logarithm for empirical KL_inf, applicable to unbounded data.
Corrects sample selection bias in empirical risk minimization using importance sampling.
problem Statistical learning with biased training data.
method Weighted empirical risk minimization using importance sampling.
result Generalization capacity preserved with estimated importance weights.
Prove non-asymptotic bounds for minimal risk in statistical learning
problem Estimating minimal risk in statistical learning
method Using concentration inequalities
result Non-asymptotic bounds for minimal risk
New measure of robustness for estimators, with tight bounds for Gaussian mean estimation.
problem Developing robust statistical estimators for datasets with noise or outliers.
method Introducing empirical sensitivity as a new robustness measure and proving lower bounds for Gaussian mean estimation.
result Empirical sensitivity bounds for optimal estimators are tight, showing obstructions on mean and variance.
New statistical test for change-point detection using relative entropy.
problem Offline change-point detection using divergence metrics.
method Study of empirical relative entropy distributions, derivation of approximations, introduction of new Berry-Esseen bounds.
result Theoretical and practical validation of relative entropy for change-point detection.
In a wide range of statistical learning problems such as ranking, clustering or metric learning among others, the risk is accurately estimated by U-statistics of degree d≥1, i.e. functionals of the training data with low variance that take the form of averages over k-tuples. From a computational perspective, …
Paper develops error rates for physics-informed learning, comparing it to data-driven methods.
problem Understanding the trade-off between soft penalties and hard constraints in PISL.
method Develops complexity-dependent error rates using the small-ball method.
result Physics-informed estimators have comparable error rates to hard constrained methods, differing only by constants.
New smoothing technique improves Wasserstein distance estimation in high dimensions.
problem Estimating statistical distances between high-dimensional distributions.
method Gaussian smoothing of p-Wasserstein distance and analysis of its asymptotic behavior. result Gaussian-smoothed p-Wasserstein distance converges at rate n−1/2, improving over n−1/d for unsmoothed distances. The paper improves generative models to avoid replicating observed examples.
problem Improving generative models to avoid replicating observed examples.
method Theoretical insights into the Wasserstein GAN, constrained to left-invertible push-forward maps, generating distributions that avoid replication and significantly deviate from the empirical distribution.
result Left-invertibility achieves this without compromising statistical optimality.
Paper uses learned summary statistics for Bayesian inference with difficult likelihood functions.
problem Difficult to obtain exact likelihood function for observation data and simulation model.
method Simulation-based inference with learned summary statistics, using Cressie-Read discrepancy criterion.
result Effective inference performed over selected sample sets of observation data.
Empirical model tackles decision problems without specifying states of the world.
problem Decision problems under uncertainty with inaccessible states of the world.
method Empirical approach using observed act--consequence pairs as model primitives.
result Optimality in empirical decision problems addressed using protocol-based empirical choice functions.
The study improves the assessment of fairness in face recognition using ROC curves and statistical guarantees.
problem Improving the assessment of fairness in face recognition systems.
method Proves asymptotic guarantees for empirical ROC curves and fairness metrics, and introduces a recentering technique to avoid bootstrap pitfalls.
result Demonstrates the practical relevance of the methods for assessing fairness in face recognition systems.
New bounds show empirical EOT adapts to simpler measure.
problem Statistical performance of empirical EOT estimators.
method Novel statistical bounds, empirical process theory, dual formulation.
result Empirical EOT and its unregularized version follow lower complexity adaptation.
The optimal approach is to theorize after examining data, not before.
problem Optimal sequencing of theory and empirical analysis for economic questions.
method Formalized a Bayesian model to trade off Darwinian and Statistical Learning.
result Post hoc theorizing is typically optimal in modern economics.
Many scientifically well-motivated statistical models in natural, engineering and environmental sciences are specified through a generative process, but in some cases it may not be possible to write down a likelihood for these models analytically. Approximate Bayesian computation (ABC) methods, which allow Bayesian inf…
Paper tackles constrained learning with non-convex losses, overcoming challenges with new approach.
problem Challenges in learning with non-convex losses and statistical constraints.
method Learning in the empirical dual domain, bounding empirical duality gap.
result Established a constrained counterpart to classical learning theory.
Boosting improves data fitting while maintaining fairness guarantees.
problem Ensuring fairness in data preprocessing.
method Boosting algorithm to learn sufficient statistics of exponential families.
result The learned distribution maintains fairness guarantees while fitting the data better.
Nyström KPCA balances computational efficiency and statistical accuracy.
problem Computational burden in large sample situations for kernel methods.
method Theoretical analysis of Nyström approximate kernel principal component analysis (KPCA).
result Nyström approximate KPCA matches statistical performance of non-approximate KPCA while being computationally beneficial.
This paper tackles constrained statistical learning problems by proposing a new approach.
problem Statistical learning problems with constraints are challenging and scarce.
method Directly tackling the constrained problem using finite dimensional parameterizations, sample averages, and duality theory.
result We bound the empirical duality gap, showing the effectiveness of the constrained formulation.
New algorithms avoid non-monotonic risk curves in statistical learning.
problem Non-monotonic behavior of risk curves in statistical learning.
method Derive risk-monotonic algorithms under weak assumptions.
result Risk monotonicity does not necessarily lead to worse excess risk rates.
We highlight a very simple statistical tool for the analysis of financial bubbles, which has already been studied in [1]. We provide extensive empirical tests of this statistical tool and investigate analytically its link with stocks correlation structure.
New statistical mechanics analysis shows edge pruning outperforms node pruning in neural networks.
problem Theoretical understanding of neural network pruning effectiveness is lacking.
method Statistical mechanics analysis of a teacher-student framework.
result DPP node pruning method is superior to other methods, but edge pruning is better overall.
A cluster tree provides a highly-interpretable summary of a density function by representing the hierarchy of its high-density clusters. It is estimated using the empirical tree, which is the cluster tree constructed from a density estimator. This paper addresses the basic question of quantifying our uncertainty by ass…
A new sequential method estimates Poisson means in streaming data, achieving optimality and efficiency.
problem Estimating Poisson means in a streaming, or online, framework.
method A quasi-Bayesian approach based on Newton's algorithm for a sequential estimate.
result Established frequentist guarantees including consistency and asymptotic optimality.
Natural and social multivariate systems are commonly studied through sets of simultaneous and time-spaced measurements of the observables that drive their dynamics, i.e., through sets of time series. Typically, this is done via hypothesis testing: the statistical properties of the empirical time series are tested again…
Efficient and robust algorithms for decentralized estimation in networks are essential to many distributed systems. Whereas distributed estimation of sample mean statistics has been the subject of a good deal of attention, computation of U-statistics, relying on more expensive averaging over pairs of observations, is…
DRIFT uses neural flows to replace distributional regression models.
problem Lack of neural network representations for distributional regression models.
method Inverse flow transformations (DRIFT) for distributional regression.
result Neural representations in DRIFT match classical statistical methods in performance.
Study finds short-term instability in financial ARCH models.
problem Short-term stability of financial ARCH models.
method Analyzes quadratic ARCH processes using historical data and empirical innovations.
result Empirical innovations have variance significantly above 1, indicating short-term instability.
This paper extends Median-of-Means to new learning problems involving pairwise comparisons.
problem Learning from pairwise comparisons in machine learning.
method Segmenting data into blocks, comparing pairs of decision rules, and declaring the winner based on majority performance.
result The Median-of-Means approach maintains robustness and performance under various sampling schemes.
The paper proves concentration inequalities for two-sample rank processes and applies them to ranking performance criteria.
problem Measuring the performance of ranking statistics between two populations.
method Proves concentration inequalities for two-sample rank processes indexed by VC classes of scoring functions.
result Generalization capacity of empirical maximizers of ranking performance criteria is investigated.
Town hall discusses AI's impact on statistics, culture, and training.
problem Adapting statistics to AI advancements and infrastructure.
method Open panel discussion and audience Q&A.
result Candid perspectives on evolving statistical practices.
Polyak step size GD reaches final radius of convergence after log iterations.
problem Statistical and computational complexities of Polyak step size GD.
method Generalized smoothness and Lojasiewicz conditions, stability of gradients.
result Polyak step size GD reaches final statistical radius of convergence after logarithmic number of iterations.
Investigates financial and economic systems using statistical mechanics and information theory.
problem Complexity, asymmetry, stochasticity, and non-linearity in financial and economic systems.
method Model-based and empirical analyses using statistical mechanics and information theory.
result Derives probability distribution functions for better understanding of financial and economic dynamics.
Paper uses SGD for solving linear inverse problems, improving empirical performance.
problem Solving statistical inverse problems in science and engineering.
method Stochastic Gradient Descent (SGD) for linear inverse problems, with smoothing techniques.
result Consistency and finite sample bounds for excess risk demonstrated.
Paper explores using bootstrap methods to improve SGD's stability and robustness.
problem Improving the stability and robustness of SGD.
method Investigates empirical bootstrap approaches for SGD from algorithmic stability and statistical robustness perspectives.
result Demonstrates construction of purely distribution-free confidence intervals using bootstrap SGD.
Introduces BPEL for EL, enhancing flexibility and using MCMC for inference.
problem Computational challenges in EL methods.
method Bayesian Penalized Empirical Likelihood (BPEL) framework with MCMC sampling.
result Enhanced flexibility and practicality of EL methods with MCMC.