New methods improve confidence set calibration in complex models.
problem Challenges in maintaining confidence set coverage in complex models.
method TRUST and TRUST++ methods using simulated data for calibration.
result Methods achieve distribution-free conditional coverage and robust inference.
Paper addresses statistical inference for GANs and minimax problems.
problem Statistical properties of GANs and minimax problems.
method Consistent estimation and confidence sets for GAN parameters.
result Confidence sets for GAN parameters contain the population solutions with desired coverage probability.
The time value of money is a critical factor not only in risk analysis, but also in insurance and financial applications. In this paper, we consider a special class of set-valued risk statistics by introducing the time value of money. In fact, the risk statistics established by this method is closer to financial realit…
Study uses online bootstrap for RL inference, showing effectiveness.
problem Statistical inference for RL parameters in online settings.
method Online bootstrap method applied to TD and GTD algorithms in RL.
result Method is distributionally consistent for policy evaluation inference.
A new framework bridges classical and machine learning methods for reliable inference from complex models.
problem Intractable likelihood functions in complex systems make classical statistics ineffective for likelihood-free inference.
method Likelihood-Free Frequentist Inference (LF2I) framework that combines classical statistics and machine learning.
result Valid confidence sets with near finite-sample validity can be constructed for any parameter value.
Study examines local extrema and crossing statistics in financial markets.
problem Understanding local extrema and crossing statistics in financial markets.
method Excursion set theory, numerical computation, theoretical prediction, clustering of geometrical measures, cross-correlation, Singular Value Decomposition.
result Excursion sets reveal statistical coherency and sensitivity to crises in financial markets.
The paper outlines future work in random sets theory.
problem Developing a theory of statistical reasoning with random sets.
method Generalizing logistic regression, probability laws, and geometric uncertainty.
result A new geometric approach to uncertainty with general random sets.
A method uses neural networks to approximate sampling distributions of test statistics.
problem Accurate modeling of p-value functions or cdfs for correct confidence set coverage.
method Uses neural networks to model the cdf of test statistics, approximating sampling distributions.
result Neural network approximations of sampling distributions are effective and simple.
The paper optimizes private data sharing by selecting statistics and using MCMC for Bayesian inference.
problem Optimizing private data sharing by selecting statistics and performing Bayesian inference.
method Promotes Fisher information for statistic selection and proposes MCMC algorithms for inference.
result The Fisher information of the privatized statistic predicts the relative performance of the statistic in Bayesian estimation.
A new statistical model uses Orlicz-Sobolev spaces with Gaussian weight.
problem Statistical modeling of infinite-dimensional probability measures.
method Affine statistical bundle on Gaussian Orlicz-Sobolev space.
result Provides tools for solving infinite-dimensional evolution problems.
In this note we prove certain necessary and sufficient conditions for the existence of an embedding of statistical manifolds. In particular, we prove that any compact smooth (C1 resp.) statistical manifold can be embedded into the space of probability measures on a finite set. As a result, we get an answer to the La…
Paper introduces statistical learning for point processes.
problem Statistical learning for point processes in general spaces.
method Combines bivariate innovations and point process cross-validation.
result Statistical learning approach outperforms state of the art.
A neural network approach unifies Lasso for variable selection.
problem Combining statistical and machine learning techniques for variable selection.
method Representing Lasso through a neural network and developing a new optimization algorithm.
result The new optimization algorithm achieves better performance than previous methods.
New bandit algorithms focus on extreme values, outperforming existing methods.
problem Optimizing decisions based on extreme values rather than expected values.
method Robust statistics-based algorithms with vanishing extremal regret.
result The proposed algorithms achieve superior performance compared to existing methods.
Twinning splits data into fast, statistically similar sets.
problem Creating statistically similar data splits for Big Data.
method Twinning is a method based on SPlit for fast, model-independent dataset splitting.
result Twinning is orders of magnitude faster than SPlit.
Unified approach for quantum and classical learning from evaluation oracles.
problem Learning from evaluation oracles in quantum and classical settings.
method Inspired by Kearns' SQ and Valiant's weak evaluation oracle, a unified framework is established.
result Characterizes query complexity for learning linear function classes and extends learnability results for quantum circuits.
Fisher width is a geometric measure of complexity on statistical manifolds.
problem Complexity measures on statistical manifolds
method Introducing Fisher width as a Fisher-geometric analogue of Gaussian width
result Fisher width retains key structural features of Gaussian width while capturing anisotropic geometric effects
Statistical Machine Learning (SML) refers to a body of algorithms and methods by which computers are allowed to discover important features of input data sets which are often very large in size. The very task of feature discovery from data is essentially the meaning of the keyword `learning' in SML. Theoretical justifi…
New framework connects online learning to statistical learning for better generalization bounds.
problem Deriving generalization bounds for statistical learning algorithms.
method Constructing an online learning game and showing a connection to statistical learning.
result Established a connection between online and statistical learning, leading to new generalization bounds.
Paper tackles reinforcement learning with complex observations and simple latent dynamics.
problem Understanding reinforcement learning with complex observations and simple latent dynamics.
method Statistical and algorithmic analysis of reinforcement learning under general latent dynamics.
result Identifies latent pushforward coverability as a condition for statistical tractability.
Natural and social multivariate systems are commonly studied through sets of simultaneous and time-spaced measurements of the observables that drive their dynamics, i.e., through sets of time series. Typically, this is done via hypothesis testing: the statistical properties of the empirical time series are tested again…
Bayesian Federated Inference improves statistical model estimation from multicenter data.
problem Combining data from different medical centers is challenging due to regulatory and logistic issues.
method Bayesian Federated Inference (BFI) framework for multicenter data.
result BFI framework infers additional features of the posterior parameter distribution, capturing more information than Federated Learning.
In this paper we address the problem of performing statistical inference for large scale data sets i.e., Big Data. The volume and dimensionality of the data may be so high that it cannot be processed or stored in a single computing node. We propose a scalable, statistically robust and computationally efficient bootstra…
Develops statistical confidence sets for multidimensional scaling.
problem Statistical uncertainty in multidimensional scaling of noisy data.
method Formal statistical framework, distributional convergence results, uniform confidence sets, bootstrap procedures.
result Construction of reliable confidence sets for latent configurations in multidimensional scaling.
Many conventional statistical procedures are extremely sensitive to seemingly minor deviations from modeling assumptions. This problem is exacerbated in modern high-dimensional settings, where the problem dimension can grow with and possibly exceed the sample size. We consider the problem of robust estimation of sparse…
Statistical guarantees for hyperparameter selection
problem Hyperparameter selection in AI systems
method Learn-then-test framework
result Provable reliability and safety
This paper develops dimension-agnostic inference methods for high-dimensional data.
problem Understanding how classical inference methods behave in high-dimensional settings.
method Using variational representations, sample splitting, and self-normalization to create a refined test statistic.
result The resulting statistic has a Gaussian limiting distribution regardless of how dimensionality scales with sample size.
Formulates mechanics for probability distributions on statistical manifold.
problem Formulating mechanics for probability distributions on statistical manifold.
method Information-geometric formulation of Classical Mechanics on statistical manifold, using dually-flat connection and Hilbert bundle structure.
result Provides coherent formalism for Lagrangian and Hamiltonian mechanics on statistical bundle.
New method reconstructs data subsets from limited published statistics.
problem Reconstructing tabular data from aggregate statistics when full datasets are not possible.
method Generates and verifies subsets of rows and columns that are guaranteed to be correct.
result Privacy violations can persist even with sparse published statistics.
The paper reviews and improves concentration inequalities for statistical inference.
problem Analyzing statistical inference in various settings with high-dimensional data.
method Review and improvement of concentration inequalities for different types of random variables and statistical measures.
result Fresh new results and improved bounds with sharper constants.
Improves inference from sparse data with hybrid summary statistics.
problem Robust simulation-based inference from limited data.
method Augment traditional summary statistics with neural network outputs to maximize mutual information.
result Improves information extraction and makes inference robust in low-data settings.
Statistical uncertainty of different filtration techniques for market network analysis is studied. Two measures of statistical uncertainty are discussed. One is based on conditional risk for multiple decision statistical procedures and another one is based on average fraction of errors. It is shown that for some import…
New method for efficient inference in large datasets.
problem Statistical inference in massive datasets.
method Combines divide-and-conquer method and empirical likelihood.
result Reduces computation burden and demonstrates effectiveness.
Efficiently learns Ising model parameters with limited statistics.
problem Learning Ising model parameters with limited sample configurations.
method Examines trade-offs between computation and observation, using Ising model as example.
result Reconstructs model parameters with statistics up to order O(γ) for ℓ1 width γ. Paper establishes statistical inference for performative predictions.
problem Dynamic influence of predictions on their targets.
method End-to-end framework for estimation and inference under performativity.
result Established central limit theorem for performative settings.
The paper tackles robust policy learning in MDPs using statistical methods.
problem Offline data-driven sequential decision making in MDPs.
method Evaluates policies using average rewards centered at policy-induced stationary distributions. Developed a statistically efficient method for estimating robust optimal policies.
result Established a rate-optimal regret bound up to a logarithmic factor.
Cookbook transforms constrained statistical inference into unconstrained problems.
problem Transforming constrained statistical inference into unconstrained problems.
method Bijective and diffeomorphisms parametrizations.
result Maintains statistical inference properties like identifiability.
We present srlearn, a Python library for boosted statistical relational models. We adapt the scikit-learn interface to this setting and provide examples for how this can be used to express learning and inference problems.
This paper advances FL algorithms for composite optimization and statistical recovery.
problem Federated learning optimization and statistical recovery in composite settings.
method Proposes Fast Federated Dual Averaging for strongly convex and smooth loss, and Multi-stage Federated Dual Averaging for restricted strongly convex and smooth loss.
result Establishes state-of-the-art iteration and communication complexity, and high probability complexity bound with linear speedup.
Hypothesis testing is one of the most common types of data analysis and forms the backbone of scientific research in many disciplines. Analysis of variance (ANOVA) in particular is used to detect dependence between a categorical and a numerical variable. Here we show how one can carry out this hypothesis test under the…
Linear statistics of random zero sets are integrals of smooth differential forms over the zero set and as such are smooth analogues of the volume of the random zero set inside a fixed domain. We derive an asymptotic expansion for the variance of linear statistics of the zero divisors of random holomorphic sections of p…
A new method for safer statistical inference after predictions.
problem Statistical inference with pseudo-outcomes from machine learning predictions.
method Prediction De-Correlated Inference (PDC) framework.
result PDC consistently outperforms supervised methods and can adapt to any model.
Unified framework for set-valued classification tackles ambiguous multi-class datasets.
problem Ambiguous multi-class datasets in modern statistics.
method Unified statistical framework encompassing various set-valued classification formulations.
result Infinite sample optimal strategies and plug-in principle for data-driven algorithms.
New test for conditional independence using kernel embeddings.
problem Testing conditional independence in high-dimensional settings.
method Analytic kernel embeddings, asymptotic distribution.
result New test outperforms existing methods in high-dimensional settings.
New method robustifies topological data analysis against outliers.
problem Outliers make topological data analysis unstable.
method Proposed a robust distance function (MoM Dist) for persistent homology.
result MoM Dist sublevel filtrations and weighted filtrations are consistent estimators in adversarial settings.
Bandit algorithms struggle with consistent performance and robustness.
problem Achieving consistent and robust performance in stochastic multi-armed bandit settings.
method Analyzing regret minimization trade-offs and proposing distribution-oblivious algorithms.
result Logarithmic regret is inconsistent and super-logarithmic regret is necessary for consistent learning.
The paper analyzes Tikhonov regularization in Hilbert scales for statistical inverse problems.
problem Statistical inverse problems in Hilbert scales with general noise.
method Tikhonov regularization scheme with conditional stability estimates and high probability error bounds.
result Explicit rates of convergence for oversmoothing and regular cases over defined regularity classes.
Flexible multi-task learning framework using summary statistics.
problem Data-sharing constraints in healthcare settings.
method Proposes a flexible multi-task learning framework utilizing summary statistics and adaptive parameter selection.
result Systematic non-asymptotic analysis and simulations demonstrate the method's performance.