Cookbook transforms constrained statistical inference into unconstrained problems.
problem Transforming constrained statistical inference into unconstrained problems.
method Bijective and diffeomorphisms parametrizations.
result Maintains statistical inference properties like identifiability.
A new statistical model uses Orlicz-Sobolev spaces with Gaussian weight.
problem Statistical modeling of infinite-dimensional probability measures.
method Affine statistical bundle on Gaussian Orlicz-Sobolev space.
result Provides tools for solving infinite-dimensional evolution problems.
We investigate artificial neural networks as a parametrization tool for stochastic inputs in numerical simulations. We address parametrization from the point of view of emulating the data generating process, instead of explicitly constructing a parametric form to preserve predefined statistics of the data. This is done…
A tractable pseudo-metric for non-parametric distributions via SPD geometry.
problem Computing distances between non-parametric probability distributions is intractable.
method Two-stage framework: projection onto parametric family, embedding into SPD matrices.
result Closed-form pseudo-metric for two-sample hypothesis testing.
Enhances selective inference for generalized lasso using parametric programming.
problem Low statistical power in selective inference for generalized lasso.
method Parametric programming to compute solution paths and identify model selection events.
result Improves selective inference power and practicality for various problems.
Examines algorithmic modeling across three cultures.
problem Tackles algorithmic modeling in different cultural contexts.
method Uses parametric regressions, interpretable algorithms, and complex algorithms.
result Extension of Leo Breiman's thesis to include cultural differences.
Theoretical guarantees for neural estimators in parametric statistics are derived.
problem Lack of theoretical guarantees for neural estimators in parametric statistics.
method Decompose risk into terms and verify assumptions for convergence.
result Derive theoretical guarantees for neural estimators.
Develops coresets for scalable multivariate distribution estimation.
problem Handling large-scale data in non-parametric or semi-parametric regression and density estimation.
method Novel coreset construction for multivariate conditional transformation models (MCTMs).
result Substantial data reduction with high log-likelihood accuracy.
A nonparametric two-sample test using a parametric integral probability metric
problem Detecting distributional differences between two independent samples
method Propose a new two-sample test statistic based on a newly introduced integral probability metric (IPM)
result Establish theoretical guarantees for the associated two-sample testing procedure
Breiman's two cultures reconciled through blending statistical thinking.
problem Tension between parametric statistical and machine learning approaches.
method Establishing a link between parametric statistical and machine learning frameworks.
result Integrated statistical thinking can bridge the gap between two cultures.
Study provides guarantees for kernel clustering under non-parametric mixtures.
problem Statistical guarantees for kernel-based clustering without strong assumptions.
method Non-parametric mixture models, kernel-based clustering, consistency guarantees.
result Necessary and sufficient separability conditions for consistent clustering recovery.
New framework for privacy-preserving statistical inference using robust statistics.
problem Privacy-preserving statistical inference with robust statistics.
method Introducing a general framework for parametric inference with differential privacy guarantees using M-estimators and test statistics.
result Demonstrated that differential privacy is weaker than robustness and can be achieved by randomizing robust M-estimators.
One of the main challenges in the parametrization of geological models is the ability to capture complex geological structures often observed in the subsurface. In recent years, generative adversarial networks (GAN) were proposed as an efficient method for the generation and parametrization of complex data, showing sta…
Proposes a private empirical bootstrap for Gaussian Differential Privacy.
problem Quantifying uncertainty in massive data under Differential Privacy.
method Gaussian Differential Private Bootstrap by Subsampling.
result Consistent and efficient private inference method.
This paper finds a unique partition of a sample space for estimating continuous distributions.
problem Estimating continuous probability distributions from finite samples.
method Equal-probability partition of the sample space using order statistics.
result The partition yields an entropy of log2(N+1) bits, providing a discrete entropy estimate.
We analyse the learning performance of Distributed Gradient Descent in the context of multi-agent decentralised non-parametric regression with the square loss function when i.i.d. samples are assigned to agents. We show that if agents hold sufficiently many samples with respect to the network size, then Distributed Gra…
Paper reproduces a kernel-based scan B-statistic for online change-point detection.
problem Continuous detection of distribution changes in online data streams.
method Efficient kernel-based scan B-statistic for online change-point detection.
result Scan B-statistic outperforms parametric methods in challenging scenarios.
The paper provides stability guarantees for non-parametric maximum likelihood estimation using statistical mechanics.
problem Non-parametric maximum likelihood estimation and Gaussian mixture models.
method Statistical mechanics analysis to establish stability guarantees for NPMLE.
result High probability upper bounds on the Kullback-Leibler divergence between NPMLE estimators and the true density.
The parametric complexity is the key quantity in the minimum description length (MDL) approach to statistical model selection. Rissanen and others have shown that the parametric complexity of a statistical model approaches a simple function of the Fisher information volume of the model as the sample size n goes to in…
Neural networks can learn relationships that traditional models cannot.
problem Identifying factors that differentiate neural networks from traditional models.
method Proving non-identifiability of neural networks compared to smooth parametric models.
result Neural networks can learn nontrivial relationships that traditional models cannot.
There are various parametric models for analyzing pairwise comparison data, including the Bradley-Terry-Luce (BTL) and Thurstone models, but their reliance on strong parametric assumptions is limiting. In this work, we study a flexible model for pairwise comparisons, under which the probabilities of outcomes are requir…
The paper tackles robust policy learning in MDPs using statistical methods.
problem Offline data-driven sequential decision making in MDPs.
method Evaluates policies using average rewards centered at policy-induced stationary distributions. Developed a statistically efficient method for estimating robust optimal policies.
result Established a rate-optimal regret bound up to a logarithmic factor.
Robust learning method minimizes risk with corrupted data.
problem Statistical learning with unknown corrupted data fraction.
method Develops a robust learning method with specified corrupted data fraction upper bound.
result Optimal weights provide robustness against corrupted data.
A graph-based method for two-sample testing across connected nodes.
problem Identifying nodes where two probability distributions differ significantly.
method Collaborative non-parametric two-sample testing (CTST) framework.
result CTST outperforms independent node tests by leveraging graph structure.
Novel nonparametric method for GLMs improves prediction and inference performance.
problem Improving prediction and inference in GLMs with minimal assumptions.
method Combines binary regression and latent variable formulations, extends parametric versions, introduces new classification statistic.
result Uniformly better prediction and inference performance over parametric formulation, especially with asymmetric data.
Paper introduces statistical learning for point processes.
problem Statistical learning for point processes in general spaces.
method Combines bivariate innovations and point process cross-validation.
result Statistical learning approach outperforms state of the art.
Researchers explore geometric dualities in statistical manifolds.
problem Understanding geometric dualities in statistical manifolds.
method Exploring the dualistic geometry of statistical manifolds, focusing on Hessian manifolds.
result Moduli space of univariate normal distributions corresponds to Siegel half-space and Siegel-Jacobi space.
Locally private methods detect changes in time series data.
problem Detecting distributional changes in time series data under local differential privacy.
method Proposed locally differentially private algorithms based on randomized response and binary mechanisms.
result Theoretical performance bounds and empirical validation of detection accuracy.
KQT-EWMA monitors multivariate data streams online with flexible and practical change detection.
problem Online monitoring of multivariate data streams for detecting changes.
method Combines Kernel-QuantTree histogram and EWMA statistic for non-parametric monitoring.
result Controls Average Run Length (ARL0) while achieving comparable detection delays.
We study the Wasserstein natural gradient in parametric statistical models with continuous sample spaces. Our approach is to pull back the L2-Wasserstein metric tensor in the probability density space to a parameter space, equipping the latter with a positive definite metric tensor, under which it becomes a Riemanni…
Study finds a method to discover causal relationships that are invariant to marginal distributions.
problem Current causal discovery methods are sensitive to marginal distributions, leading to unreliable results.
method Proposes a non-parametric estimator that marginalizes the marginals to find intrinsic causal relationships.
result The proposed method yields causal estimators competitive with current methodologies and emphasizes uncertainty.
We propose a new non parametric technique to estimate the CALL function based on the superhedging principle. Our approach does not require absence of arbitrage and easily accommodates bid/ask spreads and other market imperfections. We prove some optimal statistical properties of our estimates. As an application we firs…
Estimates neural drift for stochastic equations, improving inference on noisy data.
problem Estimating drift in stochastic differential equations with neural networks.
method Non-parametric estimation using ReLU neural networks, enforcing theoretical bounds.
result Practical method for inference on noisy and rough functional data.
Estimates conditional Brenier maps using entropic optimal transport.
problem Non-parametric estimation of conditional Brenier maps.
method Entropic optimal transport for scalable non-parametric estimation.
result Entropic optimal transport maps asymptotically converge to conditional Brenier maps.
We propose a robust estimator to improve maximum likelihood in probabilistic models.
problem Overfitting and sensitivity to noise in maximum likelihood estimation.
method Distributionally robust maximum likelihood estimator that minimizes worst-case expected log-loss.
result The robust estimator is statistically consistent and performs well in regression and classification tasks.
Interpolation improves performance in nearest neighbor algorithms without over-parametrization.
problem Achieving zero training error in deep learning without over-parametrization.
method Introduced a class of interpolated weighting schemes in nearest neighbor algorithms.
result Mild data interpolation strictly improves prediction performance and statistical stability.
The increased usage of solar energy places additional importance on forecasts of solar radiation. Solar panel power production is primarily driven by the amount of solar radiation and it is therefore important to have accurate forecasts of solar radiation. Accurate forecasts that also give information on the forecast u…
We propose a novel approach for density estimation with exponential families for the case when the true density may not fall within the chosen family. Our approach augments the sufficient statistics with features designed to accumulate probability mass in the neighborhood of the observed points, resulting in a non-para…
The paper provides a statistical decision-theoretical derivation of the Two-Stage approach for parameter estimation.
problem Theoretical justification for the Two-Stage approach in situations where likelihood is difficult to evaluate.
method Statistical decision-theoretical derivation leading to Bayesian and Minimax estimators.
result The Two-Stage approach is justified theoretically and applied to independent and identically distributed samples.
Compressive learning framework adapted for semi-parametric models.
problem Handling large datasets efficiently with semi-parametric models.
method Reformulate compressive learning framework to handle semi-parametric models, capturing their inherent topology and structure.
result Demonstrated robustness and efficiency of the framework in independent component analysis and subspace clustering.
A new principle for extrapolating regression outside training data.
problem Regression extrapolation when predictions are outside training data range.
method Data-adaptive marginal transformation and simple relationship assumption.
result Progression method offers guarantees on approximation error beyond training data range.
The paper develops efficient estimators for semi-parametric binary models in distributed computing.
problem Estimation and inference challenges in large-scale data under non-smooth objective functions.
method Proposes one-shot and multi-round divide-and-conquer estimators with adaptive kernel smoothing to relax constraints and achieve superlinear optimization error.
result Establishes quadratic convergence up to optimal statistical error rate and handles dataset heterogeneity and high-dimensional sparse parameters.
Paper proves optimality of doubly robust estimators for treatment effects.
problem Estimating treatment effects in causal inference.
method Structure-agnostic framework of statistical lower bounds, using non-parametric regression and classification oracles.
result Doubly robust estimators are statistically optimal for ATE and ATT.
Develops a simulation-based method to translate expert knowledge into prior distributions for Bayesian models.
problem Effective incorporation of expert knowledge into prior distributions for diverse model structures.
method Simulation-based stochastic gradient descent to learn hyperparameters of parametric priors from expert knowledge.
result Method is adaptable to various elicitation techniques and independent of model structure.
Develops non-parametric tests for group symmetry in data.
problem Lack of statistical tests for group symmetry in data.
method Formulates and implements non-parametric tests for distributional symmetry under specified groups.
result Develops tests for conditional invariance/equivariance and applies them to real-world data.
Deep adaptive sampling improves surrogate modeling for complex systems.
problem Statistical errors in random sampling for high-dimensional problems.
method DAS^2 method, using deep generative models to refine training sets.
result Reduces statistical errors in approximating solutions for low-regularity problems.
New test assesses probabilistic model calibration without expensive approximations.
problem Assessing calibration of probabilistic models with scores.
method Kernel Calibration Conditional Stein Discrepancy (KCCSD) test using new score-based kernels.
result Control over type-I error with improved scalability and efficiency.
We analyze numerically the training dynamics of deep neural networks (DNN) by using methods developed in statistical physics of glassy systems. The two main issues we address are (1) the complexity of the loss landscape and of the dynamics within it, and (2) to what extent DNNs share similarities with glassy systems. O…