Randomized exploration methods are more statistically efficient than optimistic methods in reinforcement learning.
problem Comparing and contrasting optimistic and randomized exploration methods in reinforcement learning.
method Analytic examples to compare optimistic and randomized approaches.
result Randomized approaches are more statistically efficient than optimistic approaches.
I introduce a new geometrical approach to thermo--statistical mechanics. Here I highlight the main physical ideas, and how do they translate into geometrical language. I contrast the present approach with previous thermo--statistical--geometrical formalisms, (pseudo-)Riemannian [Weinhold 1975; Ruppeiner 1979] as well a…
A neural network approach unifies Lasso for variable selection.
problem Combining statistical and machine learning techniques for variable selection.
method Representing Lasso through a neural network and developing a new optimization algorithm.
result The new optimization algorithm achieves better performance than previous methods.
Improves statistical and computational efficiency in sparse regression.
problem Balancing statistical accuracy and computational efficiency in data analysis.
method Proposes a unified approach to sparse and group-sparse regression.
result Shows improved performance in both statistical and computational aspects.
Flexible approach for normal approximations in geometric and topological statistics.
problem Normal approximation for complex statistics not expressible as sums of score functions.
method Flexible add-one cost operator combined with strong stabilization theory.
result Established normal approximation results for geometric and topological statistics.
New method uses neural networks to solve statistical mechanics problems.
problem Statistical mechanics of systems with finite size.
method Variational autoregressive neural networks with reinforcement learning.
result Directly computes free energy, entropy, magnetizations, and correlations.
Discovering statistically significant patterns from databases is an important challenging problem. The main obstacle of this problem is in the difficulty of taking into account the selection bias, i.e., the bias arising from the fact that patterns are selected from extremely large number of candidates in databases. In …
Modern statistical methods find the most common value in data.
problem Finding the most common value in data.
method Modern statistical methods applied to mode estimation.
result Traditional approaches to mode estimation are explored and modern methods are applied to new fields.
New approach uses statistical mechanics to explain deep learning generalization.
problem Understanding deep learning's generalization properties.
method Revisiting statistical mechanics in neural networks, introducing control parameters.
result Simple model explains overfitting, discontinuous learning, and sharp transitions.
A new method for inference without likelihood, using logistic regression.
problem Statistical inference in the absence of a likelihood function.
method Estimate the ratio of data generating and marginal distributions using logistic regression.
result Automatic selection of relevant summary statistics.
Unified approach to private statistics from empirical to population data.
problem Divided focus on empirical vs population statistics in private statistics.
method Unified methods for both types of statistics.
result Methods for empirical statistics can be applied to population statistics.
The paper provides a statistical decision-theoretical derivation of the Two-Stage approach for parameter estimation.
problem Theoretical justification for the Two-Stage approach in situations where likelihood is difficult to evaluate.
method Statistical decision-theoretical derivation leading to Bayesian and Minimax estimators.
result The Two-Stage approach is justified theoretically and applied to independent and identically distributed samples.
New approach improves AI's handling of incomplete data.
problem Improving AI's ability to work with incomplete data.
method Proposes a new likelihood-free EM algorithm for faster, more efficient inference.
result More statistically efficient than masking approach and faster than conventional EM.
Machine learning and statistical modeling complement each other in healthcare analytics.
problem Choosing between machine learning and statistical modeling for analytics challenges.
method Choosing based on problem, data, and desired outcomes.
result Machine learning and statistical modeling are complementary, using similar principles but different tools.
New method assesses neural network robustness with statistical estimates.
problem Assessing neural network robustness under input models.
method Statistical approach based on estimating the proportion of inputs violating a property.
result Provides an informative notion of network robustness, scaling to larger networks.
Adaptive data fusion boosts efficiency in multi-task optimization.
problem Multi-task non-smooth optimization in various fields.
method Adaptive data fusion approach leveraging commonalities among objectives.
result Significant improvements in sample efficiency with sharp statistical guarantees.
The paper proves statistical consistency and fairness guarantees for a plug-in algorithm.
problem Establishing statistical guarantees for fairness-aware binary classification.
method Proves statistical consistency and derives finite sample guarantees for the plug-in algorithm.
result The plug-in algorithm is statistically consistent and guarantees fairness and differential privacy.
Statistical learning theory connects to spin glass models via Rademacher complexity and replica theory.
problem Bounding generalization gap in statistical learning theory.
method Linking Rademacher complexity in statistical learning to synthetic models in statistical physics.
result Rademacher complexity is closely related to ground state energy in spin glass models.
Bayesian approach controls FDR in high-dimensional models.
problem High-dimensional variable selection and inference.
method Adapted Mirror Statistic to Bayesian framework for FDR control.
result Effective FDR control without data splitting.
New approach for distributed learning of Gaussian mixtures.
problem Large datasets distributed across different centers.
method Split-and-conquer approach with MM algorithm.
result New estimator is consistent and retains root-n consistency.
Statistical evaluation of machine learning models can lead to false positives.
problem Statistical significance in model comparison can be misleading.
method Evaluation using train, dev, and test sets with statistical significance testing.
result Statistical significance does not necessarily indicate a superior learning approach.
The relationship between statistical dependency and causality lies at the heart of all statistical approaches to causal inference. Recent results in the ChaLearn cause-effect pair challenge have shown that causal directionality can be inferred with good accuracy also in Markov indistinguishable configurations thanks to…
This paper describes a new approach to time series modeling that combines subject-matter knowledge of the system dynamics with statistical techniques in time series analysis and regression. Applications to American option pricing and the Canadian lynx data are given to illustrate this approach.
Novel framework for ML-assisted inference valid for any statistical task.
problem Limited validity of existing methods for post-prediction inference.
method Introduces PSPS framework for task-agnostic ML-assisted inference.
result Valid and efficient inference for arbitrary ML models.
Proposes an additive approximation method for multiplicative noise.
problem Limitations in existing approaches to marginalize over multiplicative errors.
method Embeds multiplicative noise in an additive error term.
result Proposed approach provides feasible error estimates.
Develops a deep learning approach for statistical arbitrage.
problem Temporal price differences between similar assets.
method Constructs arbitrage portfolios using latent asset pricing factors and a convolutional transformer for time series signals.
result High risk-adjusted returns and Sharpe ratios with optimal trading policy.
Paper introduces data-dependent SSP for private linear and logistic regression.
problem Private linear and logistic regression with better performance.
method Data-dependent sufficient statistic perturbation (SSP) for linear and logistic regression.
result Data-dependent SSP outperforms state-of-the-art methods for linear and logistic regression.
New DP framework using data truncation for efficient estimation.
problem Differential privacy in unbounded data support.
method Data truncation, exponential family distributions, maximum likelihood estimation, DP stochastic gradient descent.
result Near-optimal sample complexity for Gaussian mean and covariance estimation.
New ML method detects incomplete bid-rigging cartels.
problem Detecting incomplete bid-rigging cartels in competitive bidding.
method Combines statistical screens with machine learning.
result Algorithm outperforms existing methods in incomplete cartels.
Enhances U-statistics for semi-supervised datasets using unlabeled data.
problem Efficiently utilizing unlabeled data in semi-supervised settings.
method Semi-supervised U-statistics enhanced by unlabeled data.
result Proposed method is asymptotically Normal and more efficient than classical U-statistics.
Study compares machine learning and statistical methods for downscaling precipitation.
problem Improving accuracy of precipitation predictions using machine learning.
method Compared four statistical methods and three machine learning methods.
result Linear methods outperform non-linear approaches in capturing daily anomalies and extremes.
Chentsov's theorem proved for exponential families.
problem Characterizing the Fisher information metric in exponential families.
method Unified proof using the central limit theorem.
result Fisher information metric is the only Riemannian metric invariant under extensions and sufficient statistics.
Statistical query algorithms and low-degree tests are nearly equivalent in high-dimensional hypothesis testing.
problem High-dimensional hypothesis testing and information-computation gaps.
method Analysis of statistical query framework and low-degree polynomials.
result Statistical query algorithms and low-degree polynomials are almost equivalent in power under mild conditions.
New method improves statistical inference using machine learning-imputed data.
problem Improving statistical inference with imputed data from machine learning.
method Two-phase sampling approach for Z-estimation with ML-imputed outcomes.
result Guaranteed efficiency matching or exceeding classical inference, regardless of prediction quality.
Paper presents a statistical method for detecting adversarial inputs.
problem Adversarial attacks on deep learning models.
method Statistical approach based on per-class feature distribution comparison.
result Approach achieves good adversarial detection performance on MNIST and CIFAR-10 datasets.
A new framework for brain mapping using statistical agnostic methods.
problem Estimating brain connectivity with limited data and controlling false positives.
method Statistical Agnostic Mapping (SAM) based on concentration inequalities.
result Relieves instability and provides less conservative p-value correction.
Active inference framework improves U-statistic estimation efficiency.
problem Costly acquisition of labels for U-statistics. method Active inference framework with optimal sampling rule.
result Substantial gains in estimation efficiency over baseline methods.
Researchers explore statistical perspectives to understand GNN generalization.
problem Limited mathematical understanding of GNN performance.
method Three broad frameworks: learning theory, asymptotics, and random graph models.
result Various theoretical results and open questions identified.
New method reconstructs data subsets from limited published statistics.
problem Reconstructing tabular data from aggregate statistics when full datasets are not possible.
method Generates and verifies subsets of rows and columns that are guaranteed to be correct.
result Privacy violations can persist even with sparse published statistics.
Bayesian method improves hit identification in compound screening.
problem Identifying candidate hits from thousands to millions of compounds.
method Bayesian nonparametric modeling for cross-plate correlation and statistical strength.
result Significant improvements in hit identification sensitivity and specificity.
Differentially private learning of graphs improves on naive methods.
problem Learning discrete, undirected graphical models while preserving privacy.
method Developed a principled approach using collective graphical models within an expectation-maximization framework.
result The new method learns better models than competing approaches.
Method converts sparse systems to dense ones for statistical mechanics problems.
problem Statistical mechanics on sparse graphs
method Extracts a Feedback Vertex Set, learns variational distribution, estimates free energy.
result More accurate and faster than existing methods for sparse systems.
Research aims to bridge statistical learning to causal models in AI.
problem Challenges in machine learning and AI related to causality.
method Transition from statistical learning to causal models.
result Progress in AI may require advances in causal modeling.
Boosting improves data fitting while maintaining fairness guarantees.
problem Ensuring fairness in data preprocessing.
method Boosting algorithm to learn sufficient statistics of exponential families.
result The learned distribution maintains fairness guarantees while fitting the data better.
Unified approach connects robust statistics models for efficient mean estimation.
problem Efficient mean estimation in the presence of heavy-tailed noise and Huber contamination.
method Developed connections between Huber's epsilon-contamination model and heavy-tailed noise model, providing efficient and robust estimators.
result Simple efficient estimators are robust to both Huber contamination and heavy-tailed noise.
Landmark Ordinal Embedding improves scalability of ordinal embedding.
problem Learning low-dimensional Euclidean representations from ordinal constraints.
method Landmark-based strategy (LOE) that trades statistical efficiency for computational efficiency.
result LOE is significantly more efficient than conventional methods as the number of items grows.
Improved likelihood-free inference by localizing and refining low-dimensional approximations.
problem Poor performance of common likelihood-free methods in high-dimensional models.
method Localisation followed by refinement of low-dimensional summaries.
result Improved accuracy in marginal posteriors through localized and refined approximations.
The scalability of statistical estimators is of increasing importance in modern applications. One approach to implementing scalable algorithms is to compress data into a low dimensional latent space using dimension reduction methods. In this paper we develop an approach for dimension reduction that exploits the assumpt…