Big Data bring new opportunities to modern society and challenges to data scientists. On one hand, Big Data hold great promises for discovering subtle population patterns and heterogeneities that are not possible with small-scale data. On the other hand, the massive sample size and high dimensionality of Big Data intro…
Factor models are a class of powerful statistical models that have been widely used to deal with dependent measurements that arise frequently from various applications from genomics and neuroscience to economics and finance. As data are collected at an ever-growing scale, statistical machine learning faces some new cha…
The analysis of mixed data has been raising challenges in statistics and machine learning. One of two most prominent challenges is to develop new statistical techniques and methodologies to effectively handle mixed data by making the data less heterogeneous with minimum loss of information. The other challenge is that …
Tensor analysis tackles complex multidimensional data across fields.
problem Efficiently extracting information from high-dimensional data.
method Interdisciplinary approach combining statistics, optimization, and numerical linear algebra.
result Significant progress in tensor analysis over the last decade.
The need for new methods to deal with big data is a common theme in most scientific fields, although its definition tends to vary with the context. Statistical ideas are an essential part of this, and as a partial response, a thematic program on statistical inference, learning, and models in big data was held in 2015 i…
FedOS tackles challenges in federated learning by using open-set learning.
problem Challenges in federated learning due to data locality and privacy constraints.
method Introduces open-set learning to stabilize training in federated learning.
result Demonstrates improved model performance through open-set learning.
We here summarize our experience running a challenge with open data for musical genre recognition. Those notes motivate the task and the challenge design, show some statistics about the submissions, and present the results.
Discovering statistically significant patterns from databases is an important challenging problem. The main obstacle of this problem is in the difficulty of taking into account the selection bias, i.e., the bias arising from the fact that patterns are selected from extremely large number of candidates in databases. In …
This paper tackles constrained statistical learning problems by proposing a new approach.
problem Statistical learning problems with constraints are challenging and scarce.
method Directly tackling the constrained problem using finite dimensional parameterizations, sample averages, and duality theory.
result We bound the empirical duality gap, showing the effectiveness of the constrained formulation.
Survey of challenges and future directions in applying RL to real-world settings.
problem Challenges in deploying RL in practical settings due to limited interaction and changing environments.
method Analysis of RL system design, implementation, and continual improvement.
result Need for theory and methodology to bridge research and application gap.
Graph Neural Networks improve financial time series forecasting accuracy.
problem Forecasting univariate financial time series with statistical significance.
method Introducing the Time-Geometric model combining geometric and temporal patterns.
result Statistically significant improvements in forecasting accuracy through geometric patterns.
New algorithms tackle statistical heterogeneity in federated learning.
problem Statistical heterogeneity in distributed machine learning models.
method Introduces three novel methods: SuPerFed, AAggFF, and FedEvg.
result Mitigates statistical heterogeneity in federated learning.
CAD-DA controls anomaly detection under domain adaptation.
problem Valid statistical inference after domain adaptation.
method Conditional Selective Inference to handle domain adaptation effects.
result Valid statistical inference under domain adaptation achieved.
Federated learning poses new statistical and systems challenges in training machine learning models over distributed networks of devices. In this work, we show that multi-task learning is naturally suited to handle the statistical challenges of this setting, and propose a novel systems-aware optimization method, MOCHA,…
Machine learning and statistical modeling complement each other in healthcare analytics.
problem Choosing between machine learning and statistical modeling for analytics challenges.
method Choosing based on problem, data, and desired outcomes.
result Machine learning and statistical modeling are complementary, using similar principles but different tools.
Paper stabilizes persistent homology rank functions for statistical inference.
problem Stability issues in persistent homology rank functions.
method Derive stability results for rank functions under FDA metrics.
result Rank functions stabilize, improving statistical inference.
Enhances statistical inference using synthetic data.
problem Limited labeled data for statistical inference.
method GESPI framework that combines synthetic and real data.
result Error rate remains below a user-specified bound and decreases with synthetic data quality.
Bayesian nonparametrics adapt model complexity to diverse datasets.
problem Complex challenges across statistics, computer science, and engineering.
method Flexible Bayesian nonparametric models that adapt model complexity.
result Bayesian nonparametrics offer innovative solutions to multi-object tracking.
Both the human brain and artificial learning agents operating in real-world or comparably complex environments are faced with the challenge of online model selection. In principle this challenge can be overcome: hierarchical Bayesian inference provides a principled method for model selection and it converges on the sam…
A new method selects features for ERGMs to improve network modeling.
problem ERGM degeneracy creates unrealistic network structures.
method Stochastic step-wise feature selection to overcome computational burden and improve network accommodation.
result Enhanced ERGM modeling with more accurate interpretations.
The search for higher-order feature interactions that are statistically significantly associated with a class variable is of high relevance in fields such as Genetics or Healthcare, but the combinatorial explosion of the candidate space makes this problem extremely challenging in terms of computational efficiency and p…
Normalizing flows improve ptychography reconstruction quality and uncertainty quantification.
problem Challenges in ptychography due to large-scale nonlinear and non-convex inverse problems and photon statistics.
method Use of normalizing flows to model the posterior distribution and quantify reconstruction uncertainty.
result Normalizing flows enable better characterization and uncertainty quantification in ptychography reconstructions.
The paper tackles statistical and computational challenges in learning correlated reward models.
problem The Independence of Irrelevant Alternatives (IIA) assumption collapses human preferences into a universal utility function, leading to coarse approximations.
method The paper investigates the statistical and computational challenges of learning a correlated probit model using best-of-three preference data.
result Best-of-three preference data overcomes the limitations of pairwise preference data, allowing for more fine-grained modeling of human preferences.
New theory challenges traditional machine learning assumptions.
problem Traditional machine learning theories are critiqued.
method A new theory is proposed and discussed.
result Learning true probabilities is not equivalent to other learning goals.
Bayesian Neural Networks help quantify uncertainty in deep learning predictions.
problem Uncertainty quantification in deep learning predictions.
method Bayesian statistics applied to neural networks.
result Design, implementation, training, and evaluation of Bayesian Neural Networks.
Study enhances neural network interpretability through statistical methods.
problem Complexity and interpretability challenges in neural networks.
method Theoretical framework, statistical tests, dimensionality reduction algorithms.
result Developed bootstrapping technique and statistical tests for ANN performance.
Finding statistically significant high-order interaction features in predictive modeling is important but challenging task. The difficulty lies in the fact that, for a recent applications with high-dimensional covariates, the number of possible high-order interaction features would be extremely large. Identifying stati…
RLHF uses human feedback to train AI models, posing statistical challenges.
problem Aligning AI models with human preferences using noisy, subjective feedback.
method Supervised fine-tuning, reward modeling, policy optimization, statistical ideas.
result Statistical methods for reward function learning and policy optimization.
Interpretable ML helps discover insights from big data.
problem Validating data-driven discoveries from complex datasets.
method Statistical and machine learning techniques for interpretable models.
result Challenges in validating data-driven discoveries remain.
SFS-DA method statistically tests FS reliability under domain adaptation.
problem Feature selection reliability under domain adaptation with limited target data.
method Selective Inference framework to control false positive rate and enhance true positive rate.
result SFS-DA method controls FPR below a pre-specified level α (e.g., 0.05) while maximizing true positive rate. New methods improve confidence set calibration in complex models.
problem Challenges in maintaining confidence set coverage in complex models.
method TRUST and TRUST++ methods using simulated data for calibration.
result Methods achieve distribution-free conditional coverage and robust inference.
New method improves SBI efficiency and scalability.
problem Scalability issues in SBI methods for large datasets.
method Langevin dynamics with score matching, exploiting likelihood structure.
result Structured score network enhances statistical efficiency and scalability.
The accurate measurement of security metrics is a critical research problem because an improper or inaccurate measurement process can ruin the usefulness of the metrics, no matter how well they are defined. This is a highly challenging problem particularly when the ground truth is unknown or noisy. In contrast to the w…
Machine learning improves official statistics but needs rigorous validation.
problem Lack of methodological robustness in machine learning for official statistics.
method Total Machine Learning Error (TMLE) framework to validate ML models.
result TMLE addresses representativeness and measurement errors in ML models.
Introduces BPEL for EL, enhancing flexibility and using MCMC for inference.
problem Computational challenges in EL methods.
method Bayesian Penalized Empirical Likelihood (BPEL) framework with MCMC sampling.
result Enhanced flexibility and practicality of EL methods with MCMC.
New lattice path method for statistical inference of persistent diagrams.
problem Statistical inference on persistent diagrams.
method Lattice path representation and combinatorial enumerations.
result Topological changes observed in spike proteins of COVID-19 virus.
The paper uses statistics to improve the explainability of models.
problem Subjective human assessment of explanations and lack of theoretical guarantees.
method Leveraging statistical estimators for proper definition and evaluation of explanations.
result Statistical tools provide theoretical guarantees and evaluation metrics for explanations.
Novel framework for ML-assisted inference valid for any statistical task.
problem Limited validity of existing methods for post-prediction inference.
method Introduces PSPS framework for task-agnostic ML-assisted inference.
result Valid and efficient inference for arbitrary ML models.
Breiman's two cultures reconciled through blending statistical thinking.
problem Tension between parametric statistical and machine learning approaches.
method Establishing a link between parametric statistical and machine learning frameworks.
result Integrated statistical thinking can bridge the gap between two cultures.
Federated learning involves training statistical models over remote devices or siloed data centers, such as mobile phones or hospitals, while keeping data localized. Training in heterogeneous and potentially massive networks introduces novel challenges that require a fundamental departure from standard approaches for l…
Summary statistics of genome-wide association studies (GWAS) teach causal relationship between millions of genetic markers and tens and thousands of phenotypes. However, underlying biological mechanisms are yet to be elucidated. We can achieve necessary interpretation of GWAS in a causal mediation framework, looking to…
New strategy debiases synthetic data generated by DGMs for improved statistical inference.
problem Bias and imprecision in synthetic data generated by DGMs impede statistical convergence and inference.
method Debiasing strategy based on debiased and targeted machine learning.
result Enhanced convergence rates and accurate estimators with easily approximated variances.
Statistical framework improves LLM chatbot ranking.
problem Improving evaluation of LLM-based chatbots through pairwise comparisons.
method Factored tie model, covariance modeling, and parameter constraints.
result Substantial improvements in modeling pairwise comparison data.
Framework for automatically assessing and correcting data quality issues without domain knowledge.
problem Ensuring data quality in datasets across various domains.
method Hybrid approach combining statistical and machine learning methods.
result Effective detection and correction of missing values, duplicates, and typographical errors.
New estimators for causal effects in DAGs with hidden variables, addressing computational and statistical challenges.
problem Estimating causal effects in DAGs with hidden variables beyond traditional criteria.
method Introduces novel one-step corrected plug-in and targeted minimum loss-based estimators for causal effects in DAGs with hidden variables.
result Root-n consistent causal effect estimates with desirable statistical properties.
Using AI predictions as data can mislead inference, study shows.
problem Misleading inference when using AI predictions instead of real data.
method Characterized statistical challenges and reviewed methods for IPD.
result High predictive accuracy doesn't ensure valid inference.
Develops a statistical framework for coherent risk estimation.
problem Constructing coherent risk estimators with sound financial and statistical properties.
method Inspired by axiomatic risk measure theory, defines coherent risk estimators through robust representations linked to L-estimators. result Demonstrates that coherence of a risk measure does not necessarily carry over to its estimators and shows alternative weight structures can lead to different outcomes.
As Internet-based commerce becomes increasingly widespread, large data sets about the demand for and pricing of a wide variety of products become available. These present exciting new opportunities for empirical economic and business research, but also raise new statistical issues and challenges. In this article, we su…