Machine learning improves official statistics but needs rigorous validation.
problem Lack of methodological robustness in machine learning for official statistics.
method Total Machine Learning Error (TMLE) framework to validate ML models.
result TMLE addresses representativeness and measurement errors in ML models.
Discusses new probabilistic morphisms and geometric methods in machine and statistical learning.
problem Addressing challenges in statistical, machine, and manifold learning.
method Introduces category of probabilistic morphisms and geometric methods.
result New insights and applications in various learning fields.
Unified platform for statistical and machine learning in bioinformatics.
problem Workflow inefficiencies in using multiple tools for data analysis.
method Automated hyperparameter optimization, feature importance analysis, statistical tests.
result Accelerates biological discovery workflows with methodological soundness.
In the first half of 2018, the Federal Statistical Office of Germany (Destatis) carried out a "Proof of Concept Machine Learning" as part of its Digital Agenda. A major component of this was surveys on the use of machine learning methods in official statistics, which were conducted at selected national and internationa…
This paper examines risks and uncertainties of changing data sources in machine learning for official statistics.
problem Risks and uncertainties associated with changing data sources in machine learning for official statistics.
method An overview of risks, causes, and repercussions of changing data sources, with a checklist of measures.
result Maintaining integrity, reliability, consistency, and relevance in official statistics.
This work uses statistical mechanics to explain AI learning.
problem Understanding the statistical principles behind AI learning.
method Starting from sample concentration behaviors, the study applies statistical mechanics principles to AI and machine learning.
result Exponential families and statistical quantities are key in AI and machine learning.
This review assesses statistical and machine learning methods for coral bleaching.
problem Coral bleaching due to rising sea temperatures and environmental factors.
method Statistical and machine learning models for predicting and analyzing coral bleaching.
result Statistical and machine learning methods are crucial for effective reef management.
DPpack offers R tools for private data analysis and machine learning.
problem Ensuring privacy in statistical analysis and machine learning.
method Differential privacy mechanisms (Laplace, Gaussian, exponential).
result User-friendly implementation of privacy-preserving models.
This paper connects functional data analysis with machine learning techniques.
problem Lack of theoretical analysis for functional depths.
method Viewing functional depths as kernel mean embeddings in machine learning.
result Facilitates answers to open questions about functional depths.
Survey on importance weighting in machine learning applications.
problem Distribution shift in supervised learning.
method Weighting objective function or probability distribution based on instance importance.
result Importance weighting can guarantee desirable statistical properties in distribution shift scenarios.
Machine learning and statistical modeling complement each other in healthcare analytics.
problem Choosing between machine learning and statistical modeling for analytics challenges.
method Choosing based on problem, data, and desired outcomes.
result Machine learning and statistical modeling are complementary, using similar principles but different tools.
Statistical downscaling of global climate models (GCMs) allows researchers to study local climate change effects decades into the future. A wide range of statistical models have been applied to downscaling GCMs but recent advances in machine learning have not been explored. In this paper, we compare four fundamental st…
Research aims to bridge statistical learning to causal models in AI.
problem Challenges in machine learning and AI related to causality.
method Transition from statistical learning to causal models.
result Progress in AI may require advances in causal modeling.
Paper reviews algebraic research in machine learning theory.
problem Understanding phase transitions in machine learning models.
method Algebraic approaches in statistical mechanics.
result Algebraic methods are essential for analyzing machine learning models with singularities.
Machine learning should incorporate maximum likelihood for better estimation.
problem Lack of rigorous foundational theory in machine learning.
method Integrate maximum likelihood estimation into machine learning models.
result Foundationally rigorous machine learning models have greater practical impact.
This paper reviews quantum machine learning from NISQ to fault tolerance.
problem The challenges and opportunities in quantum machine learning.
method Comprehensive review of quantum machine learning concepts.
result Coverage of NISQ and fault-tolerant quantum computing approaches.
MegazordNet combines stats and ML for better financial time series forecasting.
problem Forecasting financial time series is challenging due to its chaotic nature.
method MegazordNet integrates statistical features with a deep learning model.
result MegazordNet outperforms single statistical and machine learning methods in S&P 500 stock price prediction.
New theory challenges traditional machine learning assumptions.
problem Traditional machine learning theories are critiqued.
method A new theory is proposed and discussed.
result Learning true probabilities is not equivalent to other learning goals.
The paper analyzes extreme temperature forecasting using machine learning models.
problem Forecasting extreme temperatures in U.S. cities.
method Auto-Regressive Integrated Moving Average, Exponential Smoothing, Multilayer Perceptrons, Gaussian Processes.
result Multilayer Perceptrons were found to be the most effective approach for forecasting extreme temperatures.
A textbook on statistical machine learning for astronomy.
problem Uncertainty quantification in astronomical data analysis.
method Bayesian inference and classical statistical methods.
result Unified framework connecting modern and traditional methods.
Improves survey sampling with unbiased machine learning methods.
problem Design-consistent model-assisted estimation lacks a general theory for machine learning.
method Proposes a subsampling Rao-Blackwell method for design-unbiased estimation.
result Yields efficiency gains over standard methods while ensuring valid estimation.
Interpretable ML helps discover insights from big data.
problem Validating data-driven discoveries from complex datasets.
method Statistical and machine learning techniques for interpretable models.
result Challenges in validating data-driven discoveries remain.
Statistical learning theory provides the theoretical basis for many of today's machine learning algorithms. In this article we attempt to give a gentle, non-technical overview over the key ideas and insights of statistical learning theory. We target at a broad audience, not necessarily machine learning researchers. Thi…
Breiman's paper sparked debate on the future of statistics and machine learning.
problem The tension between traditional statistical modeling and model-free machine learning approaches.
method Discussion of the implications of machine learning's success and the need for new inferential approaches.
result The importance of understanding 'why' and 'if' questions in machine learning is now recognized.
Solla discusses neural processing using statistical physics and Bayesian methods.
problem Understanding neural information processing through statistical physics.
method Bayesian inference, Gibbs description, Generalized Linear Models, dimensionality reduction.
result Connection between neural processing and statistical physics.
A method to select validation data from a dataset using statistical criteria.
problem Selecting a validation basis from a full dataset for machine learning model validation.
method Adopting a 'design of experiments' point of view and using statistical criteria, particularly Maximum Mean Discrepancy criteria.
result The 'support points' concept is particularly relevant for selecting validation data.
Improves statistical inference using machine learning predictions with imputed data.
problem Invalid statistical inference due to machine learning prediction errors.
method Bootstrap confidence intervals for nonuniform samples and arbitrary imputed features.
result Valid confidence intervals without assumptions on machine learning model quality.
This chapter reviews classic regression methods and their evolution to physics-informed approaches.
problem Finding relationships between variables using regression.
method Introduces traditional and physics-informed regression methods, linking them to computational science.
result Regression methods have evolved from purely statistical to incorporating physical knowledge.
Responds to critiques on tests for causal parameter confidence intervals.
problem Testing nominal confidence interval coverage for causal parameters estimated by machine learning.
method Rejoinder to critiques on nearly assumption-free tests.
result Clarifies and supports the original research's approach.
New ML method detects incomplete bid-rigging cartels.
problem Detecting incomplete bid-rigging cartels in competitive bidding.
method Combines statistical screens with machine learning.
result Algorithm outperforms existing methods in incomplete cartels.
Statistical machine learning plays an important role in modern statistics and computer science. One main goal of statistical machine learning is to provide universally consistent algorithms, i.e., the estimator converges in probability or in some stronger sense to the Bayes risk or to the Bayes decision function. Kerne…
Kempe discusses NTK approach to machine learning problems.
problem Generally unsolvable machine learning problems.
method NTK approach focusing on kernel formulations.
result Practical applications like data distillation and adversarial robustness.
New model enhances SPIM for solving low-rank combinatorial optimization and statistical learning problems.
problem Solving large-scale combinatorial optimization problems efficiently.
method Proposed a new computing model for SPIM that can handle low-rank interaction matrices.
result Demonstrated efficient learning, classification, and sampling of MNIST images using the model.
AI needs causal inference to avoid being just a correlation machine.
problem AI's inability to distinguish correlation from causation.
method Develops a unified framework connecting various causal statistical estimators and proves a Statistical Necessity Theorem for causal generalization.
result AI systems without causal grounding are brittle and biased, highlighting the need for causal statistics.
Study shows heavy-tailed distributions affect reliability of machine learning calibration statistics.
problem Reliability of calibration statistics for machine learning regression tasks is affected by heavy-tailed uncertainty and error distributions.
method Examined two calibration error estimation methods (CE and ZMS) and found ZMS to be less sensitive to heavy-tailed distributions.
result Heavy-tailed distributions make MSE and MV unreliable, but ZMS remains a reliable approach.
Method controls extrapolation in prediction profiles for statistical and machine learning models.
problem Avoiding invalid predictions due to extrapolation in prediction profiles.
method Genetic algorithm optimization over constrained factor regions.
result Optimal factor settings without constraint are often invalid and extrapolated.
Active inference uses machine learning to prioritize data labeling for more efficient statistical inference.
problem Efficiently collecting data points for statistical inference with limited labels.
method A machine learning-assisted approach that identifies uncertain data points for labeling.
result Achieves the same level of accuracy with fewer samples, resulting in smaller confidence intervals and more powerful p-values.
Unified framework for analyzing neural networks in high dimensions.
problem Understanding neural networks' efficiency in high-dimensional data.
method Statistical physics techniques, including replica method and approximate message-passing algorithms.
result Unified analysis of various machine learning architectures and tasks.
This paper reviews statistical and machine learning methods for anti-money laundering.
problem Lack of scientific literature on statistical and machine learning methods for anti-money laundering.
method Client risk profiling and suspicious behavior flagging.
result Client risk profiling involves diagnostics, while suspicious behavior flagging involves non-disclosed features and hand-crafted risk indices.
Study on kernel methods in large-scale machine learning problems.
problem Large-scale machine learning with many interacting variables.
method Mean field limit analysis of kernels and their Hilbert spaces.
result Mean field convergence of empirical and infinite-sample solutions.
We present two new statistical machine learning methods designed to learn on fully homomorphic encrypted (FHE) data. The introduction of FHE schemes following Gentry (2009) opens up the prospect of privacy preserving statistical machine learning analysis and modelling of encrypted data without compromising security con…
Establishes statistical and computational bounds for influence diagnostics.
problem Identifying influential datapoints or subsets in machine learning models.
method Finite-sample statistical bounds and computational complexity for influence functions and approximate maximum influence perturbations.
result Established statistical and computational guarantees for influence diagnostics.
This paper introduces a novel measure-theoretic theory for machine learning that does not require statistical assumptions. Based on this theory, a new regularization method in deep learning is derived and shown to outperform previous methods in CIFAR-10, CIFAR-100, and SVHN. Moreover, the proposed theory provides a the…
The relevance of data quantifies learning efficiency.
problem Understanding the statistical nature of high-dimensional, sparse data.
method Defining relevance as information content, and using it to define ideal limits of samples and learning machines.
result Maximally informative samples and optimal learning machines exhibit critical features like power-law frequency distributions and anomalously large susceptibility.
MM (majorization--minimization) algorithms are an increasingly popular tool for solving optimization problems in machine learning and statistical estimation. This article introduces the MM algorithm framework in general and via three popular example applications: Gaussian mixture regressions, multinomial logistic regre…
SMILE improves explainability of machine learning models.
problem Difficulty in understanding and trusting the conclusions of black-box machine learning models.
method Statistical Model-agnostic Interpretability with Local Explanations (SMILE).
result SMILE makes machine learning models more interpretable.
Statistical physics helps solve complex machine learning problems.
problem Large dimensional inference problems in machine learning.
method Replica symmetric level analysis and cavity methods.
result General framework for solving various problems with weak long-range interactions.
New framework assesses extreme errors in machine learning models.
problem Current validation methods fail to quantify extreme errors in high-stakes domains.
method Uses Extreme Value Theory (EVT) to estimate worst-case failures.
result Establishes EVT as a fundamental tool for assessing model reliability.