Survey of statistical queries and their applications.
problem Understanding statistical queries and their applications.
method Exploration of statistical queries model, definitions, and connections to learnability.
result Connections to learnability and applications in optimization, evolvability, and differential privacy.
Improves survey sampling with unbiased machine learning methods.
problem Design-consistent model-assisted estimation lacks a general theory for machine learning.
method Proposes a subsampling Rao-Blackwell method for design-unbiased estimation.
result Yields efficiency gains over standard methods while ensuring valid estimation.
Machine learning improves official statistics but needs rigorous validation.
problem Lack of methodological robustness in machine learning for official statistics.
method Total Machine Learning Error (TMLE) framework to validate ML models.
result TMLE addresses representativeness and measurement errors in ML models.
Survey of recent measures of association, including a new coefficient.
problem Exploring new measures of association in statistics.
method Survey and introduction of a new correlation coefficient.
result Proposed a new extension of the correlation coefficient to standard Borel spaces.
Survey on importance weighting in machine learning applications.
problem Distribution shift in supervised learning.
method Weighting objective function or probability distribution based on instance importance.
result Importance weighting can guarantee desirable statistical properties in distribution shift scenarios.
In the first half of 2018, the Federal Statistical Office of Germany (Destatis) carried out a "Proof of Concept Machine Learning" as part of its Digital Agenda. A major component of this was surveys on the use of machine learning methods in official statistics, which were conducted at selected national and internationa…
Survey on statistical inference under memory constraints.
problem Effect of memory limitations on statistical inference performance.
method Review of state-of-the-art in several canonical problems.
result Identification of fundamental building blocks and useful techniques.
Survey on using low-degree polynomials to assess statistical tasks complexity.
problem Understanding the complexity of statistical tasks using polynomial functions.
method Applying low-degree polynomials to measure the complexity of statistical tasks, including detection, recovery, and estimation.
result Low-degree polynomials provide a framework to predict and explain statistical-computational tradeoffs.
Survey of extreme value modeling techniques for insurance.
problem Modeling of insurance industry's extreme events.
method Truncation, tempering, censoring, regression techniques.
result Adapted techniques for insurance applications.
Survey on factor models and their applications in econometrics.
problem Estimating low-rank structures in high-dimensional models.
method Low-rank recovery techniques for factor model estimation.
result New insights into factor model applications in econometrics.
Survey of robust clustering methods for hotspot detection.
problem Detecting false positives in spatial hotspot mapping.
method Statistically rigorous clustering techniques.
result Survey of models and algorithms for robust clustering.
The paper provides a survey of results related to the "κ-generalized distribution", a statistical model for the size distribution of income and wealth. Topics include, among others, discussion of basic analytical properties, interrelations with other statistical distributions as well as aspects that are of special in…
In certain situations that shall be undoubtedly more and more common in the Big Data era, the datasets available are so massive that computing statistics over the full sample is hardly feasible, if not unfeasible. A natural approach in this context consists in using survey schemes and substituting the "full data" stati…
This paper surveys techniques to personalize federated learning models.
problem Personalized models outperform shared models for some clients, reducing participation.
method Surveys recent research on personalizing federated learning models.
result Personalization techniques improve model performance for individual clients.
Survey of deep learning methods for time series forecasting.
problem Improving accuracy in time series predictions across various domains.
method Analysis of common encoder and decoder designs, hybrid models, and decision support.
result Advancements in deep learning for time series forecasting.
Survey of robust statistical methods for efficient computation.
problem Efficient robust statistical methods for various forms of data contamination and heavy-tailed distributions.
method Survey and technical connections between robustness forms, showing efficient algorithms.
result Same algorithmic ideas lead to efficient estimators for robustness in different settings.
AI-assisted interviews allow respondents to describe experiences naturally, but mapping those accounts into structured survey variables is fallible.
problem Mapping AI-assisted interview responses into structured survey variables is fallible.
method Adaptive Matrix Validation (AMV) is proposed, which involves mapping responses into tabular data and using a small set of structured questions for statistical adjustment.
result The estimator calibrates mapped values using validation answers from other respondents and corrects remaining error with validation answers observed for the target respondent.
PPI uses survey sampling methods for inference, bridging ML and statistics.
problem Combining machine learning predictions with small labeled data for valid inference.
method Equivalence of PPI estimators to survey sampling methods.
result PPI estimators are algebraically equivalent to survey sampling methods.
Surveying nonparametric inference with shape constraints, past and future.
problem Statistical inference under shape constraints.
method Historical overview and future directions.
result Outlook on future research directions.
Paper extends conformal prediction to complex survey data.
problem Applying distribution-free prediction intervals to complex survey data.
method Design-based conformal prediction for non-exchangeable data.
result Empirical guarantees of finite-sample coverage for complex survey data.
Surveying joint Gaussian graphical models to identify shared structures across domains.
problem Estimating shared structures across different data sources.
method Statistical inference of joint Gaussian graphical models.
result Improved estimation power for high-dimensional data.
Survey on deep learning methods for stock market prediction.
problem Lack of comprehensive survey on deep learning methods for stock market prediction.
method Propose a novel taxonomy summarizing state-of-the-art models based on deep neural networks.
result Provide detailed statistics on datasets and evaluation metrics.
Survey on statistical learning theory for control, focusing on linear systems.
problem Applying machine learning techniques to control systems, especially linear ones.
method Adapting tools from modern high-dimensional statistics and learning theory.
result Recent advances in statistical learning theory for control, particularly for linear systems.
This survey explores causal inference in banking, finance, and insurance.
problem Explaining decisions in banking, finance, and insurance using causal inference.
method Categorizes 37 papers on causal inference applications in banking, finance, and insurance.
result Causal inference is still in its infancy in banking and insurance sectors.
Surveying random sections on Kähler manifolds, leading to metrics.
problem Understanding statistics of random sections on Kähler manifolds.
method Analyzing tensor powers of line bundles.
result Induced metrics from random sections.
Bayesian nonparametrics adapt model complexity to diverse datasets.
problem Complex challenges across statistics, computer science, and engineering.
method Flexible Bayesian nonparametric models that adapt model complexity.
result Bayesian nonparametrics offer innovative solutions to multi-object tracking.
The rapid development of computing power and efficient Markov Chain Monte Carlo (MCMC) simulation algorithms have revolutionized Bayesian statistics, making it a highly practical inference method in applied work. However, MCMC algorithms tend to be computationally demanding, and are particularly slow for large datasets…
Survey of robust streaming techniques and their relationships.
problem Challenges in robust streaming and online learning.
method Overview and survey of robust streaming techniques, unifying theorems.
result Proved the relationship between robust streaming techniques.
Online surveys have the potential to support adaptive questions, where later questions depend on earlier responses. Past work has taken a rule-based approach, uniformly across all respondents. We envision a richer interpretation of adaptive questions, which we call dynamic question ordering (DQO), where question order …
Survey on learning with graph-dependent data, deriving new generalization bounds.
problem Traditional i.i.d. data assumption fails in many real-life applications.
method Collect and analyze graph-dependent concentration bounds, derive generalization bounds.
result New generalization bounds for graph-dependent data.
A textbook on statistical machine learning for astronomy.
problem Uncertainty quantification in astronomical data analysis.
method Bayesian inference and classical statistical methods.
result Unified framework connecting modern and traditional methods.
The study introduces backward baselines to distinguish past prediction from future prediction in machine learning models.
problem Differentiating between past and future prediction in machine learning models.
method Theoretical, empirical, and normative arguments support a family of simple and efficient statistical tests called backward baselines.
result The study provides a meaningful backward baseline for auditing black-box prediction systems.
SDRF estimates complex survey designs for conditional distributions.
problem Estimating conditional distributions under complex survey designs.
method Survey-calibrated distributional random forest (SDRF) with pseudo-population bootstrap and MMD split criterion.
result Established design consistency and model consistency for survey designs.
Survey of universal portfolio techniques for minimizing investment regret.
problem Minimizing investment regret in algorithmic trading.
method Explains various universal portfolio techniques and their proofs.
result Coverage of fundamental concepts and algorithms in regret minimization.
Financial market prediction on the basis of online sentiment tracking has drawn a lot of attention recently. However, most results in this emerging domain rely on a unique, particular combination of data sets and sentiment tracking tools. This makes it difficult to disambiguate measurement and instrument effects from f…
In this short paper, we overview and extend the results of our papers cond-mat/0001432, cond-mat/0008305, and cond-mat/0103544, where we use an analogy with statistical physics to describe probability distributions of money, income, and wealth in society. By making a detailed quantitative comparison with the available …
New method improves active statistical inference by reducing noise.
problem Inaccurate uncertainty estimates in active sampling lead to noisy results.
method Robust sampling strategies that interpolate between uniform and active sampling based on uncertainty scores.
result The robust sampling ensures that the estimator is never worse than uniform sampling and usually outperforms active inference.
FOSS is an acronym for Free and Open Source Software. The FOSS 2013 survey primarily targets FOSS contributors and relevant anonymized dataset is publicly available under CC by SA license. In this study, the dataset is analyzed from a critical perspective using statistical and clustering techniques (especially multiple…
In this survey, a short introduction in the recent discovery of log-normally distributed market-technical trend data will be given. The results of the statistical evaluation of typical market-technical trend variables will be presented. It will be shown that the log-normal assumption fits better to empirical trend data…
An efficient LDP protocol for QMLE with improved practicality and theoretical guarantees.
problem Difficult implementation of existing LDP QMLE for large-scale surveys.
method Developed an alternative LDP protocol without long waiting time, high communication cost, and derivative boundedness assumptions.
result Sufficient conditions for consistency and asymptotic normality of the protocol.
This paper reviews various sampling methods from statistics and machine learning.
problem Addressing sampling methods in statistics and machine learning.
method Explains and reviews simple random sampling, bootstrapping, stratified sampling, cluster sampling, multistage sampling, network sampling, snowball sampling, and sampling from cumulative distribution function.
result Summarizes characteristics, pros, and cons of different sampling methods.
This paper provides a tutorial on Boltzmann Machines and Deep Belief Networks.
problem Understanding and applying Boltzmann Machines and Deep Belief Networks.
method Explains the structures, conditional distributions, Gibbs sampling, training methods, and deep belief networks of RBMs.
result Comprehensive overview of RBMs and DBNs, useful in various fields.
We survey some of the recent advances in mean estimation and regression function estimation. In particular, we describe sub-Gaussian mean estimators for possibly heavy-tailed data both in the univariate and multivariate settings. We focus on estimators based on median-of-means techniques but other methods such as the t…
ESRLCM clusters similar responses, more broadly than traditional models.
problem Clustering multivariate categorical data with common response patterns.
method Bayesian Equivalence Set Restricted Latent Class Model (ESRLCM).
result ESRLCM identifies clusters with similar item response probabilities.
A new method models galaxies as points in space for better analysis.
problem Limitations of binning and voxelization in galaxy surveys.
method A diffusion-based generative model for galaxy point clouds.
result Demonstrated on dark matter haloes in Quijote simulations.
Audio fingerprinting, also named as audio hashing, has been well-known as a powerful technique to perform audio identification and synchronization. It basically involves two major steps: fingerprint (voice pattern) design and matching search. While the first step concerns the derivation of a robust and compact audio si…
Online crowdsourcing provides a scalable and inexpensive means to collect knowledge (e.g. labels) about various types of data items (e.g. text, audio, video). However, it is also known to result in large variance in the quality of recorded responses which often cannot be directly used for training machine learning syst…
Recently, a number of statistical problems have found an unexpected solution by inspecting them through a "modal point of view". These include classical tasks such as clustering or regression. This has led to a renewed interest in estimation and inference for the mode. This paper offers an extensive survey of the tradi…