New method estimates model parameters from incomplete data.
problem Estimating model parameters from incomplete data.
method Variational Gibbs Inference (VGI)
result Competitive or better performance compared to existing methods.
DACE estimates covariance from compressed data, improving accuracy.
problem Estimating covariance from large, distributed data.
method Data-aware weighted sampling for unbiased estimation.
result DACE provides more accurate covariance estimation under compression.
Paper estimates spectral risk measures for insurance data with truncated and censored data.
problem Estimating spectral risk measures for insurance data with left truncation and right censoring.
method Proposes a non-parametric estimator using product limit estimator and establishes asymptotic normality.
result Proposed estimator outperforms existing methods for small k and small sample sizes.
Robust deep neural networks estimate multi-dimensional functional data robustly.
problem Estimating location function from multi-dimensional functional data robustly.
method Deep neural networks with ReLU activation, robust to outliers and model misspecification.
result Uniform convergence rates for robust deep neural network estimators.
New estimator robust to adversarial noise and data heterogeneity.
problem Sensitive to adversarial noise and poor performance with heterogeneous data.
method Distributionally robust estimator minimizing worst-case conditional expected loss over adversarial distributions.
result Efficiently finds non-parametric local estimates via convex optimization.
Optimal and safe semi-supervised learning estimator for high-dimensional data.
problem Improving regression parameter estimation with unlabeled data in high-dimensional settings.
method Established minimax lower bound, proposed optimal and safe semi-supervised estimators.
result Optimal semi-supervised estimator achieves the minimax lower bound.
New method estimates mutual information using normalizing flows.
problem Mutual information estimation in high-dimensional data.
method Normalizing flows to map data to target distributions with known MI.
result Theoretical guarantees and practical advantages demonstrated.
Paper uses Super-App data to improve income estimation models.
problem Improving accuracy of income estimation models.
method TreeSHAP method for Stochastic Gradient Boosting Interpretation.
result Alternative data from Super-Apps capture more information than traditional financial data.
Survey on mean estimation and regression for heavy-tailed data.
problem Estimating mean and regression functions in heavy-tailed distributions.
method Sub-Gaussian mean estimators, median-of-means, trimmed mean, Catoni's estimator.
result Detailed proofs for estimators in heavy-tailed settings.
TraDE uses self-attention for better density estimation of tabular and image data.
problem Improving density estimation for tabular and image data.
method Self-attention-based architecture trained with a penalized maximum likelihood objective.
result TraDE produces significantly better density estimates than existing methods.
Graph neural networks extend neural Bayes estimators to irregular spatial data.
problem Estimating parameters from irregular spatial data with computational efficiency.
method Employing graph neural networks to approximate Bayes estimators for irregular spatial data.
result Extending neural Bayes estimation to irregular spatial data with computational benefits.
Generative method avoids function estimation for data generation.
problem Challenges in function estimation for generative models.
method Deterministic point transport with gradient descent.
result Data generation possible without function estimation.
Improved SV estimator for efficient data valuation.
problem Computational inefficiency in Shapley value estimation.
method Group Testing-based SV estimator with improvements.
result Enhanced asymptotic sample complexity and insights into challenges.
Paper proposes robust estimators for heavy-tailed data with infinite variance.
problem Developing robust estimators for heavy-tailed data with infinite variance.
method Proposes two robust estimators: ridge log-truncated M-estimator and elastic net log-truncated M-estimator.
result Demonstrates robustness of log-truncated estimations over standard estimations through simulations and real data analysis.
Estimates copula density for complex data distributions.
problem Estimating copula density from observed data.
method Neural network-based copula density neural estimation (CODINE).
result Novel approach capable of modeling complex distributions.
The paper resolves the paradox of using unlabeled data for treatment effect estimation.
problem Using unlabeled data to estimate propensity scores for treatment effect estimation.
method Proposes a simple procedure to reconcile the use of estimated propensity scores with the advice to use true propensity scores.
result Direct regression may be preferable to inverse-propensity weighting in many circumstances.
RCUKF combines data-driven modeling and Bayesian estimation for accurate system state estimation.
problem Challenges in obtaining reliable process models for complex systems.
method Integrates reservoir computing with unscented Kalman filtering.
result Demonstrated effectiveness on benchmark problems and real-time vehicle trajectory estimation.
Deep learning improves causal effect estimation from complex observational data.
problem Estimating causal effects from complex observational data with low bias.
method Unified deep learning framework using multitask recurrent neural networks.
result Deep learning estimator shows lower bias in causal effect estimates.
New estimator improves mutual information estimation.
problem Estimating mutual information in data science and machine learning.
method Proposes a new estimator that uses a preliminary estimate of the data distribution.
result A preliminary estimate helps in estimating mutual information more accurately.
New data improves market impact estimation methods.
problem Improving efficiency of market impact estimation.
method Investigates the use of price trajectory data for market impact estimation.
result Estimation methods using early trade prices outperform established methods asymptotically.
Proposes a robust method for predicting missing outcomes in covariate shift adaptation.
problem Predicting missing outcomes in test data with covariate shift.
method Doubly robust estimator for covariate shift adaptation via importance weighting, incorporating an additional estimator for the regression function.
result Shows robustness against density-ratio estimation errors, maintaining consistency if either estimator is consistent.
New methods for estimating treatment effects with missing data.
problem Missing outcome data complicates estimating treatment effects.
method Proposed two de-biased machine learning estimators (mDR-learner and mEP-learner) to address under-representation.
result Oracle efficiency of the proposed estimators under reasonable conditions.
New local ID estimators based on data separability.
problem Estimating intrinsic dimensionality locally in multi-dimensional data.
method Local estimators based on concentration of measure.
result Empirical comparison with other ID estimators.
Paper tackles causal effect estimation in observational data with hidden variables.
problem Estimating causal effects in observational data with hidden confounders.
method Developed a theorem for local search to find superset of adjustment variables, proposing a data-driven algorithm.
result Proposed algorithm produces more accurate causal effect estimates than existing methods.
Robustly estimates mean in incomplete data with corrupted examples.
problem Estimating mean in data with missing values and outliers.
method Algorithms for robust estimation with optimal error guarantees in nearly-linear time.
result Information-theoretically optimal error guarantees for mean estimation.
Improved quantile estimation using semi-supervised data.
problem Quantile estimation in high-dimensional settings with limited labeled data.
method Proposes semi-supervised estimators using a flexible imputation strategy and debiasing step.
result Improved estimation accuracy compared to supervised methods, robust to misspecification.
Paper improves ML estimation from incomplete data with robust M-estimator.
problem Estimating parameters from incomplete data with improved accuracy.
method Developed a robust M-estimator and a sandwich estimator for standard errors.
result Improved estimation accuracy with smaller standard errors than ML estimates.
Paper compares LSTM and GARCH for estimating value-at-risk.
problem Estimating value-at-risk on time series with heteroscedastic dynamics.
method Uses LSTM neural networks to estimate value-at-risk compared to GARCH benchmarks.
result LSTM outperforms GARCH on real market data in terms of exception rate and mean quantile score.
We present a multi-task learning approach to jointly estimate the means of multiple independent data sets. The proposed multi-task averaging (MTA) algorithm results in a convex combination of the single-task maximum likelihood estimates. We derive the optimal minimum risk estimator and the minimax estimator, and show t…
VAE leverages MMSE channel estimation with data-driven modeling.
problem Data-driven channel estimation for wireless communications.
method Variational autoencoder (VAE) modeling of channel distribution and LMMSE approximation.
result VAE-based channel estimators approximate MMSE performance with practical training methods.
Study nonparametric covariance function estimation for noisy data.
problem Estimating covariance function from discrete noisy data in high dimensions.
method Adaptive learning-based estimators, including deep learning.
result Established oracle inequality and convergence rates for deep learning estimators.
Combines multiple OPE estimators into a more accurate and efficient estimate.
problem Offline evaluation of recommender systems using biased data.
method Meta-analysis of correlated OPE estimators, accounting for inter-estimator correlation.
result Improved statistical efficiency and accuracy in estimating policy value.
Density Estimation is one of the central areas of statistics whose purpose is to estimate the probability density function underlying the observed data. It serves as a building block for many tasks in statistical inference, visualization, and machine learning. Density Estimation is widely adopted in the domain of unsup…
Modes and ridges of the probability density function behind observed data are useful geometric features. Mode-seeking clustering assigns cluster labels by associating data samples with the nearest modes, and estimation of density ridges enables us to find lower-dimensional structures hidden in data. A key technical cha…
Paper proposes a method to estimate total variation distance for synthetic data fidelity.
problem Assessing the fidelity of synthetic data generated by AI.
method Discriminative approach to estimate total variation distance between two distributions.
result Estimation of total variation distance reduces to quantifying Bayes risk in classification.
Study examines mean estimation in high dimensions with small data.
problem Efficiently estimating mean in high-dimensional data with limited data size.
method Extensive experimentation of various mean estimation techniques.
result Developed robust methods for mean estimation with low data size.
New framework estimates graph from multimodal functional data.
problem Estimating graph from joint multimodal functional data.
method Integrative framework using partial correlation operator.
result Estimator converges to stationary point with quantifiable error.
A new method estimates the number of clusters on spherical data.
problem Estimating the number of clusters in spherical data.
method Spherical X-means (SX-means) method assuming von Mises-Fisher distributions.
result Shows the performance of SX-means in estimating the number of clusters.
Estimates population mean from user-level data with privacy, accounting for heterogeneity.
problem Heterogeneous user data with varying numbers of data points and distributions.
method Simple model of heterogeneous user data, differential privacy mechanism for estimation.
result Asymptotic optimality of the proposed estimator and general lower bounds on error.
Measuring Mutual Information (MI) between high-dimensional, continuous, random variables from observed samples has wide theoretical and practical applications. Recent work, MINE (Belghazi et al. 2018), focused on estimating tight variational lower bounds of MI using neural networks, but assumed unlimited supply of samp…
Ranked data appear in many different applications, including voting and consumer surveys. There often exhibits a situation in which data are partially ranked. Partially ranked data is thought of as missing data. This paper addresses parameter estimation for partially ranked data under a (possibly) non-ignorable missing…
Method estimates causal effects from incremental data, overcoming missing data challenges.
problem Estimating causal effects from non-stationary, incrementally available observational data.
method Continual Causal Effect Representation Learning
result Method achieves continual causal effect estimation without compromising original data.
Combines public and private data for better statistical estimation.
problem Estimating aggregate statistics from mixed data with varying privacy needs.
method Mixed estimators optimized for minimizing variance or median, using differential privacy techniques.
result Our mechanisms often outperform baseline methods in empirical tests.
New model estimates species population trends from citizen science data.
problem Interannual confounding in citizen science data.
method Double Machine Learning framework to estimate population change and propensity scores for confounding adjustment.
result Spatially detailed trend estimates from citizen science data with low error rates.
New methods estimate covariance for matrix data without assuming fixed size or specific distributions.
problem Estimating covariance for high-dimensional matrix data without distributional assumptions.
method Unified framework for bandable covariance estimation with rank one approximation, robust to heavy-tailed data.
result Proposed estimators are rate-optimal and perform well in simulations and real applications.
Estimates causal effects in Gaussian Linear SCMs with finite data.
problem Estimating causal effects from observational data with latent confounders.
method Centralized Gaussian Linear SCMs (CGL-SCMs) and EM-based estimation algorithm.
result Learned CGL-SCM parameters accurately recover causal distributions from finite observational samples.
Estimates neural drift for stochastic equations, improving inference on noisy data.
problem Estimating drift in stochastic differential equations with neural networks.
method Non-parametric estimation using ReLU neural networks, enforcing theoretical bounds.
result Practical method for inference on noisy and rough functional data.
Estimates non-parametric logistic model using case-control data and external summary info.
problem Imbalanced binary data in case-control studies.
method Two-step estimation procedure with deep neural network for functional approximation.
result Proposed estimator achieves optimal convergence rate in non-parametric regression.