Robust variable selection for high-dimensional data with missing and measurement errors.
problem Missing data and measurement errors confound data distribution.
method Exponential loss function with inverse probability weighting and additive error models.
result The Atan punishment method improves robust variable selection.
Latent diffusion improves robustness in missing data imputation.
problem Missing data imputation under MCAR corruption.
method Two-stage framework: VAE for latent feature learning, diffusion model in latent space.
result Latent diffusion maintains high quality and stability up to 50% missingness.
Merlin improves robustness of MTSF models to missing data.
problem Suboptimal forecasting performance due to unfixed missing rates in MTSF models.
method Offline knowledge distillation and multi-view contrastive learning.
result Merlin enhances robustness of MTSF models while preserving accuracy.
Proves subspace tracking with missing data and improves matrix completion.
problem Subspace tracking in the presence of missing data.
method Modified robust subspace tracking algorithm.
result Proves subspace estimates are close to true subspaces under mild assumptions.
Paper develops adaptive models for robust energy forecasting with missing data.
problem Operational models assume complete data; missing data can degrade forecast accuracy.
method Adaptive robust optimization and adversarial machine learning for missing data.
result Proposed models perform well even with short-term missing data and significantly outperform imputation with longer-term missing data.
New estimator for symmetric kernel expectations, robust to missing data.
problem Efficient estimation of symmetric kernel expectations with missing data.
method Median-of-Incomplete-U-Statistics (MIU) estimator.
result Established finite-sample concentration rate for MIU.
A novel Bayesian approach for latent variable models from mixed data with missing values.
problem Learning parameters of latent variable models from mixed (continuous and ordinal) data with missing values.
method Proposes a novel Bayesian Gaussian copula factor (BGCF) approach that is consistent and robust to violations of certain conditions.
result BGCF substantially outperforms two state-of-the-art alternative approaches in simulations and is favorable over robust maximum likelihood (MLR) in practice.
Scalable and robust TR decomposition for large-scale data with missing entries and outliers.
problem Handling large-scale tensor data with missing entries and outliers.
method Auto-weighted steepest descent method for missing entries and outliers identification, FGMC and RStS strategies.
result Outperforms existing TR decomposition methods in the presence of outliers and runs faster than robust tensor completion algorithms.
Similarity-based approaches represent a promising direction for time series analysis. However, many such methods rely on parameter tuning, and some have shortcomings if the time series are multivariate (MTS), due to dependencies between attributes, or the time series contain missing data. In this paper, we address thes…
Develops kernel machines for missing response data.
problem Missing responses in data.
method Proposes kernel machine families for handling missing responses, including doubly-robust estimators.
result Oracle inequalities and consistency proved for kernel machine estimators.
Proposes a method to evaluate classifiers with missing labels using multiple imputation.
problem Missing labels during model evaluation can introduce bias, especially in Missing Not At Random (MNAR) data.
method Develops a multiple imputation technique to estimate and provide predictive distributions for metrics like precision, recall, and ROC-AUC.
result The predictive distribution's location and shape are generally correct, even in the MNAR regime.
Robust Lasso-Zero handles missing covariates and sparse corruptions.
problem Sparse corruptions and missing covariates in sparse linear models.
method Extension of Lasso-Zero to handle sparse corruptions, with theoretical guarantees on sign recovery.
result Robust Lasso-Zero can handle missing values without specifying a parametric model.
Graphon estimation improved with extended NBS for missing subgraph data.
problem Estimating graphon from partially observed network data.
method Extended NBS algorithm applied to missing data.
result Extended NBS achieves smaller error rates than other methods.
Bayes predictor remains robust to ignorable missingness shifts.
problem Challenges in prediction with missing covariates and shifts in missingness reasons.
method Bayesian approach and different prediction methods.
result Bayes predictor remains unchanged by ignorable shifts, but robust prediction requires disregarding missingness for non-ignorable shifts.
Paper improves robust PCA for noisy, outlier, and missing data.
problem Robust PCA with noise, outliers, and missing data.
method Bridging convex and nonconvex optimization.
result Near-optimal statistical accuracy for robust PCA.
New algorithm for robust Boolean matrix factorization handles noise and missing data.
problem Robust probabilistic Boolean matrix factorization in the presence of noise and missing values.
method Probabilistic Expectation Maximization algorithm without latent factor assumptions.
result Outperforms state-of-the-art probabilistic algorithms on real data.
We tackle missing data in SBI methods and introduce a neural process approach.
problem Missing data in SBI methods can bias parameter estimation.
method We introduce a neural process approach to jointly learn imputation and inference.
result Our method provides robust inference outcomes compared to baselines.
New EM algorithm for mixtures of elliptical distributions handles missing data and outliers.
problem Missing data imputation for noisy and non-Gaussian data.
method Investigation of a new EM algorithm for mixtures of elliptical distributions.
result The proposed algorithm is robust to outliers and competitive with other methods.
LRMM learns to recommend with missing modalities, improving robustness to data sparsity and cold-start issues.
problem Learning to recommend with missing modalities and cold-start problems.
method LRMM uses modality dropout and multimodal sequential autoencoder to learn multimodal representations and impute missing modalities.
result LRMM achieves state-of-the-art performance on rating prediction tasks and is more robust to data sparsity and cold-start issues.
Generative imputation network improves missing value estimation.
problem Missing data in statistical analysis.
method Integrates generative networks with chained equations for robust multiple imputation.
result GCMI outperforms other imputation techniques in simulations and real data.
Proposes SDRG to adjust missingness in machine learning models.
problem Systemic missingness in observational data leads to biased parameter estimation.
method Introduces SDRG using two models: weight-corrected gradients and per-covariate control variates.
result Empirically demonstrates convergence in training image classifiers with missing data.
ptype infers data types robustly in real-world data.
problem Type inference fails with missing data and anomalies.
method Probabilistic robust type inference method.
result Outperforms existing methods.
DSF improves EEG model robustness to missing channels and noise.
problem Robust learning from corrupted EEG data with missing channels.
method Dynamic Spatial Filtering (DSF) as a multi-head attention module.
result DSF achieves up to 29.4% accuracy improvement over baseline models in noisy conditions.
New method calibrates crypto option prices more robustly.
problem Large bid-ask spreads and missing quotes in crypto markets.
method Designs a novel calibration procedure for crypto options.
result Calibration is more robust and accurate than standard methods.
StableDR stabilizes doubly robust learning for biased recommendation data.
problem Data missing not at random in recommender systems.
method StableDR, a stabilized doubly robust learning approach.
result StableDR achieves bounded bias, variance, and generalization error.
Trinary decision tree improves handling of missing data in machine learning.
problem Improving accuracy in decision tree algorithms when dealing with missing data.
method Introduces Trinary decision tree, which does not assume missing values contain information about the response.
result Trinary decision tree outperforms other algorithms in Missing Completely at Random settings, especially when data is only missing out-of-sample.
Estimator improves prediction with missing data in multi-environment settings.
problem Handling missing data in multi-environment settings for robust prediction.
method Derive an estimator from invariance objective under missing outcomes.
result The estimator achieves lower prediction error despite using a biased imputation model.
NeuMiss networks tackle supervised learning with missing values, offering efficient and robust predictions.
problem Challenges in supervised learning with missing values, especially when the response is a linear function of the complete data.
method Derive analytical form of optimal predictor under linearity assumption and various missing data mechanisms. Propose NeuMiss networks using multiplication by missingness indicator.
result Upper bound on Bayes risk and good predictive accuracy with independent complexity of missing data patterns.
Paper tackles HMM learning with unknown missing observation locations.
problem Learning HMMs with missing data locations unknown.
method Proposes reconstruction algorithms without structural assumptions.
result Can reconstruct process dynamics as if missing locations were known.
A new method for handling missing values in data.
problem Handling missing values in machine learning models.
method Sharing pattern submodels with sparsity-inducing regularization.
result Sharing pattern submodels provide robust predictions and maintain/improve pattern submodel performance.
Method tackles missing covariates in large-scale datasets.
problem Cross-population missing data problem in large-scale datasets.
method Augmented transfer regression learning method combining importance-weighted estimating equations and imputation terms.
result Estimator is n1/2-consistent and asymptotically normal, attaining semiparametric efficiency bound under correct specification. RTC-GTNLN model recovers traffic data from missing values and noise.
problem Simultaneous missing data and noise in traffic data.
method Gradient tensor nuclear L1-L2 norm for robust tensor completion.
result RTC-GTNLN model outperforms existing methods in complex recovery scenarios.
VAEs struggle with missing data imputation, especially for extreme values.
problem Imputation of missing data in complex, non-linear relationships.
method Investigated variational autoencoders (VAEs) for multiple imputation and improved with β-VAEs.
result β-VAEs provide better uncertainty calibration and avoid false discoveries.
In medical domain, data features often contain missing values. This can create serious bias in the predictive modeling. Typical standard data mining methods often produce poor performance measures. In this paper, we propose a new method to simultaneously classify large datasets and reduce the effects of missing values.…
Proposes a method to make statistical inferences robust in spatially dependent settings with missing at random labels.
problem Statistical inference challenges with missing at random labels and spatial dependence.
method Doubly robust estimator with cross-fit nuisances and jackknife spatial HAC variance correction.
result Asymptotically valid confidence intervals with improved finite-sample calibration.
CoIFNet unifies imputation and forecasting for robust multivariate time series prediction with missing values.
problem Pervasive missing values degrade multivariate time series forecasting accuracy.
method CoIFNet integrates imputation and forecasting through Cross-Timestep Fusion and Cross-Variate Fusion modules.
result CoIFNet achieves 24.40% improvement over state-of-the-art methods at 0.6 point (block) missing rate.
New algorithm improves GAN performance with minimal labels.
problem Improving GAN performance with little supervision.
method Intentionally corrupts generated labels to match real data statistics, trains discriminator with corrupted labels.
result Minimizing proposed loss is equivalent to minimizing true divergence between real and generated data.
SRTC model for background/foreground separation with missing pixels.
problem Background/foreground separation with missing pixels in videos.
method Smooth robust tensor completion (SRTC) model with tensor proximal alternating minimization (tenPAM).
result Global convergence guarantee for the proposed algorithm.
New model handles missing data effectively in autoregressive models.
problem Handling missing data in autoregressive models.
method Reinterpret existing models through missing data lens, introduce principled framework for incomplete datasets, active information acquisition.
result MO-ARM consistently outperforms imputation baselines across real-world benchmarks.
New algorithm tracks subspaces with missing and corrupted data, simpler and federated.
problem Subspace tracking with missing and corrupted data.
method Proposes a novel algorithm that does not assume piecewise constant subspace changes and is simpler.
result Guarantees for both subspace tracking with missing data and outliers.
New model clusters mixed-type data with missing values, improving air quality analysis.
problem Clustering mixed-type data with missing values and regime persistence.
method Statistical jump model incorporating regime persistence and handling missing data.
result Superior performance in inferring persistent air quality regimes compared to traditional methods.
Random forest (RF) missing data algorithms are an attractive approach for dealing with missing data. They have the desirable properties of being able to handle mixed types of missing data, they are adaptive to interactions and nonlinearity, and they have the potential to scale to big data settings. Currently there are …
Improved nearest neighbors for missing data in latent factor models.
problem Estimation with missing data in latent factor models.
method Doubly robust nearest neighbors (NN) that considers both row and column neighbors.
result Provides a consistent estimate and near-quadratic improvement in error.
GPCCA integrates multi-modal data with missing values, improving clustering accuracy.
problem Integrating and analyzing multi-modal data with missing values and partial observations.
method Generalized Probabilistic Canonical Correlation Analysis (GPCCA) for unsupervised multi-modal data integration and dimensionality reduction.
result GPCCA outperforms existing methods in capturing essential patterns across modalities and provides robust low-dimensional embeddings.
This paper proposes a model to learn multimodal representations robust to missing data.
problem Learning multimodal representations from heterogeneous sources of information.
method Optimizes a joint generative-discriminative objective across multimodal data and labels, factorizing representations into multimodal discriminative and modality-specific generative factors.
result The proposed model achieves state-of-the-art performance on six multimodal datasets and can reconstruct missing modalities without significant performance drop.
Paper proposes methods to handle missing values in clinical data.
problem Missing Not At Random (MNAR) data in clinical studies.
method Model-based and surrogate estimation strategies for low-rank matrix completion.
result Proposed methods can handle different missing value mechanisms robustly.
New method for fMRI missing value imputation improves robustness.
problem High frequency of missing values in fMRI data.
method Spatial and time-dependent regularization with a novel recurrent layer.
result Improved robustness against state-of-the-art alternatives.
Proposes methods to handle missing data in multi-task learning for healthcare.
problem Handling missing features in multi-task learning datasets.
method Plug-in covariance matrix estimators with LASSO and graph regularization.
result Effective for prediction and model estimation in datasets with missing data.