Paper tackles stutter detection using deep learning.
problem Identification and classification of stuttered speech.
method Uses a deep residual network with bidirectional LSTM layers.
result Achieves an average miss rate of 10.03%, outperforming state-of-the-art.
NCA improves fault detection in nonlinear processes.
problem Fault detection in nonlinear chemical processes.
method Neural Component Analysis (NCA) using feedforward neural networks with orthogonal constraints.
result NCA outperforms traditional PCA and autoencoder methods in fault detection.
Project accelerates pedestrian detection in self-driving cars using FPGA.
problem Improve pedestrian detection accuracy in self-driving vehicles.
method Training and tuning SSD, FPGA design for acceleration.
result 4x performance improvement with minimal accuracy loss.
Novel CGTF model for recommender systems and community detection from coupled graphs and tensors.
problem Lack of effective methods for analyzing multiple information repositories with graph side information.
method Coupled Graph-Tensor Factorization (CGTF) with ADMM for nonnegative factor recovery.
result CGTF model successfully detects communities even with missing graph links.
Detects missing tensor signals in a KS subspace with high probability.
problem Detecting tensor signals with many missing entities in a KS subspace.
method Projecting the signal onto the KS subspace and bounding residual energy.
result Reliable detection is possible if the missing signal cardinality exceeds KS subspace dimensions.
New method predicts positive samples with missing labels.
problem Missing labels due to response-dependent factors.
method P(U)U-O-Mixture algorithm for joint estimation.
result Non-convex algorithm leads to optimal statistical error.
New algorithm detects network outliers with missing links.
problem Identifying outliers in networks with missing links.
method Proposes a new algorithm that detects outliers and predicts missing links.
result Proves the algorithm exactly detects outliers and achieves best error for missing link prediction.
This paper tackles anomaly detection with missing values.
problem Anomaly detection methods struggle with missing data.
method Five strategies for handling missing values in test queries: mean imputation, MAP imputation, reduction, marginalization, and proportional distribution.
result MAP imputation and proportional distribution outperform other methods.
New method detects concept drift in data streams with missing values.
problem Uncertainty introduced by missing values in concept drift detection.
method Fuzzy distance estimation and histogram bin allocation.
result Fuzzy set theory improves drift detection in data with missing values.
New methods detect changes in data with missing values.
problem Detecting changes in high-dimensional data with missing values.
method Proposes three imputation methods and adapts model selection for incomplete data.
result Methods improve change point detection in scenarios with missing values.
Proposes a new Lasso method for high missing rate data.
problem Handling high-dimensional data with many missing values.
method Integrates mean imputed covariance to overcome estimation bias.
result Effective even with high missing rates, improving upon CoCoLasso.
The study proposes algorithms to minimize rating discordance in missing data.
problem Missing ratings in combined rating lists.
method Optimization models and algorithms that minimize total rating discordance.
result The proposed methods outperform state-of-the-art imputation methods in accuracy.
ECAD detects anomalies without data exchangeability, improving traffic flow detection.
problem Detecting anomalies in spatio-temporal data with missing values.
method ECAD uses conformal prediction to wrap around any regression algorithm, controlling Type-I error without data exchangeability.
result ECAD outperforms other methods in detecting anomalous traffic flow.
Hierarchical GANs reduce anomaly detection costs.
problem Balancing anomaly detection accuracy and sampling costs.
method Hierarchical GANs for nonuniform sampling and buffer zones.
result Proposed GAN-based detector outperforms baseline in detection delay and average cost of error.
This paper describes a novel approach to change-point detection when the observed high-dimensional data may have missing elements. The performance of classical methods for change-point detection typically scales poorly with the dimensionality of the data, so that a large number of observations are collected after the t…
The paper addresses bias in fraud detection models by improving label recovery in payment networks.
problem Systematic bias in chargeback labels in payment networks.
method Formalizes the observation pipeline as a sequential missing-data problem with three stages and a corruption layer. Constructs the Sequential Triply Robust (STR) estimator to correct for all four impairments simultaneously.
result Achieves the semiparametric efficiency bound and provably dominates naive chargeback-based training in mean squared error.
New algorithm improves fraud detection by analyzing financial account relationships.
problem High false positive rates and missed detections in conventional fraud detection systems.
method Personalized PageRank (PPR) algorithm to capture social dynamics of fraud.
result Integrating PPR enhances fraud detection model's predictive power.
Merlin improves robustness of MTSF models to missing data.
problem Suboptimal forecasting performance due to unfixed missing rates in MTSF models.
method Offline knowledge distillation and multi-view contrastive learning.
result Merlin enhances robustness of MTSF models while preserving accuracy.
Paper proposes anomaly detection using Eigentraces and one-class classification.
problem Detect anomalies in system call trace data for Linux OS.
method One-class classification with Eigentraces feature extraction, Radial Basis Function neural network, and Random Forest.
result High performance in detecting anomalies and normal activities.
Rating prediction is an important application, and a popular research topic in collaborative filtering. However, both the validity of learning algorithms, and the validity of standard testing procedures rest on the assumption that missing ratings are missing at random (MAR). In this paper we present the results of a us…
Optimizes variational autoencoder for detecting missing data in Mars rover transmissions.
problem Detecting missing data in Mars rover transmissions to prevent volume loss and corruption.
method Applies derivative-free optimization to tune variational autoencoder.
result Improves variational autoencoder's ability to detect missing data, aiding GDSA team.
Random Forests are reinterpreted as generative models to handle missing data and detect outliers.
problem Handling missing features and detecting outliers in Random Forests.
method Interpreting Random Forests as Probabilistic Circuits and applying marginalisation for missing data.
result GeDTs and GeFs can handle missing data and detect outliers under certain assumptions.
The study reveals fundamental limits of fraud detection in card payment networks.
problem Fraud detection in card payment networks is challenging due to structural information impairments.
method Formalized card authorization as a sequential decision problem with delayed feedback, derived minimax regret lower bound.
result Improving issuer reporting quality or reducing censorship can yield larger reductions in the regret floor than increasing model complexity.
Paper uses GNNs to efficiently detect profitable triangular arbitrage opportunities.
problem Detecting profitable triangular arbitrage opportunities in dynamic markets.
method Formulate the problem as a graph-based optimization task and use a GNN architecture to capture complex relationships.
result GNN-based method achieves higher average yield with reduced computational time compared to traditional methods.
Study improves accuracy of weather data for real-time building simulations.
problem Anomalous and missing weather data affect real-time building energy simulations.
method Introduces a framework for quality control of measured weather data using anomaly detection and neural network infilling.
result Neural Networks enhance the accuracy of data imputation compared to traditional methods.
DeepIFSAC uses attention mechanisms and contrastive learning to impute missing values in tabular data.
problem Missing values in tabular data, especially when high and not random.
method Row and column attention in a contrastive learning framework with CutMix data augmentation.
result Proposed method outperforms state-of-the-art methods for missing rates between 10% and 90% and various missing value types.
Develops panoramic gastroscopy for automatic polyp detection.
problem Missed diagnosis of gastric polyps during endoscopy.
method Panoramic reconstruction method and end-to-end multi-object detection.
result Average error of panorama less than 2 mm, polyp detection accuracy 95%, recall rate 99%.
New deep probabilistic model handles missing data in time series forecasting.
problem Handling missing data in time series forecasting.
method Combination of deep learning and probabilistic methods.
result Advantage in forecasting and novelty detection with missing data.
Estimates the probability of discovering a new type in samples from a population.
problem Estimating the missing mass of unknown type proportions in samples.
method Bayesian nonparametric tools and Good-Turing estimator for regularly varying type proportions.
result The Good-Turing estimator is rate optimal under regularly varying type proportions.
New online imputation method for mixed data improves accuracy and speed.
problem Missing value imputation in online settings for mixed data types.
method Online Gaussian copula model for imputation and change point detection.
result The model improves accuracy and speed, especially on large datasets.
Paper builds ML classifier to detect crypto-ransomware.
problem Detecting crypto-ransomware with high accuracy and low false positives.
method Behavior-based detection using input/output activities and file-content entropy. Deep-learning classifier with adversarial research and Integrated Gradient method for explanation.
result Deep-learning classifier achieves high accuracy and low false positive rate in detecting crypto-ransomware.
ELMV uses ensemble learning to handle missing values in EHR data.
problem Significant missing values in EHR data cause bias and unreliable conclusions.
method ELMV constructs multiple subsets with lower missing rates and uses a support set for ensemble learning.
result ELMV outperforms conventional methods in critical feature identification and outcome prediction.
Recommending items to users is a challenging task due to the large amount of missing information. In many cases, the data solely consist of ratings or tags voluntarily contributed by each user on a very limited subset of the available items, so that most of the data of potential interest is actually missing. Current ap…
We present an automatic classification method for astronomical catalogs with missing data. We use Bayesian networks, a probabilistic graphical model, that allows us to perform inference to pre- dict missing values given observed data and dependency relationships between variables. To learn a Bayesian network from incom…
New MMD estimators detect differences in missing paired data.
problem Handling missing data in matched pairs with complex distributions.
method Maximum mean discrepancy (MMD) estimators for complex data with missing values.
result Valid and consistent estimators detect differences in data distributions.
BlockEcho method improves imputation of block-wise missing data.
problem Block-wise missing data reduces interpolation capability and predictive power.
method Integrates Matrix Factorization (MF) within Generative Adversarial Networks (GAN) to retain long-distance inter-element relationships.
result Superior performance on public datasets across three domains, especially at higher missing rates.
Study tests uniformity of categorical data against missing-ball alternatives, finding chi-squared test outperforms.
problem Testing uniformity of categorical data against missing-ball alternatives.
method Characterizes minimax risk, uses collisions and chi-squared test, reduces to structured subset of alternatives.
result Minimax test outperforms chi-squared test under least favorable alternative.
IFGAN uses feature-specific GANs for missing value imputation.
problem Missing value imputation in data mining.
method Feature-specific Generative Adversarial Networks (GAN).
result IFGAN outperforms state-of-the-art algorithms in various missing conditions.
The paper analyzes the benefit-cost ratio for feature selection in machine learning.
problem Tackling the challenge of distinguishing relevant features from noise in feature selection.
method Simulation study with different cost and data settings to analyze the benefit-cost ratio.
result The benefit-cost ratio can overemphasize cheap noise features in scenarios with large cost differences and small effect sizes.
Random Forest outperforms other IDS algorithms in smart grids.
problem Security vulnerabilities in smart grids.
method Comparison of four data mining algorithms (Random Forest, SVM, Neural Network, and KNN) in detecting attacks.
result Random Forest outperforms other algorithms in terms of detection accuracy and efficiency.
The paper analyzes k k k -means clustering for missing data, proving statistical guarantees under MCAR.
problem Statistical guarantees for k k k -means clustering with missing data, especially under Missing Completely at Random (MCAR). method Established n \sqrt{n} n -excess risk bound and consistency of cluster centers under general missing mechanisms; derived n \sqrt{n} n -convergence rate and asymptotic normality for MCAR. result Achieving n \sqrt{n} n -rate and converging to true cluster centers requires distinct true cluster centers in every dimension under MCAR. GRU-D detects age-specific missing patterns in vital signs.
problem Temporal missingness in clinical time series data.
method Gated recurrent unit with decay mechanisms (GRU-D) trained on MIMIC-IV vital signs.
result GRU-D achieves AUROC 0.780 and AUPRC 0.810 on bootstrapped data.
New method for tensor classification with missing data.
problem Handling incomplete tensor data in high-dimensional classification.
method High-dimensional tensor linear discriminant analysis with TGMM and Tensor LDA-MD.
result Established convergence rates and minimax optimal bounds for misclassification rate.
PyPOTS simplifies machine learning on time series with missing data.
problem Handling missing data in time series analysis.
method Unified interface for imputation, forecasting, anomaly detection, classification, and clustering.
result Robust and scalable Python toolkit for multivariate partially-observed time series.
Study compares traditional and machine learning methods for handling missing data in longitudinal studies.
problem Handling missing Not at Random (MNAR) and nonnormal data in longitudinal research.
method Monte Carlo simulations to assess six missing data techniques.
result FIML is most effective for MNAR data, while TSRE excels for MAR data.
Robots learn new skills from demonstrations, using active learning to detect missing information.
problem Detecting missing information during skill generalization and transitioning to new tasks.
method Novel active learning algorithm based on deep generative models and metric learning in latent spaces.
result Smooth trajectories generated by asking for additional demonstrations when non-smooth transitions are detected.
Unsupervised method detects surgical site infections from blood samples.
problem Detecting surgical site infections from blood samples without labeled data.
method Powerful kernels for multivariate time series with missing data handling.
result Framework shows superior performance compared to baselines.
New method for matrix completion under complex missing data patterns.
problem Matrix completion with complex missing data patterns.
method Estimate the probability matrix of observation via low-rank matrix estimation and use inverse probabilities weighting to complete the target matrix.
result Optimal asymptotic convergence rates for observation probabilities and target matrix estimation.