The paper introduces the Banzhaf value for robust data valuation in machine learning, addressing stochastic model performance.
problem Inconsistent data value rankings due to model performance noise.
method Introduces the Banzhaf value and Maximum Sample Reuse (MSR) principle for efficient estimation.
result The Banzhaf value outperforms other semivalues in robust data valuation.
The Shapley value method calculates data contributions efficiently.
problem Valuing data contributions fairly among multiple contributors.
method Utilizing the Shapley value, a game-theoretic approach, with efficient algorithms.
result Efficient algorithms approximate the Shapley value for data valuation.
The paper introduces Absolute Shapley Value to handle negative contributions in machine learning model training.
problem Negative marginal contributions in machine learning model training.
method Investigates three philosophies: Original Shapley Value, Zero Shapley Value, and Absolute Shapley Value.
result Absolute Shapley Value significantly outperforms other definitions in evaluating data importance.
A new STAR framework models integer-valued data with flexible distributions.
problem Modeling integer-valued data with flexibility and accuracy.
method Simultaneously Transforming and Rounding (STAR) a continuous-valued process.
result STAR framework designs a new BART model for integer-valued data with impressive predictive accuracy.
Enhances data valuation by integrating global and local statistical properties.
problem Insufficient consideration of global and local statistical properties in data valuation methods.
method Proposes a method that fuses global and local statistical properties into regularization terms for Shapley value estimation and dynamic data valuation.
result Demonstrates improved performance and efficiency of data valuation methods through integration of global and local statistical properties.
RDIS fills missing values in time series data explicitly.
problem Missing values in time series data.
method Random Drop Imputation with Self-training.
result RDIS achieves competitive results on real-world datasets.
New methods for ordinal classification of interval-valued data and functional data.
problem Ordinal classification of interval-valued data and functional data.
method Six ordinal classifiers are proposed, including parametric, binary decomposition, logistic regression, distance-based, k-nearest-neighbor, kernel PCA, and random forest methods.
result Considering ordering and interval-valued information improves the accuracy of ordinal classification.
A complex-valued convolutional network (convnet) implements the repeated application of the following composition of three operations, recursively applying the composition to an input vector of nonnegative real numbers: (1) convolution with complex-valued vectors followed by (2) taking the absolute value of every entry…
DeepIFSAC uses attention mechanisms and contrastive learning to impute missing values in tabular data.
problem Missing values in tabular data, especially when high and not random.
method Row and column attention in a contrastive learning framework with CutMix data augmentation.
result Proposed method outperforms state-of-the-art methods for missing rates between 10% and 90% and various missing value types.
A new framework assigns values to data points considering their distribution.
problem Limited applicability of data Shapley to points outside the fixed data set.
method Proposes distributional Shapley, defining point value in context of data distribution.
result Distributional Shapley values are stable under data point and distribution perturbations.
Imputation-free method learns tabular data with missing values using transformer.
problem Machine learning on tabular data with missing values often leads to unreliable outcomes due to synthetic imputation.
method Incremental attention learning (IFIAL) using transformer with attention masks.
result IFIAL outperforms state-of-the-art methods in 17 diverse tabular data sets.
Developing an explainable outlier detection method for interval-valued data using Shapley value-based approach.
problem Outlier detection in interval-valued data.
method Proposed a novel approach based on Shapley value for interval-valued data.
result Fine-grained interpretation of outliers with variable contributions.
IFGAN uses feature-specific GANs for missing value imputation.
problem Missing value imputation in data mining.
method Feature-specific Generative Adversarial Networks (GAN).
result IFGAN outperforms state-of-the-art algorithms in various missing conditions.
CVNN outperforms RVNN on non-circular data.
problem Classifying complex-valued data with statistical dependence.
method Comparison of CVNN and RVNN on non-circular data.
result CVNN outperforms RVNN in accuracy and generalization.
Many data mining and data analysis techniques operate on dense matrices or complete tables of data. Real-world data sets, however, often contain unknown values. Even many classification algorithms that are designed to operate with missing values still exhibit deteriorated accuracy. One approach to handling missing valu…
Paper assesses the market value of sharing privacy-protected smart meter data.
problem Value of sharing privacy-protected smart meter data between consumers and load serving entities.
method Discounted differential privacy model, ANN-based load forecasting, optimal procurement problem.
result Significant value in sharing smart meter data while retaining individual consumer privacy.
ELMV uses ensemble learning to handle missing values in EHR data.
problem Significant missing values in EHR data cause bias and unreliable conclusions.
method ELMV constructs multiple subsets with lower missing rates and uses a support set for ensemble learning.
result ELMV outperforms conventional methods in critical feature identification and outcome prediction.
New method forecasts values and timing in irregular time series.
problem Forecasting values and timing in sparse, irregularly sampled multivariate time series.
method Proposes a novel approach for forecasting values and timing in irregular time series.
result Successfully forecasts values and timing in irregular time series.
Proposes a normalization technique for manifold valued data.
problem Instability in optimization for manifold valued data.
method Develops a general normalization technique for manifold valued data.
result Demonstrates performance gain in synthetic and real datasets.
Paper improves conformal prediction for imprecise training data.
problem Applying conformal prediction to partially labeled data.
method Generalizes conformal prediction for set-valued training and calibration data.
result Validates the proposed method and shows it outperforms baselines.
Develops a new metric to equitably value data for machine learning models.
problem Equitable valuation of individual data in machine learning predictions.
method Data Shapley framework, Monte Carlo and gradient-based methods.
result Data Shapley uniquely satisfies properties of equitable data valuation.
New method predicts y distributions from imperfect data.
problem Predicting y from imperfect data (discrete, truncated, censored).
method Optimal transformations to estimate p(y|x).
result Estimates location, scale, and shape of y distribution.
Gaussian Processes improve missing value imputation in datasets.
problem Handling missing values in large datasets.
method Sparse Gaussian Processes combined with stochastic variational inference.
result MGP significantly outperforms other imputation methods.
This paper proposes a new method to fairly value data in federated learning.
problem Fairly valuing decentralized data contributions in federated learning.
method Variant of Shapley value (federated Shapley value) that is efficient and respects the order of data contributions.
result The federated Shapley value can reflect the real utility of data sources and enhance system robustness, security, and efficiency.
New insights on Shapley value precision for tabular data predictions.
problem Precision of Shapley value explanations for individual observations.
method Conditional Shapley value estimation methods for tabular data.
result Shapley value explanations are less precise for outer observations.
Study uses Open Banking data to estimate customer value, showing potential 21% increase.
problem Limited CLV estimation using single-entity data.
method Introduces PCLV framework using Open Banking data for comprehensive customer value estimation.
result Open Banking data can estimate PCLV per competitor, showing a 21.06% increase over Actual CLV.
Framework for joint learning of tasks on dementia data with missing values.
problem Lack of multi-task learning, handling time-dependent data, and missing values in dementia forecasting.
method Proposes SSHIBA model using Bayesian variational inference for imputation and combined information from different views.
result SSHIBA model outperforms baselines in predicting diagnosis, ventricle volume, and clinical scores in dementia.
Analyzes premium data of Indian non-life insurers, finding GEV distribution best fits Lognormal and GEV extremes.
problem Modeling premiums of non-life insurance companies in India.
method Empirical analysis using Lognormal, GEV, and GPD distributions.
result Generalized Extreme Value distribution best fits premium data for ten Indian non-life insurers.
Paper improves matrix-valued data classification using nonparametric LDA.
problem Classification of matrix-valued data in neuroimaging and signal processing.
method Nonparametric LDA based on NPMLE for vectorized and scaled matrices.
result Improves classification performance across various data structures.
DVRL uses RL to estimate data value for machine learning tasks.
problem Adaptive learning of data value for machine learning tasks.
method Meta learning framework with reinforcement learning for data value estimation.
result DVRL yields superior data value estimates compared to alternative methods.
Framework reconstructs missing spatio-temporal data for extreme value prediction.
problem Predicting extreme values from incomplete spatio-temporal data.
method Convolutional deep neural networks and autoencoder-like models for conditional sampling.
result Framework produces accurate reconstructions of missing data for extremal values.
New algorithms handle missing data to improve fairness in machine learning.
problem Missing values in data can lead to unfair outcomes in machine learning models.
method Developed scalable and adaptive algorithms to handle missing values while preserving predictive information.
result Our adaptive algorithms consistently achieve higher fairness and accuracy than standard impute-then-classify methods.
Techniques such as clusterization, neural networks and decision making usually rely on algorithms that are not well suited to deal with missing values. However, real world data frequently contains such cases. The simplest solution is to either substitute them by a best guess value or completely disregard the missing va…
Estimates peeking effects in p-values to correct bias.
problem Data peeking biases reported p-values downward.
method Develops mechanisms to estimate running extrema of test statistics.
result Corrects bias in p-values due to peeking.
Optimal clustering handles missing values without imputation.
problem Missing values complicate clustering algorithms in biomedical studies.
method Integrates missing value mechanism into optimal clustering framework.
result Superior performance compared to other clustering approaches.
DVWU framework improves model performance by considering data value heterogeneity.
problem Existing machine unlearning algorithms ignore data value heterogeneity, potentially degrading model performance.
method Data Value-Weighted Unlearning (DVWU) framework that integrates data values into the unlearning process.
result DVWU achieves superior predictive performance and robustness compared to conventional unlearning approaches.
New Shapley values reveal non-linear feature dependencies.
problem Understanding non-linear dependencies in machine learning models.
method Model-independent Shapley values using non-parametric measures of dependence.
result Model-independent Shapley values can uncover non-linear dependencies.
Extends Fisher's Discriminant Analysis for interval-valued data.
problem Classifying entities represented by intervals and histograms.
method Adapts Fisher's Discriminant Analysis using Moore's interval arithmetic and Mallows' distance.
result Discriminant directions for interval-valued data are numerically maximized.
In medical domain, data features often contain missing values. This can create serious bias in the predictive modeling. Typical standard data mining methods often produce poor performance measures. In this paper, we propose a new method to simultaneously classify large datasets and reduce the effects of missing values.…
Complex-valued neural networks improve seismic data analysis by preserving phase information.
problem Low-frequency aliasing in seismic data due to discarded phase information.
method Developed complex-valued deep convolutional networks to leverage phase information in deterministic physical data.
result Complex-valued networks outperform real-valued networks in training and inference from deterministic physical data.
This paper develops a pricing model for data assets from the buyer's perspective.
problem Insufficient research on pricing data assets from the buyer's perspective.
method Develops a pricing model based on the informational value of data assets from the buyer's perspective, using an implicit function derived from value functions in investment-consumption problems under ambiguity markets.
result Derives general expressions and explicit pricing formulas for data assets under various conditions.
New method maps global value chains at product level from trade data.
problem Lack of detailed product-level value chain information in existing datasets.
method Machine learning and trade theory applied to international trade data.
result Approximate product-level value chain information inferred from trade patterns.
The paper studies OPE with missing data, showing bias under nonignorable missingness and proposing a solution.
problem Estimating value of a target policy from logged data with missingness.
method Investigates OPE with monotone missingness, proposes an IPW value estimator, and conducts statistical inference.
result Value estimates remain unbiased under ignorable missingness but can be biased under nonignorable missingness.
The uncertainty or the variability of the data may be treated by considering, rather than a single value for each data, the interval of values in which it may fall. This paper studies the derivation of basic description statistics for interval-valued datasets. We propose a geometrical approach in the determination of s…
Consistent supervised learning with missing values is possible using imputation or specialized models.
problem Predicting with missing values in both training and testing data.
method Two approaches: imputing with a constant and using a predictor for complete observations through multiple imputation. Decision trees can handle missing values naturally.
result Imputing with a constant can be consistent when missing values are not informative.
This paper calculates data valuation for nearest neighbor models efficiently.
problem Distributing payment for training ML models based on data contributions.
method Defined 'relative value of data' via Shapley value for fairness and decentralizability. Developed algorithms for exact and approximate computation of Shapley values for nearest neighbor models.
result Exact computation of Shapley values for nearest neighbor models in O(N log N) time, significantly faster than previous methods.
New asymptotic e-values improve inference by eliminating data-dependent scaling inefficiency.
problem Data-dependent scaling inefficiency in existing asymptotic e-values.
method Drawing on Bentkus's near-optimal concentration inequalities, introduce Bentkus-type asymptotic e-values.
result Bentkus-type asymptotic e-values consistently deliver sharper inference than existing alternatives.
Paper derives fast algorithms for interpreting machine learning data contributions.
problem Efficiently quantify individual data contributions in machine learning.
method Developed analytic expressions for Distributional Shapley values for common ML tasks.
result New algorithms estimate DShapley values up to several orders of magnitude faster.