Fisher score is one of the most widely used supervised feature selection methods. However, it selects each feature independently according to their scores under the Fisher criterion, which leads to a suboptimal subset of features. In this paper, we present a generalized Fisher score to jointly select features. It aims …
A new method assigns anomaly scores to features for better interpretation.
problem Interpreting anomaly scores from feature attributions.
method Proposes a characteristic function to attribute anomaly scores using Shapley value.
result Demonstrates the potential utility of the proposed attribution methods.
New method disentangles feature importance scores in machine learning.
problem Misinterpretation of feature importance scores due to interactions and dependencies.
method Derive DIP (Disentangled Importance) decomposition of feature importance scores.
result DIP decomposition uniquely separates standalone contributions from interactions and dependencies.
This paper improves random feature sampling using empirical leverage scores.
problem Optimizing the number of features for kernel approximation and supervised learning.
method Uses empirical leverage scores to optimize feature sampling.
result Empirical sampling of random features using leverage scores outperforms vanilla Monte Carlo sampling.
MLS improves feature selection for imbalanced data.
problem Machine learning challenges with imbalanced high-dimensional data.
method Introduces Marginal Laplacian Score (MLS) for better feature selection.
result MLS improves performance on synthetic and public datasets.
Optimal scoring framework for kernel classification with feature selection.
problem Two-group classification problem.
method Optimal scoring framework, structured sparsity using weighted kernels, automated parameter selection.
result Superior classification performance compared to existing nonparametric classifiers.
We revisit the problem of feature selection in linear discriminant analysis (LDA), that is, when features are correlated. First, we introduce a pooled centroids formulation of the multiclass LDA predictor function, in which the relative weights of Mahalanobis-transformed predictors are given by correlation-adjusted t…
Two new algorithms improve feature importance scoring for graph-structured data.
problem Efficiently scoring feature importance for structured data.
method Developed two linear complexity algorithms for instancewise feature importance scoring.
result Our methods compare favorably with other feature importance scoring methods.
Three RFF-based methods for nonlinear causal discovery in mixed data.
problem Nonlinear causal discovery in mixed data with computational constraints.
method FFML, TRFF, and FFCI methods for score-based, constraint-based, and hybrid causal discovery.
result FFML and TRFF methods provide complementary performance in causal discovery.
Deconfounding scores improve causal effect estimation with weak overlap.
problem Challenges in causal treatment effect estimation due to weak overlap in high-dimensional data.
method Propose deconfounding scores to preserve identification and target estimation while improving overlap.
result Prognostic scores are overlap-optimal under a broad family of generalized linear models with Gaussian features.
The article warns against assuming normality in machine learning, especially when feature vectors are not normally distributed.
problem The assumption of normality in machine learning models can be misleading, especially when feature vectors are not normally distributed.
method The article provides a mathematical counterexample and experiments to illustrate the risks of assuming normality.
result Prudence is needed when assuming normality in machine learning models, particularly when feature vectors are not normally distributed.
Paper compares ML methods for credit scoring, highlighting feature selection and scaling impacts.
problem Determining default risk in credit scoring models.
method Eight ML methods (SVM, Naive Bayes, DT, RF, XGBoost, KNN, MLP, LR) with feature selection and scaling.
result Feature selection and scaling improve model performance in credit scoring.
Directly compute classification by learning features with class scores.
problem Classification efficiency and accuracy on various datasets.
method PCA for feature encoding, supervised learning model with encoder-decoder structure.
result Effective classification performance on multiple datasets.
Study shows how diffusion models learn on low-dimensional manifolds.
problem Learning efficiency of diffusion models on manifolds.
method Analyzes denoising score matching with random feature neural networks.
result Sample complexity scales linearly with intrinsic dimension, not ambient dimension.
DeepSleepNet uses CNN and LSTM to score sleep stages from raw EEG data.
problem Automatic sleep stage scoring using raw EEG data.
method Deep learning model using CNN for time-invariant features and LSTM for transition rules.
result DeepSleepNet achieves similar accuracy to state-of-the-art methods on different EEG datasets.
TRIP detects unreliable feature importance scores in random forests.
problem Unreliable feature importance scores in random forests due to model extrapolation.
method Develops TRIP (Test for Reliable Interpretation via Permutation) to detect unreliable permutation feature importance scores.
result TRIP reliably detects unreliable permutation feature importance scores in high-dimensional settings.
Proposes CDTD, a diffusion model for mixed-type tabular data.
problem Adapting diffusion models to mixed-type tabular data.
method Score matching and score interpolation for continuous features, adaptive noise schedules for categorical features.
result Consistently outperforms state-of-the-art models in mixed-type tabular data.
Deconfounding scores improve causal effect estimation with weak overlap.
problem Poor overlap in treatment and control groups makes causal effect estimators brittle.
method Introduces feature representations that improve overlap without introducing bias.
result Deconfounding scores satisfy a zero-covariance condition that is identifiable in observed data.
Develops fair feature importance scores for tree-based models to interpret fairness.
problem Ensuring fairness in machine learning models, especially tree-based ones.
method Inspired by decision trees, proposes a novel fair feature importance score based on mean decrease in group bias.
result Valid interpretations of fairness for tree-based ensembles and surrogates of other ML systems.
Approach for selecting features by discarding nuisance and correlated ones.
problem Large datasets with correlated and nuisance features.
method Laplacian score criterion, autoencoder architecture, concrete layer.
result Outperforms similar approaches in clustering performance.
Regularizes attention scores in vision transformers using bootstrapping.
problem Noisy and diffused attention maps in ViT limit interpretability.
method Statistical learning techniques, bootstrapping of attention scores.
result Improves shrinkage and sparsity of attention scores.
A scoring method for driving safety using trajectory data.
problem Managing traffic safety through driver behaviors and violations.
method Extract driving habits and violations from trajectories, train a model, score drivers.
result Proves the effectiveness of the scoring method using traffic simulation.
Feature learning forms the cornerstone for tackling challenging learning problems in domains such as speech, computer vision and natural language processing. In this paper, we consider a novel class of matrix and tensor-valued features, which can be pre-trained using unlabeled samples. We present efficient algorithms f…
A new method selects features for clustering without labels.
problem Identifying meaningful features in large datasets.
method Differentiable unsupervised feature selection using a gated Laplacian.
result The method improves clustering performance in noisy data.
Feature learning forms the cornerstone for tackling challenging learning problems in domains such as speech, computer vision and natural language processing. In this paper, we consider a novel class of matrix and tensor-valued features, which can be pre-trained using unlabeled samples. We present efficient algorithms f…
FUJI scores similarity of ranked lists more robustly.
problem Improving similarity assessment of ranked lists.
method Integrates a membership function into Jaccard index for better rank consideration.
result More stable and accurate similarity estimates.
Research creates a machine learning model for predicting TAVI patient mortality.
problem Lack of robust risk scores for TAVI patients.
method Gradient boosting on decision trees, feature analysis and selection, model validation.
result Model outperforms existing risk scores with AUC of 0.83.
Bayesian encoding improves lead scoring for WeWork using conjugate models.
problem Encoding high-cardinality categorical features for machine learning.
method Conjugate Bayesian models for categorical features, ensemble learning.
result AUC improved from 0.87 to 0.97 for WeWork's lead scoring engine.
Paper uses PSM to improve fake news detection generalizability.
problem Confounding variables in fake news features.
method Propensity Score Matching (PSM) to select features.
result Generalizability of fake news detection methods improved significantly.
The paper investigates how irrelevant features affect clustering performance.
problem The challenge of identifying relevant features in unsupervised clustering tasks.
method Investigation of clustering performance with added irrelevant features.
result Different types of irrelevant features impact clustering outcomes differently.
Proposes a new scoring function for linear classifiers to improve object positioning in feature space.
problem Lack of information about relative positions of recognized objects in feature space.
method Calculates a scoring function based on object distance from decision boundary and class centroid.
result Demonstrates effectiveness of the proposed method compared to other ensemble algorithms on multiple datasets.
Evaluates explanations of LTR models using decision paths and compares their accuracy.
problem Challenges in evaluating local explanations of LTR models due to lack of ground truth feature importance scores.
method Focuses on tree-based LTR models, extracts ground truth feature importance scores using decision paths, and compares them with explanation techniques.
result Explanation accuracy varies depending on the model and data point.
Score function estimators improve k-subset sampling efficiency.
problem Efficiently sampling k-subsets in machine learning tasks. method Revisit score function estimators, using discrete Fourier transform and control variates.
result Efficient and unbiased gradient estimates for k-subset sampling. Nonparametric IPSS selects features with false discovery control.
problem Feature selection in high-dimensional data with theoretical false discovery control.
method Integrated Path Stability Selection (IPSS) applied to nonparametric feature importance scores.
result IPSS accurately controls false discovery rate and detects more true positives than existing methods.
EHBOS enhances HBOS by capturing feature interactions, improving anomaly detection.
problem Limited ability of HBOS to detect anomalies in datasets with feature interactions.
method Incorporates two-dimensional histograms to capture feature pair dependencies.
result EHBOS outperforms HBOS on datasets with critical feature interactions, achieving notable improvements in ROC AUC.
This paper investigates two feature-scoring criteria that make use of estimated class probabilities: one method proposed by \citet{shen} and a complementary approach proposed below. We develop a theoretical framework to analyze each criterion and show that both estimate the spread (across all values of a given feature)…
TDA improves FX clustering quality over traditional methods.
problem Capturing complex currency co-movements in FX markets.
method Topological Data Analysis (TDA) compared to traditional statistical methods on monthly FX returns.
result TDA-based clustering yields more compact and well-separated clusters.
A faster method for density estimation using denoising score matching with random Fourier features.
problem Intractability of normalizing constant in density estimation.
method Denoising Score Matching with Random Fourier Features.
result The method is computationally faster and scalable to complex high-dimensional data.
Cluster LOCO: A model-agnostic feature importance score for interpreting cluster outputs
problem Interpreting and auditing cluster outputs
method Cluster LOCO (Leave-One-Covariate-Out)
result More reliably recovers informative features than existing methods
XGBoost fails to accurately identify relevant features, while interpretable methods do.
problem Accurately identifying relevant features in black-box models like XGBoost.
method Comparison of variable importance methods (CART, Optimal Trees, XGBoost, SHAP) across various experiments.
result Interpretable methods outperform black-box models in feature selection accuracy.
Proposes a new method to identify important input features using maximally invariant data perturbation.
problem Lack of formal mathematical definitions for feature scoring in complex machine learning models.
method Formulates the problem as linear programming to find the maximally invariant data perturbation.
result Identifies relevant parts of images effectively, distinguishing important input features.
A novel approach combines feature importance scores with deep learning for forex price prediction.
problem Improving forex price prediction using deep learning models.
method Feature importance recap combined with stacking models.
result Proper feature selection significantly improves model performance.
Unified analysis improves random Fourier features for kernel methods.
problem Pessimistic theoretical bounds on random Fourier features.
method Unified risk analysis for squared error and Lipschitz loss.
result Improved bounds on number of features for convergence.
A method for ranking features in multi-label classification using Markov Networks.
problem Feature ranking in multi-label classification.
method Markov Networks, Ising model, score statistic.
result The proposed method outperforms conventional approaches on artificial and real datasets.
MIAEAD detects anomalies in mixed data types.
problem Challenges of heterogeneity in feature subsets for anomaly detection.
method Multiple-Input Variational Auto-Encoder (MIVAE) for simultaneous feature subset anomaly scoring.
result MIVAE outperforms conventional methods and state-of-the-art unsupervised models in AUC score.
Paper proposes hybrid approach for transparent credit scoring models.
problem Lack of transparency in machine learning models limits their use in regulated environments.
method Post-hoc interpretation of black-box models guides feature selection, followed by training glass-box models.
result Reduces feature usage from 106 to 10 while maintaining comparable performance.
Interpretable machine learning uncovers ESG's explanatory power on equity returns across sectors and capitalizations.
problem Explaining equity returns beyond market factors using ESG data.
method Interpretable machine learning models, cross-validation scheme, random company-wise validation.
result Gradient boosting models explain unaccounted price returns, with ESG data outperforming basic fundamental features.
The study examines how the number of noise samples affects diffusion models' performance.
problem Understanding the balance between generalization and memorization in diffusion models.
method Theoretical analysis and empirical experiments with Denoising Score Matching (DSM) using random features.
result Precise expressions for test and train errors under specific conditions reveal the mechanisms of generalization and memorization.