New algorithm mitigates bias in subset selection with noisy protected attributes.
problem Mitigating bias in subset selection when protected attributes are noisy.
method Formulated a denoised selection problem and developed a linear-programming based approximation algorithm.
result The approach can produce fairer subsets despite noisy protected attributes.
Greedy PIG adapts integrated gradients for better feature attribution.
problem Interpreting deep learning model predictions.
method Unified discrete optimization framework for feature attribution and selection.
result Greedy PIG improves feature attribution on various tasks.
The paper proposes a method to improve data analysis by considering multiple subsets of attributes (views) to enhance geometric information.
problem Distortion of distance metrics in high-dimensional data analysis.
method Partitioning attributes into multiple subsets (views) and using consensus between views to extract geometric information.
result Enhanced geometric information from multiple views improves data analysis.
Algorithm uncovers latent attribute graph from molecular data.
problem Learning latent representations and interpreting them for limited data.
method Perturbation experiments on latent codes of a generative autoencoder.
result Effective graphical model of latent codes and attributes.
Bayesian approach scores influential training examples for model predictions.
problem Enhance interpretability and safety of machine learning models.
method Formulate TDA as a Bayesian information-theoretic problem, scoring subsets by information loss.
result Method aligns with classical influence scores while promoting diversity for subsets.
New algorithm tackles subgroup fairness in AI with multiple sensitive attributes.
problem Heavy computational burdens and data sparsity in subgroup fairness for multiple sensitive attributes.
method Doubly Regressing Adversarial learning (DRAF) for subgroup fairness, focusing on subgroups with sufficient sample sizes and marginal fairness.
result DRAF algorithm reduces a surrogate fairness gap for supIPM with less computation than directly reducing supIPM.
PSI models and infers feature attributions efficiently and accurately.
problem Modeling and inferring feature attributions in flexible predictive models.
method Probabilistic Shapley inference (PSI) framework using latent random variables and a masking-based neural network architecture.
result PSI learns feature attribution distributions centered at Shapley values, revealing meaningful uncertainty.
Tab-Shapley identifies top-k anomalies in tabular data quality insights.
problem Challenges in identifying anomalies in unlabeled tabular datasets.
method Cooperative game theory using Shapley values to quantify attribute contributions.
result Efficiently identifies top-k tabular data quality insights using closed-form Shapley values.
Private release of sensitive data enables fair learning.
problem Learning fair predictors with restricted sensitive data.
method Private release of sensitive demographic data, adapting non-discriminatory learners.
result The approach provides theoretical guarantees on performance for fair predictors.
Estimates KL divergence with fairness considerations for sub-populations.
problem Fairly estimate KL divergence between distributions considering sub-populations.
method Proposes multi-group attribution for KL divergence estimation, derived from multi-calibration.
result Shows multi-group attribution provides better KL divergence estimates conditioned on sub-populations.
The paper formalizes feature attribution to address inconsistent definitions and evaluate methods.
problem Inconsistent definitions of feature relevance in feature attribution.
method Formalization based on relaxed functional dependence, extended to instance-wise setting.
result State-of-the-art methods often fail to verify necessary properties for candidate selection.
In-Run Data Shapley offers efficient data attribution for large-scale models.
problem Existing data attribution methods are computationally intensive and cannot target specific models.
method In-Run Data Shapley, which efficiently attributes data contributions to a specific model without re-training.
result In-Run Data Shapley achieves significant efficiency, enabling data attribution for pretraining models.
New method explains anomalies in multivariate time series data.
problem Understanding and explaining anomalies in multivariate time series data.
method Counterfactual reasoning applied to MDI-detected anomalous intervals.
result Our method accurately identifies and explains anomalies in various extreme events.
The outlying property detection problem is the problem of discovering the properties distinguishing a given object, known in advance to be an outlier in a database, from the other database objects. In this paper, we analyze the problem within a context where numerical attributes are taken into account, which represents…
Enhances apparel attribute recognition with a two-layer ensemble method.
problem Improving accuracy in apparel attributes classification using deep neural networks.
method Proposes a two-layer mixture framework combining bagging and boosting for ensemble learning.
result The proposed method outperforms individual models and ensemble methods.
This research simplifies computation of feature attribution methods under certain conditions.
problem Computational complexity of feature attribution methods, especially power indices.
method Identifying conditions for polynomial computation and introducing new indices.
result Conditions for efficient computation of feature attribution methods are identified.
Scalable learning of nodal attributes in large graphs with privacy.
problem Scalability and privacy in learning nodal attributes of large graphs.
method Multikernel-based approach for real-time evaluation of nodal attributes.
result Real-time evaluation of nodal attributes without re-solving the problem over all nodes.
The paper explains neural network predictions of psychological attributes using heatmaps.
problem Explaining predictions of psychological attributes from face photographs using deep neural networks.
method Transfer learning with age and gender prediction models, novel explanation method using heatmaps.
result The explanation method provides insights into the base model's features and aptitude for transfer learning.
Develops methods to measure and reduce fairness in datasets with limited protected attribute labels.
problem Measuring and reducing fairness in datasets with limited protected attribute labels.
method Proposes methods to estimate fairness metrics and train models to limit fairness violations using probabilistic protected attribute labels.
result Our methods provide tighter bounds on true disparity and effectively reduce fairness violations with lesser fairness-accuracy trade-offs.
Study shows data attribution methods are sensitive to hyperparameters, making tuning costly.
problem Hyperparameter sensitivity in data attribution methods makes tuning impractical.
method Theoretical analysis and lightweight procedure for selecting regularization value without retraining.
result Proposes a lightweight procedure for selecting regularization value without model retraining.
Detects domain shifts in datasets using interpretable feature subspaces.
problem Detecting subtle differences in dataset probability distributions.
method Localised density anomaly detection in high-dimensional feature spaces.
result Extracts interpretable feature subspaces for domain shifts.
K-Metamodes clusters security data without converting categorical attributes.
problem Clustering heterogeneous security data sets with categorical and numerical attributes.
method Frequency-based distance function for ensemble-based k-modes clustering, adapted feature discretisation.
result Higher effectiveness compared to previous methods on public security data sets.
Efficiently classifies binary labels with XOR queries, even under noisy conditions.
problem Binary classification with unknown labels using XOR queries.
method Effective query type and an efficient inference algorithm for noisy conditions.
result Achieves information-theoretic limit on optimal number of queries.
PredDiff measures prediction changes while marginalizing features, offering new insights into interaction effects.
problem Understanding interaction effects in black-box models.
method Model-agnostic, local attribution method based on probability theory.
result Introduced a new measure for interaction effects between arbitrary feature subsets.
New credit attribution methods for machine learning models using relaxed stability guarantees.
problem Ensuring proper attribution in generative models trained on existing works.
method Proposed new definitions of stability that allow for non-stable processing of a subset of datapoints with permission.
result Extended well-studied stability notions and provided a comprehensive characterization of learnability.
New algorithm selects features more efficiently than existing methods.
problem Feature selection and attribute reduction in machine learning.
method Hybridizes B-SFLA with FRDD for efficient feature selection.
result Significantly outperforms other metaheuristic methods in feature selection and classification accuracy.
The paper proposes a unified taxonomy for biclustering methods.
problem Lack of a unified taxonomy for biclustering methods.
method Using concept lattices and attribute exploration to build a taxonomy.
result A unified taxonomy for biclustering methods.
In this paper we analyze a budgeted learning setting, in which the learner can only choose and observe a small subset of the attributes of each training example. We develop efficient algorithms for ridge and lasso linear regression, which utilize the geometry of the data by a novel data-dependent sampling scheme. When …
Study evaluates consistency of feature attribution in deep learning for multi-omics data.
problem Challenges in interpretability of deep learning models in biological research.
method Investigation of Shapley Additive Explanations (SHAP) on multi-view deep learning models applied to multi-omics data.
result SHAP rankings are sensitive to architecture and random initialization, suggesting caution.
HC test measures word-frequency similarity for authorship attribution.
problem Identifying the author of a document based on word-frequency patterns.
method Adapting Higher Criticism (HC) to compare word-frequency tables.
result HC identifies characteristic words of the author, unaffected by topic structure.
The paper defines fairness in AI using intersectionality, proving guarantees and providing an algorithm.
problem Fairness in AI systems, especially considering multiple overlapping protected attributes.
method Intersectional framework, proving guarantees, learning algorithm.
result Intersectional fairness criteria behave sensibly for any subset of protected attributes.
Identifies minimal training subset to flip a prediction.
problem Flipping predictions in machine learning models.
method Extended influence function for relabeling minimal subset.
result Relabeling fewer than 2% of training points can flip a prediction.
New framework analyzes pre-stock jump trading behaviors using multivariate time series analysis.
problem Understanding micro-trading behaviors before stock price jumps.
method Multivariate time series analysis considering temporal information.
result Identifies highly informative attributes for predicting price jumps.
A new method for fairness in machine learning without explicit protected attributes.
problem Ensuring fairness in machine learning without explicit protected attributes.
method Example-based approach using paired consistency as a fairness regularizer.
result Demonstrated effectiveness on the Income Census dataset.
LLM extracts actionable insights from customer reviews.
problem Extracting actionable insights from customer reviews.
method Large language model approach distinguishing perceptual attributes from actionable features.
result High consistency and predictive validity of LLM insights compared to human coders.
Framework for multi-view redescription mining overcomes limitations of existing approaches.
problem Challenges in revealing non-trivial associations between different subsets of attributes (views).
method Memory-efficient, extensible framework using multi-target regression or multi-label classification algorithms.
result Framework can generate redescriptions from multiple views, improving over existing two-view approaches.
Nodes in real world networks often have class labels, or underlying attributes, that are related to the way in which they connect to other nodes. Sometimes this relationship is simple, for instance nodes of the same class are may be more likely to be connected. In other cases, however, this is not true, and the way tha…
GraphSAC detects anomalies in large graphs by sampling and filtering node subsets.
problem Vulnerability of holistic anomaly detection methods to compromised nodal attributes and network links.
method Randomly draws subsets of nodes, filters out contaminated sets, and uses SSL to estimate nominal label distributions.
result GraphSAC provides performance guarantees and is scalable to large graphs.
Differentiable Masking reveals how neural models make decisions across layers.
problem Intractable and expensive approximate search for input relevance in deep models.
method Differentiable Masking learns to mask inputs while maintaining differentiability.
result Reveals how decisions are formed across network layers in BERT models.
ExCIR provides efficient, consistent, and scalable explainability for complex models.
problem Complex models lack transparency and require efficient, stable, and scalable explainability methods.
method ExCIR uses correlation-aware feature attribution with robust centering and groupwise aggregation.
result ExCIR delivers trustworthy agreement with global baselines and full model rankings, reduces runtime, and scales to large datasets.
MISS identifies influential subsets in ML models.
problem Capturing collective influence of training samples.
method Analyzed and compared influence-based greedy heuristics and adaptive versions.
result Adaptive heuristics can better capture sample interactions.
SAM adds semantic attributes to language models for better interpretation and style variation.
problem Improving text interpretation and style variation in language models.
method SAM includes document attributes, scores them, and embeds them into the model's input space.
result SAM generates interpretable texts and shows superior performance on various datasets.
New algorithm reduces regret in dynamic assortment selection.
problem Dynamic assortment selection with consumer choice modeling.
method Optimistic algorithm with convex relaxation.
result Regret bound of O ( d T + κ ) O(\sqrt{dT} + κ) O ( d T + κ ) , improving over existing methods. Proposes a supervised VAE to reveal model invariances for interpretability.
problem Understanding and interpreting complex supervised models.
method Supervised variational auto-encoders (VAEs) with latent space invariances.
result Reveals model invariances through sampling nuisance dimensions.
AttGAN edits facial attributes by changing only what you want, preserving details.
problem Facial attribute editing with preservation of details.
method Encoder-decoder architecture with attribute classification and reconstruction learning.
result Outperforms state-of-the-arts on realistic attribute editing with preserved details.
Rotation forest is superior to other classifiers for problems with continuous features.
problem Classifying problems with real-valued features.
method Empirical comparison of classifiers from three families: SVM, tree-based ensembles, and neural networks.
result Rotation forest is significantly more accurate than competing techniques on average.
Proposes a Taylor framework to unify and analyze attribution methods.
problem Lack of a unified guideline for feature contribution assignment in machine learning models.
method Introduces a Taylor attribution framework to model the attribution problem and reformulates fourteen mainstream methods.
result Empirically validates the Taylor reformulations and reveals a positive correlation between performance and principles followed.
End-to-end models classify composers from musical scores.
problem Classifying composers from musical scores using machine learning.
method Pooled and convolutional architectures for feature extraction.
result Models achieve high accuracy on a large corpus of scores.