New attribution model boosts ad bidding efficiency.
problem Inefficiency of standard bidding policies in ad exchanges.
method Developed and applied an attribution model within the bidder.
result Average bid increased after incorporating attribution model.
In-Run Data Shapley offers efficient data attribution for large-scale models.
problem Existing data attribution methods are computationally intensive and cannot target specific models.
method In-Run Data Shapley, which efficiently attributes data contributions to a specific model without re-training.
result In-Run Data Shapley achieves significant efficiency, enabling data attribution for pretraining models.
A-DOGE embeds attributed graphs efficiently using density of states.
problem Efficiently represent node-attributed graphs with few numerical features.
method A-DOGE uses density of states to blend topology and attributes, leveraging efficient approximation algorithms.
result A-DOGE achieves competitive performance with modern supervised GNNs while being significantly faster.
New method clusters graph topology and attributes efficiently.
problem Clustering attributed graphs with complex topology-attribute relationships.
method Symmetric NMF with PU learning for non-linear projection.
result Outperforms existing methods in clustering quality.
GGDA simplifies DA for large models, speeding up attribution by up to 50x.
problem Computational intensity of existing DA methods limits their applicability to large-scale models.
method Generalized Group Data Attribution (GGDA) framework attributing to groups of training points.
result GGDA achieves up to 50x speedups over standard DA methods while maintaining effectiveness.
Efficiently explains model outputs using HSIC, a dependence measure.
problem Efficiently explain model outputs for various architectures.
method HSIC, RKHS, Reproducing Kernel Hilbert Spaces, black-box attribution.
result Up to 8 times faster than previous methods while maintaining fidelity.
PSI models and infers feature attributions efficiently and accurately.
problem Modeling and inferring feature attributions in flexible predictive models.
method Probabilistic Shapley inference (PSI) framework using latent random variables and a masking-based neural network architecture.
result PSI learns feature attribution distributions centered at Shapley values, revealing meaningful uncertainty.
Efficiently learns halfspaces with malicious noise, near-optimal label complexity.
problem Learning s-sparse halfspaces under malicious label noise. method Active learning algorithm with instance reweighting and empirical risk minimization.
result Near-optimal label complexity of O(slog4d/ε) and noise tolerance Ω(ε). This paper tackles efficient clustering of moderate dimensional data from dual observation and attribute spaces.
problem Clustering high-dimensional data, especially when dimensionality is moderate to small.
method Develops an efficient clustering processing pipeline using dual spaces of observations and attributes.
result Established an effective method for clustering in moderate dimensional data.
New algorithm learns PTFs with noisy data efficiently.
problem Learning low-degree PTFs with noisy data efficiently.
method Structural result and novel robust Chow vector estimation.
result PAC learns PTFs with nasty noise using efficient samples.
SIM-Shapley improves SV approximation efficiency and stability.
problem High computational costs of Shapley value methods in high-dimensional settings.
method Stochastic Iterative Momentum for Shapley Value Approximation (SIM-Shapley).
result Reduced computation time by up to 85% while maintaining feature attribution quality.
DeepGL learns hierarchical graph representations from attributed graphs.
problem Learning deep node and edge representations from large attributed graphs.
method Derives base features, learns multi-layered hierarchical graph representation, leverages previous layer outputs, supports attributed graphs, learns interpretable features, and is space-efficient.
result DeepGL learns relational functions that generalize across-networks and is effective for across-network transfer learning tasks.
Improved graph clustering with modularity and coarsening for attributes and communities.
problem Inaccurate community detection and computational inefficiency in graph clustering.
method Integrates coarsening and modularity maximization, using a loss function with log-determinant, smoothness, and modularity components.
result Superior clustering outcomes, proven consistent under DC-SBM, and efficient algorithm integration with GNNs and VGAEs.
Generative model controls text attributes for realistic sentences.
problem Challenges in generating natural language sentences with desired attributes.
method Combines variational auto-encoders and holistic attribute discriminators for semantic structure imposition.
result Effective generation of realistic sentences with desired attributes.
A framework for hypothesis testing on attributed graphs using sampling.
problem Statistical testing on graph data, especially large attributed graphs.
method Sampling-based framework with PHASE and PHASEopt for accurate and efficient hypothesis testing.
result PHASE and PHASEopt improve accuracy and efficiency of hypothesis testing in attributed graphs.
New methods solve sparse linear regression with limited attribute observation.
problem Sparse linear regression with limited attribute observation.
method Stochastic gradient methods using hard thresholding and adaptive combination of exploration and exploitation.
result Achieves sample complexity of O(1/ε) for error ε under restricted eigenvalue condition.
Embed nodes with multi-scale attributes for robust network analysis.
problem Capturing complex node attributes across different scales.
method Multi-scale attributed node embedding (AE & MUSAE) using Skip-gram approach.
result Proves node-feature mutual information is implicitly factorized by embeddings.
Bayesian approach uses node attributes to improve incomplete relational data.
problem Improving community detection and link prediction in incomplete relational data.
method Bayesian probabilistic approach incorporating binary node attributes for both directed and undirected networks, using efficient Gibbs sampling.
result State-of-the-art link prediction results, especially with highly incomplete data.
New method combines hypergraph structure and node attributes for better community detection.
problem Improving community detection in hypergraphs with node attributes.
method Developed a principled model that learns from data to combine higher-order interactions and node attributes.
result Strong performance in hyperedge prediction and community detection, especially when attributes are informative.
Proposes methods for local clustering in attributed graphs.
problem Finding a single cluster concentrated on a specific region in a graph.
method Introduces Graph Unimodality (GU) and Attribute Unimodality (AU) measures, and LOCLU algorithm to optimize Compactness score.
result Local cluster detected by LOCLU concentrates on the region of interest and exhibits unimodal data distribution.
Hierarchical CNNs improve image recognition with fewer parameters.
problem Difficulty in analyzing deep neural networks.
method Structured deep convolutional networks with progressively higher dimensional attributes learned from data.
result Hierarchical networks achieve comparable precision to state-of-the-art networks with fewer parameters.
GLACE embeds large-scale attributed graphs effectively, preserving structure and attributes.
problem Uncertainty and complexity in large-scale attributed graphs.
method Gaussian embeddings for scalable and efficient graph embedding.
result GLACE outperforms state-of-the-art methods on multiple graph analysis tasks.
Nonparametric method measures influence of training images on diffusion model outputs.
problem Quantifying influence of individual training examples on diffusion model outputs.
method Patch-level similarity between generated and training images, using optimal score function.
result Strong attribution performance, matching gradient-based approaches and outperforming baselines.
This work introduces a fair learning method for diverse sensitive attributes.
problem Fairness in supervised learning with complex sensitive attributes.
method Neural network with a simple random sampler for fairness penalties.
result The method improves fairness and utility on benchmark data.
Study on list learning with noisy data, showing limits and some learnable cases.
problem Learning from noisy data in a list learning context.
method Inspired by coding theory, extends list learning model to study sparse conjunctions and parities/majors.
result Sparse conjunctions can be efficiently list learned under certain conditions, but parities and majors cannot be efficiently learned.
FairICP addresses equalized odds fairness for multiple sensitive attributes.
problem Equalized odds fairness for multiple sensitive attributes.
method Adversarial learning with inverse conditional permutation.
result Promotes equalized odds under complex, multi-dimensional sensitive attributes.
Improved scalability and interpretability in training data attribution.
problem Identifying which training data drives specific behaviors, especially unintended ones.
method Leveraging interpretable structures within the model to attribute model behavior to semantic directions, not individual test examples.
result Simple probe-based attribution methods are first-order approximations of Concept Influence that achieve comparable performance while being over an order-of-magnitude faster.
Study on protecting sensitive properties of datasets during analysis.
problem Ensuring privacy of sensitive properties in datasets.
method Proposes definitions and mechanisms for attribute privacy using the Pufferfish framework.
result Developed efficient and inefficient mechanisms for attribute privacy.
EmbNum learns numerical attribute representations without distributional assumptions.
problem Semantic labeling of numerical values with unknown distributions.
method Neural numerical embedding model (EmbNum) for deep metric learning.
result EmbNum significantly outperforms state-of-the-art methods for numerical attribute semantic labeling.
Attributing forecast gaps to component models in complex model suites
problem Attributing forecast gaps between model-suite forecasts and realized outcomes
method Formalizing walk analysis and adapting order-independent attribution frameworks
result Deriving efficient formulas for elementwise and vectorized gap attribution
XAI-Bench releases synthetic datasets for evaluating feature attribution methods.
problem Evaluating and comparing feature attribution methods is challenging.
method Released synthetic datasets and benchmarking library.
result Efficiently evaluates feature attribution methods across various metrics.
Many real world network problems often concern multivariate nodal attributes such as image, textual, and multi-view feature vectors on nodes, rather than simple univariate nodal attributes. The existing graph estimation methods built on Gaussian graphical models and covariance selection algorithms can not handle such d…
TRAK traces model predictions to training data efficiently.
problem Inefficiency in data attribution methods for large-scale models.
method TRAK: a new data attribution method that is both effective and computationally tractable.
result TRAK matches the performance of methods requiring thousands of models with just a handful.
In this paper we analyze a budgeted learning setting, in which the learner can only choose and observe a small subset of the attributes of each training example. We develop efficient algorithms for ridge and lasso linear regression, which utilize the geometry of the data by a novel data-dependent sampling scheme. When …
The paper tackles attributing forecast gaps in complex model suites.
problem Attributing forecast gaps to individual component models in complex model suites.
method Formalized walk analysis, adapted LMDI and Shapley value approaches.
result Developed efficient formulas for gap attribution in practical portfolio-scale examples.
Concept Relation Discovery and Innovation Enabling Technology (CORDIET), is a toolbox for gaining new knowledge from unstructured text data. At the core of CORDIET is the C-K theory which captures the essential elements of innovation. The tool uses Formal Concept Analysis (FCA), Emergent Self Organizing Maps (ESOM) and…
New attribution method for neural networks using causal principles.
problem Understanding the impact of features on neural network outputs.
method Viewing neural networks as Structural Causal Models, computing causal effects efficiently.
result Efficient computation of feature impacts on neural network outputs.
SISR improves feature attribution in complex payoff schemes.
problem Distorted feature attributions due to non-additive payoff functions and high-dimensional feature spaces.
method Sparse Isotonic Shapley Regression (SISR) learns a monotonic transformation to restore additivity and enforces L0 sparsity.
result SISR achieves strong support recovery and stable attributions across various payoff schemes.
Efficiently learns sparse halfspaces with minimal label queries.
problem PAC active learning of sparse linear classifiers.
method Attribute-efficient algorithm with sublinear label complexity.
result Achieves label complexity of O(t · polylog(d, 1/ε)) under certain assumptions.
DGBO optimizes attributed graphs efficiently.
problem Optimizing graphs with rich contextual features.
method Deep graph neural network surrogate for scalable Bayesian optimization.
result DGBO scales linearly with observations and outperforms state-of-the-art methods.
Develops GNNs for incomplete graphs, improving learning from missing node attributes.
problem Learning from incomplete graphs with missing node attributes.
method Introduces PaGNNs with novel partial aggregation functions for incomplete graph data.
result Demonstrates effectiveness and efficiency of PaGNNs on various datasets.
Proposes a method to quantify and explain deep learning model uncertainties.
problem Deep learning model predictions are sensitive to perturbations and adversarial attacks.
method Gradient-based uncertainty attribution method to identify problematic regions and propose mitigation strategies.
result Proposed UA-Backprop method achieves competitive accuracy and efficiency compared to existing methods.
A method for trust evaluation of devices in human-device coexistence systems.
problem Efficient trust evaluation of devices in systems with diverse physical and social attributes.
method Canonical correlation analysis-enhanced hypergraph self-supervised learning (HSLCCA).
result The proposed HSLCCA method significantly outperforms baseline algorithms in identifying trusted devices.
New method attributes feature uncertainty in ML models using cooperative game theory.
problem Lack of feature-level uncertainty attribution in explainable AI.
method Proposes a novel, model-agnostic uncertainty attribution method using cooperative game theory and conformal prediction.
result Demonstrates improved runtime efficiency and practical utility in real-world applications.
Study shows data attribution methods are sensitive to hyperparameters, making tuning costly.
problem Hyperparameter sensitivity in data attribution methods makes tuning impractical.
method Theoretical analysis and lightweight procedure for selecting regularization value without retraining.
result Proposes a lightweight procedure for selecting regularization value without model retraining.
Improved decision tree algorithm for more accurate data classification.
problem ID3's tendency to choose attributes with many values.
method Divide attributes into groups, apply selection measure 5, and recursively refine until good classification is achieved.
result Proposed algorithm classifies data sets more accurately and efficiently.
Paper learns latent and hierarchical structures in CDMs from data.
problem Jointly learning latent and hierarchical structures in CDMs from observed data.
method Penalized likelihood approach for selecting attributes and estimating structures; EM and latent structure recovery algorithms.
result Good performance demonstrated by simulation and real data applications.
GLSR-VAE enhances VAE latent space for better attribute control.
problem Fine-tuning VAE latent space for continuous data attributes.
method Geodesic Latent Space Regularization (GLSR) for VAEs.
result Controls latent space changes to reflect data attributes.