New methods solve sparse linear regression with limited attribute observation.
problem Sparse linear regression with limited attribute observation.
method Stochastic gradient methods using hard thresholding and adaptive combination of exploration and exploitation.
result Achieves sample complexity of O(1/ε) for error ε under restricted eigenvalue condition.
We consider the most common variants of linear regression, including Ridge, Lasso and Support-vector regression, in a setting where the learner is allowed to observe only a fixed number of attributes of each example at training time. We present simple and efficient algorithms for these problems: for Lasso and Ridge reg…
Algorithm uncovers latent attribute graph from molecular data.
problem Learning latent representations and interpreting them for limited data.
method Perturbation experiments on latent codes of a generative autoencoder.
result Effective graphical model of latent codes and attributes.
This paper tackles efficient clustering of moderate dimensional data from dual observation and attribute spaces.
problem Clustering high-dimensional data, especially when dimensionality is moderate to small.
method Develops an efficient clustering processing pipeline using dual spaces of observations and attributes.
result Established an effective method for clustering in moderate dimensional data.
WGNN learns graph representations from incomplete attribute data.
problem Missing node attributes in graphs.
method WGNN learns node representations from decomposed attribute matrices and uses Wasserstein space for message passing.
result WGNN outperforms existing methods in node classification tasks with missing attribute data.
New approach uses causal reasoning to address fairness issues.
problem Fairness criteria based on observational data are limited and unreliable.
method Shifts focus from observational criteria to causal reasoning.
result Formalizes why and when observational criteria fail.
Infer-AVAE infers missing user attributes from incomplete data using a novel adversarial approach.
problem Incomplete user attributes in social networks.
method Infer-AVAE combines MLP and GNNs with adversarial training to infer missing attributes.
result Infer-AVAE outperforms baselines by 7.0% in accuracy on real-world datasets.
Unified framework for analyzing machine learning model attributions.
problem Lack of a general and theoretical framework for understanding attribution methods.
method Proposes a Taylor attribution framework to unify and analyze seven mainstream attribution methods.
result Established three principles for good attribution and empirically validated the Taylor reformulations.
Concept modulation models unify identifiability and extrapolation in conditional latent variable models.
problem Reliable generalization in conditional latent variable models
method Concept modulation models (CMMs) with structure AoΛoCoX result Lifts identifiability to conditional settings and controls extrapolation through attribute potentials.
Develops methods to measure and reduce fairness in datasets with limited protected attribute labels.
problem Measuring and reducing fairness in datasets with limited protected attribute labels.
method Proposes methods to estimate fairness metrics and train models to limit fairness violations using probabilistic protected attribute labels.
result Our methods provide tighter bounds on true disparity and effectively reduce fairness violations with lesser fairness-accuracy trade-offs.
Embed nodes with multi-scale attributes for robust network analysis.
problem Capturing complex node attributes across different scales.
method Multi-scale attributed node embedding (AE & MUSAE) using Skip-gram approach.
result Proves node-feature mutual information is implicitly factorized by embeddings.
In this paper we consider classes of models that have been recently developed for quantitative finance that involve modelling a highly complex multivariate, multi-attribute stochastic process known as the Limit Order Book (LOB). The LOB is the primary data structure recorded each day intra-daily for all assets on every…
Paper studies fairness postprocessing with imperfect attribute information.
problem Ensuring fairness with imperfect protected attribute information.
method Equalized odds postprocessing method with imperfect attribute information.
result Conditions on perturbation ensure reduced bias in classifier.
New method explains visual models by altering features causally.
problem Inability of pure observational data to compute reliable feature effects.
method Causal Counterfactuals and intervened causal models.
result Computes counterfactuals to show model reactions to feature changes.
Detects adversarial examples with feature attribution differences.
problem Easily fooled by small adversarial perturbations in deep neural networks.
method Thresholding a scale estimate of feature attribution scores.
result Superior performance in distinguishing adversarial examples.
Paper develops methods for fair insurance pricing without direct access to sensitive attributes.
problem Fairness in insurance pricing with restricted access to sensitive attributes.
method Develops statistical methods for estimating discrimination-free premiums using privatized sensitive attributes.
result The proposed methods enable fair insurance pricing while respecting privacy and regulatory constraints.
New framework for fair ranking with noisy protected attributes.
problem Errors in socially-salient attributes undermine fairness guarantees.
method Modeling perturbations in protected attributes and incorporating probabilistic information.
result Framework provides provable guarantees on fairness and utility.
Anomaly detection method separates contextual from behavioral attributes.
problem Detect anomalies in data without labeled examples.
method Uses joint deep variational generative models.
result Robust to anomalous or novel contextual attributes.
New research on Shapley values for feature attribution in machine learning, considering model vs. data fidelity.
problem Controversy in connecting machine learning models to coalitional games, differing approaches.
method Investigates two approaches: interventional vs. observational conditional expectation Shapley values for linear models.
result The choice between model and data fidelity depends on the specific application.
DGBO optimizes attributed graphs efficiently.
problem Optimizing graphs with rich contextual features.
method Deep graph neural network surrogate for scalable Bayesian optimization.
result DGBO scales linearly with observations and outperforms state-of-the-art methods.
Unified approach for conversational recommendation by integrating attributes and items.
problem Cold-start users' real-time personalization in online recommendation.
method Seamlessly unifies attributes and items in Thompson Sampling framework for interactive decision-making.
result Conversational Thompson Sampling (ConTS) outperforms existing methods in success rate and conversation turns.
New method separates graph structure from node attributes to recover lost signal.
problem Standard representation learning on attributed graphs merges incompatible metric spaces, leading to geometrically flawed alignment.
method Custom variational autoencoder that separates manifold learning from structural alignment.
result Transforms geometric conflict into interpretable structural descriptor, uncovering connectivity patterns and anomalies.
Generative model controls text attributes for realistic sentences.
problem Challenges in generating natural language sentences with desired attributes.
method Combines variational auto-encoders and holistic attribute discriminators for semantic structure imposition.
result Effective generation of realistic sentences with desired attributes.
Robustly detects and attributes climate change impacts under interventions.
problem Detect and attribute climate change impacts from observations robustly.
method Supervised learning with anchor regression for robust predictions under interventions.
result CO2 forcing can be robustly predicted from temperature patterns under strong solar forcing interventions.
Proposes a Taylor framework to unify and analyze attribution methods.
problem Lack of a unified guideline for feature contribution assignment in machine learning models.
method Introduces a Taylor attribution framework to model the attribution problem and reformulates fourteen mainstream methods.
result Empirically validates the Taylor reformulations and reveals a positive correlation between performance and principles followed.
New algorithm disentangles latent space in GANs using video sequences.
problem Learning disentangled latent spaces in GANs without supervision.
method Adversarial training with video sequences, modifying standard GAN algorithm.
result Disentangled latent space into content and motion attributes.
New method uses neural networks to identify sources from limited data in complex systems.
problem Identifying sources from noisy and limited data in high-dimensional systems.
method Calibrating deep neural network surrogates to ensemble simulations and using Bayesian optimization for source identification.
result Reliable source identification with uncertainty quantification using limited data and auxiliary processes.
ETNA model predicts demographic attributes from transactions.
problem Predicting demographic attributes from transaction histories with limited data.
method Embedding Transformation Network with Attention (ETNA) model.
result ETNA model outperforms previous models on all tasks.
New analysis improves accuracy of Newton step and influence function data attributions.
problem Improving accuracy of data attribution methods for logistic regressions.
method Introducing a new analysis of Newton Step and Influence Function data attribution methods for convex learning problems.
result Proved asymptotically tight error bounds for Newton Step and Influence Function data attribution methods.
New smart contract mechanisms evade traditional AML systems by decoupling transaction roles.
problem Current AML systems fail to track economic value migration in composable smart contracts.
method Introduce PEB separation and state-mediated value migration to demonstrate how traditional tracing fails.
result Transfer-layer observation is incomplete and causally ambiguous in composable smart contracts.
Proposes VCLANC for attributed network clustering using node and attribute embeddings.
problem Lack of mutual affinity exploitation between nodes and attributes in graph convolution.
method Dual variational auto-encoders for node and attribute embeddings, Gaussian mixture model priors, mutual distance and clustering assignment hardening losses.
result Demonstrates effectiveness on real-world attributed network datasets.
In this paper, we address the problem of conditional modality learning, whereby one is interested in generating one modality given the other. While it is straightforward to learn a joint distribution over multiple modalities using a deep multimodal architecture, we observe that such models aren't very effective at cond…
Proposes m-GCRF for semi-supervised regression on graphs with missing data.
problem Handling missing labels in partially observed temporal attributed graphs.
method Marginalized Gaussian Conditional Random Fields (m-GCRF) for semi-supervised learning.
result Consistently more accurate than alternative models in experiments.
New method clusters graph topology and attributes efficiently.
problem Clustering attributed graphs with complex topology-attribute relationships.
method Symmetric NMF with PU learning for non-linear projection.
result Outperforms existing methods in clustering quality.
Consider observation data, comprised of n observation vectors with values on a set of attributes. This gives us n points in attribute space. Having data structured as a tree, implied by having our observations embedded in an ultrametric topology, offers great advantage for proximity searching. If we have preprocessed d…
While state-of-the-art kernels for graphs with discrete labels scale well to graphs with thousands of nodes, the few existing kernels for graphs with continuous attributes, unfortunately, do not scale well. To overcome this limitation, we present hash graph kernels, a general framework to derive kernels for graphs with…
New method combines hypergraph structure and node attributes for better community detection.
problem Improving community detection in hypergraphs with node attributes.
method Developed a principled model that learns from data to combine higher-order interactions and node attributes.
result Strong performance in hyperedge prediction and community detection, especially when attributes are informative.
GGDA simplifies DA for large models, speeding up attribution by up to 50x.
problem Computational intensity of existing DA methods limits their applicability to large-scale models.
method Generalized Group Data Attribution (GGDA) framework attributing to groups of training points.
result GGDA achieves up to 50x speedups over standard DA methods while maintaining effectiveness.
New model generates unseen attribute combinations from limited data.
problem Lack of generalization in deep generative models for unseen attribute combinations.
method Introduces multilinear latent conditioning to capture multiplicative interactions.
result Demonstrates effectiveness on MNIST, Fashion-MNIST, and CelebA datasets.
Improved 3D ECG feature attributions for clinical interpretation.
problem Lack of interpretability in deep learning models for 12-lead ECG analysis.
method Cross-modal mapping of feature attributions from 12-lead ECG models onto CineECG 3D space.
result Mapped feature attributions yield higher Dice scores than standard 12-lead attributions.
We interpret black box predictive models using causal attribution.
problem Interpreting models trained using machine learning in high-stakes applications.
method Estimate causal effects of model inputs on output using observational data.
result Effective interpretation of black box predictive models via causal attribution.
In-Run Data Shapley offers efficient data attribution for large-scale models.
problem Existing data attribution methods are computationally intensive and cannot target specific models.
method In-Run Data Shapley, which efficiently attributes data contributions to a specific model without re-training.
result In-Run Data Shapley achieves significant efficiency, enabling data attribution for pretraining models.
Adaptive graph convolution improves attributed graph clustering performance.
problem Joint modeling of graph structures and node attributes is challenging.
method Adaptive graph convolution that captures global cluster structure and selects appropriate order for different graphs.
result Empirical results show our method compares favorably with state-of-the-art methods.
Study on list learning with noisy data, showing limits and some learnable cases.
problem Learning from noisy data in a list learning context.
method Inspired by coding theory, extends list learning model to study sparse conjunctions and parities/majors.
result Sparse conjunctions can be efficiently list learned under certain conditions, but parities and majors cannot be efficiently learned.
Nonparametric method measures influence of training images on diffusion model outputs.
problem Quantifying influence of individual training examples on diffusion model outputs.
method Patch-level similarity between generated and training images, using optimal score function.
result Strong attribution performance, matching gradient-based approaches and outperforming baselines.
Proposes DAPr framework to learn feature importance from prior knowledge.
problem Ensuring meaningful feature attributions in deep models.
method Jointly learns feature importance from prior knowledge and biases models to rely on important features.
result Improves model generalization and provides new interpretation methods.
Simple graph representation outperforms complex methods in graph classification.
problem Graph classification and representation learning on graphs.
method Developed a simple yet meaningful graph representation and tested its effectiveness.
result Simple graph representation achieves similar performance to state-of-the-art methods for non-attributed graph classification.
New method explains time series classification by assessing causal effects.
problem Understanding machine learning model decisions in time series classification.
method Model-agnostic causal attribution method using diffusion models.
result Causal attributions differ from associational ones, highlighting risks.