New method identifies subgroups in censored data.
problem Identifying meaningful patterns in heterogeneous populations.
method Combining inverse probability weighting, M-estimation, and concave pairwise fusion penalization.
result Robust approach for censored data under heterogeneous AFT models.
Sparse GFA identifies disease factors in FTD subgroups.
problem Heterogeneity in neurological disorders hinders understanding and treatment.
method Sparse Group Factor Analysis (GFA) with regularised horseshoe priors.
result Identified latent disease factors differentially expressed in FTD subgroups.
Develops algorithm to find subgroups with different treatment effects in HIV patients.
problem Estimating treatment effects in EHR data with challenges like time-varying confounding.
method SDLD algorithm combining generalized interaction tree and longitudinal targeted maximum likelihood estimation.
result Identifies subgroups of HIV patients at higher risk of weight gain with dolutegravir-containing ARTs.
Model patching closes subgroup performance gaps in skin cancer classification.
problem Inconsistent model performance on specific subgroups of a class.
method Two-stage framework that models subgroup features and learns semantic transformations, followed by data augmentation.
result Reductions in robust error of up to 33% relative to best baseline on benchmark datasets.
Causal Interaction Trees identify treatment subgroup effects in observational data.
problem Identifying subgroups with enhanced treatment effects in observational studies.
method Extending Classification and Regression Trees with subgroup-specific treatment effect estimators.
result The proposed algorithms enhance treatment effect heterogeneity in subgroups.
Proposes a method to learn fair predictors for multiple subgroups with limited data.
problem Fairness and accuracy issues in learning from multiple subgroups with limited data.
method Formulates a bilevel objective to learn subgroup-specific predictors and a fair predictor that is close to all of them.
result The method effectively controls group sufficiency and generalization error, improving fairness and accuracy.
Proposes a method to select features for subgroup datasets with systematic missing data.
problem Feature selection for datasets with subgroup structure and systematic missing data.
method Develops a heterogeneous graph neural network to propagate information between feature-subgroup-target variable connections.
result Demonstrates improved feature selection performance and scalability.
We analyse an issue when comparing survival curves between two subgroups. We show that there is a direct relationship between estimates of subgroups' survival at a time point and positive and negative predictive values in the binary classification settings. Our findings present a case where current methods of comparing…
The paper proposes a method to find subgroups with significant treatment effects in noisy data.
problem Estimating the causal effects of interventions on noisy outcomes.
method A machine-learning method specifically optimized for finding subgroups with significant effects, designed to maximize the probability of obtaining a statistically significant positive treatment effect.
result The proposed method yields higher power in detecting subgroups affected by the treatment compared to standard tree-based tools.
We discover subgroups for Cox model survival analysis, improving model accuracy.
problem Finding interpretable subsets of data where Cox model is highly accurate.
method Developed new metrics (EPE, CRS) and algorithms to solve subgroup discovery problem.
result Our methods improve model fit and recover known nonlinearities in data.
Proposes a new method for finding non-redundant, standout subgroups in numeric datasets.
problem Mining large numbers of redundant subgroups in numeric datasets.
method Dispersion-aware problem formulation based on MDL principle for subgroup set discovery.
result Empirically demonstrates SSD++ returns outstanding subgroup lists.
New framework shows diverse training data improves subgroup and overall performance.
problem Lack of understanding how diverse training data affects subgroup and overall performance.
method Casts data collection as part of the learning process, analyzes dataset compositions, and guides dataset design.
result Diverse representation in training data improves subgroup and overall performance.
Proposes a method to identify subgroup structure and estimate covariate effects for multivariate response data.
problem Identifying subgroup structure and estimating covariate effects in multivariate response data.
method Joint heterogeneity and reduced-rank learning framework using rank-constrained pairwise fusion penalization.
result Established the asymptotic properties of the estimators and proposed a predictive information criterion for rank selection.
New methods improve subgroup analysis in trials with limited data.
problem Limited sample sizes in subgroup analyses of randomized controlled trials.
method Two TMLEs that borrow information from non-subgroup participants.
result Improved precision in subgroup-specific treatment effect estimates.
A new method clusters heterogeneous subgroups for accurate causal learning.
problem Diverse causal relationships across different time spans, regions, or strategies.
method Nonlinear Causal Kernel Clustering
result Reduction in prediction error through enhanced causal learning.
We consider high-dimensional regression over subgroups of observations. Our work is motivated by biomedical problems, where disease subtypes, for example, may differ with respect to underlying regression models, but sample sizes at the subgroup-level may be limited. We focus on the case in which subgroup-specific model…
Adapts to shifts in latent subgroup distributions without labeled target data.
problem Adapting to domain shifts when latent subgroup distributions differ.
method Uses concept and proxy variables from source domain, and unlabeled target data.
result Optimal target predictor can be identified and estimated.
Robust subgroup discovery finds non-redundant, statistically significant subgroups.
problem Finding interpretable, robust subgroups from data.
method Formulated subgroup lists for univariate and multivariate targets, used MDL principle and greedy heuristic SSD++.
result SSD++ outperforms previous methods in quality and size of subgroup lists.
The paper tackles fairness in forecasting and learning linear dynamical systems.
problem Under-representation bias in training data for multiple subgroups.
method Introducing subgroup-fair and instant-fair learning of LDS from multiple trajectories of varying lengths, using hierarchies of convexifications of non-commutative polynomial optimisation problems.
result Empirical results show both the beneficial impact of fairness considerations on statistical performance and encouraging effects of exploiting sparsity on run time.
CAPITAL algorithm identifies optimal patient subgroups for better treatment.
problem Identify maximum number of patients benefiting from better treatment.
method Constrained Policy Tree Search (CAPITAL) algorithm to find optimal subgroup selection rule (SSR).
result Maximizes the number of patients with enhanced treatment effects.
Chiseling finds valid subgroups interactively, improving on existing methods.
problem Finding valid subgroups with inferential guarantees in regression and causal inference.
method Interactive subgroup refinement with inferential validity guarantees.
result Chiseling identifies better subgroups than existing methods with inferential guarantees.
New algorithm tackles subgroup fairness in AI with multiple sensitive attributes.
problem Heavy computational burdens and data sparsity in subgroup fairness for multiple sensitive attributes.
method Doubly Regressing Adversarial learning (DRAF) for subgroup fairness, focusing on subgroups with sufficient sample sizes and marginal fairness.
result DRAF algorithm reduces a surrogate fairness gap for supIPM with less computation than directly reducing supIPM.
We present a novel subset scan method to detect if a probabilistic binary classifier has statistically significant bias -- over or under predicting the risk -- for some subgroup, and identify the characteristics of this subgroup. This form of model checking and goodness-of-fit test provides a way to interpretably detec…
Posterior conformal prediction improves prediction interval validity for subgroups.
problem Marginal and conditional prediction interval validity for subgroups.
method Modeling conditional nonconformity score distribution as a mixture of cluster distributions.
result PCP produces tighter prediction intervals, especially for well-represented clusters.
A new method improves AI fairness assessment by estimating performance across intersectional subgroups.
problem Limited evaluation of AI systems across intersectional subgroups due to small sample sizes.
method Structured regression approach to disaggregated evaluation.
result Our method yields more accurate performance estimates, especially for small subgroups.
Improves fairness in machine learning by adding underrepresented group data.
problem Machine learning biases across subgroups due to under-representation or societal biases.
method Data augmentation via pairwise mixup across subgroups to balance subpopulations.
result Achieves fair outcomes with robust if not improved accuracy.
The growing capability and accessibility of machine learning has led to its application to many real-world domains and data about people. Despite the benefits algorithmic systems may bring, models can reflect, inject, or exacerbate implicit and explicit societal biases into their outputs, disadvantaging certain demogra…
In this paper, we investigate the effect of machine learning based anonymization on anomalous subgroup preservation. In particular, we train a binary classifier to discover the most anomalous subgroup in a dataset by maximizing the bias between the group's predicted odds ratio from the model and observed odds ratio fro…
Method learns representation invariant to subgroup support using adversarial matching.
problem Representational bias in training data leads to spurious correlations.
method Adversarial support-matching using semi-supervised clustering.
result Improves generalizability to unseen sources.
Improved linear regression for diverse data batches.
problem Learning from multiple heterogeneous data sources.
method Gradient-based algorithm for different input distributions.
result Significant reduction in the number of required batches and sample size.
cMCA uses contrastive learning to identify latent subgroups in political party data.
problem Identifying latent subgroups within political party data.
method Contrastive learning applied to multiple correspondence analysis (MCA).
result cMCA identifies latent subgroups not seen by traditional methods.
A subgroup of a Kac-Moody group is called bounded if it is contained in the intersection of two finite type parabolic subgroups of opposite signs. In this paper, we study the isomorphisms between Kac-Moody groups over arbitrary fields of cardinality at least 4, which preserve the set of bounded subgroups. We show that …
This study evaluates subgroup analysis methods for time-to-event outcomes in randomized controlled trials.
problem Identifying subgroups of good responders in non-significant randomized controlled trials.
method Evaluation of several subgroup analysis algorithms for time-to-event outcomes using synthetic and semi-synthetic data.
result Provides a new synthetic and semi-synthetic data generation process and an open-source Python package for benchmarking.
DDGroup identifies subgroups with uniform linear relationships.
problem Heterogeneous effects of covariates in linear models.
method Data-driven method to identify subgroups with uniform linear relationships.
result DDGroup can discover subgroups with improved performance.
D3M debiases models by selectively removing problematic examples.
problem Model failures on underrepresented subgroups.
method Isolates and removes specific training examples that cause failures.
result Efficiently trains debiased classifiers with minimal example removal.
This paper studies three aspects around dimension datum: (1), a generalization of the dimension datum, which we call the tau-dimension datum; (2), dimension data of disconnected subgroups; (3), compactness of isospectral sets of normal homogeneous spaces.
Cluster analysis methods are used to identify homogeneous subgroups in a data set. In biomedical applications, one frequently applies cluster analysis in order to identify biologically interesting subgroups. In particular, one may wish to identify subgroups that are associated with a particular outcome of interest. Con…
In subgroup discovery, also known as supervised pattern mining, discovering high quality one-dimensional subgroups and refinements of these is a crucial task. For nominal attributes, this is relatively straightforward, as we can consider individual attribute values as binary features. For numerical attributes, the task…
Proposes a method to detect anomalies in multi-subgroup normal data.
problem Anomaly detection with limited labeled anomalies and multi-subgroup normal data.
method Learn multi-normal prototypes with deep embedding clustering and contrastive learning. Estimate the likelihood of unlabeled samples being normal during training.
result Superior performance compared to state-of-the-art methods on various datasets.
R2P method identifies homogeneous and heterogeneous subgroups for better treatment effect estimation.
problem Current subgroup analysis methods are weak in identifying homogeneous and heterogeneous subgroups and lack confidence estimates.
method R2P uses an arbitrary ITE estimator and quantifies uncertainty robustly.
result R2P produces more homogeneous and heterogeneous partitions than other methods.
MOB-dS uses permutation to correct for dependency in discrete survival data.
problem Identifying subgroups in discrete event time data with potential spurious results.
method Model-based recursive partitioning (MOB) with modified data matrix and permutation test.
result MOB-dS controls type I error rate better than standard MOB for discrete survival data.
Efficient neural network invariant to symmetry subgroups.
problem Designing neural networks invariant to symmetry subgroups for computational efficiency.
method A new G-invariant transformation module and multi-layer perceptron. result The proposed architecture is computationally and memory efficient, and universal.
New risk measures control subgroup imbalances, improving PAC-Bayesian bounds.
problem Insufficient risk bounds for subgroup imbalances in data.
method Introduce constrained f-entropic risk measures and derive PAC-Bayesian bounds.
result First disintegrated PAC-Bayesian guarantees beyond standard risks.
The paper introduces moment multicalibration for estimating uncertainty across subgroups.
problem Ensuring fairness and accurate uncertainty estimation in predictions across different subgroups.
method Develops a method for multicalibration of higher moments, enabling point predictions and interval estimation.
result Moment multicalibration allows for valid prediction intervals that are fair across various subgroups.
GLMM trees identify subgroups with different growth patterns in longitudinal data.
problem Identifying subgroups with distinct growth trajectories in longitudinal studies.
method Extended GLMM trees for longitudinal data.
result Extended GLMM trees outperform other methods in accuracy and speed.
In this paper, we compute the subgroup distortion of all finitely generated subgroups of all finitely generated 3-manifold groups, and the subgroup distortion in this case can only be linear, quadratic, exponential and double exponential. It turns out that the subgroup distortion of a subgroup of a 3-manifold group is …
GAME improves matrix completion by considering subgroup-specific latent structures.
problem Heterogeneous data with overlapping categories, smoothing away subgroup-specific variation.
method Group-Aware Matrix Estimation (GAME) with overlapping nuclear-norm penalties.
result GAME outperforms global low-rank estimators in structured missingness regimes.
Regular subgroups of SL3(R) are identified and ruled out.
problem Identifying and characterizing regular subgroups of SL3(R).
method Using Kapovich–Leeb–Porti and Guichard–Wienhard divergent subgroups criteria, and Oh's results.
result Regular subgroups of SL3(R) are precisely lattices in minimal horospherical subgroups.