Chiseling finds valid subgroups interactively, improving on existing methods.
problem Finding valid subgroups with inferential guarantees in regression and causal inference.
method Interactive subgroup refinement with inferential validity guarantees.
result Chiseling identifies better subgroups than existing methods with inferential guarantees.
Proposes a method to select features for subgroup datasets with systematic missing data.
problem Feature selection for datasets with subgroup structure and systematic missing data.
method Develops a heterogeneous graph neural network to propagate information between feature-subgroup-target variable connections.
result Demonstrates improved feature selection performance and scalability.
Selective regression allows abstention to improve fairness criteria.
problem Selective regression can exacerbate disparities between subgroups.
method Proposes new fairness criteria and two approaches to mitigate performance disparity.
result Proposed fairness criteria ensures performance improvement for every subgroup with reduced coverage.
CAPITAL algorithm identifies optimal patient subgroups for better treatment.
problem Identify maximum number of patients benefiting from better treatment.
method Constrained Policy Tree Search (CAPITAL) algorithm to find optimal subgroup selection rule (SSR).
result Maximizes the number of patients with enhanced treatment effects.
D3M debiases models by selectively removing problematic examples.
problem Model failures on underrepresented subgroups.
method Isolates and removes specific training examples that cause failures.
result Efficiently trains debiased classifiers with minimal example removal.
Optimizes subgroup selection in clinical trials.
problem Identifying regions in feature space where a regression function exceeds a threshold.
method Formulates subgroup selection as constrained optimisation, determining minimax optimal rate for regret.
result Determines the minimax optimal rate for regret in sample size and Type I error probability.
This paper presents an original approach for jointly fitting survival times and classifying samples into subgroups. The Coxlogit model is a generalized linear model with a common set of selected features for both tasks. Survival times and class labels are here assumed to be conditioned by a common risk score which depe…
The paper proposes a method to find subgroups with significant treatment effects in noisy data.
problem Estimating the causal effects of interventions on noisy outcomes.
method A machine-learning method specifically optimized for finding subgroups with significant effects, designed to maximize the probability of obtaining a statistically significant positive treatment effect.
result The proposed method yields higher power in detecting subgroups affected by the treatment compared to standard tree-based tools.
Unified framework for variable selection in model-based clustering with missing data.
problem Challenges in identifying relevant variables and handling missing data in model-based clustering.
method Unified framework incorporating a data-driven penalty matrix and a mechanism for missingness modeling.
result Achieves both asymptotic consistency and selection consistency in the presence of missing data.
Proposes a method to identify subgroup structure and estimate covariate effects for multivariate response data.
problem Identifying subgroup structure and estimating covariate effects in multivariate response data.
method Joint heterogeneity and reduced-rank learning framework using rank-constrained pairwise fusion penalization.
result Established the asymptotic properties of the estimators and proposed a predictive information criterion for rank selection.
New framework for interpreting disaggregated fairness evaluations using causal models.
problem Misinterpretation of disaggregated fairness evaluations due to data representativeness and selection bias.
method Causal graphical models to characterize fairness properties and metric stability under different data generating processes.
result Disaggregated evaluations are unreliable without explicit assumptions regarding bias mechanisms.
WHOMP optimizes randomized controlled trials by minimizing subgroup bias.
problem Minimizing subgroup bias in randomized controlled trials.
method Wasserstein Homogeneity Partition (WHOMP) method.
result WHOMP optimally minimizes type I and type II errors in trials.
We discover subgroups for Cox model survival analysis, improving model accuracy.
problem Finding interpretable subsets of data where Cox model is highly accurate.
method Developed new metrics (EPE, CRS) and algorithms to solve subgroup discovery problem.
result Our methods improve model fit and recover known nonlinearities in data.
A spherical topological manifold of dimension n-1 forms a prototile on its cover, the (n-1)-sphere. The tiling is generated by the fixpoint-free action of the group of deck transformations. By a general theorem, this group is isomorphic to the first homotopy group. Multiplicity and selection rules appear in the form of…
Study on rotating surfaces in 4D space with matrices.
problem Understanding rotational surfaces in pseudo-Euclidean 4-space.
method Defined hyperbolic and elliptic rotational surfaces using curves and matrices in 4D semi-Euclidean space.
result Generated rotated surfaces using specific rotation matrices.
The identification of predictive biomarkers from a large scale of covariates for subgroup analysis has attracted fundamental attention in medical research. In this article, we propose a generalized penalized regression method with a novel penalty function, for enforcing the hierarchy structure between the prognostic an…
Given two possible treatments, there may exist subgroups who benefit greater from one treatment than the other. This problem is relevant to the field of marketing, where treatments may correspond to different ways of selling a product. It is similarly relevant to the field of public policy, where treatments may corresp…
ShapShift explains shifts in model predictions due to data distribution changes.
problem Prediction shifts caused by changes in input distribution.
method Subgroup Conditional Shapley Values applied to decision trees and ensembles.
result Simple, faithful, and near-complete explanations of prediction shifts across model classes.
Method provides statistical guarantees for identifying subgroups in ML studies.
problem Bias and noise in estimating conditional average treatment effects (CATE).
method Develops uniform confidence bands (GATES) for estimating group average treatment effects (GATEs).
result Identifies subgroups with statistical guarantees, regardless of effect size.
ADHAM provides interpretable survival analysis for healthcare.
problem Limited interpretability in deep learning survival models.
method Additive Deep Hazard Analysis Mixtures (ADHAM) with latent subgroup structure.
result ADHAM offers interpretable insights into exposure-outcome associations.
New feature selection methods improve uplift modeling accuracy.
problem Overfitting and poor interpretability in feature selection for uplift models.
method Explicitly designed feature selection methods inspired by statistics and information theory.
result Proposed methods outperform traditional feature selection methods in uplift modeling.
DCEM algorithm reduces bias in machine learning models trained on selective labels.
problem Bias in machine learning models trained on selective labels.
method Disparate Censorship Expectation-Maximization (DCEM) algorithm.
result DCEM improves bias mitigation without sacrificing discriminative performance.
New method resolves causal heterogeneity by defining a resolution profile.
problem Causal subgroup analyses often oversimplify heterogeneity into a small number of groups.
method Introduces a resolution profile as a functional of the causal feature law, using Bayesian-bootstrap inference.
result Shows that the resolution profile is a continuous path with discontinuities at knots, providing integer-valued subgroup numbers.
New holistic approach measures sample-level adversarial vulnerability for trustworthy systems.
problem Inherent bias in adversarial attacks across subgroups.
method Combining high-frequency feature reliance and sample-distance to decision boundary.
result Holistic approach improves adversarial vulnerability estimation and system trustworthiness.
PCS-UQ framework improves uncertainty quantification for machine learning models.
problem Ensuring trustworthy uncertainty quantification for machine learning models in high-stakes domains.
method PCS-UQ framework based on Predictability, Computability, and Stability principles, integrating prediction-checking, bootstrap samples, and multiplicative calibration.
result PCS-UQ maintains target coverage while outperforming or matching conformal methods in interval width and subgroup coverage.
Estimates statistical power for cluster analysis in biomedical research.
problem Lack of established methods to compute a priori statistical power for cluster analysis.
method Simulation studies varying subgroup size, number, separation, and covariance structure.
result Sufficient statistical power achieved with small samples (N=20-30) for large effect sizes.
In this paper, we propose a new deep feature selection method based on deep architecture. Our method uses stacked auto-encoders for feature representation in higher-level abstraction. We developed and applied a novel feature learning approach to a specific precision medicine problem, which focuses on assessing and prio…
A/B testing improves marketing decisions by selecting effective stratification variables.
problem Improving the sensitivity of A/B testing through stratified sampling.
method Designing an algorithm to select a subset of stratification variables for variance reduction.
result The subset selection method outperforms other variance reduction techniques in A/B testing.
CRL approach improves understanding of heterogeneous treatment effects in complex diseases.
problem Estimating heterogeneous treatment effects in complex diseases.
method Causal rule learning (CRL) workflow consisting of rule discovery, selection, and analysis.
result CRL outperforms other methods in providing interpretable estimates of HTE.
A statistical toolbox for analyzing model performance in medical imaging.
problem Analyzing model performance by patient and recording properties, especially in medical imaging.
method Selection of appropriate performance metrics, correction of multiple comparisons, and finding interesting subgroups.
result Enables rigorous assessment of model performance for potential subgroup disparities.
Identifying homogeneous subgroups of variables can be challenging in high dimensional data analysis with highly correlated predictors. We propose a new method called Hexagonal Operator for Regression with Shrinkage and Equality Selection, HORSES for short, that simultaneously selects positively correlated variables and…
A homogeneous Riemannian manifold (M=G/K,g) is called a space with homogeneous geodesics or a G-g.o. space if every geodesic γ(t) of M is an orbit of a one-parameter subgroup of G, that is γ(t)=exp(tX)⋅o, for some non zero vector X in the Lie algebra of G. We give an exposition on the subject, …
In this paper, we compute the subgroup distortion of all finitely generated subgroups of all finitely generated 3-manifold groups, and the subgroup distortion in this case can only be linear, quadratic, exponential and double exponential. It turns out that the subgroup distortion of a subgroup of a 3-manifold group is …
Federated online learning for streaming data with privacy and efficiency.
problem Analyzing continuous, heterogeneous data streams in a privacy-preserving manner.
method Personalized models for each data source, subgroup assumption, penalized renewable estimation, proximal gradient descent.
result Effective model for distributed multi-source streaming data analysis with privacy and efficiency.
Regular subgroups of SL3(R) are identified and ruled out.
problem Identifying and characterizing regular subgroups of SL3(R).
method Using Kapovich–Leeb–Porti and Guichard–Wienhard divergent subgroups criteria, and Oh's results.
result Regular subgroups of SL3(R) are precisely lattices in minimal horospherical subgroups.
Study on braid group quotients by congruence subgroups.
problem Understanding the image of congruence subgroups in GL(n,Z).
method Characterization through symplectic congruence subgroups.
result Open problem solved: image of congruence subgroups in GL(n,Z).
The paper explores geometric finiteness in mapping class groups and constructs new examples of these subgroups.
problem Understanding geometric finiteness in mapping class groups and constructing new examples.
method Examined several constructions of subgroups and determined conditions for geometric finiteness.
result Provides new examples of parabolically geometrically finite and reducibly geometrically finite subgroups.
Proposes a semi-supervised K-Means algorithm for better feature selection.
problem Data clustering with unknown feature quality and limited labelled data.
method Combines unsupervised sparse clustering and semi-supervised learning with labelled data.
result The algorithm identifies informative features and maintains high performance.
Proposes a new method for finding non-redundant, standout subgroups in numeric datasets.
problem Mining large numbers of redundant subgroups in numeric datasets.
method Dispersion-aware problem formulation based on MDL principle for subgroup set discovery.
result Empirically demonstrates SSD++ returns outstanding subgroup lists.
Proves Congruence Subgroup Property for two types of groups.
problem Proving Congruence Subgroup Property for specific groups.
method Elementary proof of Johnson filtration and geometric subsurface inclusions.
result Proves Congruence Subgroup Property for nilpotent quotients and subsurface subgroups.
New method constructs non-quasiconvex subgroups in hyperbolic groups.
problem Creating non-quasiconvex subgroups in hyperbolic groups.
method Using Stallings-like techniques on right-angled Coxeter groups (RACGs).
result Explicit examples of non-quasiconvex subgroups constructed.
Characterizes knotted subgroups of Lie groups and provides examples.
problem Defining and understanding knotted subgroups of Lie groups.
method Geometric equivalence, one-parameter subgroups, infinitesimal elements, canonical forms, spectrum analysis.
result Completely classified knotted subgroups of SL(2,R) and SL(3,R).
Study subgroups of pro-p PD^3 groups, finding specific conditions.
problem Characterize subgroups of pro-p PD^3 groups. method Analyzes properties of subnormal and finitely presented subgroups.
result Conditions on subgroups of pro-p PD^3 groups. No hyperbolic group can have an infinite chain of free subgroups of fixed rank.
problem Infinite ascending chains of free subgroups in hyperbolic groups.
method Proof by contradiction and properties of hyperbolic groups.
result Hyperbolic groups do not contain strictly ascending chains of free quasiconvex subgroups of constant rank.
Robust subgroup discovery finds non-redundant, statistically significant subgroups.
problem Finding interpretable, robust subgroups from data.
method Formulated subgroup lists for univariate and multivariate targets, used MDL principle and greedy heuristic SSD++.
result SSD++ outperforms previous methods in quality and size of subgroup lists.
Sparse GFA identifies disease factors in FTD subgroups.
problem Heterogeneity in neurological disorders hinders understanding and treatment.
method Sparse Group Factor Analysis (GFA) with regularised horseshoe priors.
result Identified latent disease factors differentially expressed in FTD subgroups.
Let N be at least 4. We prove that every injective homomorphism from the Torelli subgroup into Out(FN) differs from the inclusion by a conjugation in Out(FN). This applies more generally to the following subgroups: every finite-index subgroup of Out(FN) (recovering a theorem of Farb and Handel); every subgro…
New lattices in higher dimensions have dense surface subgroups.
problem Finding dense subgroups in higher-dimensional arithmetic lattices.
method Exhibited nonuniform arithmetic lattices in SO(n,1).
result Contain Zariski-dense surface subgroups.