Chiseling finds valid subgroups interactively, improving on existing methods.
problem Finding valid subgroups with inferential guarantees in regression and causal inference.
method Interactive subgroup refinement with inferential validity guarantees.
result Chiseling identifies better subgroups than existing methods with inferential guarantees.
WHOMP optimizes randomized controlled trials by minimizing subgroup bias.
problem Minimizing subgroup bias in randomized controlled trials.
method Wasserstein Homogeneity Partition (WHOMP) method.
result WHOMP optimally minimizes type I and type II errors in trials.
Optimizes subgroup selection in clinical trials.
problem Identifying regions in feature space where a regression function exceeds a threshold.
method Formulates subgroup selection as constrained optimisation, determining minimax optimal rate for regret.
result Determines the minimax optimal rate for regret in sample size and Type I error probability.
The paper proposes a method to find subgroups with significant treatment effects in noisy data.
problem Estimating the causal effects of interventions on noisy outcomes.
method A machine-learning method specifically optimized for finding subgroups with significant effects, designed to maximize the probability of obtaining a statistically significant positive treatment effect.
result The proposed method yields higher power in detecting subgroups affected by the treatment compared to standard tree-based tools.
A/B testing improves marketing decisions by selecting effective stratification variables.
problem Improving the sensitivity of A/B testing through stratified sampling.
method Designing an algorithm to select a subset of stratification variables for variance reduction.
result The subset selection method outperforms other variance reduction techniques in A/B testing.
New framework for interpreting disaggregated fairness evaluations using causal models.
problem Misinterpretation of disaggregated fairness evaluations due to data representativeness and selection bias.
method Causal graphical models to characterize fairness properties and metric stability under different data generating processes.
result Disaggregated evaluations are unreliable without explicit assumptions regarding bias mechanisms.
The paper controls the geometry of surface subgroups in specific Kleinian groups.
problem Understanding the geometry of surface subgroups in specific Kleinian groups.
method Finding surface subgroups that are quasi-conformally conjugate to finite index subgroups of a genus-2 quasi-Fuchsian group.
result The existence of surface subgroups that are K-quasiconformally conjugate to finite index subgroups of a genus-2 quasi-Fuchsian group. Proposes a method to select features for subgroup datasets with systematic missing data.
problem Feature selection for datasets with subgroup structure and systematic missing data.
method Develops a heterogeneous graph neural network to propagate information between feature-subgroup-target variable connections.
result Demonstrates improved feature selection performance and scalability.
This study evaluates subgroup analysis methods for time-to-event outcomes in randomized controlled trials.
problem Identifying subgroups of good responders in non-significant randomized controlled trials.
method Evaluation of several subgroup analysis algorithms for time-to-event outcomes using synthetic and semi-synthetic data.
result Provides a new synthetic and semi-synthetic data generation process and an open-source Python package for benchmarking.
Adaptive Prespecification improves precision in randomized trials.
problem Selecting optimal covariates for precision in randomized trials.
method Adaptive Prespecification using V-fold cross-validation and influence curve-squared loss function.
result Substantial gains in precision, equivalent to 20-43% reductions in sample size for the same power.
We propose a sliding surface for systems on the Lie group SO(3)×R3 . The sliding surface is shown to be a Lie subgroup. The reduced-order dynamics along the sliding subgroup have an almost globally asymptotically stable equilibrium. The sliding surface is used to design a sliding-mode controller for t…
Selective regression allows abstention to improve fairness criteria.
problem Selective regression can exacerbate disparities between subgroups.
method Proposes new fairness criteria and two approaches to mitigate performance disparity.
result Proposed fairness criteria ensures performance improvement for every subgroup with reduced coverage.
CAPITAL algorithm identifies optimal patient subgroups for better treatment.
problem Identify maximum number of patients benefiting from better treatment.
method Constrained Policy Tree Search (CAPITAL) algorithm to find optimal subgroup selection rule (SSR).
result Maximizes the number of patients with enhanced treatment effects.
D3M debiases models by selectively removing problematic examples.
problem Model failures on underrepresented subgroups.
method Isolates and removes specific training examples that cause failures.
result Efficiently trains debiased classifiers with minimal example removal.
New risk measures control subgroup imbalances, improving PAC-Bayesian bounds.
problem Insufficient risk bounds for subgroup imbalances in data.
method Introduce constrained f-entropic risk measures and derive PAC-Bayesian bounds.
result First disintegrated PAC-Bayesian guarantees beyond standard risks.
New methods improve subgroup analysis in trials with limited data.
problem Limited sample sizes in subgroup analyses of randomized controlled trials.
method Two TMLEs that borrow information from non-subgroup participants.
result Improved precision in subgroup-specific treatment effect estimates.
This chapter covers different approaches to policy evaluation for assessing the causal effect of a treatment or intervention on an outcome of interest. As an introduction to causal inference, the discussion starts with the experimental evaluation of a randomized treatment. It then reviews evaluation methods based on se…
This paper presents an original approach for jointly fitting survival times and classifying samples into subgroups. The Coxlogit model is a generalized linear model with a common set of selected features for both tasks. Survival times and class labels are here assumed to be conditioned by a common risk score which depe…
Study controllability of diffeomorphisms of simple polytopes.
problem Controllability of diffeomorphisms of simple polytopes.
method Lie group structure and controllability results for diffeomorphisms of simple polytopes.
result Identity component of the diffeomorphism group is generated by the exponential image.
We define a new condition on relatively hyperbolic Dehn filling which allows us to control the behavior of a relatively quasiconvex subgroups which need not be full. As an application, in combination with a recent result of Cooper and Futer, we provide a new proof of the virtual fibering of non-compact finite-volume hy…
Randomized Controlled Trials (RCTs) are the gold standard for comparing the effectiveness of a new treatment to the current one (the control). Most RCTs allocate the patients to the treatment group and the control group by uniform randomization. We show that this procedure can be highly sub-optimal (in terms of learnin…
Paper proposes a sparse synthetic control method to select important predictors.
problem Choosing and weighting predictors affects synthetic control estimator performance.
method Sparse synthetic control procedure that penalizes predictors, derived in a linear factor model.
result Sparse synthetic control achieves lower bias and better post-treatment performance.
Private variable selection method controls FDR with simulations showing reasonable power.
problem Performing variable selection with privacy constraints.
method Private knockoff filter using Gaussian and Laplace mechanisms.
result Achieves controlled false discovery rate (FDR) in variable selection.
In variable or graph selection problems, finding a right-sized model or controlling the number of false positives is notoriously difficult. Recently, a meta-algorithm called Stability Selection was proposed that can provide reliable finite-sample control of the number of false positives. Its benefits were demonstrated …
Proposes mCS for multivariate selection with FDR control.
problem Selecting high-quality candidates from multivariate datasets.
method Introduces regional monotonicity and multivariate nonconformity scores.
result Significantly improves selection power with FDR control.
FlowSelect uses normalizing flows to control FDR in feature selection.
problem Controlled feature selection with knockoffs often fails to control false discovery rate (FDR).
method FlowSelect uses normalizing flows for accurate feature modeling and a novel MCMC-based p-value calculation to enforce knockoff properties.
result FlowSelect consistently controls FDR and demonstrates greater power compared to competing methods.
Nonparametric IPSS selects features with false discovery control.
problem Feature selection in high-dimensional data with theoretical false discovery control.
method Integrated Path Stability Selection (IPSS) applied to nonparametric feature importance scores.
result IPSS accurately controls false discovery rate and detects more true positives than existing methods.
ACS is an interactive framework for model-free selection with guaranteed error control.
problem Model-free selection with rigorous error control.
method Adaptive conformal selection with human-in-the-loop data exploration and new information incorporation.
result ACS provides concrete selection algorithms for various goals, including model update/selection, diversified selection, and incorporating new data.
Controlled K-theory is used to show that algebraic K-theory of virtually abelian groups is described by an assembly map defined using possibly-infinite hyperelementary subgroups. The Farrell-Jones summand (coming from infinite subgroups) is parameterized by the rational projective space of the group, and a reduced …
New method learns from subgroup feedback in complex systems.
problem Optimizing complex systems with heterogeneous components.
method Decomposed Gaussian Process (GP) regression and optimization algorithm.
result Proved lower variance and improved accuracy in subgroup feedback.
OptCS optimizes model selection after conformal inference, controlling FDR and power loss.
problem Challenges in model selection for conformal inference, especially when limited labeled data and many model choices are available.
method OptCS framework that allows valid statistical testing after flexible data-driven model optimization, using novel multiple testing procedures.
result Valid conformal p-values constructed despite substantial data reuse, maintaining FDR control.
CAP algorithm controls FCR in online selective prediction.
problem Online predictive tasks with temporal multiplicity and FCR control.
method CAP framework with adaptive pick rule and calibration set construction.
result CAP achieves exact selection-conditional coverage guarantee and FCR control.
Proposes a method to learn fair predictors for multiple subgroups with limited data.
problem Fairness and accuracy issues in learning from multiple subgroups with limited data.
method Formulates a bilevel objective to learn subgroup-specific predictors and a fair predictor that is close to all of them.
result The method effectively controls group sufficiency and generalization error, improving fairness and accuracy.
Agents acting in the natural world aim at selecting appropriate actions based on noisy and partial sensory observations. Many behaviors leading to decision mak- ing and action selection in a closed loop setting are naturally phrased within a control theoretic framework. Within the framework of optimal Control Theory, o…
Develops methods to select informative conformal prediction sets with FCR control.
problem Selecting informative prediction sets with FCR control in supervised learning.
method Unified framework for informative conformal prediction sets with FCR control.
result First procedures providing FCR control for informative prediction sets.
Minimizes indecisions in selective classification to control misclassification rates.
problem Controlling misclassification rates in high-risk scenarios.
method Using indecisions to control misclassification rates, even below Bayes optimal.
result Control of misclassification rates to any user-specified level, even below Bayes optimal.
Simulation study evaluates causal ML models under confounding violations.
problem Assessing conditional exchangeability in causal machine learning models.
method Simulation study with varying confounding, sample size, and NCO structures.
result Causal ML models fail to recover true treatment effect heterogeneity under violations of conditional exchangeability.
Algorithm identifies interpretable subgroups with elevated treatment effects.
problem Estimating high-dimensional, uninterpretable CATE results.
method Rule sets summarizing CATE estimates, optimizing subgroup size and effect size.
result Frontier of Pareto optimal rule sets for subgroup identification.
New method reduces memory usage for high-dimensional variable selection.
problem Scalability issues in high-dimensional variable selection, especially in genomics.
method Adaptive sampling of null features to eliminate dummy matrix materialization.
result Reduces memory and runtime by several orders of magnitude while preserving FDR control.
T-Rex selector selects variables fast and controls FDR in high-dimensional data.
problem Variable selection in high-dimensional data with FDR control.
method Fused solutions of early terminated random experiments.
result FDR control at target level with high variable selection power.
Selective inference framework for CART trees to control error rates and coverage.
problem Inference on CART trees does not control Type 1 error rates and coverage.
method Selective inference framework conditioning on tree estimation, efficient algorithms.
result Proposes tests and intervals for CART trees with selective error control.
New method improves reliability of selecting individuals based on predicted treatment effects.
problem Reliability of selecting individuals based on predicted conditional average treatment effects (CATE) is unreliable.
method Denoised Conformal Alignment, combining proxy errors, variance estimation, and Benjamini-Hochberg selection.
result Significantly improved power in selecting individuals while maintaining false discovery rate control.
We introduce the notion of controlled Floyd separation between geodesic rays starting at the identity in a finitely generated group G. Two such geodesic rays are said to be Floyd separated with respect to quasigeodesics if the (Floyd) length of c-quasigeodesics (for fixed but arbitrary c) joining points on the geodesic…
The paper introduces a method to control false splits in tree-based data aggregation.
problem Identifying the correct subgroups to treat as a single entity in tree-based data.
method Introduces the 'false split rate' and proposes a multiple hypothesis testing algorithm for tree-based aggregation.
result The proposed algorithm controls the false split rate, demonstrating its effectiveness on stock volatility and taxi fare data.
Unified framework for variable selection in model-based clustering with missing data.
problem Challenges in identifying relevant variables and handling missing data in model-based clustering.
method Unified framework incorporating a data-driven penalty matrix and a mechanism for missingness modeling.
result Achieves both asymptotic consistency and selection consistency in the presence of missing data.
Proposes a method to identify subgroup structure and estimate covariate effects for multivariate response data.
problem Identifying subgroup structure and estimating covariate effects in multivariate response data.
method Joint heterogeneity and reduced-rank learning framework using rank-constrained pairwise fusion penalization.
result Established the asymptotic properties of the estimators and proposed a predictive information criterion for rank selection.
Paper solves a control problem with robust methods.
problem Monotone mean-variance problems with stochastic coefficients.
method Finding saddle point through BSDEs with unbounded coefficients.
result Optimal control and value match mean-variance problems.
If a class of finitely generated groups Curly(G) is closed under isometric amalgamations along free subgroups, then every G in Curly(G) can be quasi-isometrically embedded in a group Hat(G) in Curly(G) that has no proper subgroups of finite index. Every compact, connected, non-positively curved space X admits an isomet…