Chiseling finds valid subgroups interactively, improving on existing methods.
problem Finding valid subgroups with inferential guarantees in regression and causal inference.
method Interactive subgroup refinement with inferential validity guarantees.
result Chiseling identifies better subgroups than existing methods with inferential guarantees.
Posterior conformal prediction improves prediction interval validity for subgroups.
problem Marginal and conditional prediction interval validity for subgroups.
method Modeling conditional nonconformity score distribution as a mixture of cluster distributions.
result PCP produces tighter prediction intervals, especially for well-represented clusters.
We analyse an issue when comparing survival curves between two subgroups. We show that there is a direct relationship between estimates of subgroups' survival at a time point and positive and negative predictive values in the binary classification settings. Our findings present a case where current methods of comparing…
New method combines randomization tests and flexible models for valid inference without splitting data.
problem Valid inference in randomized panel experiments with complex effect heterogeneity.
method Model-assisted randomization tests that estimate unsigned CATE from residualized outcomes.
result CATE-assisted tests control Type I error and achieve higher power than alternatives.
AI-assisted interviews allow respondents to describe experiences naturally, but mapping those accounts into structured survey variables is fallible.
problem Mapping AI-assisted interview responses into structured survey variables is fallible.
method Adaptive Matrix Validation (AMV) is proposed, which involves mapping responses into tabular data and using a small set of structured questions for statistical adjustment.
result The estimator calibrates mapped values using validation answers from other respondents and corrects remaining error with validation answers observed for the target respondent.
Algorithm identifies interpretable subgroups with elevated treatment effects.
problem Estimating high-dimensional, uninterpretable CATE results.
method Rule sets summarizing CATE estimates, optimizing subgroup size and effect size.
result Frontier of Pareto optimal rule sets for subgroup identification.
Novel strategy benchmarks observational studies against randomized trials.
problem Benchmarking observational studies for treatment effect bias.
method Statistical test for null hypothesis of treatment effect difference.
result Valid lower bound on maximum bias strength for any subgroup.
The paper introduces moment multicalibration for estimating uncertainty across subgroups.
problem Ensuring fairness and accurate uncertainty estimation in predictions across different subgroups.
method Develops a method for multicalibration of higher moments, enabling point predictions and interval estimation.
result Moment multicalibration allows for valid prediction intervals that are fair across various subgroups.
WHOMP optimizes randomized controlled trials by minimizing subgroup bias.
problem Minimizing subgroup bias in randomized controlled trials.
method Wasserstein Homogeneity Partition (WHOMP) method.
result WHOMP optimally minimizes type I and type II errors in trials.
CAPITAL algorithm identifies optimal patient subgroups for better treatment.
problem Identify maximum number of patients benefiting from better treatment.
method Constrained Policy Tree Search (CAPITAL) algorithm to find optimal subgroup selection rule (SSR).
result Maximizes the number of patients with enhanced treatment effects.
New method detects strong calibration in ML models, even for small poorly calibrated subgroups.
problem Auditing machine learning models for strong calibration is difficult, especially for small poorly calibrated subgroups.
method Reorder observations by expected residuals and use changepoint detection for score-based cumulative sum (CUSUM) test.
result The proposed adaptive CUSUM test consistently achieved higher power and more than doubled power in auditing mortality risk prediction models.
The paper introduces a model to measure ASR fairness, addressing key issues.
problem Measuring fairness in ASR systems for different subgroups.
method Mixed-effects Poisson regression to control nuisance factors and handle unobserved heterogeneity.
result The method effectively addresses WER gaps among subgroups and is flexible for practical analyses.
The paper proves an ascending chain condition for subgroups in hyperbolic and graph 3-manifolds.
problem Proving an ascending chain condition for subgroups in specific types of 3-manifolds.
method Uses profinite techniques and geometric proofs for hyperbolic and graph manifolds.
result Established the ascending chain condition for free subgroups of constant rank in closed hyperbolic and graph 3-manifolds.
Personalized medicine aims at identifying best treatments for a patient with given characteristics. It has been shown in the literature that these methods can lead to great improvements in medicine compared to traditional methods prescribing the same treatment to all patients. Subgroup identification is a branch of per…
This paper presents an original approach for jointly fitting survival times and classifying samples into subgroups. The Coxlogit model is a generalized linear model with a common set of selected features for both tasks. Survival times and class labels are here assumed to be conditioned by a common risk score which depe…
We apply a study of orders in quaternion algebras, to the differential geometry of Riemann surfaces. The least length of a closed geodesic on a hyperbolic surface is called its systole, and denoted syspi_1. P. Buser and P. Sarnak constructed Riemann surfaces X whose systole behaves logarithmically in the genus g(X). Th…
M-learner estimates treatment effects in mediation models with subgroup identification.
problem Estimating heterogeneous treatment effects in mediation models.
method Four-step procedure: compute conditional effects, construct distance matrix, apply tSNE and K-means clustering, refine clusters.
result Validates robustness and effectiveness in real-world dataset.
Method provides statistical guarantees for identifying subgroups in ML studies.
problem Bias and noise in estimating conditional average treatment effects (CATE).
method Develops uniform confidence bands (GATES) for estimating group average treatment effects (GATEs).
result Identifies subgroups with statistical guarantees, regardless of effect size.
LCMQR improves prediction intervals by adapting to local heteroscedasticity.
problem Efficient and adaptive prediction intervals for local heteroscedasticity.
method LCMQR combines multi-quantile information with kernel-based localization.
result LCMQR constructs tighter intervals than prior methods, especially in heterogeneous environments.
New method resolves causal heterogeneity by defining a resolution profile.
problem Causal subgroup analyses often oversimplify heterogeneity into a small number of groups.
method Introduces a resolution profile as a functional of the causal feature law, using Bayesian-bootstrap inference.
result Shows that the resolution profile is a continuous path with discontinuities at knots, providing integer-valued subgroup numbers.
Simulation study evaluates causal ML models under confounding violations.
problem Assessing conditional exchangeability in causal machine learning models.
method Simulation study with varying confounding, sample size, and NCO structures.
result Causal ML models fail to recover true treatment effect heterogeneity under violations of conditional exchangeability.
Proposes a tool to contrast global vs personalized models in clinical prediction.
problem Balancing global vs personalized models in clinical prediction.
method Localized regression approach using autoencoder for dimension reduction.
result Identification of patient subgroups where global models fall short.
The study uses transfer learning to compare surgical outcomes across racial/ethnic subgroups.
problem Difficulty in comparing surgical outcomes due to racial/ethnic and geographic differences.
method Causal inference framework and transfer learning to incorporate data from multiple populations.
result Racial and ethnic differences in surgical outcomes are found, with non-Hispanic Black patients experiencing wide variability.
A statistical toolbox for analyzing model performance in medical imaging.
problem Analyzing model performance by patient and recording properties, especially in medical imaging.
method Selection of appropriate performance metrics, correction of multiple comparisons, and finding interesting subgroups.
result Enables rigorous assessment of model performance for potential subgroup disparities.
Given a constant magnetic field on Euclidean space Rp determined by a skew-symmetric (p×p) matrix Θ, and a Zp-invariant probability measure μ on the disorder set Σ which is by hypothesis a Cantor set, where the action is assumed to be minimal, the corresponding Integrated Density…
Using certain Thom spectra appearing in the study of cobordism categories, we show that the odd half of the Miller-Morita-Mumford classes on the mappping class group of a surface with negative Euler characteristic vanish in integral cohomology when restricted to the handlebody subgroup. This is a special case of a more…
The paper breaks down AUC into cluster-level components for better model diagnostics.
problem Global AUC masks weaknesses in specific subpopulations, leading to financial or operational risks.
method Formal decomposition of AUC into intra- and inter-cluster components, comparing with other performance metrics.
result Allows practitioners to evaluate and diagnose model performance within and across clusters.
We explicitly compute the lower algebraic K-theory of the split three-dimensional crystallographic groups; i.e., the groups G that act properly and cocompactly on three-dimensional Euclidean space by isometries, such that the natural map from G to O(3) is a split injection onto its image. There are 73 split three-dimen…
New method reduces variance in subpopulation model performance estimates.
problem High variance in subpopulation performance metrics for small groups.
method Using an evaluation model to form model-based metric (MBM) estimates.
result MBMs produce more accurate and lower variance estimates for small subpopulations.
In this paper, we compute the subgroup distortion of all finitely generated subgroups of all finitely generated 3-manifold groups, and the subgroup distortion in this case can only be linear, quadratic, exponential and double exponential. It turns out that the subgroup distortion of a subgroup of a 3-manifold group is …
Machine learning models predict depression risk based on various factors.
problem Identifying individuals at greatest risk for depression.
method Random Effects/Expectation Maximization (RE-EM) trees and Mixed Effects Random Forest (MERF) algorithms.
result Machine learning models accurately predict depression severity and identify key predictors.
Regular subgroups of SL3(R) are identified and ruled out.
problem Identifying and characterizing regular subgroups of SL3(R).
method Using Kapovich–Leeb–Porti and Guichard–Wienhard divergent subgroups criteria, and Oh's results.
result Regular subgroups of SL3(R) are precisely lattices in minimal horospherical subgroups.
Study on braid group quotients by congruence subgroups.
problem Understanding the image of congruence subgroups in GL(n,Z).
method Characterization through symplectic congruence subgroups.
result Open problem solved: image of congruence subgroups in GL(n,Z).
The paper explores geometric finiteness in mapping class groups and constructs new examples of these subgroups.
problem Understanding geometric finiteness in mapping class groups and constructing new examples.
method Examined several constructions of subgroups and determined conditions for geometric finiteness.
result Provides new examples of parabolically geometrically finite and reducibly geometrically finite subgroups.
Proposes a new method for finding non-redundant, standout subgroups in numeric datasets.
problem Mining large numbers of redundant subgroups in numeric datasets.
method Dispersion-aware problem formulation based on MDL principle for subgroup set discovery.
result Empirically demonstrates SSD++ returns outstanding subgroup lists.
Proves Congruence Subgroup Property for two types of groups.
problem Proving Congruence Subgroup Property for specific groups.
method Elementary proof of Johnson filtration and geometric subsurface inclusions.
result Proves Congruence Subgroup Property for nilpotent quotients and subsurface subgroups.
New method constructs non-quasiconvex subgroups in hyperbolic groups.
problem Creating non-quasiconvex subgroups in hyperbolic groups.
method Using Stallings-like techniques on right-angled Coxeter groups (RACGs).
result Explicit examples of non-quasiconvex subgroups constructed.
Enhanced survival trees improve computational efficiency and inference.
problem Censored failure time data and variable selection bias.
method Improved splitting procedure, intersected validation, fused regularization, and bootstrap-based bias correction.
result Valid confidence intervals for median survival times.
Characterizes knotted subgroups of Lie groups and provides examples.
problem Defining and understanding knotted subgroups of Lie groups.
method Geometric equivalence, one-parameter subgroups, infinitesimal elements, canonical forms, spectrum analysis.
result Completely classified knotted subgroups of SL(2,R) and SL(3,R).
Study subgroups of pro-p PD^3 groups, finding specific conditions.
problem Characterize subgroups of pro-p PD^3 groups. method Analyzes properties of subnormal and finitely presented subgroups.
result Conditions on subgroups of pro-p PD^3 groups. New holistic approach measures sample-level adversarial vulnerability for trustworthy systems.
problem Inherent bias in adversarial attacks across subgroups.
method Combining high-frequency feature reliance and sample-distance to decision boundary.
result Holistic approach improves adversarial vulnerability estimation and system trustworthiness.
No hyperbolic group can have an infinite chain of free subgroups of fixed rank.
problem Infinite ascending chains of free subgroups in hyperbolic groups.
method Proof by contradiction and properties of hyperbolic groups.
result Hyperbolic groups do not contain strictly ascending chains of free quasiconvex subgroups of constant rank.
Robust subgroup discovery finds non-redundant, statistically significant subgroups.
problem Finding interpretable, robust subgroups from data.
method Formulated subgroup lists for univariate and multivariate targets, used MDL principle and greedy heuristic SSD++.
result SSD++ outperforms previous methods in quality and size of subgroup lists.
Sparse GFA identifies disease factors in FTD subgroups.
problem Heterogeneity in neurological disorders hinders understanding and treatment.
method Sparse Group Factor Analysis (GFA) with regularised horseshoe priors.
result Identified latent disease factors differentially expressed in FTD subgroups.
Let N be at least 4. We prove that every injective homomorphism from the Torelli subgroup into Out(FN) differs from the inclusion by a conjugation in Out(FN). This applies more generally to the following subgroups: every finite-index subgroup of Out(FN) (recovering a theorem of Farb and Handel); every subgro…
New lattices in higher dimensions have dense surface subgroups.
problem Finding dense subgroups in higher-dimensional arithmetic lattices.
method Exhibited nonuniform arithmetic lattices in SO(n,1).
result Contain Zariski-dense surface subgroups.
For a finitely generated group, there are two recent generalizations of the notion of a quasiconvex subgroup of a word-hyperbolic group, namely a stable subgroup and a Morse or strongly quasiconvex subgroup. Durham and Taylor defined stability and proved stability is equivalent to convex cocompactness in mapping class …
A new algorithm COVA-FC improves subgroup-fair clustering efficiency.
problem Challenges in making cluster assignments independent of sensitive attributes in subgroups.
method Defining a subgroup-fairness gap, deriving a covariance-based surrogate, and introducing a continuous relaxation for efficient optimization.
result COVA-FC achieves competitive cost-fairness trade-offs and improves computational efficiency.