A new method identifies sub-populations in unlabelled heterogeneous data by accounting for co-features.
problem Estimating sub-populations in unlabelled heterogeneous data with co-features.
method Mixture of Conditional Gaussian Graphical Models (CGGM) with penalized EM algorithm.
result The method successfully identifies sub-populations disrupted by co-features.
Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.
problem Underrepresentation of minorities in precision medicine research leads to underperforming risk prediction models.
method Two-way federated transfer learning strategy integrating diverse populations and healthcare institutions.
result Improves risk prediction models for underrepresented populations, reducing performance gaps.
Machine learning approaches have been effective in predicting adverse outcomes in different clinical settings. These models are often developed and evaluated on datasets with heterogeneous patient populations. However, good predictive performance on the aggregate population does not imply good performance for specific …
New method stops experiments early for harm in diverse groups.
problem Early stopping of experiments for harmful treatment effects in diverse populations.
method Causal machine learning approach (CLASH) for early stopping.
result CLASH effectively stops experiments early for harmful treatment effects in diverse groups.
Develops robust learning methods for datasets with sub-populations.
problem Robust performance and generalization to unseen testing populations in datasets with sub-populations.
method Min-max-regret (MMR) formulation for distribution-free robust hierarchical model.
result Empirical MMR enjoys regret guarantees on training and unseen testing populations.
A method for inferring motility models and heterogeneity from particle trajectories.
problem Understanding motility patterns from discrete trajectory data of biological agents.
method Maximum likelihood approach for second-order Langevin models with population heterogeneity.
result The proposed method outperforms alternative approaches for short trajectories.
Estimates personalized policies robust to shifts in target populations.
problem Estimating policies that perform well in diverse target populations.
method Develops methods for estimating robust policies considering shifts in outcomes and characteristics.
result Welfare-maximizing policies are robust to certain shifts in potential outcomes.
Kernel measures similarity of nonlinear causal structures in heterogeneous populations.
problem Learning causal structure in populations with diverse underlying structures.
method Distance covariance-based kernel for measuring similarity of causal structures.
result Kernel enables clustering of homogeneous subpopulations for causal structure learning.
LEARNER improves low-rank matrix estimation using source population data.
problem Improving low-rank matrix estimation in target populations with diverse data sources.
method LEARNER uses similarity in latent spaces between source and target populations to enhance estimation.
result LEARNER often outperforms benchmark methods, especially with higher signal-to-noise ratios in the source population.
Method improves robustness and generalizability of CATE estimation.
problem Lack of external validity in site-specific models for diverse populations.
method Minimax-regret framework with robust optimization.
result Interpretable closed-form solution for generalizable CATE model.
Estimates population mean from user-level data with privacy, accounting for heterogeneity.
problem Heterogeneous user data with varying numbers of data points and distributions.
method Simple model of heterogeneous user data, differential privacy mechanism for estimation.
result Asymptotic optimality of the proposed estimator and general lower bounds on error.
We introduce a general framework for estimation of inverse covariance, or precision, matrices from heterogeneous populations. The proposed framework uses a Laplacian shrinkage penalty to encourage similarity among estimates from disparate, but related, subpopulations, while allowing for differences among matrices. We p…
This paper addresses external validity bias in causal inference.
problem Estimating causal effects in a target population.
method Synthesis of approaches for generalizability and transportability, including tests for heterogeneity of treatment effects and differences between study and target populations.
result Framework for addressing external validity bias in causal inference.
New method estimates treatment effects across different populations.
problem Estimating treatment effects across populations with changing distributions.
method SBRL-HAP framework combining balancing and independence regularizers with hierarchical attention.
result Significant improvement in HTE estimation across out-of-distribution populations.
Paper proposes robust method to detect risk heterogeneity across ethnic groups.
problem Detecting risk heterogeneity across ethnic groups in ICU studies.
method Proposes a robust framework using Neyman orthogonality for inference.
result Demonstrates improved inferential stability and reduced bias compared to standard methods.
New method resolves causal heterogeneity by defining a resolution profile.
problem Causal subgroup analyses often oversimplify heterogeneity into a small number of groups.
method Introduces a resolution profile as a functional of the causal feature law, using Bayesian-bootstrap inference.
result Shows that the resolution profile is a continuous path with discontinuities at knots, providing integer-valued subgroup numbers.
New method identifies subgroups in censored data.
problem Identifying meaningful patterns in heterogeneous populations.
method Combining inverse probability weighting, M-estimation, and concave pairwise fusion penalization.
result Robust approach for censored data under heterogeneous AFT models.
Proposes a new method to handle data heterogeneity in causal inference.
problem Challenges of collaborating between different data centers due to heterogeneity.
method Collaborative inverse propensity score weighting estimator to adjust for distribution shift.
result Significant improvements over traditional meta-analysis methods when dealing with increased heterogeneity.
Performing inference on data obtained through observational studies is becoming extremely relevant due to the widespread availability of data in fields such as healthcare, education, retail, etc. Furthermore, this data is accrued from multiple homogeneous subgroups of a heterogeneous population, and hence, generalizing…
Method learns cell interaction rules from individual trajectories.
problem Inferring interaction rules from heterogeneous cellular data.
method WSINDy for second order IPSs, learning individual cell models.
result Efficiently identifies different species and best-fit models for each.
A new model for heterogeneous populations optimizes consumption and investment over short horizons.
problem Optimizing consumption and investment in economies with a heterogeneous population over short time periods.
method Continuous-time general equilibrium framework with Brownian flow on a type space, solving vanishing-horizon problems under relative-income criteria.
result Existence and characterization of short-horizon Duesenberry equilibrium, with sharp asset-pricing implications.
Develops a method to estimate personalized treatment regimes from summary statistics.
problem Estimating optimal treatment regimes for a target population when individual-level data is unavailable.
method A weighting framework that tailors a treatment regime for the target population using summary statistics.
result Consistent and asymptotically normal estimator for optimal treatment regimes.
The paper proposes a mixture model with segmentation for heterogeneous functional data.
problem Heterogeneity in time and population for functional data.
method Mixture model with segmentation of time, maximum likelihood estimator, EM algorithm with dynamic programming.
result The method is consistent and identifiable, and illustrated on simulated and real datasets.
CRE discovers interpretable subgroups with heterogeneous treatment effects.
problem Identifying subgroups with notable treatment effect heterogeneity.
method Causal Rule Ensemble (CRE) using an ensemble-of-trees approach.
result CRE offers interpretable decision rules and high stability in subgroup discovery.
CLSB models system dynamics from cross-sectional data with population-level regularization.
problem Challenges in modeling system dynamics from limited cross-sectional samples and heterogeneous individual behaviors.
method Introduces CLSB framework for learning dynamics, regularized for population-level temporal variations.
result Empirically superior in single-cell sequencing data analyses, e.g., simulating cell development and drug response.
Study explains mortgage burnout using Cox hazard models.
problem Understanding burnout in mortgage pools.
method Modeling mortgage prepayment using Cox hazard processes.
result Observed pool hazard is a survival-weighted mean of individual hazards with a selection term.
In this work, we consider hypothesis testing and anomaly detection on datasets where each observation is a weighted network. Examples of such data include brain connectivity networks from fMRI flow data, or word co-occurrence counts for populations of individuals. Current approaches to hypothesis testing for weighted n…
A new method reduces preference distortion in LLM alignment.
problem Vulnerability of traditional LLM alignment methods to human preference heterogeneity.
method Sign Estimator: A simple, provably consistent, and efficient estimator using binary classification loss.
result Substantially reduces preference distortion over a panel of simulated personas.
Bayesian approach learns nonparametric mixture components from heterogeneous data.
problem Realistic modeling of heterogeneous data populations with nonparametric mixture components.
method Bayesian nonparametric modeling using Dirichlet process mixture priors.
result Posterior contraction rates for component densities are nearly polynomial, improving over deconvolution methods.
Study optimizes data collection from biased, costly sources to minimize risk.
problem Estimating population means and group-conditional means from multiple sources with varying costs and biases.
method Develops a sampling plan that maximizes effective sample size, paired with a post-stratification estimator.
result Achieves budgeted minimax optimal risk for estimating population means and group-conditional means.
Bayesian Federated Inference combines local data analyses to estimate regression models.
problem Estimating accurate parameters with limited data from different centers.
method Bayesian Federated Inference (BFI) for pooling local data analyses.
result Excellent performance of BFI methodology shown in real-life examples.
This article analyzes the problem of estimating the time until an event occurs, also known as survival modeling. We observe through substantial experiments on large real-world datasets and use-cases that populations are largely heterogeneous. Sub-populations have different mean and variance in their survival rates requ…
Proposes a generalized causal tree for handling multiple treatments in uplift modeling.
problem Handling multiple treatments in uplift modeling.
method Generalizes causal tree algorithm to handle multiple discrete and continuous-valued treatments.
result Demonstrates improved performance over existing methods in experiments and real data examples.
In this paper we model the problem of learning preferences of a population as an active learning problem. We propose an algorithm can adaptively choose pairs of items to show to users coming from a heterogeneous population, and use the obtained reward to decide which pair of items to show next. We provide computational…
Conventional survival analysis approaches estimate risk scores or individualized time-to-event distributions conditioned on covariates. In practice, there is often great population-level phenotypic heterogeneity, resulting from (unknown) subpopulations with diverse risk profiles or survival distributions. As a result, …
Study optimal investment in large populations of competitive, heterogeneous agents.
problem Maximizing utility in a large, interacting agent system with relative performance concerns.
method Analyzes stochastic utility maximization game in finite and infinite agent settings, using graphon models and backward stochastic differential equations.
result Convergence of Nash equilibria and optimal utilities from finite to infinite agent models under specific conditions.
Study shows income inequality increases with city size, affecting only the wealthiest deciles.
problem Understanding income inequality in urban areas.
method Urban scaling analysis of total income scaling in population percentiles.
result Income in the poorest decile does not increase with city size, while the wealthiest deciles show superlinear scaling.
Clustering and community detection with multiple graphs have typically focused on aligned graphs, where there is a mapping between nodes across the graphs (e.g., multi-view, multi-layer, temporal graphs). However, there are numerous application areas with multiple graphs that are only partially aligned, or even unalign…
Novel framework identifies pump-specific deterioration rates using Bayesian hierarchical hazard modeling and causal discovery.
problem Challenges in asset management due to heterogeneous deterioration rates in pump equipment.
method Bayesian hierarchical hazard modeling with causal discovery, GPU-accelerated No-U-Turn Sampling (NUTS), and DirectLiNGAM.
result Identified striking heterogeneity in deterioration rates, with negative effects 400 times larger than positive effects.
Model infers diffusion networks from heterogeneous cascade data.
problem Understanding and predicting diffusion processes in interconnected populations.
method Double mixture directed graph model with layer-specific constraints.
result Convex formulation allows for statistical and computational guarantees.
The paper proposes methods to identify and sample from mixtures of Mallows models for top-k rankings.
problem Identifying and sampling from mixtures of Mallows models for top-k rankings in a heterogeneous population.
method Efficient sampling algorithms and identifiability proofs for both components of the mixture.
result The identifiability and learnability of the Mallows components' parameters in the mixture.
Response time improves alignment with diverse human preferences.
problem Standard aggregation of feedback ignores heterogeneity and anonymity.
method Augmenting feedback with response time data and modeling decisions with DDM.
result Estimator of heterogeneous preferences converges to true average preference.
New model explains price dynamics of Bitcoin with psychological factors.
problem Understanding price variations in cryptocurrency markets with psychological factors.
method Extended agent-based model with heterogeneous psychological parameters.
result Model shows diverse dynamics based on psychological correlation.
Proposes a convex model for mixed logit to handle individual heterogeneity.
problem Non-convex optimization in mixed logit models for individual heterogeneity.
method Sparse and low-rank decomposition for convex formulation.
result Convex formulation avoids simulation-based approximation and unstable model interpretation.
Paper proposes a method to optimize policies for diverse individuals using heterogeneous data.
problem Learning optimal policies for a heterogeneous population from pre-collected data.
method Individualized offline policy optimization framework for heterogeneous MDPs.
result The proposed P4L algorithm achieves a fast rate of average regret.
Bayesian meta-learning improves health prediction models across similar diseases.
problem Inter- and intra-task variability in healthcare predictions due to disease heterogeneity and patient differences.
method Bayesian meta-learning approach that models task similarity to mitigate negative transfer and improve generalizability.
result Significant generalizability improvements in stroke prediction tasks using electronic health record data.
ScoreFusion fuses multiple diffusion models to enhance generative modeling of a target population.
problem Enhancing generative modeling of a target population with limited data.
method ScoreFusion uses KL barycenters of auxiliary populations and recasts the learning problem as score matching in denoising diffusion.
result ScoreFusion achieves a dimension-free sample complexity bound in total variation distance.
Study evaluates MRIQC pipeline's generalization on large multi-center datasets.
problem Generalizing MRI QC methods to new, unseen data.
method Evaluated MRIQC preprocessing steps on ABIDE and CATI datasets.
result Model without preprocessing yields best results on unseen data.