Proposes Population Difference Criterion for visually observed subpopulation differences.
problem Statistical significance of visually observed subpopulation differences in high-dimensional and high-signal contexts.
method Balanced permutation approach and bootstrap confidence interval for quantifying uncertainty.
result Balanced permutation approach is more powerful in high-signal contexts.
Proposes a stability evaluation criterion for learning models using distributional perturbations.
problem Ensuring reliable deployment of learning models in out-of-sample environments.
method Uses optimal transport discrepancy with moment constraints to quantify minimal perturbation required for model deterioration.
result Validates the practical utility of the stability evaluation criterion across various real-world applications.
In this work we use an inelastic scattering process of particles to propose a model able to reproduce the salient features of the wealth distribution in an economy by including taxes to each trading process and redistributing that collected among the population according to a given criterion. Additionally, we show that…
We present a new approach for mitigating unfairness in learned classifiers. In particular, we focus on binary classification tasks over individuals from two populations, where, as our criterion for fairness, we wish to achieve similar false positive rates in both populations, and similar false negative rates in both po…
Stable random variables are motivated by the central limit theorem for densities with (potentially) unbounded variance and can be thought of as natural generalizations of the Gaussian distribution to skewed and heavy-tailed phenomenon. In this paper, we introduce stable graphical (SG) models, a class of multivariate st…
New criterion assesses cluster separability for validation.
problem Validating cluster analysis results and determining the number of clusters.
method Distinguishability criterion, combined loss function-based framework.
result Validated cluster configurations and determined the number of clusters.
The paper offers simple, near-optimal algorithms for multi-group learning.
problem Learning predictors within subgroups of a population, addressing fairness and hidden stratification.
method Studies the structure of solutions and provides simple, near-optimal algorithms.
result Simple and near-optimal algorithms for multi-group learning.
fAux tests individual fairness without domain knowledge or out-of-domain samples.
problem Testing for individual fairness in machine learning models.
method Derivative comparison between model predictions and auxiliary model predictions.
result Effectively identifies discrimination on synthetic and real-world datasets.
Investigates optimal pension policies in PAYG systems with forward utility and ageing population.
problem Optimal investment and pension policies in PAYG systems with sustainability and adequacy constraints.
method Non-zero volatility forward CRRA utilities, closed-form optimal policies, detailed numerical analysis.
result Characterization of optimal policies and detailed impact analysis under various scenarios.
The paper proposes a new method for comparing logistic regression models across different populations.
problem Comparing logistic regression models across sub-populations can lead to misleading results.
method Develops a cascading set of equivalence tests for logistic regression models, addressing coding, predictions, and overall accuracy.
result Equivalence testing incentivizes accurate inference and avoids perverse incentives from significance tests.
Stacked regressions improve predictive accuracy by combining estimators.
problem Improve predictive accuracy in regression models.
method Analogous to least-squares, learn combination weights by minimizing regularized empirical risk with nonnegativity constraint.
result The stacked estimator has strictly smaller population risk than the best single estimator, especially when signal-to-noise ratio is small.
Unified framework for estimating indirect effects in observational studies with unmeasured confounding.
problem Challenges in evaluating indirect effects due to unmeasured confounding and unethical exposures.
method Developed a unified identification and estimation framework using proximal causal inference.
result Unified identification and estimation of PIIE and causal effect of an intervening variable in settings with pervasive unmeasured confounding.
New decompositions misattribute differences between populations, even when outcomes are identical.
problem Misattribution of differences between populations using common functional decompositions.
method Extending the Kitagawa-Oaxaca-Blinder decomposition to nonlinear functional decompositions.
result Functional ANOVA and Accumulated Local Effects can misattribute differences even when outcomes are identical in two populations.
Bayesian networks learn sub-population differences from data.
problem Inference from a single network structure can be misleading when data populations are heterogeneous.
method A mixture of Bayesian networks where component probabilities depend on individual characteristics.
result Identifies both network structures and demographic predictors of sub-population membership.
Improved average distance classifier for HDLSS settings with multiple population differences.
problem Poor performance of average distance classifier in HDLSS settings with location and scale differences.
method Proposed transformations to the average distance classifier to handle multiple population differences.
result The proposed classifiers perform well even when populations differ in other aspects than location and scale.
In this work, we systematically investigate mean field games and mean field type control problems with multiple populations using a coupled system of forward-backward stochastic differential equations of McKean-Vlasov type stemming from Pontryagin's stochastic maximum principle. Although the same cost functions as well…
The method of covariate adjustment is often used for estimation of population average treatment effects in observational studies. Graphical rules for determining all valid covariate adjustment sets from an assumed causal graphical model are well known. Restricting attention to causal linear models, a recent article der…
New method combines regional HIV prevention trial data without sharing individual patient info.
problem Regional differences in HIV prevention efficacy, privacy concerns, and data sharing limitations.
method Federated learning approach that combines site-specific estimators via L1-regularization.
result Improved precision in estimating region-specific survival curves.
We investigate the problem of testing whether d random variables, which may or may not be continuous, are jointly (or mutually) independent. Our method builds on ideas of the two variable Hilbert-Schmidt independence criterion (HSIC) but allows for an arbitrary number of variables. We embed the d-dimensional joint …
Analyzes how bias evolves in SGD training across different data sub-populations.
problem Understanding bias formation during machine learning training.
method Analytical description of SGD dynamics in a teacher-student setup with Gaussian-mixture model.
result Different sub-populations influence bias at different timescales, revealing shifting classifier preferences.
SDRF estimates complex survey designs for conditional distributions.
problem Estimating conditional distributions under complex survey designs.
method Survey-calibrated distributional random forest (SDRF) with pseudo-population bootstrap and MMD split criterion.
result Established design consistency and model consistency for survey designs.
While machine learning is rapidly being developed and deployed in health settings such as influenza prediction, there are critical challenges in using data from one environment in another due to variability in features; even within disease labels there can be differences (e.g. "fever" may mean something different repor…
Population attributes are essential in health for understanding who the data represents and precision medicine efforts. Even within disease infection labels, patients can exhibit significant variability; "fever" may mean something different when reported in a doctor's office versus from an online app, precluding direct…
A new and an enriched JPEG algorithm is provided for identifying redundancies in a sequence of irregular noisy data points which also accommodates a reference-free criterion function. Our main contribution is by formulating analytically (instead of approximating) the inverse of the transpose of JPEGwavelet transform wi…
Optimizes pension mix of PAYGO, EET, and individual savings.
problem Balancing PAYGO, EET, and individual savings in funded pension schemes.
method Solves a Nash equilibrium between pension participants and government, considering age-dependent preferences and optimal asset allocation.
result Identifies critical ages and optimal contribution rates for maximizing overall utility.
CEDA analyzes large categorical datasets using tree geometry and binary codes.
problem Analyzing large categorical datasets with extreme-K samples. method CEDA uses tree geometry and binary codes to analyze categorical data.
result CEDA discovers patterns and evaluates their reliability in large categorical datasets.
New methods resolve conflicting treatment effect estimates in health tech assessments.
problem Conflicting conclusions from different sponsors analyzing the same data.
method Arbitrated indirect treatment comparisons (ArMAIC) targeting a common target population.
result Estimates treatment effects in a common target population, resolving the MAIC paradox.
PSC classifier improves HDLSS classification on class-imbalanced data.
problem Classification on high-dimension low-sample-size data with class imbalance.
method Population Structure-learned Classifier (PSC) maximizing inter-class and intra-class scatter matrices.
result PSC outperforms state-of-the-art methods on IHDLSS.
Estimates personalized policies robust to shifts in target populations.
problem Estimating policies that perform well in diverse target populations.
method Develops methods for estimating robust policies considering shifts in outcomes and characteristics.
result Welfare-maximizing policies are robust to certain shifts in potential outcomes.
Develops certificates for local population-risk increments using cross-fitted ridge calibration.
problem Certifying local population-risk increments in statistical models.
method Cross-fitted ridge calibration for linear feature classes, separating Taylor fluctuations and remainders.
result Certifies measurable updates from the same sample with penalties dependent on empirical geometry.
The paper automates policy learning for nonlinear welfare criteria using machine learning and debiasing techniques.
problem Learning optimal policies from observational data with nonlinear welfare criteria.
method Modeling a nonlinear welfare criterion with a utility function, estimating propensity scores with machine learning, and using sieve approximations and cross-validation for model selection.
result The proposed policy learning method satisfies oracle inequalities, providing theoretical guarantees on performance.
A clustering method for multivariate populations with similar dependence structures.
problem Grouping populations with similar dependence structures.
method Orthogonal projection coefficients of density copulas estimated from populations.
result Clusters of populations with similar dependence structures.
Unified approach to private statistics from empirical to population data.
problem Divided focus on empirical vs population statistics in private statistics.
method Unified methods for both types of statistics.
result Methods for empirical statistics can be applied to population statistics.
Generalizing empirical findings to new environments, settings, or populations is essential in most scientific explorations. This article treats a particular problem of generalizability, called "transportability", defined as a license to transfer information learned in experimental studies to a different population, on …
Machine learning approaches have been effective in predicting adverse outcomes in different clinical settings. These models are often developed and evaluated on datasets with heterogeneous patient populations. However, good predictive performance on the aggregate population does not imply good performance for specific …
Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.
problem Underrepresentation of minorities in precision medicine research leads to underperforming risk prediction models.
method Two-way federated transfer learning strategy integrating diverse populations and healthcare institutions.
result Improves risk prediction models for underrepresented populations, reducing performance gaps.
We introduce a new criterion to determine the order of an autoregressive model fitted to time series data. It has the benefits of the two well-known model selection techniques, the Akaike information criterion and the Bayesian information criterion. When the data is generated from a finite order autoregression, the Bay…
The perennial problem of "how many clusters?" remains an issue of substantial interest in data mining and machine learning communities, and becomes particularly salient in large data sets such as populational genomic data where the number of clusters needs to be relatively large and open-ended. This problem gets furthe…
In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we develop an inference method that is resilient to sampling biases and is able to …
Kernel measures similarity of nonlinear causal structures in heterogeneous populations.
problem Learning causal structure in populations with diverse underlying structures.
method Distance covariance-based kernel for measuring similarity of causal structures.
result Kernel enables clustering of homogeneous subpopulations for causal structure learning.
New criterion improves predictive evaluation in weighted inference scenarios.
problem Improving predictive evaluation in scenarios with different likelihoods for estimation and evaluation.
method Developed the posterior covariance information criterion (PCIC) to handle weighted likelihood inference.
result PCIC is asymptotically unbiased for quasi-Bayesian generalization error in weighted inference.
Paper develops a two-population model to assess longevity basis risk.
problem Mismatch between hedger's liability and hedging instrument causes longevity basis risk.
method Develops a two-population mortality model using Lee-Carter model and renewal process.
result Proposed model provides significant risk reduction when mortality jumps and sampling risk are considered.
Fiber simplifies RL and population-based methods for distributed training.
problem Challenges in RL and population-based methods, including frequent interaction with simulations and dynamic scaling.
method Introducing Fiber, a scalable distributed computing framework.
result Significantly expands accessibility of large-scale parallel computation.
A new model for heterogeneous populations optimizes consumption and investment over short horizons.
problem Optimizing consumption and investment in economies with a heterogeneous population over short time periods.
method Continuous-time general equilibrium framework with Brownian flow on a type space, solving vanishing-horizon problems under relative-income criteria.
result Existence and characterization of short-horizon Duesenberry equilibrium, with sharp asset-pricing implications.
The Ward error sum of squares hierarchical clustering method has been very widely used since its first description by Ward in a 1963 publication. It has also been generalized in various ways. However there are different interpretations in the literature and there are different implementations of the Ward agglomerative …
We introduce principal differences analysis (PDA) for analyzing differences between high-dimensional distributions. The method operates by finding the projection that maximizes the Wasserstein divergence between the resulting univariate populations. Relying on the Cramer-Wold device, it requires no assumptions about th…
A new stopping criterion for active learning based on deterministic generalization bounds.
problem Determining the optimal stopping point for active learning when data acquisition is costly.
method The proposed stopping criterion is based on the difference in expected generalization errors and hypothesis testing, derived from PAC-Bayesian theory.
result The proposed stopping criterion effectively stops active learning by combining an upper bound with a statistical test.
In systems biomedicine, an experimenter encounters different potential sources of variation in data such as individual samples, multiple experimental conditions, and multi-variable network-level responses. In multiparametric cytometry, which is often used for analyzing patient samples, such issues are critical. While c…