New findings show that multi-armed bandit methods can't guarantee correct population selection with finite samples.
problem Selecting the best population or probability distribution from many with unknown means.
method Adapting sequential methods from stochastic multi-armed bandit literature.
result Infinite number of samples are required for correct population selection with finite probability.
Framework improves policy generalizability under biased training data.
problem Learning policies that generalize to a target population from biased training data.
method Characterizes sample selection bias using a selection variable, optimizes minimax value over uncertainty set, derives efficient algorithm.
result Policies generalize to target population, outperform standard methods.
Proposes a method to learn population and subject-specific brain connectivity networks.
problem Quantifying inter-subject variability in brain connectivity networks.
method Mixed Neighborhood Selection to estimate population and subject-specific graphical models.
result Identifies regions or sub-networks with heterogeneity across subjects.
Aims to describe neural network training dynamics using two-time-scale models.
problem Lack of a general mathematical description of neural network training.
method Introduces a theoretical framework based on two-time-scale population dynamics.
result Derives selection-mutation equations and effective fitness for hyperparameters.
New method predicts model performance under selection bias in healthcare.
problem Selection bias limits model generalizability in healthcare.
method Proposes a novel upper bound method for estimating model performance.
result Validates and demonstrates the practical utility of the method.
New method finds minimum in noisy data, useful for model selection.
problem Finding the index of the minimum value in noisy observations.
method Developed an asymptotically normal test statistic integrating cross-validation and differential privacy.
result Achieves a favorable bias-variance trade-off in practical scenarios.
New method estimates treatment effects across different populations.
problem Estimating treatment effects across populations with changing distributions.
method SBRL-HAP framework combining balancing and independence regularizers with hierarchical attention.
result Significant improvement in HTE estimation across out-of-distribution populations.
Bayesian method identifies causal sets across populations without graph knowledge.
problem Transporting causal information across populations without causal graph knowledge.
method Combines observational and experimental data to identify s-admissible backdoor sets.
result Proves asymptotic convergence and corrects transportability bias in simulations.
The paper explores how to select data points for optimal learning performance.
problem Optimizing data selection for empirical risk minimizers.
method Fixing a learning rule and focusing on optimizing the training data selection.
result Achieving performance comparable to training on the entire population with a small subset of data points.
Feature selection predicts immune state changes in RA mouse model.
problem Predicting the immune state change after RA immunotherapy.
method Feature selection algorithms applied to mouse CIA model data.
result Selected features predict both T cell markers and treatment efficacy.
Study merges datasets to improve AI model performance.
problem Improving machine learning models with diverse datasets.
method Developed an algorithm using oracle inequality and data-driven estimators.
result Algorithm reduces population loss with high probability.
Lapse-supported life insurance exacerbates adverse selection risks.
problem Lapse-supported life insurance increases adverse selection costs.
method Modeling 'Term to 100' contracts and analyzing three methods of managing lapse surplus.
result Adverse selection losses can be almost unlimited under certain conditions.
Study explains mortgage burnout using Cox hazard models.
problem Understanding burnout in mortgage pools.
method Modeling mortgage prepayment using Cox hazard processes.
result Observed pool hazard is a survival-weighted mean of individual hazards with a selection term.
New methods reveal consensus and dissensus in network partitions.
problem Degenerate community detection methods often yield multiple competing answers.
method Comprehensive set of methods to characterize and summarize complex populations of partitions.
result It is not possible to obtain a consistent answer from point estimates when the distribution is heterogeneous.
Study finds incorporating fairness in healthcare models doesn't improve performance or net benefit.
problem Addressing health inequities in healthcare through algorithmic fairness.
method Empirical case study using models to estimate atherosclerotic cardiovascular disease risk.
result Incorporating fairness considerations into model training objective does not improve model performance or net benefit.
System exposes study population descriptions in clinical guidelines.
problem Challenges in understanding applicability of clinical guidelines.
method Developed an ontology-enabled prototype system using SIO.
result Allows medical practitioners to better understand study populations.
Framework estimates precision matrices for heterogeneous populations.
problem Estimating precision matrices in populations with subpopulations.
method Laplacian shrinkage penalty, ADMM algorithm, hierarchical clustering.
result Consistent estimation of precision matrices in heterogeneous populations.
The study optimizes bandwidth for nonparametric modal clustering.
problem Optimizing bandwidth for nonparametric modal clustering.
method Asymptotic analysis of density-based partitions and bandwidth selection.
result Asymptotic approximation of a metric for partition distance.
Optimizes biomarker selection for cost-effective treatment rules.
problem Incorporating multiple biomarkers in treatment selection rules can be costly and reduce model performance.
method Developed procedures for estimating linear and nonlinear combinations of biomarkers using 0-norm penalized weighted classification.
result Demonstrated the importance of feature selection and marker cost in treatment selection rules.
Active learning method for neural population dynamics using optogenetics.
problem Efficiently selecting neurons to stimulate for identifying neural population dynamics.
method Developed active learning procedure for low-rank regression to determine informative photostimulation patterns.
result Demonstrated a two-fold reduction in data required for predictive power using low-rank linear dynamical systems model.
A new evolutionary algorithm improves k-means clustering by recombining the entire population.
problem Optimizing the k-means clustering problem, especially in non-convex cases.
method Recombinator-k-means uses stochastic recombination with a reweighting mechanism.
result Recombinator-k-means outperforms standard genetic algorithms in optimization objective.
Optimizes subgroup selection in clinical trials.
problem Identifying regions in feature space where a regression function exceeds a threshold.
method Formulates subgroup selection as constrained optimisation, determining minimax optimal rate for regret.
result Determines the minimax optimal rate for regret in sample size and Type I error probability.
Selection mechanisms impact market volatility in evolving markets.
problem Determining how selection mechanisms affect market volatility in evolving markets.
method Used a population of evolving zero-intelligence agents and a frequent batch auction price-discovery mechanism to analyze the role of selection mechanisms.
result Local fitness-proportionate selection mechanisms correlate with high correlation between risk-aversion and volatility, while quantile-based selection mechanisms show less correlation.
FAQ efficiently evaluates LLMs with statistical guarantees using adaptive query selection.
problem Efficiently evaluating many LLMs on a large suite of benchmarks is expensive.
method FAQ uses Bayesian factor models, adaptive sampling, and proactive active inference to select queries.
result FAQ delivers up to 5x effective sample size gains over baselines, matching CI width with fewer queries.
New framework for inference with LAR, explaining variable contributions and providing stopping rules.
problem LAR's lack of well-understood termination point and basic behavioral properties.
method Developed a novel framework for inference with LAR, providing new mathematical properties and stopping rules.
result LAR estimates of non-zero population correlations have independent normal distributions for inference, and zero-valued correlations have a non-normal joint distribution.
A new method corrects for bias in selecting the best candidate.
problem Bias in selecting the best candidate leads to invalid conclusions.
method Zoom correction, flexible for parametric and nonparametric settings.
result Valid inference on the winner is possible, even with selection bias.
Optimizes kernel discrepancies by selecting subsets efficiently.
problem Improving kernel discrepancies for QMC methods.
method Introduces a novel subset selection algorithm for kernel discrepancies.
result Efficiently generates low-discrepancy samples from various distributions.
Unified method for learning from selectively labeled data.
problem Classification with selectively labeled data from multiple decision-makers.
method Unified cost-sensitive learning (UCL) approach.
result Unified method for robust classification in selective labeling.
Microarray cancer gene expression data comprise of very high dimensions. Reducing the dimensions helps in improving the overall analysis and classification performance. We propose two hybrid techniques, Biogeography - based Optimization - Random Forests (BBO - RF) and BBO - SVM (Support Vector Machines) with gene ranki…
Study shows improving information content in risk scores can reduce algorithmic discrimination.
problem Algorithmic prediction systems may discriminate against underrepresented populations.
method Investigates the role of information in fair prediction, introduces refinement concept.
result Improving information content through refinements improves downstream selection rules and reduces disparity in treatment.
We study ranking quantilized mean-field games to select top-performing agents.
problem Selecting top-performing agents in competitive scenarios.
method Developed two formulations: target-based and threshold-based, and provided analytic and semi-explicit solutions.
result Analytic and semi-explicit solutions for quantilized mean-field consistency conditions.
Study on optimal trading in a finite population with market frictions and asymmetric information.
problem Optimal trading in a finite population with market frictions and asymmetric information.
method Investigates stochastic differential games with asymmetric information and market frictions, proving existence and uniqueness of Nash and Stackelberg-Nash equilibria.
result Existence and uniqueness of Nash and Stackelberg-Nash equilibria in both unconstrained and constrained trading scenarios.
A controller optimizes sampling from unknown distributions to maximize a score function.
problem Optimizing sampling from unknown distributions to maximize a score function.
method Uniformly Fast (UF) sampling policies and UCB policy.
result UCB policy is asymptotically optimal and achieves the lower bound for sub-optimal activations.
Study on Brazilian income distribution using PNAD data.
problem Analyzing income inequality in Brazil.
method Used semi-empirical model to reconcile Pareto and Boltzmann-Gibbs distributions, calculated multiple inequality measures.
result Decreasing income inequality over 2001-2014 period, especially for historically disadvantaged groups.
Paper proposes a method to estimate coverage in data streams.
problem Estimating coverage in data streams with limited storage.
method Modified CVM algorithm for estimating coverage in streaming settings.
result The method efficiently estimates coverage in data streams.
aLTT selects hyperparameters efficiently with statistical guarantees.
problem Statistical validity and efficiency in hyperparameter selection.
method Sequential data-dependent multiple hypothesis testing with early termination.
result Reduces testing rounds while maintaining statistical validity.
Rodent hippocampal population codes represent important spatial information about the environment during navigation. Several computational methods have been developed to uncover the neural representation of spatial topology embedded in rodent hippocampal ensemble spike activity. Here we extend our previous work and pro…
The paper reveals a spinning top geometry in real-world games.
problem Understanding the structure of real-world games.
method Developed a spinning top geometric model and used Nash clustering.
result Real-world games exhibit a spinning top structure with transitive and non-transitive dimensions.
Proposes a resampling method to compare uplift models with uncertainty.
problem Uncertainty in estimating uplift curves when full population data is unavailable.
method Two-step sampling procedure and resampling-based approach.
result Validates the proposed method through simulations and real data applications.
Validates conformal prediction for network data under non-uniform sampling.
problem Validity of conformal prediction for network data under non-representative sampling.
method Interprets sampling mechanisms as selection rules, studies validity conditional on selection events, uses permutation invariance and joint exchangeability.
result Finite-sample validity of conformal prediction for certain selection events and asymptotic validity for random walk sampling.
The paper addresses learning from biased data samples.
problem Selection bias in training data affects machine learning outcomes.
method Develops a method to create a nearly debiased training population from biased samples.
result The learning rate is nearly as good as in unbiased scenarios.
MarkerMap selects key genes for cell type analysis in single-cell RNA-seq.
problem Selecting informative genes from large single-cell RNA-seq datasets is challenging and computationally intensive.
method MarkerMap is a generative model that identifies minimal gene sets explaining cell type variability.
result MarkerMap outperforms existing methods in both supervised and unsupervised marker selection.
Proposes a method to select fair performance metrics through metric elicitation.
problem Choosing fair performance metrics in multiclass classification with multiple sensitive groups.
method Metric elicitation strategy that requires only relative preference feedback and is robust to noise.
result Elicits group-fair performance metrics for multiclass classification problems.
Optimizes tensor rank selection for neural network compression.
problem Finding optimal tensor rank for regression models.
method Analyzes population expressions for training-testing discrepancy under Gaussian design.
result Optimal rank minimizes prediction error and aligns with cross-validation.
Operator calculus for population-based optimization provides a unified framework for analyzing convergence of various methods.
problem Convergence analysis of population-based optimization methods
method Introduce an operator calculus for describing composite mean-field algorithms as compositions of elementary operators acting on probability measures.
result Establish a modular Lyapunov principle for certifying exponential decay of state-space Lyapunov function and search errors.
Social learning can make financial markets inefficient, but individual learning can fix this.
problem Inefficiencies in financial markets due to social learning.
method Study of the Minority Game model with social and individual learning mechanisms.
result Individual learning can rescue a population from the inefficiencies caused by social learning.
Heavy-tailed distributions are frequently used to enhance the robustness of regression and classification methods to outliers in output space. Often, however, we are confronted with "outliers" in input space, which are isolated observations in sparsely populated regions. We show that heavy-tailed stochastic processes (…
Paper tackles fair low-rank approximation and column subset selection.
problem Minimize loss over sub-populations in machine learning.
method Developed algorithms for fair low-rank approximation and fair column subset selection.
result Achieved polynomial time algorithms for fair low-rank approximation.