Infinitesimal boosting converges to a deterministic process in large sample limit.
problem Characterizing the asymptotic behavior of infinitesimal gradient boosting in large sample sizes.
method Proving convergence to a deterministic process using large sample theory and differential equations.
result The test error decreases over time in the population limit.
Aims to describe neural network training dynamics using two-time-scale models.
problem Lack of a general mathematical description of neural network training.
method Introduces a theoretical framework based on two-time-scale population dynamics.
result Derives selection-mutation equations and effective fitness for hyperparameters.
EPC curriculum improves MARL performance as agent population grows.
problem Challenges in learning good policies for large multi-agent systems.
method Evolutionary Population Curriculum (EPC) for scaling MARL.
result EPC consistently outperforms baselines as agent population increases.
Evaluating the computational reproducibility of data analysis pipelines has become a critical issue. It is, however, a cumbersome process for analyses that involve data from large populations of subjects, due to their computational and storage requirements. We present a method to predict the computational reproducibili…
New method pools labels from similar data items to improve learning from small samples.
problem Learning from small, human-annotated samples with potential disagreement among annotators.
method Proposes neighborhood-based pooling for sharing labels across similar data items.
result Improves learning from small, noisy samples by pooling labels from similar items.
Many modern data analysis problems involve inferences from streaming data. However, streaming data is not easily amenable to the standard probabilistic modeling approaches, which assume that we condition on finite data. We develop population variational Bayes, a new approach for using Bayesian modeling to analyze strea…
A powerful approach for understanding neural population dynamics is to extract low-dimensional trajectories from population recordings using dimensionality reduction methods. Current approaches for dimensionality reduction on neural data are limited to single population recordings, and can not identify dynamics embedde…
Study equilibrium consumption habits in a large population using mean field games.
problem Equilibrium consumption under external habit formation in a large population.
method Formulated and solved mean field games for linear and multiplicative habit formation preferences, constructed approximate Nash equilibria for large n-player games.
result Characterized mean field equilibrium strategies and derived financial implications.
Study multiple-population games using McKean-Vlasov equations.
problem Mean field games and control problems with multiple populations.
method Coupled forward-backward SDEs and Pontryagin's principle.
result Existence of mean field equilibria under various cooperation scenarios.
A new method simulates large, diverse populations of learning agents evolving in games.
problem Limited scalability and efficiency of Multi-Agent Reinforcement Learning.
method Parallelizable implementation of Policy Gradient and Opponent-Learning Awareness for evolutionary simulations.
result Simulated large, diverse populations of learning agents evolve under various strategies.
Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.
problem Underrepresentation of minorities in precision medicine research leads to underperforming risk prediction models.
method Two-way federated transfer learning strategy integrating diverse populations and healthcare institutions.
result Improves risk prediction models for underrepresented populations, reducing performance gaps.
In order to investigate the breast cancer prediction problem on the aging population with the grades of DCIS, we conduct a tree augmented naive Bayesian network experiment trained and tested on a large clinical dataset including consecutive diagnostic mammography examinations, consequent biopsy outcomes and related can…
Framework generates precise synthetic populations for scalable modeling.
problem Generating accurate synthetic populations without personal data.
method Constraint-programming framework encoding aggregated statistics and structural relations.
result Exact control of demographic profiles without requiring microdata.
Method tackles missing covariates in large-scale datasets.
problem Cross-population missing data problem in large-scale datasets.
method Augmented transfer regression learning method combining importance-weighted estimating equations and imputation terms.
result Estimator is n1/2-consistent and asymptotically normal, attaining semiparametric efficiency bound under correct specification. With the development of high-throughput technologies, principal component analysis (PCA) in the high-dimensional regime is of great interest. Most of the existing theoretical and methodological results for high-dimensional PCA are based on the spiked population model in which all the population eigenvalues are equal ex…
CrowdLLM uses LLMs and generative models to create diverse digital populations.
problem Lack of diversity and accuracy in digital populations created by LLMs.
method Integrates pretrained LLMs and generative models to enhance diversity and fidelity.
result CrowdLLM achieves promising performance in accuracy and distributional fidelity.
This paper studies the sample complexity of searching over multiple populations. We consider a large number of populations, each corresponding to either distribution P0 or P1. The goal of the search problem studied here is to find one population corresponding to distribution P1 with as few samples as possible. The main…
Fiber simplifies RL and population-based methods for distributed training.
problem Challenges in RL and population-based methods, including frequent interaction with simulations and dynamic scaling.
method Introducing Fiber, a scalable distributed computing framework.
result Significantly expands accessibility of large-scale parallel computation.
Study shows neural networks outperform traditional methods in speaker identification.
problem Open-set speaker identification with large populations.
method Discriminative neural networks compared to Gaussian mixture models.
result Multi-class neural networks outperform traditional methods for large speaker populations.
Study uses MFG approach to model equilibrium pricing with market clearing condition.
problem Continuous asset pricing with market clearing condition.
method Mean field game approach to solve forward-backward SDEs of McKean-Vlasov type.
result Net order flow converges to zero in large N-limit with specified conditions.
Insiders camouflage trading to balance wealth and stealth, avoiding legal penalties.
problem Legal penalties and insider trading among liquidity traders.
method Kyle-type model with a diverse spectrum of prosecution schemes.
result Existence and uniqueness of equilibria for large populations, with a stealth index revealing trading scale.
In systems biomedicine, an experimenter encounters different potential sources of variation in data such as individual samples, multiple experimental conditions, and multi-variable network-level responses. In multiparametric cytometry, which is often used for analyzing patient samples, such issues are critical. While c…
The emergence of complex life on Earth is often attributed to the arms race that ensued from a huge number of organisms all competing for finite resources. We present an artificial intelligence research environment, inspired by the human game genre of MMORPGs (Massively Multiplayer Online Role-Playing Games, a.k.a. MMO…
New model estimates species population trends from citizen science data.
problem Interannual confounding in citizen science data.
method Double Machine Learning framework to estimate population change and propensity scores for confounding adjustment.
result Spatially detailed trend estimates from citizen science data with low error rates.
Consider the problem of finding a population or a probability distribution amongst many with the largest mean when these means are unknown but population samples can be simulated or otherwise generated. Typically, by selecting largest sample mean population, it can be shown that false selection probability decays at an…
Paper uses neural networks to calibrate Lee-Carter models for multiple populations.
problem Calibrating Lee-Carter models for multiple populations with neural networks.
method Developed neural network architectures to fit Lee-Carter and Poisson Lee-Carter models simultaneously.
result Smooth and less sensitive parameter estimates, improved forecasting performance.
In lowest unique bid auctions, N players bid for an item. The winner is whoever places the \emph{lowest} bid, provided that it is also unique. We use a grand canonical approach to derive an analytical expression for the equilibrium distribution of strategies. We then study the properties of the solution as a function…
System exposes study population descriptions in clinical guidelines.
problem Challenges in understanding applicability of clinical guidelines.
method Developed an ontology-enabled prototype system using SIO.
result Allows medical practitioners to better understand study populations.
The perennial problem of "how many clusters?" remains an issue of substantial interest in data mining and machine learning communities, and becomes particularly salient in large data sets such as populational genomic data where the number of clusters needs to be relatively large and open-ended. This problem gets furthe…
The paper proposes a new method for comparing logistic regression models across different populations.
problem Comparing logistic regression models across sub-populations can lead to misleading results.
method Develops a cascading set of equivalence tests for logistic regression models, addressing coding, predictions, and overall accuracy.
result Equivalence testing incentivizes accurate inference and avoids perverse incentives from significance tests.
Population attributes are essential in health for understanding who the data represents and precision medicine efforts. Even within disease infection labels, patients can exhibit significant variability; "fever" may mean something different when reported in a doctor's office versus from an online app, precluding direct…
New models extrapolate false alarms in ASV without new data.
problem Reliable extrapolation of false alarm rates in ASV without new speaker data.
method Generative models in ASV score space for arbitrary systems.
result Models accurately extrapolate false alarm rates for large speaker populations.
Paper studies optimal tracking portfolio in mean field game of large fund competition.
problem Optimal tracking portfolio in large fund competition with relative performance benchmark.
method Formulated mean field game problem, established existence of mean field equilibrium using PDE approach, constructed approximate Nash equilibrium.
result Existence of mean field equilibrium and consistency condition verified.
New method generates geolocated synthetic populations from real data.
problem Generating synthetic populations with explicit geographic coordinates.
method Mapping coordinates into a latent space using Normalizing Flows (NF), then combining with other features in a Variational Autoencoder (VAE).
result NF+VAE architecture outperforms existing methods in generating geolocated synthetic populations.
We introduce an algorithmic method for population anomaly detection based on gaussianization through an adversarial autoencoder. This method is applicable to detection of `soft' anomalies in arbitrarily distributed highly-dimensional data. A soft, or population, anomaly is characterized by a shift in the distribution o…
A framework is presented for unsupervised learning of representations based on infomax principle for large-scale neural populations. We use an asymptotic approximation to the Shannon's mutual information for a large neural population to demonstrate that a good initial approximation to the global information-theoretic o…
NPE trains neural networks to approximate posterior distributions in SIR models from final outcome data.
problem Computational challenges in Bayesian inference for SIR models with final outcome data.
method Neural posterior estimation (NPE) using a logNormal posterior approximated by a neural network.
result NPE accurately recovers reference posteriors across various population sizes and transmission regimes.
MFM integrates multiple evolving populations using Wasserstein manifold flows.
problem Learning dynamics of multiple interacting populations evolving over time.
method Meta Flow Matching (MFM) integrates vector fields on Wasserstein manifold using amortized flow models and GNN embeddings.
result MFM improves prediction of individual treatment responses on multi-patient single-cell drug screen data.
We study the top-K ranking problem where the goal is to recover the set of top-K ranked items out of a large collection of items based on partially revealed preferences. We consider an adversarial crowdsourced setting where there are two population sets, and pairwise comparison samples drawn from one of the populat…
Improves flu prediction by blending environment and population info.
problem Challenges in using data from one environment in another due to feature variability and population subgroup differences.
method Population-aware hierarchical Bayesian domain adaptation framework with multiple invariant components.
result Model improves flu prediction in new environments with unlabelled data.
Improves Gaussian process factor models for multi-population recordings.
problem Cubic runtime scaling with trial length and group number limits application to large-scale recordings.
method Two approximate approaches: inducing variables and frequency domain.
result Achieved orders of magnitude speed-up with minimal statistical performance impact.
Much recent research aims to identify evidence for Drug-Drug Interactions (DDI) and Adverse Drug reactions (ADR) from the biomedical scientific literature. In addition to this "Bibliome", the universe of social media provides a very promising source of large-scale data that can help identify DDI and ADR in ways that ha…
Develops an equilibrium model for securities pricing in a mixed cooperative and non-cooperative market.
problem Equilibrium pricing of securities in a market with cooperative and non-cooperative agents.
method Conditional extended mean-field control for cooperative agents, mean-field model for both cooperative and non-cooperative agents.
result Existence of a unique equilibrium for both finite-agent and mean-field models under certain conditions.
Study laws of large numbers in online classification, determining optimal regret bounds.
problem Understanding how sequential sampling affects online learning and classification.
method Characterized online learnable classes and determined optimal regret bounds using Littlestone's dimension.
result Optimal regret bounds in online learning are determined, resolving open questions.
Study on order book dynamics with uniform catastrophes, explaining volatility and trends.
problem Understanding volatility and trends in financial markets with different types of liquidity.
method Stochastic models and population processes with uniform catastrophes.
result Law of large numbers, central limit theorem, and large deviations proved for the model.
PAVI speeds up VI for large-scale studies by sharing parameterization across i.i.d. variables.
problem Challenges in Bayesian inference for large population studies with many latent parameters.
method Designing plate-amortized variational inference (PAVI) to share parameterization across i.i.d. variables.
result Significant speedup in training large-scale hierarchical variational distributions.
MutaGAN predicts mutations of evolving protein populations using GANs.
problem Predicting mutations in evolving protein populations.
method Generative adversarial networks (GANs) with recurrent neural networks (RNNs).
result MutaGAN generates complete protein sequences with mutations.
PAVI speeds up Bayesian inference for large datasets.
problem Challenges in Bayesian inference for large population studies.
method Plate-amortized Variational Inference (PAVI) that shares parameterization across i.i.d. variables.
result Significant speedup in training variational distributions, orders of magnitude faster.