Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2725448161,088 · Jun 202019922001200920182026
48 results for large population studies

Method predicts computational reproducibility of large population studies data analysis pipelines.

problem Difficulty in evaluating reproducibility of large population studies due to computational and storage requirements.
method Formulated as collaborative filtering process with constraints on training set construction.
result One sampling method, 'Random File Numbers (Uniform)', predicts reproducibility with good accuracy.

Study equilibrium consumption habits in a large population using mean field games.

problem Equilibrium consumption under external habit formation in a large population.
method Formulated and solved mean field games for linear and multiplicative habit formation preferences, constructed approximate Nash equilibria for large n-player games.
result Characterized mean field equilibrium strategies and derived financial implications.

Study shows neural networks outperform traditional methods in speaker identification.

problem Open-set speaker identification with large populations.
method Discriminative neural networks compared to Gaussian mixture models.
result Multi-class neural networks outperform traditional methods for large speaker populations.

Study multiple-population games using McKean-Vlasov equations.

problem Mean field games and control problems with multiple populations.
method Coupled forward-backward SDEs and Pontryagin's principle.
result Existence of mean field equilibria under various cooperation scenarios.

Many modern data analysis problems involve inferences from streaming data. However, streaming data is not easily amenable to the standard probabilistic modeling approaches, which assume that we condition on finite data. We develop population variational Bayes, a new approach for using Bayesian modeling to analyze strea…

2015-07-19abs ↗pdf ↗

This paper studies the sample complexity of searching over multiple populations. We consider a large number of populations, each corresponding to either distribution P0 or P1. The goal of the search problem studied here is to find one population corresponding to distribution P1 with as few samples as possible. The main…

2012-09-06abs ↗pdf ↗

Study uses MFG approach to model equilibrium pricing with market clearing condition.

problem Continuous asset pricing with market clearing condition.
method Mean field game approach to solve forward-backward SDEs of McKean-Vlasov type.
result Net order flow converges to zero in large N-limit with specified conditions.

Framework generates precise synthetic populations for scalable modeling.

problem Generating accurate synthetic populations without personal data.
method Constraint-programming framework encoding aggregated statistics and structural relations.
result Exact control of demographic profiles without requiring microdata.

Method tackles missing covariates in large-scale datasets.

problem Cross-population missing data problem in large-scale datasets.
method Augmented transfer regression learning method combining importance-weighted estimating equations and imputation terms.
result Estimator is n1/2n^{1/2}-consistent and asymptotically normal, attaining semiparametric efficiency bound under correct specification.

Infinitesimal boosting converges to a deterministic process in large sample limit.

problem Characterizing the asymptotic behavior of infinitesimal gradient boosting in large sample sizes.
method Proving convergence to a deterministic process using large sample theory and differential equations.
result The test error decreases over time in the population limit.

Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.

problem Underrepresentation of minorities in precision medicine research leads to underperforming risk prediction models.
method Two-way federated transfer learning strategy integrating diverse populations and healthcare institutions.
result Improves risk prediction models for underrepresented populations, reducing performance gaps.

Aims to describe neural network training dynamics using two-time-scale models.

problem Lack of a general mathematical description of neural network training.
method Introduces a theoretical framework based on two-time-scale population dynamics.
result Derives selection-mutation equations and effective fitness for hyperparameters.

Paper studies optimal tracking portfolio in mean field game of large fund competition.

problem Optimal tracking portfolio in large fund competition with relative performance benchmark.
method Formulated mean field game problem, established existence of mean field equilibrium using PDE approach, constructed approximate Nash equilibrium.
result Existence of mean field equilibrium and consistency condition verified.

PAVI speeds up VI for large-scale studies by sharing parameterization across i.i.d. variables.

problem Challenges in Bayesian inference for large population studies with many latent parameters.
method Designing plate-amortized variational inference (PAVI) to share parameterization across i.i.d. variables.
result Significant speedup in training large-scale hierarchical variational distributions.

The paper proposes a new method for comparing logistic regression models across different populations.

problem Comparing logistic regression models across sub-populations can lead to misleading results.
method Develops a cascading set of equivalence tests for logistic regression models, addressing coding, predictions, and overall accuracy.
result Equivalence testing incentivizes accurate inference and avoids perverse incentives from significance tests.

New model estimates species population trends from citizen science data.

problem Interannual confounding in citizen science data.
method Double Machine Learning framework to estimate population change and propensity scores for confounding adjustment.
result Spatially detailed trend estimates from citizen science data with low error rates.

Study on order book dynamics with uniform catastrophes, explaining volatility and trends.

problem Understanding volatility and trends in financial markets with different types of liquidity.
method Stochastic models and population processes with uniform catastrophes.
result Law of large numbers, central limit theorem, and large deviations proved for the model.

A game environment simulates competition among many agents for resources.

problem Understanding large-scale multiagent interactions and resource competition.
method Developed a persistent, massively multiplayer AI environment.
result Population size affects the development of skillful behaviors and niche differentiation.

Study laws of large numbers in online classification, determining optimal regret bounds.

problem Understanding how sequential sampling affects online learning and classification.
method Characterized online learnable classes and determined optimal regret bounds using Littlestone's dimension.
result Optimal regret bounds in online learning are determined, resolving open questions.

We study the top-KK ranking problem where the goal is to recover the set of top-KK ranked items out of a large collection of items based on partially revealed preferences. We consider an adversarial crowdsourced setting where there are two population sets, and pairwise comparison samples drawn from one of the populat…

2016-02-15abs ↗pdf ↗

NPE trains neural networks to approximate posterior distributions in SIR models from final outcome data.

problem Computational challenges in Bayesian inference for SIR models with final outcome data.
method Neural posterior estimation (NPE) using a logNormal posterior approximated by a neural network.
result NPE accurately recovers reference posteriors across various population sizes and transmission regimes.

New method pools labels from similar data items to improve learning from small samples.

problem Learning from small, human-annotated samples with potential disagreement among annotators.
method Proposes neighborhood-based pooling for sharing labels across similar data items.
result Improves learning from small, noisy samples by pooling labels from similar items.

Study shows finite agent equilibrium converges to mean-field limit in asset pricing.

problem Asset pricing equilibrium in markets with finite vs infinite agents.
method Existence of finite agent equilibrium and strong convergence to mean-field limit.
result Finite agent equilibrium converges to mean-field limit under suitable conditions.

A new method simulates large, diverse populations of learning agents evolving in games.

problem Limited scalability and efficiency of Multi-Agent Reinforcement Learning.
method Parallelizable implementation of Policy Gradient and Opponent-Learning Awareness for evolutionary simulations.
result Simulated large, diverse populations of learning agents evolve under various strategies.

Study proves convergence of subgradients for optimal transport-based objectives.

problem Ensuring statistical consistency and optimization stability in transport-based models.
method Proves graphical convergence of subdifferentials to the subdifferential of the population objective.
result Standard subgradient methods consistently approach stationary points of the population-level problem.

Estimates population parameters from limited trials per individual.

problem Estimating parameters from Bernoulli trials with limited data per individual.
method Maximum Likelihood Estimation (MLE) for optimally estimating the underlying distribution.
result MLE achieves optimal error bounds of O(1t)\mathcal{O}(\frac{1}{t}) for t<clogNt < c\log{N} and O(1tlogN)\mathcal{O}(\frac{1}{\sqrt{t\log N}}) for larger tt.

Improves flu prediction by blending environment and population info.

problem Challenges in using data from one environment in another due to feature variability and population subgroup differences.
method Population-aware hierarchical Bayesian domain adaptation framework with multiple invariant components.
result Model improves flu prediction in new environments with unlabelled data.

Fiber simplifies RL and population-based methods for distributed training.

problem Challenges in RL and population-based methods, including frequent interaction with simulations and dynamic scaling.
method Introducing Fiber, a scalable distributed computing framework.
result Significantly expands accessibility of large-scale parallel computation.

In stochastic optimization, the population risk is generally approximated by the empirical risk. However, in the large-scale setting, minimization of the empirical risk may be computationally restrictive. In this paper, we design an efficient algorithm to approximate the population risk minimizer in generalized linear …

2016-11-21abs ↗pdf ↗

Study explores fairness in loan decisions using dynamic modeling.

problem Fairness constraints do not always benefit disadvantaged groups.
method Continuous population state representation using Beta distribution; model of population dynamics under lending decisions.
result Optimal lender behavior can lead to unfair outcomes, but common fairness constraints cause convergence to the same equilibrium.

Model captures context-dependent neural correlations using Poisson mixtures.

problem Capturing context-dependent noise correlations in neural populations.
method Conditional finite mixtures of Poisson distributions, cross-validation for dimensionality, EM algorithm.
result Model successfully captures stimulus-dependent correlations in V1 neuron responses.

Consider the problem of finding a population or a probability distribution amongst many with the largest mean when these means are unknown but population samples can be simulated or otherwise generated. Typically, by selecting largest sample mean population, it can be shown that false selection probability decays at an…

2015-07-16abs ↗pdf ↗

Paper uses neural networks to calibrate Lee-Carter models for multiple populations.

problem Calibrating Lee-Carter models for multiple populations with neural networks.
method Developed neural network architectures to fit Lee-Carter and Poisson Lee-Carter models simultaneously.
result Smooth and less sensitive parameter estimates, improved forecasting performance.

The study of networks leads to a wide range of high dimensional inference problems. In many practical applications, one needs to draw inference from one or few large sparse networks. The present paper studies hypothesis testing of graphs in this high-dimensional regime, where the goal is to test between two populations…

2017-07-04abs ↗pdf ↗