Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,694 papers · 148 categories

Trend · papers per month

72144215287 · Jun 202019922001200920172026
48 results for population inference

Many modern data analysis problems involve inferences from streaming data. However, streaming data is not easily amenable to the standard probabilistic modeling approaches, which assume that we condition on finite data. We develop population variational Bayes, a new approach for using Bayesian modeling to analyze strea…

2015-07-19abs ↗pdf ↗

Bayesian predictive inference analyzes a dataset to make predictions about new observations. When a model does not match the data, predictive accuracy suffers. We develop population empirical Bayes (POP-EB), a hierarchical framework that explicitly models the empirical population distribution as part of Bayesian analys…

2014-11-02abs ↗pdf ↗

This paper addresses external validity bias in causal inference.

problem Estimating causal effects in a target population.
method Synthesis of approaches for generalizability and transportability, including tests for heterogeneity of treatment effects and differences between study and target populations.
result Framework for addressing external validity bias in causal inference.

New method bypasses global fit for LISA's Galactic binaries, extracting population parameters directly.

problem Disentangling LISA's Galactic binary sources from backgrounds in a computationally intensive process.
method Simulation-based approach using normalizing flow to infer population parameters.
result Direct inference of population parameters from LISA's frequency strain series.

Efficiently estimates variable importance in prediction tasks using Shapley values.

problem Valid statistical inference on the importance of variables in prediction tasks.
method Randomly sampling feature subsets to estimate Shapley Population Variable Importance Measure (SPVIM) efficiently.
result The proposed estimator converges at an asymptotically optimal rate and can construct valid confidence intervals and hypothesis tests.

To understand how rich dynamics emerge in neural populations, we require models exhibiting a wide range of activity patterns while remaining interpretable in terms of connectivity and single-neuron dynamics. However, it has been challenging to fit such mechanistic spiking networks at the single neuron scale to empirica…

2019-10-03abs ↗pdf ↗

A method for inferring motility models and heterogeneity from particle trajectories.

problem Understanding motility patterns from discrete trajectory data of biological agents.
method Maximum likelihood approach for second-order Langevin models with population heterogeneity.
result The proposed method outperforms alternative approaches for short trajectories.

In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we develop an inference method that is resilient to sampling biases and is able to …

2019-02-26abs ↗pdf ↗

The paper develops methods to identify stable associations across multiple studies.

problem Identifying stable associations across multiple studies with possible distributional shifts.
method Modeling heterogeneous multi-source data with multiple high-dimensional regressions and devising a novel sampling method for valid confidence intervals of maximin effects.
result Significant maximin effects indicate stable associations that can be generalized to target populations.

NPE trains neural networks to approximate posterior distributions in SIR models from final outcome data.

problem Computational challenges in Bayesian inference for SIR models with final outcome data.
method Neural posterior estimation (NPE) using a logNormal posterior approximated by a neural network.
result NPE accurately recovers reference posteriors across various population sizes and transmission regimes.

New framework for inference with LAR, explaining variable contributions and providing stopping rules.

problem LAR's lack of well-understood termination point and basic behavioral properties.
method Developed a novel framework for inference with LAR, providing new mathematical properties and stopping rules.
result LAR estimates of non-zero population correlations have independent normal distributions for inference, and zero-valued correlations have a non-normal joint distribution.

Bayesian Federated Inference combines local data analyses to estimate regression models.

problem Estimating accurate parameters with limited data from different centers.
method Bayesian Federated Inference (BFI) for pooling local data analyses.
result Excellent performance of BFI methodology shown in real-life examples.

New method uses SBI to infer magnetorotational properties of isolated pulsars.

problem Constrain magnetorotational properties of isolated Galactic radio pulsars.
method Combines population synthesis with SBI to model neutron star birth and evolution.
result Inferred μlogB=13.100.10+0.08μ_{\log B} = 13.10^{+0.08}_{-0.10}, σlogB=0.450.05+0.05σ_{\log B} = 0.45^{+0.05}_{-0.05} for lognormal distributions.

This work learns models for population dynamics using variational methods and higher-order quadrature.

problem Modeling population dynamics of physical systems with stochastic and mean-field effects.
method Variational problem to infer gradient fields, combining Monte Carlo sampling with higher-order quadrature rules.
result Accurate prediction of population dynamics over a wide range of parameters.

Causal processes in biomedicine may contain cycles, evolve over time or differ between populations. However, many graphical models cannot accommodate these conditions. We propose to model causation using a mixture of directed cyclic graphs (DAGs), where the joint distribution in a population follows a DAG at any single…

2019-01-28abs ↗pdf ↗

Estimates causal effect of managed care plans on NYC Medicaid spending.

problem Generalizing causal estimates to a target population not well-represented by randomized studies.
method Conditional cross-design synthesis estimators combining randomized and observational data.
result Estimates causal effect of managed care plans on health care spending.

Framework generates precise synthetic populations for scalable modeling.

problem Generating accurate synthetic populations without personal data.
method Constraint-programming framework encoding aggregated statistics and structural relations.
result Exact control of demographic profiles without requiring microdata.

A new method for learning gradient flows from population dynamics.

problem Reconstructing population dynamics from limited data.
method Residual approach to enforce continuity equations, combining with data-fitting divergence.
result Demonstrated state-of-the-art performance across trajectory inference benchmarks.

New methods improve causal inference generalization using trial and observational data.

problem Limited trial data makes generalizing causal inferences to target populations statistically infeasible.
method Develops algorithms that combine trial and observational data to estimate complex nuisance functions.
result Improves generalization of causal inferences when the additional observational study is high-quality.

This work compares regularization and constrained inference for label constraints in machine learning.

problem Improving model performance with label constraints in machine learning.
method Comparison of regularization and constrained inference strategies.
result Constrained inference reduces population risk by correcting model violations, while regularization narrows the generalization gap but introduces bias.

A body of recent work in modeling neural activity focuses on recovering low-dimensional latent features that capture the statistical structure of large-scale neural populations. Most such approaches have focused on linear generative models, where inference is computationally tractable. Here, we propose fLDS, a general …

2016-05-26abs ↗pdf ↗

Bayesian networks learn sub-population differences from data.

problem Inference from a single network structure can be misleading when data populations are heterogeneous.
method A mixture of Bayesian networks where component probabilities depend on individual characteristics.
result Identifies both network structures and demographic predictors of sub-population membership.

We develop amortized population Gibbs (APG) samplers, a class of scalable methods that frames structured variational inference as adaptive importance sampling. APG samplers construct high-dimensional proposals by iterating over updates to lower-dimensional blocks of variables. We train each conditional proposal by mini…

2019-11-04abs ↗pdf ↗

The Collective Graphical Model (CGM) models a population of independent and identically distributed individuals when only collective statistics (i.e., counts of individuals) are observed. Exact inference in CGMs is intractable, and previous work has explored Markov Chain Monte Carlo (MCMC) and MAP approximations for le…

2014-05-20abs ↗pdf ↗

New method infers population dynamics from snapshots using path space optimization.

problem Recover dynamics of a population from its temporal marginals.
method Grid-free algorithm using Schrödinger bridges coupled via noisy gradient descent in mean-field limit.
result Global convergence to min-entropy estimator with end-to-end theoretical guarantees.

PAVI speeds up VI for large-scale studies by sharing parameterization across i.i.d. variables.

problem Challenges in Bayesian inference for large population studies with many latent parameters.
method Designing plate-amortized variational inference (PAVI) to share parameterization across i.i.d. variables.
result Significant speedup in training large-scale hierarchical variational distributions.

Treatment recommendations within Clinical Practice Guidelines (CPGs) are largely based on findings from clinical trials and case studies, referred to here as research studies, that are often based on highly selective clinical populations, referred to here as study cohorts. When medical practitioners apply CPG recommend…

2019-07-09abs ↗pdf ↗

The paper proposes a new method for comparing logistic regression models across different populations.

problem Comparing logistic regression models across sub-populations can lead to misleading results.
method Develops a cascading set of equivalence tests for logistic regression models, addressing coding, predictions, and overall accuracy.
result Equivalence testing incentivizes accurate inference and avoids perverse incentives from significance tests.