Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

8.3%16.7%25.0%33.3% · Jan 199319922001200920172026
48 results for population estimation

Estimates Mozambique's population using remote sensing and microcensus data.

problem Lack of frequent population estimation due to censuses lacking spatio-temporal resolution.
method Combines remote sensing, microcensus data, and transfer learning with publicly available datasets.
result Population predictions improve with footprint area estimation using transfer learning.

A new data set helps estimate continental-scale population distributions.

problem Lack of comprehensive, publicly available data for population estimation.
method Comprehensive data set combining satellite imagery and open-source data.
result Provides a valuable resource for developing population estimation methods.

Develops a method to estimate personalized treatment regimes from summary statistics.

problem Estimating optimal treatment regimes for a target population when individual-level data is unavailable.
method A weighting framework that tailors a treatment regime for the target population using summary statistics.
result Consistent and asymptotically normal estimator for optimal treatment regimes.

Given samples from a distribution, how many new elements should we expect to find if we continue sampling this distribution? This is an important and actively studied problem, with many applications ranging from unseen species estimation to genomics. We generalize this extrapolation and related unseen estimation proble…

2017-07-12abs ↗pdf ↗

Estimates KL divergence with fairness considerations for sub-populations.

problem Fairly estimate KL divergence between distributions considering sub-populations.
method Proposes multi-group attribution for KL divergence estimation, derived from multi-calibration.
result Shows multi-group attribution provides better KL divergence estimates conditioned on sub-populations.

LEARNER improves low-rank matrix estimation using source population data.

problem Improving low-rank matrix estimation in target populations with diverse data sources.
method LEARNER uses similarity in latent spaces between source and target populations to enhance estimation.
result LEARNER often outperforms benchmark methods, especially with higher signal-to-noise ratios in the source population.

New model estimates species population trends from citizen science data.

problem Interannual confounding in citizen science data.
method Double Machine Learning framework to estimate population change and propensity scores for confounding adjustment.
result Spatially detailed trend estimates from citizen science data with low error rates.

Estimates unknown population sizes using the hypergeometric distribution.

problem Estimating discrete distributions with unknown population sizes and category sizes.
method Proposes a novel solution using the hypergeometric likelihood, accounting for a data generating process with a latent variable.
result Empirically demonstrates superior performance in estimating population sizes and learning latent spaces compared to other methods.

New method estimates treatment effects across different populations.

problem Estimating treatment effects across populations with changing distributions.
method SBRL-HAP framework combining balancing and independence regularizers with hierarchical attention.
result Significant improvement in HTE estimation across out-of-distribution populations.

Efficiently estimates variable importance in prediction tasks using Shapley values.

problem Valid statistical inference on the importance of variables in prediction tasks.
method Randomly sampling feature subsets to estimate Shapley Population Variable Importance Measure (SPVIM) efficiently.
result The proposed estimator converges at an asymptotically optimal rate and can construct valid confidence intervals and hypothesis tests.

This paper addresses external validity bias in causal inference.

problem Estimating causal effects in a target population.
method Synthesis of approaches for generalizability and transportability, including tests for heterogeneity of treatment effects and differences between study and target populations.
result Framework for addressing external validity bias in causal inference.

Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.

problem Underrepresentation of minorities in precision medicine research leads to underperforming risk prediction models.
method Two-way federated transfer learning strategy integrating diverse populations and healthcare institutions.
result Improves risk prediction models for underrepresented populations, reducing performance gaps.

Estimates causal effect of managed care plans on NYC Medicaid spending.

problem Generalizing causal estimates to a target population not well-represented by randomized studies.
method Conditional cross-design synthesis estimators combining randomized and observational data.
result Estimates causal effect of managed care plans on health care spending.

We introduce a general framework for estimation of inverse covariance, or precision, matrices from heterogeneous populations. The proposed framework uses a Laplacian shrinkage penalty to encourage similarity among estimates from disparate, but related, subpopulations, while allowing for differences among matrices. We p…

2016-01-02abs ↗pdf ↗

New methods resolve conflicting treatment effect estimates in health tech assessments.

problem Conflicting conclusions from different sponsors analyzing the same data.
method Arbitrated indirect treatment comparisons (ArMAIC) targeting a common target population.
result Estimates treatment effects in a common target population, resolving the MAIC paradox.

Paired estimation of change in parameters of interest over a population plays a central role in several application domains including those in the social sciences, epidemiology, medicine and biology. In these domains, the size of the population under study is often very large, however, the number of observations availa…

2019-11-28abs ↗pdf ↗

A new method identifies sub-populations in unlabelled heterogeneous data by accounting for co-features.

problem Estimating sub-populations in unlabelled heterogeneous data with co-features.
method Mixture of Conditional Gaussian Graphical Models (CGGM) with penalized EM algorithm.
result The method successfully identifies sub-populations disrupted by co-features.

Paper uses neural networks to calibrate Lee-Carter models for multiple populations.

problem Calibrating Lee-Carter models for multiple populations with neural networks.
method Developed neural network architectures to fit Lee-Carter and Poisson Lee-Carter models simultaneously.
result Smooth and less sensitive parameter estimates, improved forecasting performance.

The problem of population recovery refers to estimating a distribution based on incomplete or corrupted samples. Consider a random poll of sample size nn conducted on a population of individuals, where each pollee is asked to answer dd binary questions. We consider one of the two polling impediments: (a) in lossy pop…

2017-02-18abs ↗pdf ↗

SDRF estimates complex survey designs for conditional distributions.

problem Estimating conditional distributions under complex survey designs.
method Survey-calibrated distributional random forest (SDRF) with pseudo-population bootstrap and MMD split criterion.
result Established design consistency and model consistency for survey designs.

Paper tackles estimating individual treatment effects from observational data.

problem Estimating the difference between outcomes with and without treatment from single observation.
method Formulated as inference from hidden variables, uses a model of four causal populations, proposes ECM algorithm.
result ECM algorithm provides better performance compared to baseline methods on synthetic and real-world data.

New method constructs synthetic treatment groups without mean exchangeability assumption.

problem Violations of mean exchangeability assumption in randomized controlled trials.
method Weighted mixture of treatment groups from source populations, minimizing conditional maximum mean discrepancy.
result Asymptotic normality of synthetic treatment group estimator established.

Aims to describe neural network training dynamics using two-time-scale models.

problem Lack of a general mathematical description of neural network training.
method Introduces a theoretical framework based on two-time-scale population dynamics.
result Derives selection-mutation equations and effective fitness for hyperparameters.

Estimates class prior for unlabeled data using kernel embedding.

problem Estimating class prior in PU learning scenario where only positive and full population samples are available.
method Direct estimator based on distribution matching and kernel embedding in Reproducing Kernel Hilbert Space.
result Asymptotic consistency and explicit deviation bound for the estimator.

CTGAN synthesizes population data for travel behavior simulation.

problem Synthesizing population data for agent-based transportation modeling.
method Composite Travel Generative Adversarial Network (CTGAN).
result Consistent and accurate generation of synthetic populations with tabular and sequential mobility data.

Method tackles missing covariates in large-scale datasets.

problem Cross-population missing data problem in large-scale datasets.
method Augmented transfer regression learning method combining importance-weighted estimating equations and imputation terms.
result Estimator is n1/2n^{1/2}-consistent and asymptotically normal, attaining semiparametric efficiency bound under correct specification.

Meta learning of optimal classifier error rates allows an experimenter to empirically estimate the intrinsic ability of any estimator to discriminate between two populations, circumventing the difficult problem of estimating the optimal Bayes classifier. To this end we propose a weighted nearest neighbor (WNN) graph es…

2017-10-31abs ↗pdf ↗

NPE trains neural networks to approximate posterior distributions in SIR models from final outcome data.

problem Computational challenges in Bayesian inference for SIR models with final outcome data.
method Neural posterior estimation (NPE) using a logNormal posterior approximated by a neural network.
result NPE accurately recovers reference posteriors across various population sizes and transmission regimes.

Consider the problem of finding a population or a probability distribution amongst many with the largest mean when these means are unknown but population samples can be simulated or otherwise generated. Typically, by selecting largest sample mean population, it can be shown that false selection probability decays at an…

2015-07-16abs ↗pdf ↗

Develops a weighting framework to generalize ITRs from source to target populations.

problem Challenges in generalizing ITRs from a source population to a target population with differing characteristics.
method A robust sample weighting framework using a reproducing kernel Hilbert space to balance covariates and improve ITR learning methods.
result Improves ITR estimation for the target population compared to other weighting methods.

Entropy regularization improves interpretability of probabilistic clustering models.

problem Bayesian nonparametric mixture models often produce unbalanced cluster frequencies.
method Interpreting the posterior as penalized likelihood, entropy regularization reduces sparsely-populated clusters.
result The proposed entropy-regularized estimator enhances interpretability without sacrificing computational convenience.

Though machine learning algorithms excel at minimizing the average loss over a population, this might lead to large discrepancies between the losses across groups within the population. To capture this inequality, we introduce and study a notion we call maximum weighted loss discrepancy (MWLD), the maximum (weighted) d…

2019-06-08abs ↗pdf ↗