Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

65130194259 · Jun 202019922001200920172026
48 results for population size

Estimates unknown population sizes using the hypergeometric distribution.

problem Estimating discrete distributions with unknown population sizes and category sizes.
method Proposes a novel solution using the hypergeometric likelihood, accounting for a data generating process with a latent variable.
result Empirically demonstrates superior performance in estimating population sizes and learning latent spaces compared to other methods.

Polyak step size GD reaches final radius of convergence after log iterations.

problem Statistical and computational complexities of Polyak step size GD.
method Generalized smoothness and Lojasiewicz conditions, stability of gradients.
result Polyak step size GD reaches final statistical radius of convergence after logarithmic number of iterations.

We study the top-KK ranking problem where the goal is to recover the set of top-KK ranked items out of a large collection of items based on partially revealed preferences. We consider an adversarial crowdsourced setting where there are two population sets, and pairwise comparison samples drawn from one of the populat…

2016-02-15abs ↗pdf ↗

This paper highlights the size-dependency of income distributions, i.e. the income distribution curves versus the population of a country systematically. By using the generalized Lotka-Volterra model to fit the empirical income data in the United States during 1996-2007, we found an important parameter λλ can scale wi…

2010-12-09abs ↗pdf ↗

New method pools labels from similar data items to improve learning from small samples.

problem Learning from small, human-annotated samples with potential disagreement among annotators.
method Proposes neighborhood-based pooling for sharing labels across similar data items.
result Improves learning from small, noisy samples by pooling labels from similar items.

Direct measurements of Gini coefficients by conventional arithmetic calculations are a poor estimator, even if paradoxically, they include the entire population, as because of super-additivity they cannot lend themselves to comparisons between units of different size, and intertemporal analyses are vitiated by the popu…

2015-10-16abs ↗pdf ↗

In an adaptive population which models financial markets and distributed control, we consider how the dynamics depends on the diversity of the agents' initial preferences of strategies. When the diversity decreases, more agents tend to adapt their strategies together. This change in the environment results in dynamical…

2006-09-26abs ↗pdf ↗

Study shows finite agent equilibrium converges to mean-field limit in asset pricing.

problem Asset pricing equilibrium in markets with finite vs infinite agents.
method Existence of finite agent equilibrium and strong convergence to mean-field limit.
result Finite agent equilibrium converges to mean-field limit under suitable conditions.

A government has to finance a risk for its population. It shares the charges among the population with a fixed scale based on economic criteria. Various organisms have to collect and to redistribute fairly the subsidies. Under these conditions, when the size of the organisms is varied, the distribution's laws of the cr…

2014-01-13abs ↗pdf ↗

Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.

problem Underrepresentation of minorities in precision medicine research leads to underperforming risk prediction models.
method Two-way federated transfer learning strategy integrating diverse populations and healthcare institutions.
result Improves risk prediction models for underrepresented populations, reducing performance gaps.

Developing a scientific understanding of cities in a fast urbanizing world is essential for planning sustainable urban systems. Recently, it was shown that income and wealth creation follow increasing returns, scaling superlinearly with city size. We study scaling of per capita incomes for separate census defined incom…

2015-09-03abs ↗pdf ↗

NPE trains neural networks to approximate posterior distributions in SIR models from final outcome data.

problem Computational challenges in Bayesian inference for SIR models with final outcome data.
method Neural posterior estimation (NPE) using a logNormal posterior approximated by a neural network.
result NPE accurately recovers reference posteriors across various population sizes and transmission regimes.

Study shows income inequality increases with city size, affecting only the wealthiest deciles.

problem Understanding income inequality in urban areas.
method Urban scaling analysis of total income scaling in population percentiles.
result Income in the poorest decile does not increase with city size, while the wealthiest deciles show superlinear scaling.

New estimates for the population risk are established for two-layer neural networks. These estimates are nearly optimal in the sense that the error rates scale in the same way as the Monte Carlo error rates. They are equally effective in the over-parametrized regime when the network size is much larger than the size of…

2018-10-15abs ↗pdf ↗

Estimates population size using capture-recapture designs with binary indicators.

problem Estimating population size from capture-recapture data with binary indicators.
method Proposes a modern method using undersmoothed lasso model to estimate the target parameter of interest.
result The choice of constraint on the K-dimensional distribution significantly impacts the value of the estimand.

In this paper, we consider the problem of partitioning a small data sample drawn from a mixture of kk product distributions. We are interested in the case that individual features are of low average quality γγ, and we want to use as few of them as possible to correctly partition the sample. We analyze a spectral tech…

2007-06-25abs ↗pdf ↗

This paper studies the landscape of empirical risk of deep neural networks by theoretically analyzing its convergence behavior to the population risk as well as its stationary points and properties. For an ll-layer linear neural network, we prove its empirical risk uniformly converges to its population risk at the rat…

2017-05-19abs ↗pdf ↗

Develops an equilibrium model for securities pricing in a mixed cooperative and non-cooperative market.

problem Equilibrium pricing of securities in a market with cooperative and non-cooperative agents.
method Conditional extended mean-field control for cooperative agents, mean-field model for both cooperative and non-cooperative agents.
result Existence of a unique equilibrium for both finite-agent and mean-field models under certain conditions.

Paper improves privacy-preserving optimization rates for convex functions.

problem Differentially private stochastic convex optimization.
method Algorithmic improvements for convex and strongly convex functions under TNC and non-negative loss.
result Excess population risk bounds for DP-SCO are faster than previous results.

We develop amortized population Gibbs (APG) samplers, a class of scalable methods that frames structured variational inference as adaptive importance sampling. APG samplers construct high-dimensional proposals by iterating over updates to lower-dimensional blocks of variables. We train each conditional proposal by mini…

2019-11-04abs ↗pdf ↗

Study optimizes data collection from biased, costly sources to minimize risk.

problem Estimating population means and group-conditional means from multiple sources with varying costs and biases.
method Develops a sampling plan that maximizes effective sample size, paired with a post-stratification estimator.
result Achieves budgeted minimax optimal risk for estimating population means and group-conditional means.

The classical asymptotic theory for parametric MM-estimators guarantees that, in the limit of infinite sample size, the excess risk has a chi-square type distribution, even in the misspecified case. We demonstrate how self-concordance of the loss allows to characterize the critical sample size sufficient to guarantee …

2018-10-16abs ↗pdf ↗

Machine learning's predictive power is limited by sample size, as shown by the Limits-to-Learning Gap.

problem The limitations of machine learning in approximating true data-generating processes.
method Characterization of a universal lower bound (LLG) quantifying the discrepancy between empirical fit and population benchmark.
result Standard ML approaches can substantially understate true predictability in financial data.

In statistical connectomics, the quantitative study of brain networks, estimating the mean of a population of graphs based on a sample is a core problem. Often, this problem is especially difficult because the sample or cohort size is relatively small, sometimes even a single subject. While using the element-wise sampl…

2016-09-06abs ↗pdf ↗

The problem of population recovery refers to estimating a distribution based on incomplete or corrupted samples. Consider a random poll of sample size nn conducted on a population of individuals, where each pollee is asked to answer dd binary questions. We consider one of the two polling impediments: (a) in lossy pop…

2017-02-18abs ↗pdf ↗

New model estimates species population trends from citizen science data.

problem Interannual confounding in citizen science data.
method Double Machine Learning framework to estimate population change and propensity scores for confounding adjustment.
result Spatially detailed trend estimates from citizen science data with low error rates.

In clinical and neuroscientific studies, systematic differences between two populations of brain networks are investigated in order to characterize mental diseases or processes. Those networks are usually represented as graphs built from neuroimaging data and studied by means of graph analysis methods. The typical mach…

2015-11-19abs ↗pdf ↗

The accumulation of individual fitness or wealth is modelled as a population game in which pairs of individuals are recurrently and randomly matched to play a game over a resource. In addition, all individuals have random access to a constant background resource, and their fitness or wealth depreciates over time. For b…

2017-07-04abs ↗pdf ↗

USP test improves on Pearson's chi-squared and GG-test for independence.

problem Deficiencies in Pearson's chi-squared and GG-test for independence.
method USP test based on UU-statistic estimator of population dependence measure.
result USP test controls size, handles small cell counts, and detects minimal violations of independence.

We consider the problem of approximating the set of eigenvalues of the covariance matrix of a multivariate distribution (equivalently, the problem of approximating the "population spectrum"), given access to samples drawn from the distribution. The eigenvalues of the covariance of a distribution contain basic informati…

2016-01-30abs ↗pdf ↗