Estimates unknown population sizes using the hypergeometric distribution.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Estimates population profile from small random samples.
We are losing biodiversity at an unprecedented scale and in many cases, we do not even know the basic data for the species. Traditional methods for wildlife monitoring are inadequate. Development of new computer vision tools enables the use of images as the source of information about wildlife. Social media is the rich…
Most real-world networks are too large to be measured or studied directly and there is substantial interest in estimating global network properties from smaller sub-samples. One of the most important global properties is the number of vertices/nodes in the network. Estimating the number of vertices in a large network i…
Direct measurements of Gini coefficients by conventional arithmetic calculations are a poor estimator, even if paradoxically, they include the entire population, as because of super-additivity they cannot lend themselves to comparisons between units of different size, and intertemporal analyses are vitiated by the popu…
Improved online prediction with guaranteed coverage.
New estimates for the population risk are established for two-layer neural networks. These estimates are nearly optimal in the sense that the error rates scale in the same way as the Monte Carlo error rates. They are equally effective in the over-parametrized regime when the network size is much larger than the size of…
Estimates population size using capture-recapture designs with binary indicators.
A clustering method for multivariate populations with similar dependence structures.
NPE trains neural networks to approximate posterior distributions in SIR models from final outcome data.
New model estimates species population trends from citizen science data.
Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.
Meta learning of optimal classifier error rates allows an experimenter to empirically estimate the intrinsic ability of any estimator to discriminate between two populations, circumventing the difficult problem of estimating the optimal Bayes classifier. To this end we propose a weighted nearest neighbor (WNN) graph es…
Study shows finite agent equilibrium converges to mean-field limit in asset pricing.
Polyak step size GD reaches final radius of convergence after log iterations.
We study the top- ranking problem where the goal is to recover the set of top- ranked items out of a large collection of items based on partially revealed preferences. We consider an adversarial crowdsourced setting where there are two population sets, and pairwise comparison samples drawn from one of the populat…
TMLE improves IPM estimation for ecological population dynamics.
The paper improves methods for estimating set size using samples.
This study analyzes LTS in sparse models with finite sample error bounds.
Paired estimation of change in parameters of interest over a population plays a central role in several application domains including those in the social sciences, epidemiology, medicine and biology. In these domains, the size of the population under study is often very large, however, the number of observations availa…
The problem of population recovery refers to estimating a distribution based on incomplete or corrupted samples. Consider a random poll of sample size conducted on a population of individuals, where each pollee is asked to answer binary questions. We consider one of the two polling impediments: (a) in lossy pop…
Study optimizes data collection from biased, costly sources to minimize risk.
GNMC reduces XCSF population size while preserving function approximation and policy accuracy.
The classical asymptotic theory for parametric -estimators guarantees that, in the limit of infinite sample size, the excess risk has a chi-square type distribution, even in the misspecified case. We demonstrate how self-concordance of the loss allows to characterize the critical sample size sufficient to guarantee …
Sample measures of top centile contributions to the total (concentration) are downward biased, unstable estimators, extremely sensitive to sample size and concave in accounting for large deviations. It makes them particularly unfit in domains with power law tails, especially for low values of the exponent. These estima…
This research provides theoretical guarantees for hyperparameter estimation in complex network dynamical systems.
We propose a general semi-supervised inference framework focused on the estimation of the population mean. As usual in semi-supervised settings, there exists an unlabeled sample of covariate vectors and a labeled sample consisting of covariate vectors along with real-valued responses ("labels"). Otherwise, the formulat…
The study assesses external validity by evaluating worst-case treatment effects across subpopulations.
In statistical connectomics, the quantitative study of brain networks, estimating the mean of a population of graphs based on a sample is a core problem. Often, this problem is especially difficult because the sample or cohort size is relatively small, sometimes even a single subject. While using the element-wise sampl…
The only input to attain the portfolio weights of global minimum variance portfolio (GMVP) is the covariance matrix of returns of assets being considered for investment. Since the population covariance matrix is not known, investors use historical data to estimate it. Even though sample covariance matrix is an unbiased…
This paper highlights the size-dependency of income distributions, i.e. the income distribution curves versus the population of a country systematically. By using the generalized Lotka-Volterra model to fit the empirical income data in the United States during 1996-2007, we found an important parameter can scale wi…
USP test improves on Pearson's chi-squared and -test for independence.
We consider the problem of approximating the set of eigenvalues of the covariance matrix of a multivariate distribution (equivalently, the problem of approximating the "population spectrum"), given access to samples drawn from the distribution. The eigenvalues of the covariance of a distribution contain basic informati…
Estimates Mozambique's population using remote sensing and microcensus data.
Clustering and community detection with multiple graphs have typically focused on aligned graphs, where there is a mapping between nodes across the graphs (e.g., multi-view, multi-layer, temporal graphs). However, there are numerous application areas with multiple graphs that are only partially aligned, or even unalign…
Optimizes sampling in continuous domains by adjusting search distribution.
New method pools labels from similar data items to improve learning from small samples.
New method improves NN performance across various settings.
The paper reveals a spinning top geometry in real-world games.
A new data set helps estimate continental-scale population distributions.
Develops a method to estimate personalized treatment regimes from summary statistics.
A robust conformal method for set estimation using non-conformity scores.
This paper proposes the k-generalized distribution as a model for describing the distribution and dispersion of income within a population. Formulas for the shape, moments and standard tools for inequality measurement - such as the Lorenz curve and the Gini coefficient - are given. A method for parameter estimation is …
Given samples from a distribution, how many new elements should we expect to find if we continue sampling this distribution? This is an important and actively studied problem, with many applications ranging from unseen species estimation to genomics. We generalize this extrapolation and related unseen estimation proble…
Insiders camouflage trading to balance wealth and stealth, avoiding legal penalties.
In an adaptive population which models financial markets and distributed control, we consider how the dynamics depends on the diversity of the agents' initial preferences of strategies. When the diversity decreases, more agents tend to adapt their strategies together. This change in the environment results in dynamical…
Estimates KL divergence with fairness considerations for sub-populations.
LEARNER improves low-rank matrix estimation using source population data.