Estimates unknown population sizes using the hypergeometric distribution.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Polyak step size GD reaches final radius of convergence after log iterations.
We study the top- ranking problem where the goal is to recover the set of top- ranked items out of a large collection of items based on partially revealed preferences. We consider an adversarial crowdsourced setting where there are two population sets, and pairwise comparison samples drawn from one of the populat…
Estimates population profile from small random samples.
GNMC reduces XCSF population size while preserving function approximation and policy accuracy.
This paper highlights the size-dependency of income distributions, i.e. the income distribution curves versus the population of a country systematically. By using the generalized Lotka-Volterra model to fit the empirical income data in the United States during 1996-2007, we found an important parameter can scale wi…
Improved online prediction with guaranteed coverage.
We are losing biodiversity at an unprecedented scale and in many cases, we do not even know the basic data for the species. Traditional methods for wildlife monitoring are inadequate. Development of new computer vision tools enables the use of images as the source of information about wildlife. Social media is the rich…
Optimizes sampling in continuous domains by adjusting search distribution.
New method pools labels from similar data items to improve learning from small samples.
The paper reveals a spinning top geometry in real-world games.
A clustering method for multivariate populations with similar dependence structures.
Direct measurements of Gini coefficients by conventional arithmetic calculations are a poor estimator, even if paradoxically, they include the entire population, as because of super-additivity they cannot lend themselves to comparisons between units of different size, and intertemporal analyses are vitiated by the popu…
Most real-world networks are too large to be measured or studied directly and there is substantial interest in estimating global network properties from smaller sub-samples. One of the most important global properties is the number of vertices/nodes in the network. Estimating the number of vertices in a large network i…
Insiders camouflage trading to balance wealth and stealth, avoiding legal penalties.
In an adaptive population which models financial markets and distributed control, we consider how the dynamics depends on the diversity of the agents' initial preferences of strategies. When the diversity decreases, more agents tend to adapt their strategies together. This change in the environment results in dynamical…
In lowest unique bid auctions, players bid for an item. The winner is whoever places the \emph{lowest} bid, provided that it is also unique. We use a grand canonical approach to derive an analytical expression for the equilibrium distribution of strategies. We then study the properties of the solution as a function…
Study shows finite agent equilibrium converges to mean-field limit in asset pricing.
We investigate the problems of identity and closeness testing over a discrete population from random samples. Our goal is to develop efficient testers while guaranteeing Differential Privacy to the individuals of the population. We describe an approach that yields sample-efficient differentially private testers for the…
A government has to finance a risk for its population. It shares the charges among the population with a fixed scale based on economic criteria. Various organisms have to collect and to redistribute fairly the subsidies. Under these conditions, when the size of the organisms is varied, the distribution's laws of the cr…
Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.
Developing a scientific understanding of cities in a fast urbanizing world is essential for planning sustainable urban systems. Recently, it was shown that income and wealth creation follow increasing returns, scaling superlinearly with city size. We study scaling of per capita incomes for separate census defined incom…
NPE trains neural networks to approximate posterior distributions in SIR models from final outcome data.
We present, solve and numerically simulate a simple model that describes the consequences of increased longevity on fertility rates, population growth and the distribution of wealth in developed societies. We look at the consequences of the repeated use of life extension techniques and show that they represent a novel …
Study shows income inequality increases with city size, affecting only the wealthiest deciles.
New estimates for the population risk are established for two-layer neural networks. These estimates are nearly optimal in the sense that the error rates scale in the same way as the Monte Carlo error rates. They are equally effective in the over-parametrized regime when the network size is much larger than the size of…
This paper presents a study on discriminative artificial neural network classifiers in the context of open-set speaker identification. Both 2-class and multi-class architectures are tested against the conventional Gaussian mixture model based classifier on enrolled speaker sets of different sizes. The performance evalu…
Estimates population size using capture-recapture designs with binary indicators.
Copula-based method generates synthetic populations from marginal distributions.
In this paper, we consider the problem of partitioning a small data sample drawn from a mixture of product distributions. We are interested in the case that individual features are of low average quality , and we want to use as few of them as possible to correctly partition the sample. We analyze a spectral tech…
This paper studies the landscape of empirical risk of deep neural networks by theoretically analyzing its convergence behavior to the population risk as well as its stationary points and properties. For an -layer linear neural network, we prove its empirical risk uniformly converges to its population risk at the rat…
Develops an equilibrium model for securities pricing in a mixed cooperative and non-cooperative market.
Paper improves privacy-preserving optimization rates for convex functions.
We introduce an evolutionary algorithm called recombinator--means for optimizing the highly non-convex kmeans problem. Its defining feature is that its crossover step involves all the members of the current generation, stochastically recombining them with a repurposed variant of the -means++ seeding algorithm. Th…
We develop amortized population Gibbs (APG) samplers, a class of scalable methods that frames structured variational inference as adaptive importance sampling. APG samplers construct high-dimensional proposals by iterating over updates to lower-dimensional blocks of variables. We train each conditional proposal by mini…
Study optimizes data collection from biased, costly sources to minimize risk.
The classical asymptotic theory for parametric -estimators guarantees that, in the limit of infinite sample size, the excess risk has a chi-square type distribution, even in the misspecified case. We demonstrate how self-concordance of the loss allows to characterize the critical sample size sufficient to guarantee …
We introduce gym-city, a Reinforcement Learning environment that uses SimCity 1's game engine to simulate an urban environment, wherein agents might seek to optimize one or a combination of any number of city-wide metrics, on gameboards of various sizes. We focus on population, and analyze our agents' ability to genera…
Machine learning's predictive power is limited by sample size, as shown by the Limits-to-Learning Gap.
In statistical connectomics, the quantitative study of brain networks, estimating the mean of a population of graphs based on a sample is a core problem. Often, this problem is especially difficult because the sample or cohort size is relatively small, sometimes even a single subject. While using the element-wise sampl…
This study analyzes LTS in sparse models with finite sample error bounds.
The problem of population recovery refers to estimating a distribution based on incomplete or corrupted samples. Consider a random poll of sample size conducted on a population of individuals, where each pollee is asked to answer binary questions. We consider one of the two polling impediments: (a) in lossy pop…
New model estimates species population trends from citizen science data.
In clinical and neuroscientific studies, systematic differences between two populations of brain networks are investigated in order to characterize mental diseases or processes. Those networks are usually represented as graphs built from neuroimaging data and studied by means of graph analysis methods. The typical mach…
Many of the recent triumphs in machine learning are dependent on well-tuned hyperparameters. This is particularly prominent in reinforcement learning (RL) where a small change in the configuration can lead to failure. Despite the importance of tuning hyperparameters, it remains expensive and is often done in a naive an…
The accumulation of individual fitness or wealth is modelled as a population game in which pairs of individuals are recurrently and randomly matched to play a game over a resource. In addition, all individuals have random access to a constant background resource, and their fitness or wealth depreciates over time. For b…
USP test improves on Pearson's chi-squared and -test for independence.
We consider the problem of approximating the set of eigenvalues of the covariance matrix of a multivariate distribution (equivalently, the problem of approximating the "population spectrum"), given access to samples drawn from the distribution. The eigenvalues of the covariance of a distribution contain basic informati…