Combines cost-sensitive and Neyman-Pearson paradigms for better binary classification.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We investigate the problems of identity and closeness testing over a discrete population from random samples. Our goal is to develop efficient testers while guaranteeing Differential Privacy to the individuals of the population. We describe an approach that yields sample-efficient differentially private testers for the…
The paper examines the stability of binary choice models using Gini index and scoring indicators.
We address the problem of correcting group discriminations within a score function, while minimizing the individual error. Each group is described by a probability density function on the set of profiles. We first solve the problem analytically in the case of two populations, with a uniform bonus-malus on the zones whe…
SDRF estimates complex survey designs for conditional distributions.
In this work, we systematically investigate mean field games and mean field type control problems with multiple populations using a coupled system of forward-backward stochastic differential equations of McKean-Vlasov type stemming from Pontryagin's stochastic maximum principle. Although the same cost functions as well…
In clinical and neuroscientific studies, systematic differences between two populations of brain networks are investigated in order to characterize mental diseases or processes. Those networks are usually represented as graphs built from neuroimaging data and studied by means of graph analysis methods. The typical mach…
We show that model compression can improve the population risk of a pre-trained model, by studying the tradeoff between the decrease in the generalization error and the increase in the empirical risk with model compression. We first prove that model compression reduces an information-theoretic bound on the generalizati…
In this work, we consider hypothesis testing and anomaly detection on datasets where each observation is a weighted network. Examples of such data include brain connectivity networks from fMRI flow data, or word co-occurrence counts for populations of individuals. Current approaches to hypothesis testing for weighted n…
Infinitesimal boosting converges to a deterministic process in large sample limit.
GNMC reduces XCSF population size while preserving function approximation and policy accuracy.
Paired estimation of change in parameters of interest over a population plays a central role in several application domains including those in the social sciences, epidemiology, medicine and biology. In these domains, the size of the population under study is often very large, however, the number of observations availa…
The paper explores why a specific type of predictor works well in noisy data.
Estimates population profile from small random samples.
D2SRM solves complex PDEs using deep learning.
New TTP framework fuses control arms while controlling Type-I error.
Meta learning of optimal classifier error rates allows an experimenter to empirically estimate the intrinsic ability of any estimator to discriminate between two populations, circumventing the difficult problem of estimating the optimal Bayes classifier. To this end we propose a weighted nearest neighbor (WNN) graph es…
The paper analyzes how data augmentation affects the test error in regression models.
Estimating the largest community in a mixed population via sequential sampling.
Direct measurements of Gini coefficients by conventional arithmetic calculations are a poor estimator, even if paradoxically, they include the entire population, as because of super-additivity they cannot lend themselves to comparisons between units of different size, and intertemporal analyses are vitiated by the popu…
Study shows more frequent communication in FL reduces model's generalization power.
Optimal a priori estimates are derived for the population risk, also known as the generalization error, of a regularized residual network model. An important part of the regularized model is the usage of a new path norm, called the weighted path norm, as the regularization term. The weighted path norm treats the skip c…
Novel graph theory for neural networks improves understanding of their structure and performance.
In many statistical problems, a more coarse-grained model may be suitable for population-level behaviour, whereas a more detailed model is appropriate for accurate modelling of individual behaviour. This raises the question of how to integrate both types of models. Methods such as posterior regularization follow the id…
Many modern data analysis problems involve inferences from streaming data. However, streaming data is not easily amenable to the standard probabilistic modeling approaches, which assume that we condition on finite data. We develop population variational Bayes, a new approach for using Bayesian modeling to analyze strea…
Study evaluates different mathematical models for three case studies using statistical fitting.
In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we develop an inference method that is resilient to sampling biases and is able to …
The paper proposes a new method for comparing logistic regression models across different populations.
This study improves audit sampling by using sequential procedures with statistical guarantees.
We study the total least squares (TLS) problem that generalizes least squares regression by allowing measurement errors in both dependent and independent variables. TLS is widely used in applied fields including computer vision, system identification and econometrics. The special case when all dependent and independent…
Proposes a federated transfer learning method to improve precision medicine models for underrepresented populations.
Aims to describe neural network training dynamics using two-time-scale models.
Copula-based method generates synthetic populations from marginal distributions.
Develops a fair classifier for deep learning models.
New estimates for the population risk are established for two-layer neural networks. These estimates are nearly optimal in the sense that the error rates scale in the same way as the Monte Carlo error rates. They are equally effective in the over-parametrized regime when the network size is much larger than the size of…
Robust data fusion via subsampling improves model performance for rare data types.
With a growing interest in using non-representative samples to train prediction models for numerous outcomes it is necessary to account for the sampling design that gives rise to the data in order to assess the generalized predictive utility of a proposed prediction rule. After learning a prediction rule based on a non…
Measures of wealth and production have been found to scale superlinearly with the population of a city. Therefore, it makes economic sense for humans to congregate together in dense settlements. A recent model of population dynamics showed that population growth can become superexponential due to the superlinear scalin…
New findings on optimization landscape of Toeplitz covariance estimation.
The paper develops a minimax optimal method for high-dimensional regression using auxiliary data.
While studying response trajectory, often the population of interest may be diverse enough to exist distinct subgroups within it and the longitudinal change in response may not be uniform in these subgroups. That is, the timeslope and/or influence of covariates in longitudinal profile may vary among these different sub…
New algorithm for countable bandits with optimal regret.
New methods improve inference after prediction without strong model assumptions.
New bounds on machine learning model generalization error moments.
Develops asymptotic theory for deep Cox models to enable valid inference.
New method estimates treatment effects across different populations.
Policy mirror ascent achieves Nash equilibrium in mean field games without a population generative model.
RKD improves clustering in semi-supervised learning with limited labels.