Proposes PQ, a more precise Bayesian quantifier for prevalence estimation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper discusses thresholds and bounds for accuracy in binary classification systems.
Measures policy-violating content prevalence with ML-assisted sampling and LLM labeling.
Paper explores how unsupervised learning can be understood through linear algebra concepts.
Point estimation of class prevalences in the presence of data set shift has been a popular research topic for more than two decades. Less attention has been paid to the construction of confidence and prediction intervals for estimates of class prevalences. One little considered question is whether or not it is necessar…
Bayesian method corrects bias in imbalanced datasets.
This study connects prevalence and machine learning for diagnostic testing.
Study finds AUC is most consistent across different prevalence in binary classification.
Bayesian methods improve group testing for identifying infected patients.
The paper proves prevalent existence and partially determines moduli space of area-minimizing surfaces with fractal singular sets.
Develops RES metrics for stable rare-event forecasting evaluation.
New method adapts to structural shifts in graph data for better label prevalence estimation.
Estimates disease prevalence using non-ignorable missing data in health surveys.
The estimation of class prevalence, i.e., the fraction of a population that belongs to a certain class, is a very useful tool in data analytics and learning, and finds applications in many domains such as sentiment analysis, epidemiology, etc. For example, in sentiment analysis, the objective is often not to estimate w…
The Centers for Disease Control and Prevention (CDC) coordinates a labor-intensive process to measure the prevalence of autism spectrum disorder (ASD) among children in the United States. Random forests methods have shown promise in speeding up this process, but they lag behind human classification accuracy by about 5%…
Reconstruction error is a prevalent score used to identify anomalous samples when data are modeled by generative models, such as (variational) auto-encoders or generative adversarial networks. This score relies on the assumption that normal samples are located on a manifold and all anomalous samples are located outside…
HistNetQ improves quantification tasks by optimizing loss functions and eliminating label requirements.
Novel unsupervised scheme for highly imbalanced and overlapping datasets.
New model maps malaria prevalence across Kenya's changing administrative boundaries.
Investors in Bitcoin exhibit the disposition effect, selling winners and holding losers.
New method uses Transformers for flu forecasting.
Novel framework detects CKD in diabetic patients using sparse EHR representations.
We use methods from network science to analyze corruption risk in a large administrative dataset of over 4 million public procurement contracts from European Union member states covering the years 2008-2016. By mapping procurement markets as bipartite networks of issuers and winners of contracts we can visualize and de…
Plasmodium falciparum malaria still poses one of the greatest threats to human life with over 200 million cases globally leading to half-million deaths annually. Of these, 90% of cases and of the mortality occurs in sub-Saharan Africa, mostly among children. Although malaria prediction systems are central to the 2016-2…
Proposes a method to improve rare event prediction in healthcare.
Paper speeds up topological signal identification and cycle matching.
We propose a discrete surface theory in that unites the most prevalent versions of discrete special parametrizations. This theory encapsulates a large class of discrete surfaces given by a Lax representation and, in particular, the one-parameter associated families of constant curvature surfaces. The theo…
The quantification problem consists of determining the prevalence of a given label in a target population. However, one often has access to the labels in a sample from the training population but not in the target population. A common assumption in this situation is that of prior probability shift, that is, once the la…
Model compression has been widely adopted to obtain light-weighted deep neural networks. Most prevalent methods, however, require fine-tuning with sufficient training data to ensure accuracy, which could be challenged by privacy and security issues. As a compromise between privacy and performance, in this paper we inve…
Extends diffusion models to handle exponential family distributions for inverse problems.
New method speeds up deep learning optimization.
Geometry-aware KDE model improves multiclass quantification.
Continuous Sweep improves binary quantifier performance.
In dynamic topic modeling, the proportional contribution of a topic to a document depends on the temporal dynamics of that topic's overall prevalence in the corpus. We extend the Dynamic Topic Model of Blei and Lafferty (2006) by explicitly modeling document level topic proportions with covariates and dynamic structure…
Optimizes classification algorithms with bounds on error rates.
This article addresses persistent tangles. These are tangles whose presence in a knot diagram forces that diagram to be knotted. We provide new methods for constructing persistent tangles. Our techniques rely mainly on the existence of non-trivial colorings for the tangles in question. Our main result in this article i…
New conformal prediction methods for long-tailed classification problems.
Classification is the task of predicting the class labels of objects based on the observation of their features. In contrast, quantification has been defined as the task of determining the prevalences of the different sorts of class labels in a target dataset. The simplest approach to quantification is Classify & Count…
We show uniqueness of cylindrical blowups for mean curvature flow in all dimension and all codimension. Cylindrical singularities are known to be the most important; they are the most prevalent in any codimension. Mean curvature flow in higher codimension is a nonlinear parabolic system where many of the methods used f…
In this short note we extend some of the recent results on matrix completion under the assumption that the columns of the matrix can be grouped (clustered) into subspaces (not necessarily disjoint or independent). This model deviates from the typical assumption prevalent in the literature dealing with compression and r…
Study shows almost complex structures with certain tensor properties are prevalent.
Analyzes securitization impacts on monetary and fiscal policies.
New proof shows rapid mixing for random walks on nilmanifolds.
Estimates proportions of LLM-generated text in mixed documents.
This paper describes a simple framework for structured sparse recovery based on convex optimization. We show that many structured sparsity models can be naturally represented by linear matrix inequalities on the support of the unknown parameters, where the constraint matrix has a totally unimodular (TU) structure. For …
Non-atomic arbitrage exploits price differences on Ethereum and other blockchains, accounting for over 10% of Ethereum's block value.
Matrix approximation is a common tool in machine learning for building accurate prediction models for recommendation systems, text mining, and computer vision. A prevalent assumption in constructing matrix approximations is that the partially observed matrix is of low-rank. We propose a new matrix approximation model w…
The modern data analyst must cope with data encoded in various forms, vectors, matrices, strings, graphs, or more. Consequently, statistical and machine learning models tailored to different data encodings are important. We focus on data encoded as normalized vectors, so that their "direction" is more important than th…