Proposes PQ, a more precise Bayesian quantifier for prevalence estimation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper proves prevalent existence and partially determines moduli space of area-minimizing surfaces with fractal singular sets.
The paper discusses thresholds and bounds for accuracy in binary classification systems.
Measures policy-violating content prevalence with ML-assisted sampling and LLM labeling.
Paper explores how unsupervised learning can be understood through linear algebra concepts.
Point estimation of class prevalences in the presence of data set shift has been a popular research topic for more than two decades. Less attention has been paid to the construction of confidence and prediction intervals for estimates of class prevalences. One little considered question is whether or not it is necessar…
Bayesian method corrects bias in imbalanced datasets.
This study connects prevalence and machine learning for diagnostic testing.
The estimation of class prevalence, i.e., the fraction of a population that belongs to a certain class, is a very useful tool in data analytics and learning, and finds applications in many domains such as sentiment analysis, epidemiology, etc. For example, in sentiment analysis, the objective is often not to estimate w…
Study finds AUC is most consistent across different prevalence in binary classification.
Paper speeds up topological signal identification and cycle matching.
Bayesian methods improve group testing for identifying infected patients.
HistNetQ improves quantification tasks by optimizing loss functions and eliminating label requirements.
This article addresses persistent tangles. These are tangles whose presence in a knot diagram forces that diagram to be knotted. We provide new methods for constructing persistent tangles. Our techniques rely mainly on the existence of non-trivial colorings for the tangles in question. Our main result in this article i…
Proposes a method to improve rare event prediction in healthcare.
Analyzes securitization impacts on monetary and fiscal policies.
Geometry-aware KDE model improves multiclass quantification.
Develops RES metrics for stable rare-event forecasting evaluation.
New method adapts to structural shifts in graph data for better label prevalence estimation.
Estimates disease prevalence using non-ignorable missing data in health surveys.
New conformal prediction methods for long-tailed classification problems.
The Centers for Disease Control and Prevention (CDC) coordinates a labor-intensive process to measure the prevalence of autism spectrum disorder (ASD) among children in the United States. Random forests methods have shown promise in speeding up this process, but they lag behind human classification accuracy by about 5%…
New method uses momentum to converge in DC optimization with small batches.
Reconstruction error is a prevalent score used to identify anomalous samples when data are modeled by generative models, such as (variational) auto-encoders or generative adversarial networks. This score relies on the assumption that normal samples are located on a manifold and all anomalous samples are located outside…
Fidel-TS creates a new benchmark for time series forecasting models.
CDSSL improves representation quality by integrating linear and nonlinear dependencies.
Novel unsupervised scheme for highly imbalanced and overlapping datasets.
New method targets vaccines for new variants using Thompson sampling.
New model maps malaria prevalence across Kenya's changing administrative boundaries.
Investors in Bitcoin exhibit the disposition effect, selling winners and holding losers.
New method uses Transformers for flu forecasting.
Develops a model for RNA-seq data clustering.
Several applications of Reinforcement Learning suffer from instability due to high variance. This is especially prevalent in high dimensional domains. Regularization is a commonly used technique in machine learning to reduce variance, at the cost of introducing some bias. Most existing regularization techniques focus o…
Graph Convolutional Networks (GCNs) have recently been shown to be quite successful in modeling graph-structured data. However, the primary focus has been on handling simple undirected graphs. Multi-relational graphs are a more general and prevalent form of graphs where each edge has a label and direction associated wi…
Novel framework detects CKD in diabetic patients using sparse EHR representations.
We use methods from network science to analyze corruption risk in a large administrative dataset of over 4 million public procurement contracts from European Union member states covering the years 2008-2016. By mapping procurement markets as bipartite networks of issuers and winners of contracts we can visualize and de…
Plasmodium falciparum malaria still poses one of the greatest threats to human life with over 200 million cases globally leading to half-million deaths annually. Of these, 90% of cases and of the mortality occurs in sub-Saharan Africa, mostly among children. Although malaria prediction systems are central to the 2016-2…
Computer Vision and machine learning methods were previously used to reveal screen presence of genders in TV and movies. In this work, using head pose, gender detection, and skin color estimation techniques, we demonstrate that the gender disparity in TV in a South Asian country such as Bangladesh exhibits unique chara…
Determining risk contributions of unit exposures to portfolio-wide economic capital is an important task in financial risk management. Computing risk contributions involves difficulties caused by rare-event simulations. In this study, we address the problem of estimating risk contributions when the total risk is measur…
Proposes a deep neural network for event intensity estimation.
Time series data are prevalent in electronic health records, mostly in the form of physiological parameters such as vital signs and lab tests. The patterns of these values may be significant indicators of patients' clinical states and there might be patterns that are unknown to clinicians but are highly predictive of s…
New framework explains normalizing flows' power and limitations.
Recent advances in computing power and the potential to make more realistic assumptions due to increased flexibility have led to the increased prevalence of simulation models in economics. While models of this class, and particularly agent-based models, are able to replicate a number of empirically-observed stylised fa…
Time series of graphs are increasingly prevalent in modern data and pose unique challenges to visual exploration and pattern extraction. This paper describes the development and application of matrix factorizations for exploration and time-varying community detection in time-evolving graph sequences. The matrix factori…
Improves SSL with doubly robust estimation of unlabeled class distribution.
Cluster analysis is a fundamental tool for pattern discovery of complex heterogeneous data. Prevalent clustering methods mainly focus on vector or matrix-variate data and are not applicable to general-order tensors, which arise frequently in modern scientific and business applications. Moreover, there is a gap between …
We consider the problem of demixing a sequence of source signals from the sum of noisy bilinear measurements. It is a generalized mathematical model for blind demixing with blind deconvolution, which is prevalent across the areas of dictionary learning, image processing, and communications. However, state-of- the-art c…
We propose a discrete surface theory in that unites the most prevalent versions of discrete special parametrizations. This theory encapsulates a large class of discrete surfaces given by a Lax representation and, in particular, the one-parameter associated families of constant curvature surfaces. The theo…