BayCANN uses ANN to speed up Bayesian calibration in health sciences.
problem Bayesian calibration's practical and computational burdens in health decision sciences.
method BayCANN trains an ANN metamodel to calibrate parameters probabilistically, comparing accuracy and speed to direct Bayesian calibration.
result BayCANN is more accurate and faster than direct Bayesian calibration methods.
Proposes PQ, a more precise Bayesian quantifier for prevalence estimation.
problem Uncertainty quantification in prevalence estimation.
method Bayesian quantification methods, focusing on precision and coverage.
result PQ provides more precise and well-calibrated uncertainty quantification.
The paper discusses thresholds and bounds for accuracy in binary classification systems.
problem The accuracy of binary classification systems and its dependence on prevalence.
method Analyzing the precision-prevalence curve and negative predictive value-prevalence curve to find thresholds and bounds.
result Thresholds (φe and φn) bound various accuracy metrics (Fβ, F1, FM, MCC) and the ratio of maximum accuracy to prevalence. Measures policy-violating content prevalence with ML-assisted sampling and LLM labeling.
problem Accurate measurement of content violations that are often rare and costly to label.
method Design-based measurement system using ML-assisted probability sampling and LLM labeling.
result Produces unbiased prevalence estimates with confidence intervals and dashboard drilldowns.
Paper explores how unsupervised learning can be understood through linear algebra concepts.
problem Understanding unsupervised learning through linear algebra concepts.
method Introducing the concept of linearly independent populations and using them to solve for prevalence values.
result Unsupervised learning can be realized as a generalization of supervised learning.
Point estimation of class prevalences in the presence of data set shift has been a popular research topic for more than two decades. Less attention has been paid to the construction of confidence and prediction intervals for estimates of class prevalences. One little considered question is whether or not it is necessar…
Bayesian method corrects bias in imbalanced datasets.
problem Prevalence bias in machine learning datasets.
method Bayesian risk minimization framework, bias-corrected loss function.
result Corrected loss function improves model performance.
This study connects prevalence and machine learning for diagnostic testing.
problem Uncertainty quantification in machine learning for diagnostic tests.
method Developed a numerical homotopy algorithm to estimate classification boundaries and quantify uncertainty.
result The proposed method stabilizes uncertainty quantification in machine learning for diagnostic tests.
Study finds AUC is most consistent across different prevalence in binary classification.
problem Consistency of model evaluation metrics across varying prevalence in binary classification.
method Analysis of 156 data scenarios with 18 metrics, 5 models, and a random guess model.
result AUC has the smallest variance in evaluating individual models and ranking of models.
Bayesian methods improve group testing for identifying infected patients.
problem Identifying infected patients from group testing results with false positives.
method Bayesian inference and belief propagation algorithm, combined with expectation-maximization method.
result True-positive rate improved by considering credible intervals.
The paper proves prevalent existence and partially determines moduli space of area-minimizing surfaces with fractal singular sets.
problem Existence and moduli space of area-minimizing surfaces with fractal singular sets.
method Proof of prevalent existence, determination of moduli space, refinement of strata.
result Sharp results on moduli space and refinement of strata, showing fractal singularities do not completely dissolve under generic perturbations.
Develops RES metrics for stable rare-event forecasting evaluation.
problem Challenges in evaluating forecasts of rare events.
method Rare-event-stable (RES) metrics designed to maintain stable thresholds under extreme rarity.
result RES metrics maintain stable thresholds, consistent model rankings, and near-complete prevalence invariance.
New method adapts to structural shifts in graph data for better label prevalence estimation.
problem Structural shifts in graph data affect label prevalence estimation.
method Importance sampling variant of KDEy quantification approach.
result Adapts to structural shifts and outperforms standard approaches.
Estimates disease prevalence using non-ignorable missing data in health surveys.
problem Estimating disease prevalence in non-representative samples with non-ignorable missing data.
method Connects auxiliary proxy variable framework to label shift setting, uses high-dimensional covariates without generative models.
result Fails to account for non-ignorable missingness can lead to significant misestimations.
The estimation of class prevalence, i.e., the fraction of a population that belongs to a certain class, is a very useful tool in data analytics and learning, and finds applications in many domains such as sentiment analysis, epidemiology, etc. For example, in sentiment analysis, the objective is often not to estimate w…
The Centers for Disease Control and Prevention (CDC) coordinates a labor-intensive process to measure the prevalence of autism spectrum disorder (ASD) among children in the United States. Random forests methods have shown promise in speeding up this process, but they lag behind human classification accuracy by about 5%…
Reconstruction error is a prevalent score used to identify anomalous samples when data are modeled by generative models, such as (variational) auto-encoders or generative adversarial networks. This score relies on the assumption that normal samples are located on a manifold and all anomalous samples are located outside…
HistNetQ improves quantification tasks by optimizing loss functions and eliminating label requirements.
problem Quantification of class prevalence in bags of examples.
method Permutation-invariant Histograms and deep neural networks.
result HistNetQ outperforms other quantification methods and optimizes custom loss functions.
Novel unsupervised scheme for highly imbalanced and overlapping datasets.
problem Highly imbalanced and overlapping classes in medical datasets.
method Unsupervised domain adaptation scheme based on Quantification.
result High quality results for Quantification and Domain Adaptation.
New model maps malaria prevalence across Kenya's changing administrative boundaries.
problem Mapping disease prevalence with changing administrative boundaries.
method Combines deep learning and MCMC with aggVAE for disease mapping.
result Solves the change-of-support problem in disease surveillance.
Investors in Bitcoin exhibit the disposition effect, selling winners and holding losers.
problem The disposition effect in cryptoassets, specifically Bitcoin.
method Using transaction data from cryptoasset exchanges, the study investigated Bitcoin investors' behavior.
result Bitcoin investors exhibit the disposition effect, with intensity varying over time.
New method uses Transformers for flu forecasting.
problem Forecasting influenza-like illness trends.
method Transformer-based machine learning models with self-attention.
result Forecasting results are competitive with state-of-the-art methods.
Novel framework detects CKD in diabetic patients using sparse EHR representations.
problem Early detection of CKD in diabetic patients.
method Sparse longitudinal representations of EHR data.
result Proposed model achieves higher predictive performance than baselines.
We use methods from network science to analyze corruption risk in a large administrative dataset of over 4 million public procurement contracts from European Union member states covering the years 2008-2016. By mapping procurement markets as bipartite networks of issuers and winners of contracts we can visualize and de…
Plasmodium falciparum malaria still poses one of the greatest threats to human life with over 200 million cases globally leading to half-million deaths annually. Of these, 90% of cases and of the mortality occurs in sub-Saharan Africa, mostly among children. Although malaria prediction systems are central to the 2016-2…
Proposes a method to improve rare event prediction in healthcare.
problem Rare event classification in healthcare with low prevalence labels.
method Variational disentanglement approach to semi-parametric learning.
result Outperforms existing alternatives in mortality prediction on COVID-19 cohort.
Paper speeds up topological signal identification and cycle matching.
problem Efficiently identifying and matching topological signals across datasets.
method Cohomological approach to persistent homology computation.
result Significantly faster performance on large-scale datasets.
We propose a discrete surface theory in R3 that unites the most prevalent versions of discrete special parametrizations. This theory encapsulates a large class of discrete surfaces given by a Lax representation and, in particular, the one-parameter associated families of constant curvature surfaces. The theo…
The quantification problem consists of determining the prevalence of a given label in a target population. However, one often has access to the labels in a sample from the training population but not in the target population. A common assumption in this situation is that of prior probability shift, that is, once the la…
Model compression has been widely adopted to obtain light-weighted deep neural networks. Most prevalent methods, however, require fine-tuning with sufficient training data to ensure accuracy, which could be challenged by privacy and security issues. As a compromise between privacy and performance, in this paper we inve…
Extends diffusion models to handle exponential family distributions for inverse problems.
problem Intractability of likelihood score for non-Gaussian observations.
method Evidence trick to approximate likelihood score for exponential family distributions.
result Effective Bayesian inference on complex Poisson processes and malaria prevalence prediction.
New method speeds up deep learning optimization.
problem Scalable second-order optimization for deep learning.
method Second-order optimization with algorithmic and numerical improvements.
result Significant convergence and wall-clock time improvements.
Geometry-aware KDE model improves multiclass quantification.
problem Accurately estimating class prevalence for label shift adaptation.
method Log-ratio representations and Aitchison geometry for compositional data, shrinkage regularization.
result Competitive with state-of-the-art quantifiers, often improving over standard KDE-based baselines.
Continuous Sweep improves binary quantifier performance.
problem Estimating class prevalence in datasets.
method Parametric binary quantifier inspired by Median Sweep, using parametric class distributions and mean of Adjusted Count estimates.
result Continuous Sweep outperforms other quantifiers in simulations and empirical data analysis.
In dynamic topic modeling, the proportional contribution of a topic to a document depends on the temporal dynamics of that topic's overall prevalence in the corpus. We extend the Dynamic Topic Model of Blei and Lafferty (2006) by explicitly modeling document level topic proportions with covariates and dynamic structure…
Optimizes classification algorithms with bounds on error rates.
problem Bounding uncertainties in classifier outputs for diagnostic testing.
method Set-theoretic and probabilistic arguments to derive uniform error bounds.
result Optimal partition minimizes the largest Gershgorin radius of the confusion matrix.
This article addresses persistent tangles. These are tangles whose presence in a knot diagram forces that diagram to be knotted. We provide new methods for constructing persistent tangles. Our techniques rely mainly on the existence of non-trivial colorings for the tangles in question. Our main result in this article i…
New conformal prediction methods for long-tailed classification problems.
problem Rare classes are systematically omitted in existing conformal prediction methods.
method Introduced a new conformal score function and a new interpolation procedure.
result Smoothly trade off set size and class-conditional coverage.
Classification is the task of predicting the class labels of objects based on the observation of their features. In contrast, quantification has been defined as the task of determining the prevalences of the different sorts of class labels in a target dataset. The simplest approach to quantification is Classify & Count…
We show uniqueness of cylindrical blowups for mean curvature flow in all dimension and all codimension. Cylindrical singularities are known to be the most important; they are the most prevalent in any codimension. Mean curvature flow in higher codimension is a nonlinear parabolic system where many of the methods used f…
In this short note we extend some of the recent results on matrix completion under the assumption that the columns of the matrix can be grouped (clustered) into subspaces (not necessarily disjoint or independent). This model deviates from the typical assumption prevalent in the literature dealing with compression and r…
Study shows almost complex structures with certain tensor properties are prevalent.
problem Characterizing almost complex structures with specific tensor properties.
method Analyzes the space of almost complex structures on compact manifolds.
result The space of almost complex structures with rank at least k Nijenhuis tensor is either empty or dense in each component.
Analyzes securitization impacts on monetary and fiscal policies.
problem Impact of securitization on monetary and fiscal policies.
method Develops optimal conditions, identifies constraints, introduces new decision models.
result Identifies constraints and interactions of securitization with capital-reserve requirements.
New proof shows rapid mixing for random walks on nilmanifolds.
problem Proving rapid mixing for random walks on nilmanifolds.
method Proved rapid mixing for almost all random walks generated by m translations on nilmanifolds under mild assumptions.
result For several classical classes of nilmanifolds, m=2 suffices for rapid mixing.
Estimates proportions of LLM-generated text in mixed documents.
problem Estimating the proportion of text generated by a pre-specified LLM in mixed documents.
method Developed estimators for two observation regimes: full observation and pivotal reduction, and established sample complexity bounds.
result Full observation estimators require fewer samples than pivotal reduction estimators.
This paper describes a simple framework for structured sparse recovery based on convex optimization. We show that many structured sparsity models can be naturally represented by linear matrix inequalities on the support of the unknown parameters, where the constraint matrix has a totally unimodular (TU) structure. For …
Non-atomic arbitrage exploits price differences on Ethereum and other blockchains, accounting for over 10% of Ethereum's block value.
problem Price differences on decentralized exchanges and centralized exchanges lead to MEV.
method Analyzed non-atomic arbitrage on Ethereum's largest DEXes, identifying its prevalence and impact.
result More than 10% of Ethereum's block value is attributed to non-atomic arbitrage, involving over $132 billion.
Matrix approximation is a common tool in machine learning for building accurate prediction models for recommendation systems, text mining, and computer vision. A prevalent assumption in constructing matrix approximations is that the partially observed matrix is of low-rank. We propose a new matrix approximation model w…