In the practice of point prediction, it is desirable that forecasters receive a directive in the form of a statistical functional, such as the mean or a quantile of the predictive distribution. When evaluating and comparing competing forecasts, it is then critical that the scoring function used for these purposes be co…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A framework for sensitivity measures using scoring functions.
Observational cohort studies with oversampled exposed subjects are typically implemented to understand the causal effect of a rare exposure. Because the distribution of exposed subjects in the sample differs from the source population, estimation of a propensity score function (i.e., probability of exposure given basel…
Identification and scoring functions are statistical tools to assess the calibration and the relative performance of risk measure estimates, e.g., in backtesting. A risk measures is called identifiable (elicitable) it it admits a strict identification function (strictly consistent scoring function). We consider measure…
A statistical functional, such as the mean or the median, is called elicitable if there is a scoring function or loss function such that the correct forecast of the functional is the unique minimizer of the expected score. Such scoring functions are called strictly consistent for the functional. The elicitability of a …
Robustifies elicitable functionals to handle small distribution misspecifications.
Improves conformal prediction by combining multiple score functions and optimizing weights.
This paper considers fair probabilistic binary classification where the outputs of primary interest are predicted probabilities, commonly referred to as scores. We formulate the problem of transforming scores to satisfy fairness constraints that are linear in conditional means of scores while minimizing a cross-entropy…
Predictions are issued on the basis of certain information. If the forecasting mechanisms are correctly specified, a larger amount of available information should lead to better forecasts. For point forecasts, we show how the effect of increasing the information set can be quantified by using strictly consistent scorin…
The debate of what quantitative risk measure to choose in practice has mainly focused on the dichotomy between Value at Risk (VaR) -- a quantile -- and Expected Shortfall (ES) -- a tail expectation. Range Value at Risk (RVaR) is a natural interpolation between these two prominent risk measures, which constitutes a trad…
We address the problem of learning vector representations for entities and relations in Knowledge Graphs (KGs) for Knowledge Base Completion (KBC). This problem has received significant attention in the past few years and multiple methods have been proposed. Most of the existing methods in the literature use a predefin…
Novel estimator reduces diffusion model variance.
Paper explores connections between loss functions and consistency in binary classification and regression.
Unified score and distance-based GoF tests for model adequacy.
Paper introduces new loss functions for multi-class abstention learning.
Gini index needs auto-calibration for consistent decision-making.
CTM improves diffusion model sampling quality with efficient ODE traversal.
In statistical analysis, measuring a score of predictive performance is an important task. In many scientific fields, appropriate scores were tailored to tackle the problems at hand. A proper score is a popular tool to obtain statistically consistent forecasts. Furthermore, a mathematical characterization of the proper…
The paper analyzes elicitability of return risk measures and their scoring functions.
Proposes methods to estimate posterior probability and propensity score functions without assuming constant propensity score.
Topological anomaly scores predict return curves in S&P 500 stocks
Score matching is a popular method for estimating unnormalized statistical models. However, it has been so far limited to simple, shallow models or low-dimensional data, due to the difficulty of computing the Hessian of log-density functions. We show this difficulty can be mitigated by projecting the scores onto random…
We give a new consistent scoring function for structure learning of Bayesian networks. In contrast to traditional approaches to scorebased structure learning, such as BDeu or MDL, the complexity penalty that we propose is data-dependent and is given by the probability that a conditional independence test correctly show…
Develops a robust method for image reconstruction from limited data.
EnScale learns to downscale climate models efficiently, capturing both spatial and temporal consistency.
MAS scores cluster size consistency from points, robust to label changes.
Classifies intrinsically linked tournaments by their score sequences.
Score matching fails for general point processes, a new estimator improves accuracy.
This work extends diffusion models to function space for better generative modeling.
This guide clarifies techniques for assessing and comparing model calibration and performance.
New scoring rules compare probabilistic top lists in classification.
A new method optimizes anomaly scoring from score distribution to improve AD performance.
The paper examines the consistency of item embeddings in recommendation systems.
Generative Adversarial Networks (GANs) are known to be difficult to train, despite considerable research effort. Several regularization techniques for stabilizing training have been proposed, but they introduce non-trivial computational overheads and interact poorly with existing techniques like spectral normalization.…
A new method for training diffusion models using likelihood matching.
Graph-based semi-supervised learning is one of the most popular methods in machine learning. Some of its theoretical properties such as bounds for the generalization error and the convergence of the graph Laplacian regularizer have been studied in computer science and statistics literatures. However, a fundamental stat…
WS diffusion models handle anisotropic Gaussian noise better than conventional methods.
Paper efficiently infers differential parameters in time-varying models using time score matching.
Learning how to rank multivariate unlabeled observations depending on their degree of abnormality/novelty is a crucial problem in a wide range of applications. In practice, it generally consists in building a real valued "scoring" function on the feature space so as to quantify to which extent observations should be co…
CNP improves few-shot learning for docking scores in molecular datasets.
Knowing when a classifier's prediction can be trusted is useful in many applications and critical for safely using AI. While the bulk of the effort in machine learning research has been towards improving classifier performance, understanding when a classifier's predictions should and should not be trusted has received …
Boosting method for causal SEMs from observational data.
A faster method for density estimation using denoising score matching with random Fourier features.
CCE improves anomaly detection metrics by measuring both confidence and consistency.
Proposes measures for uncertainty quantification using proper scoring rules.
New score-based methods identify causal structures with latent variables.
Recent work has increased the performance of Generative Adversarial Networks (GANs) by enforcing a consistency cost on the discriminator. We improve on this technique in several ways. We first show that consistency regularization can introduce artifacts into the GAN samples and explain how to fix this issue. We then pr…
We give a new consistent scoring function for structure learning of Bayesian networks. In contrast to traditional approaches to score-based structure learning, such as BDeu or MDL, the complexity penalty that we propose is data-dependent and is given by the probability that a conditional independence test correctly sho…