Hierarchical framework for model evaluation on leaderboards
problem Uncertainty and variability in model performance across tasks
method Hierarchical framework with task-level and leaderboard-level rank prediction intervals
result Statistically valid and informative model rank intervals
VLM judges rank well but score poorly; task difficulty and annotation quality affect interval width.
problem VLMs as judges lack reliability indicators in multimodal evaluations.
method Conformal prediction using score-token log-probabilities.
result Evaluation uncertainty is task-dependent, affecting interval width and reliability.
We construct compactifications for median spaces with compact intervals, generalising Roller boundaries of CAT(0) cube complexes. Examples of median spaces with compact intervals include all finite rank median spaces and all proper median spaces of infinite rank. Our methods also work for general median algebra…
Formulates a Dueling Bandits problem for eliciting Kemeny rankings.
problem Eliciting preferences to find a Kemeny ranking.
method Formulates the problem as a Dueling Bandits problem, considering sampling with and without replacement.
result Approximation bounds and algorithms for finding PAC solutions with sample complexity.
Let M be a geometrically finite rank one locally symmetric manifolds. We prove that the spectrum of the Laplace operator on M is finite in a small interval which is optimal.
Atlas models are systems of Ito processes with parameters that depend on rank. We show that the parameters of a simple Atlas model can be identified by measuring the variance of the top-ranked process for different sampling intervals.
A framework for quantifying uncertainty in feature importance values.
problem Stable interpretation of feature importance values in machine learning models.
method A novel method based on pairwise comparisons of feature importance values to produce confidence intervals for feature ranks.
result The method produces simultaneous confidence intervals for feature ranks, enabling selection of top-k important features.
The paper improves ranking by integrating covariates and sparse intrinsic scores.
problem Ranking items with incomplete preference scores explained by covariates.
method Extends BTL model with covariate information and sparse intrinsic scores, using penalized MLE.
result Developed debiased estimator for penalized MLE with distributional properties.
Paper compares Bayesian and de-biased estimators for low-rank matrix completion.
problem Predict missing entries in partially observed matrices.
method Bayesian and de-biased estimators comparison.
result De-biased estimator performs similarly to Bayesian estimators but is more stable and can outperform in small samples.
A Riemannian manifold M has higher hyperbolic rank if every geodesic has a perpendicular Jacobi field making sectional curvature -1 with the geodesic. If in addition, the sectional curvatures of M lie in the interval [−1,−41], and M is closed, we show that M is a locally symmetric space of rank one. This…
In this paper, we propose exact passive-aggressive (PA) online algorithms for learning to rank. The proposed algorithms can be used even when we have interval labels instead of actual labels for examples. The proposed algorithms solve a convex optimization problem at every trial. We find exact solution to those optimiz…
The paper ranks items based on top choices in multiway comparisons.
problem Ranking items based on top choices in multiway comparisons.
method Uniform sampling scheme, statistical rates of convergence, asymptotic normality, maximum likelihood estimator, Gaussian multiplier bootstrap.
result Proposed inference framework for ranking items through maximum pairwise difference statistic.
Proposes a new matrix factorization model for interval-valued matrices.
problem Matrix factorization for matrices with entries in a given interval.
method Bounded simplex-structured matrix factorization (BSSMF) with fast algorithm for missing data.
result BSSMF provides a unique decomposition under certain conditions.
This paper presents a data set describing the evolution of results in the Portuguese Parliamentary Elections of October 6th 2019. The data spans a time interval of 4 hours and 25 minutes, in intervals of 5 minutes, concerning the results of the 27 parties involved in the electoral event. The data set is tailored f…
Bayesian framework improves LLM evaluation stability and transparency.
problem Pass@k and avg@N are unstable and misleading for LLMs.
method Bayesian evaluation with posterior estimates and credible intervals.
result Posterior-based evaluation yields stable and transparent rankings.
We consider cohomogeneity one homogeneous disk bundles and adress the question when these admit a nonnegatively curved invariant metric with normal collar, i.e., such that near the boundary the metric is the product of an interval and a normal homogeneous space. If such a bundle is not (the quotient of) a trivial bundl…
The paper examines how gradient descent stabilizes low-rank matrix factorization in noisy conditions.
problem Stability of low-rank implicit regularization in perturbed deep matrix factorization.
method Derives spectral conditions for gradient descent to exhibit a low-rank phase in noiseless settings and analyzes perturbed dynamics.
result Gradient descent converges to a low-rank solution under perturbation, with explicit dependence on perturbation size.
A common problem in machine learning is to rank a set of n items based on pairwise comparisons. Here ranking refers to partitioning the items into sets of pre-specified sizes according to their scores, which includes identification of the top-k items as the most prominent special case. The score of a given item is defi…
Study uncovers statistical optimality of nonconvex tensor completion methods.
problem Estimating a low-rank tensor from incomplete and corrupted observations.
method Two-stage estimation algorithm for nonconvex optimization.
result Nonconvex tensor completion achieves optimal ℓ2 accuracy. Given a Kaehler group G and a primitive class φ∈H1(G;Z), we show that the rank gradient of (G;φ) is zero if and only if Ker φ is finitely generated. Using this approach, we give a quick proof of the fact (originally due to Napier and Ramachandran) that Kaehler groups are not properly ascending or descending…
Generative AI reduces IR evaluation costs but introduces errors; this work provides reliable CIs.
problem Generating relevance annotations using AI introduces errors that affect IR evaluation metrics.
method Proposes two methods: prediction-powered inference and conformal risk control to place reliable CIs around IR metrics.
result Proposed methods accurately capture both variance and bias in evaluation based on AI-generated annotations.
Unified framework for statistical inference of low-rank tensors.
problem Statistical inference for tensors in high-dimensional data.
method Unified framework using debiasing and tangent space projection.
result Achieves asymptotic normality and minimax-optimal confidence intervals.
Low-rank framework for task-specific LLM ranking from sparse comparisons.
problem Challenges in reliable task-specific ranking of LLMs under sparse, imbalanced comparisons.
method Low-rank modeling of task-by-model ability matrix, max-norm accurate estimator, task-wise top-K recovery guarantees, uncertainty quantification framework.
result Improves sample efficiency and produces tighter, better-calibrated ranking certificates.
Noisy matrix completion aims at estimating a low-rank matrix given only partial and corrupted entries. Despite substantial progress in designing efficient estimation algorithms, it remains largely unclear how to assess the uncertainty of the obtained estimates and how to perform statistical inference on the unknown mat…
LLMs overestimate stock returns and are less accurate at predicting extreme outcomes.
problem Behavioral biases in LLMs' stock return forecasts.
method Comparison of LLM forecasts with crowd-sourced estimates and historical data.
result LLMs overestimate stock returns and are less accurate at predicting extreme outcomes.
Automatic detection of anomalies in space- and time-varying measurements is an important tool in several fields, e.g., fraud detection, climate analysis, or healthcare monitoring. We present an algorithm for detecting anomalous regions in multivariate spatio-temporal time-series, which allows for spotting the interesti…
PLUMAGE improves large model training efficiency and stability.
problem Accelerator memory and networking constraints during large model training.
method Probabilistic Low rank Unbiased Minimum Variance Gradient Estimator (PLUMAGE) that resolves bias and variance issues.
result PLUMAGE reduces training loss by 28% on average across the GLUE benchmark.
The probability that a user will click a search result depends both on its relevance and its position on the results page. The position based model explains this behavior by ascribing to every item an attraction probability, and to every position an examination probability. To be clicked, a result must be both attracti…
Estimates causal effect using proxies in multi-domain settings.
problem Estimating causal effect in settings with unobserved confounders across domains.
method Proposes estimation techniques using proxy variables for discrete or categorical data.
result Proves identifiability and consistency of causal effect estimation.
We consider sequential or active ranking of a set of n items based on noisy pairwise comparisons. Items are ranked according to the probability that a given item beats a randomly chosen item, and ranking refers to partitioning the items into sets of pre-specified sizes according to their scores. This notion of ranking …
Paper develops inference methods for low-rank tensors without debiasing.
problem Statistical inference for low-rank tensor models.
method Two-iteration alternating minimization for asymptotic distribution.
result Asymptotic distributions and confidence regions for singular subspaces.
A new framework evaluates LLMs by considering judge reliability.
problem Evaluating LLMs without ground truth labels can lead to biased results.
method Introduces judge-specific discrimination parameters and estimates model quality and judge reliability.
result Improves agreement with human preferences and produces calibrated uncertainty quantification.
Convex PCA improves Euclidean PCA for convex data subsets.
problem Improving PCA for convex data subsets.
method Developed new theoretical results and a numerical implementation for finite dimensional convex PCA.
result Finite dimensional convex PCA approximates Wasserstein GPCA and ranked compositional data.
In this note we show that for the group G = U(N) the space of Hecke modifications of a rank N vector bundle over a Riemann surface C coincides with the moduli space of solutions of certain non-abelian vortex equations over C . Through the recent work of Kapustin and Witten this then leads to an isomorphism between the …
New algorithm samples from Ising models efficiently, even with outliers.
problem Sampling from Ising models with general interaction matrices.
method Combines MCMC and variational inference techniques.
result First polynomial time sampling algorithms for low-rank Ising models.
The paper introduces a framework to select efficient datasets for preserving model rankings.
problem Efficient evaluation of machine learning models on small, representative datasets.
method Bootstrap aggregation, clustering, design criteria, random baselines, and greedy farthest-first (FAFI).
result Several selection strategies improve rank preservation compared to random subsets, especially in time series classification.
TripleSurv improves survival analysis by ranking samples with time-adaptive adjustments.
problem Modeling censored time-to-event data with high accuracy and robustness.
method Introduces a time-adaptive coordinate loss function to rank samples and calibrate robustness.
result TripleSurv outperforms state-of-the-art methods on various survival datasets.
The paper develops methods to infer membership probabilities and rank network nodes using the DCMM model.
problem Understanding the latent structure of network data, especially in mixed-membership models.
method Degree-Corrected Mixed Membership (DCMM) model, novel finite-sample expansion, asymptotic distributions, confidence intervals, multiplier bootstrap method.
result Valid inference on membership probabilities and node rankings, quantifying uncertainty.
Enhances VAR model estimation using transfer learning.
problem Estimating high-dimensional VAR models with temporal dependencies.
method Transfer learning for VAR models with low-rank and sparse structures.
result Theoretical guarantees for model parameter consistency and informative set selection.
Gaussian and bootstrap methods improve ATE estimator accuracy.
problem Improving the accuracy of Average Treatment Effect (ATE) estimators.
method Gaussian approximation and bootstrap procedures.
result Precise bounds on ATE estimator accuracy quantifying key parameters.
Bayesian principles improve neural additive models for better feature selection and uncertainty.
problem Lack of calibrated uncertainties and feature selection in neural additive models.
method Augmenting NAMs with Bayesian principles to provide credible intervals, feature selection, and interaction ranking.
result Improved performance on tabular datasets and real-world medical tasks.
This paper computes exact posterior distributions of mixture weights in hierarchical Bayesian models.
problem Uncertainty in class membership or data-generating processes in heterogeneous data.
method Exact marginalization of mixture weights using dynamic programming and FFT for two components, and joint dynamic program for K >= 3 components.
result Exact posterior distributions of mixture weights are finite mixtures of Beta distributions, providing credible intervals and per-observation local false-discovery rates.
Consider the problem of estimating a low-rank matrix when its entries are perturbed by Gaussian noise. If the empirical distribution of the entries of the spikes is known, optimal estimators that exploit this knowledge can substantially outperform simple spectral approaches. Recent work characterizes the asymptotic acc…
New online method for statistical inference with matrix context in decision-making.
problem Statistical inference in decision-making with matrix context.
method Proposes a fully online procedure to conduct statistical inference with adaptive data collection, handling low-rank structure.
result Establishes asymptotic normality of debiased estimators and proves validity of confidence intervals.
New method improves compatibility of risk stratification models without sacrificing accuracy.
problem Compatibility issues arise when updating clinical machine learning models.
method Proposes rank-based compatibility measure and new loss function.
result Increased compatibility of models by 0.019 with no loss in discriminative performance.
Having a regression model, we are interested in finding two-sided intervals that are guaranteed to contain at least a desired proportion of the conditional distribution of the response variable given a specific combination of predictors. We name such intervals predictive intervals. This work presents a new method to fi…
Marginal Structural Models (MSM) are the most popular models for causal inference from time-series observational data. However, they have two main drawbacks: (a) they do not capture subject heterogeneity, and (b) they only consider fixed time intervals and do not scale gracefully with longer intervals. In this work, we…
Deep learning improves PV generation quantile forecasting.
problem Accurate probabilistic forecasting of PV generation.
method Developed an encoder-decoder deep learning model for multi-output quantile PV forecasting.
result The model improves forecast quality and computational efficiency.