Develops a statistical framework to measure uncertainty in model rankings based on human preferences.
problem Uncertainty in model rankings based on human preferences due to mismatch between human and model preferences.
method Statistical framework using pairwise comparisons by humans and models to provide rank-sets for each model.
result Rank-sets constructed using only pairwise comparisons by strong models often do not cover the true ranking of human preferences.
We introduce a probabilistic framework for quantifying the semantic similarity between two groups of embeddings. We formulate the task of semantic similarity as a model comparison task in which we contrast a generative model which jointly models two sentences versus one that does not. We illustrate how this framework c…
Statistical framework improves LLM chatbot ranking.
problem Improving evaluation of LLM-based chatbots through pairwise comparisons.
method Factored tie model, covariance modeling, and parameter constraints.
result Substantial improvements in modeling pairwise comparison data.
Study on manifolds with density using modified Hessians for curvature comparison.
problem Developing comparison geometry on manifolds with density.
method Modified Hessian approach based on weighted sectional curvature framework.
result Derivation of Hessian comparison and shape operator comparison theorems.
Binary feedback outperforms ordinal comparisons in ranking recovery.
problem Challenges the conventional wisdom that ordinal comparisons offer richer information.
method Proposes a parametric framework for modeling ordinal paired comparisons, binarizing ordinal data, and proving faster convergence rates for binary comparisons.
result Binarizing ordinal data significantly improves ranking recovery accuracy and exhibits a substantial performance gap.
Enhances AI models with human feedback for noisy data.
problem Improving AI model alignment with human feedback in noisy environments.
method Two-stage SL+LHF framework connecting machine learning with human feedback.
result The LNCA ratio identifies conditions for SL+LHF superiority over pure SL.
In this study, a pairwise comparison matrix is generalized to the case when coefficients create Lie group G, non necessarily abelian. A necessary and sufficient criterion for pairwise comparisons matrices to be consistent is provided. Basic criteria for finding a nearest consistent pairwise comparisons matrix (extend…
New comparison shows differences in how value is incorporated in AIF and CAI.
problem Clarifying the relationship between Active Inference and Control-as-Inference.
method Formal comparison of AIF and CAI frameworks.
result Primary difference is how value is incorporated into generative models.
New method improves ABC for Bayesian model comparison.
problem Comparing complex models with observed data.
method Approximate Bayesian Computation with posterior density estimation.
result Efficiently assigns high posterior probabilities to ground-truth models.
Existing ordinal embedding methods usually follow a two-stage routine: outlier detection is first employed to pick out the inconsistent comparisons; then an embedding is learned from the clean data. However, learning in a multi-stage manner is well-known to suffer from sub-optimal solutions. In this paper, we propose a…
The paper proves new comparison theorems for sub-Laplacian in foliations with minimal leaves.
problem Proving comparison theorems for sub-Laplacian in Riemannian foliations with minimal leaves.
method Using Riemannian foliations with minimal leaves, the paper proves comparison theorems for the sub-Laplacian.
result The comparison theorems yield a Bonnet-Myers type theorem, stochastic completeness, and Lipschitz regularization property for the sub-Riemannian semigroup.
Volume comparison theorem for rank 1 symmetric spaces proved.
problem Volume comparison for symmetric spaces of non-compact type.
method Normalized Ricci--DeTurck flow to analyze volume functional and derive monotonicity properties.
result Volume comparison theorem established for rank 1 symmetric spaces of non-compact type.
Learning from triplet comparison data has been extensively studied in the context of metric learning, where we want to learn a distance metric between two instances, and ordinal embedding, where we want to learn an embedding in an Euclidean space of the given instances that preserves the comparison order as well as pos…
Paper proposes a new method to compare classifiers across multiple datasets.
problem Comparing classifiers over multiple datasets with multiple criteria.
method Adopting decision theory, the paper introduces generalized stochastic dominance for ranking classifiers.
result Generalized stochastic dominance can be used to rank classifiers and statistically tested.
Enhances Bayesian model comparison with a probabilistic framework for meta-uncertainty.
problem Uncertainty in posterior model probabilities (PMPs) when derived from finite data.
method Develops a fully probabilistic approach to quantify and represent meta-uncertainty over PMPs.
result Demonstrates utility in various BMC contexts, including regression, MCMC, and neural networks.
Clustering is one of the most universal approaches for understanding complex data. A pivotal aspect of clustering analysis is quantitatively comparing clusterings; clustering comparison is the basis for many tasks such as clustering evaluation, consensus clustering, and tracking the temporal evolution of clusters. In p…
Study embeds PC matrices into Grassmannian manifold for geometric interpretation.
problem Understanding algebraic consistency of pairwise comparisons matrices.
method Leverages Plücker coordinates and geometric interpretation of Grassmannian manifold.
result Algebraic consistency condition is equivalent to geometric consistency in G(2,n). Introduces AMLB, an open benchmark for AutoML frameworks.
problem Challenges in comparing AutoML frameworks.
method Open benchmark with 9 AutoML frameworks, 71 classification, 33 regression tasks, multi-faceted analysis, Bradley-Terry trees.
result Differences in AutoML frameworks' performance and trade-offs.
Proves curvature comparison theorem for manifolds with conical singularities.
problem Comparing scalar mean curvature of manifolds with conical singularities.
method Uses Dirac operator and index theory to prove curvature comparison theorem.
result Proves curvature comparison theorem without knowing the index of the twisted Dirac operator.
GNNRank uses neural networks to learn global rankings from competition match data.
problem Learning global rankings from pairwise comparisons in directed graphs.
method Proposes GNNRank, a trainable GNN-based framework with digraph embedding and new objectives.
result GNNRank achieves competitive and superior performance compared to baselines.
This study provides benchmarks for different implementations of LSTM units between the deep learning frameworks PyTorch, TensorFlow, Lasagne and Keras. The comparison includes cuDNN LSTMs, fused LSTM variants and less optimized, but more flexible LSTM implementations. The benchmarks reflect two typical scenarios for au…
Novel method for Bayesian model comparison using deep learning.
problem Comparing complex models in science with intractable likelihood functions.
method Simulation-based, purely deep learning approach that amortizes model fitting costs.
result Achieves excellent results in accuracy, calibration, and efficiency.
New framework estimates treatment effects based on preferences.
problem Estimating treatment effects with flexible outcomes.
method Preference-based Conditional Treatment Effect (CPTE) framework.
result CPTE provides interpretable targets and new identifiability conditions.
BSD is a Bayesian framework for analyzing neural spectral data.
problem Challenges in statistical analysis and group-level comparisons of neural power spectra.
method Bayesian Spectral Decomposition (BSD) for parametric models of neural spectra.
result BSD outperforms existing methods in model selection and parameter estimation.
We address the classical problem of hierarchical clustering, but in a framework where one does not have access to a representation of the objects or their pairwise similarities. Instead, we assume that only a set of comparisons between objects is available, that is, statements of the form "objects i and j are more …
Study finds lower bounds for energy on fibred manifolds using fiberwise symmetrization.
problem Finding lower bounds for energy functionals on fibred manifolds.
method Established a framework for fiberwise symmetrization to find lower bounds.
result Proved a comparison theorem for the first eigenvalue of the Laplacian on warped product manifolds.
We treat the vakonomic dynamics with general constraints within a new geometric framework which will be appropriate to study optimal control problems. We compare our formulation with Vershik-Gershkovich one in the case of linear constraints. We show how nonholonomic mechanics also admits a new geometrical description w…
Low-rank framework for task-specific LLM ranking from sparse comparisons.
problem Challenges in reliable task-specific ranking of LLMs under sparse, imbalanced comparisons.
method Low-rank modeling of task-by-model ability matrix, max-norm accurate estimator, task-wise top-K recovery guarantees, uncertainty quantification framework.
result Improves sample efficiency and produces tighter, better-calibrated ranking certificates.
While deep neural networks have become the go-to approach in computer vision, the vast majority of these models fail to properly capture the uncertainty inherent in their predictions. Estimating this predictive uncertainty can be crucial, for example in automotive applications. In Bayesian deep learning, predictive unc…
Develops a hypothesis testing framework for generalized Thurstone models.
problem Determining whether pairwise comparison data fits a generalized Thurstone model.
method Introduces separation distance and derives upper and lower bounds for testing.
result Critical threshold for testing depends on observation graph topology and scales as Θ((nk)−1/2) for complete graphs. MIRA scores assess conditional distribution accuracy using joint samples.
problem Assessing the accuracy of candidate conditional distributions.
method Analytic expression for Mira score based on equal probability mass regions.
result Mira enables Bayesian model comparison by quantifying alignment with true process.
Paper tackles ranking items with a semi-random comparison graph and a monotone adversary.
problem Ranking items based on pairwise comparisons from a semi-random comparison graph with a monotone adversary.
method Developed a weighted maximum likelihood estimator (MLE) and an SDP-based approach to reweight the semi-random graph.
result Achieves near-optimal sample complexity, up to a log^2(n) factor, for identifying the top-K preferred items.
Paper proposes a new framework to compare trading strategies by accounting for market conditions.
problem Lack of information on how trading strategy performance varies with market conditions.
method Uses a GAMLSS/ZAGA framework to model the Adjusted Information Ratio (IR∗) for a SVMP and BH strategy across 146 folds of the S&P 500. result Dominance of SVMP over BH is conditional on market regime, as shown by differences in expected IR∗ and its variance. A novel method compares 3D point clouds using information geometry.
problem Comparing 3D point clouds in machine learning applications.
method Interprets point clouds as probability density functions on a statistical manifold, using GMM and Modified Symmetric KL divergence.
result Demonstrates effectiveness through various case studies.
Pref-SHAP explains preferences using Shapley values.
problem Challenging problem of preference explanation in machine learning.
method Shapley value-based model explanation framework for pairwise comparison data.
result Richer and more insightful explanations obtained over baseline.
The paper develops a method to estimate consumer preferences from observed rankings.
problem Estimating consumer preferences from partial ranking information.
method Interpreting observed rankings as pairwise comparisons, modeling latent utility, and correcting for selection bias.
result The method improves recommendation performance, especially for previously unconsumed products.
A framework for quantifying uncertainty in feature importance values.
problem Stable interpretation of feature importance values in machine learning models.
method A novel method based on pairwise comparisons of feature importance values to produce confidence intervals for feature ranks.
result The method produces simultaneous confidence intervals for feature ranks, enabling selection of top-k important features.
GWCA analyzes cross-graph correlations for movie retrieval.
problem Cross heterogeneous graph comparison in movie retrieval.
method Spectral graph filtering, Wasserstein metric learning, generalized eigenvalue decomposition.
result Surprise consistency in learning processes and closed-form solution.
New method uses model comparison signals to improve LLM evaluation accuracy.
problem Limited benchmark sizes and model stochasticity in evaluating LLMs' mathematical reasoning.
method Combines standard labeled outcomes with model comparison signals to design a statistically efficient evaluation framework.
result Semiparametric estimator achieves the semiparametric efficiency bound and substantially improves ranking accuracy.
Feature selection methods are usually evaluated by wrapping specific classifiers and datasets in the evaluation process, resulting very often in unfair comparisons between methods. In this work, we develop a theoretical framework that allows obtaining the true feature ordering of two-dimensional sequential forward feat…
New depth function for partial orders helps compare machine learning algorithms.
problem Comparing machine learning algorithms using non-standard data types.
method Adapted simplicial depth to partial orders, using ufg depth for comparison.
result Demonstrates promising variety of analysis approaches based on ufg methods.
A new notion of stochastic ordering is introduced to compare multivariate stochastic risk models with respect to extreme portfolio losses. In the framework of multivariate regular variation comparison criteria are derived in terms of ordering conditions on the spectral measures, which allows for analytical or numerical…
Modeling preference rankings with salient features to explain irrational choices.
problem Estimating rankings from noisy pairwise comparisons with irrational choices.
method Salient feature preference model with maximum likelihood estimation.
result Strong performance of maximum likelihood estimation on synthetic and real data.
Unified framework to bridge human and LLM judgments.
problem Systematic discrepancies between human and LLM evaluations.
method Latent human preference score and linear transformations of covariates.
result Higher agreement with human ratings and exposure of systematic gaps.
ModelDiff compares learning algorithms by identifying feature transformations.
problem Comparing models trained with different learning algorithms.
method ModelDiff uses datamodels framework to find distinguishing feature transformations.
result ModelDiff can compare models trained with/without data augmentation, pre-training, and different SGD hyperparameters.
A new numerical framework simplifies elastic surface matching and comparison.
problem Challenging problem in surface comparison and matching in computer vision.
method Relaxing the geodesic boundary constraint using a varifold fidelity metric.
result Flexibility to deal with arbitrary topologies and sampling patterns, scalability to large meshes.
Statistical inference using pairwise comparison data is an effective approach to analyzing large-scale sparse networks. In this paper, we propose a general framework to model the mutual interactions in a network, which enjoys ample flexibility in terms of model parametrization. Under this setup, we show that the maximu…
In recent years, an active field of research has developed around automated machine learning (AutoML). Unfortunately, comparing different AutoML systems is hard and often done incorrectly. We introduce an open, ongoing, and extensible benchmark framework which follows best practices and avoids common mistakes. The fram…