CV inference can be invalid for relatively unstable model comparisons.
problem The validity of cross-validation for model comparison is questioned when models are relatively unstable.
method The study proves that simple, individually stable models can generate relatively unstable comparisons, invalidating CV inference.
result The Lasso and soft-thresholding generate relatively unstable comparisons, invalidating CV inferences.
Comparison data arises in many important contexts, e.g. shopping, web clicks, or sports competitions. Typically we are given a dataset of comparisons and wish to train a model to make predictions about the outcome of unseen comparisons. In many cases available datasets have relatively few comparisons (e.g. there are on…
New method improves ABC for Bayesian model comparison.
problem Comparing complex models with observed data.
method Approximate Bayesian Computation with posterior density estimation.
result Efficiently assigns high posterior probabilities to ground-truth models.
We study the problem of ranking from crowdsourced pairwise comparisons. Answers to pairwise tasks are known to be affected by the position of items on the screen, however, previous models for aggregation of pairwise comparisons do not focus on modeling such kind of biases. We introduce a new aggregation model factorBT …
New proposed models are often compared to state-of-the-art using statistical significance testing. Literature is scarce for classifier comparison using metrics other than accuracy. We present a survey of statistical methods that can be used for classifier comparison using precision, accounting for inter-precision corre…
Develops a statistical framework to measure uncertainty in model rankings based on human preferences.
problem Uncertainty in model rankings based on human preferences due to mismatch between human and model preferences.
method Statistical framework using pairwise comparisons by humans and models to provide rank-sets for each model.
result Rank-sets constructed using only pairwise comparisons by strong models often do not cover the true ranking of human preferences.
Automates model comparison in probabilistic programming.
problem Manual derivations for model comparison are error-prone and time-consuming.
method Message passing on a Forney-style factor graph with a custom mixture node.
result Automates Bayesian model averaging, selection, and combination.
Paper establishes statistical inference for pairwise comparison models.
problem Statistical inference for pairwise comparison models when the number of subjects diverges.
method Identifies Fisher information matrix as a weighted graph Laplacian for asymptotic normality.
result Near-optimal asymptotic normality result for maximum likelihood estimator.
Test log-likelihood comparisons can be misleading.
problem Misinterpretation of test log-likelihood in model comparison.
method Simple examples of model comparison and forecast accuracy.
result Test log-likelihood does not always correlate with model accuracy.
Rank regression from pairwise comparisons requires many comparisons to accurately learn model parameters.
problem Learning model parameters for rank regression from noisy pairwise comparisons.
method Uniform random pairwise comparisons to estimate model parameters with a given accuracy.
result Learning model parameters requires a number of comparisons proportional to dNlog3N/ε2. Paper presents a new insurance model equation for diverse structures.
problem Handling diverse insurance models with a single equation.
method Developed a canonical model construction and stochastic backward equations.
result Comparison theorems for different models follow from the new equation.
SC improves robustness in model comparison for misspecified models.
problem Model misspecification challenges in amortized Bayesian inference.
method Parameter posterior-based methods augmented with SC training.
result SC improves robustness under model misspecification.
Binary feedback outperforms ordinal comparisons in ranking recovery.
problem Challenges the conventional wisdom that ordinal comparisons offer richer information.
method Proposes a parametric framework for modeling ordinal paired comparisons, binarizing ordinal data, and proving faster convergence rates for binary comparisons.
result Binarizing ordinal data significantly improves ranking recovery accuracy and exhibits a substantial performance gap.
A comparison theorem for the isoperimetric profile on the universal cover of surfaces evolving by normalised Ricci flow is proven. For any initial metric, a model comparison is constructed that initially lies below the profile of the initial metric and which converges to the profile of the constant curvature metric. Th…
We introduce a probabilistic framework for quantifying the semantic similarity between two groups of embeddings. We formulate the task of semantic similarity as a model comparison task in which we contrast a generative model which jointly models two sentences versus one that does not. We illustrate how this framework c…
New model for pairwise comparisons without stochastic transitivity.
problem Suboptimal performance of models assuming stochastic transitivity in real-world scenarios.
method Proposes a general family of statistical models using a skew-symmetric matrix.
result Achieves minimax-rate optimality and adapts to data sparsity.
Pairwise comparison data arises in many domains, including tournament rankings, web search, and preference elicitation. Given noisy comparisons of a fixed subset of pairs of items, we study the problem of estimating the underlying comparison probabilities under the assumption of strong stochastic transitivity (SST). We…
Study metric learning from limited preference comparisons, showing how low-dimensional structure can still reveal metric information.
problem Learning metric from limited pairwise preference comparisons.
method Ideal point model, divide-and-conquer approach for low-dimensional structure.
result Metric can be jointly identified even with limited comparisons when items exhibit low-dimensional structure.
A faster algorithm for ranking from pairwise comparisons.
problem Efficiently ranking individuals or objects from pairwise comparisons.
method An alternative and simpler iterative algorithm for ranking that converges faster.
result The new algorithm is over 100 times faster in some cases.
This paper aims to present a general idea of method comparison of Credit Scoring techniques. Any scorecard can be made in various methods based on variable transformations in the logistic regression model. To make a comparison and come up with the proof that one technique is better than another is a big challenge due t…
Optimizes identifying top-k items from comparisons with minimal comparisons.
problem Finding the top-k items from pairwise comparisons with a fixed error rate.
method Developed an asymptotically optimal algorithm using primal-dual procedure and adaptive comparison allocation.
result Proves the algorithm is asymptotically optimal for top-k identification.
Study on deep neural networks for reward modeling with pairwise comparison data.
problem Reward modeling with deep neural networks in non-parametric settings.
method Established a non-asymptotic regret bound for deep reward estimators, introduced a margin-type condition.
result Improved regret bound for deep reward estimators, highlighting the importance of clear human beliefs.
There is increasing interest in learning algorithms that involve interaction between human and machine. Comparison-based queries are among the most natural ways to get feedback from humans. A challenge in designing comparison-based interactive learning algorithms is coping with noisy answers. The most common fix is to …
New method uses model comparison signals to improve LLM evaluation accuracy.
problem Limited benchmark sizes and model stochasticity in evaluating LLMs' mathematical reasoning.
method Combines standard labeled outcomes with model comparison signals to design a statistically efficient evaluation framework.
result Semiparametric estimator achieves the semiparametric efficiency bound and substantially improves ranking accuracy.
Quantifies Schur's theorem for curves in CAT(k) spaces.
problem Quantifying Schur's comparison theorem for curves in CAT(k) spaces.
method Comparison formula for curves in model planes, curvature measures, moment arm, and Reshetnyak's theorem.
result Sharpens and extends classical arm and bow lemmas and Riemannian analogues.
New model accounts for scale variation and noise in pairwise comparisons.
problem Nonreciprocal pairwise comparisons in decision analysis.
method Additive model with structured matrix and random perturbation.
result Explicit estimators and probability assessments of admissible ranking regions.
Graph comparison ties to Alexandrov's theorems.
problem Graph comparison conditions on metric spaces.
method Proof of Alexandrov's implications from graph comparisons.
result Complete description of graphs with trivial comparisons.
Evidence Networks simplify Bayesian model comparison for complex models.
problem Bayesian model comparison challenges with intractable likelihoods or priors.
method Loss functions and neural networks for fast, amortized estimation of Bayes factors.
result Evidence Networks provide accurate and scalable Bayes factor estimation.
We present a number of related comparison results, which allow to compare moment explosion times, moment generating functions and critical moments between rough and non-rough Heston models of stochastic volatility. All results are based on a comparison principle for certain non-linear Volterra integral equations. Our u…
Statistical framework improves LLM chatbot ranking.
problem Improving evaluation of LLM-based chatbots through pairwise comparisons.
method Factored tie model, covariance modeling, and parameter constraints.
result Substantial improvements in modeling pairwise comparison data.
We prove comparison theorems for the sub-Riemannian distortion coefficients appearing in interpolation inequalities. These results, which are equivalent to a sub-Laplacian comparison theorem for the sub-Riemannian distance, are obtained by introducing a suitable notion of sub-Riemannian Bakry-Émery curvature. The model…
Self-consistency improves the accuracy of model comparison methods.
problem Improving the accuracy of model comparison methods when simulation models are misspecified.
method Supplement traditional simulation-based training with a self-consistency loss on unlabeled real data.
result Self-consistency training improves model comparison accuracy, especially in open-world scenarios.
The paper develops a method to estimate consumer preferences from observed rankings.
problem Estimating consumer preferences from partial ranking information.
method Interpreting observed rankings as pairwise comparisons, modeling latent utility, and correcting for selection bias.
result The method improves recommendation performance, especially for previously unconsumed products.
In this paper, we propose an active learning algorithm and models which can gradually learn individual's preference through pairwise comparisons. The active learning scheme aims at finding individual's most preferred choice with minimized number of pairwise comparisons. The pairwise comparisons are encoded into probabi…
Paper tackles noisy comparison oracle for robust clustering algorithms.
problem Finding robust clustering algorithms under noisy comparison oracle.
method Develops algorithms for k-center clustering and agglomerative hierarchical clustering using noisy comparison oracle.
result Proves robust algorithms achieve good approximation guarantees with high probability.
Paper quantifies uncertainty in pairwise comparison models.
problem Uncertainty quantification in sparse Bradley-Terry-Luce models.
method Unified proof strategy for MLE and spectral estimator.
result Sharp and uniform non-asymptotic expansions for estimators.
Paper compares different models for time-to-event analysis.
problem Comparing models for time-to-event analysis.
method Experimental comparison of semi-parametric, parametric, and machine learning models.
result Models' performance evaluated using concordance index.
Proves curvature comparison for Riemannian bands in low dimensions.
problem Curvature comparison in Riemannian bands with lower bounds.
method Uses warped products over scalar-flat manifolds with log-concave warping.
result Scalar and mean curvature comparison results proven.
We analyze training dynamics in Gaussian mixture models using a comparison theorem.
problem Analyzing training algorithms with Gaussian mixture data.
method Applying a Gaussian comparison theorem to a specific family of training algorithms.
result Validated dynamic mean-field expressions and provided iterative refinement schemes.
MIRA scores assess conditional distribution accuracy using joint samples.
problem Assessing the accuracy of candidate conditional distributions.
method Analytic expression for Mira score based on equal probability mass regions.
result Mira enables Bayesian model comparison by quantifying alignment with true process.
Enhances AI models with human feedback for noisy data.
problem Improving AI model alignment with human feedback in noisy environments.
method Two-stage SL+LHF framework connecting machine learning with human feedback.
result The LNCA ratio identifies conditions for SL+LHF superiority over pure SL.
In this work, we will verify some comparison results on Kahler manifolds. They are complex Hessian comparison for the distance function from a closed complex submanifold of a Kahler manifold with holomorphic bisectional curvature bounded below by a constant, eigenvalue comparison and volume comparison in terms of scala…
The paper extends volume comparison results to total σ_l-curvature.
problem Comparing total σ_l-curvature with σ_k-curvature.
method Volume comparison theorem extension to σ-curvature comparison.
result Comparison holds for metrics close to strictly stable positive Einstein metrics.
Paper extends curvature estimates to new tensor types.
problem Mean curvature and volume comparison estimates for integral generalized quasi-Einstein tensors.
method Extends existing comparison results to new tensor types.
result Global diameter estimates derived from comparison results.
A new comparison theorem for geometric spaces.
problem Geometric space comparison theorems.
method Relative form of Toponogov comparison theorem.
result New geometric space comparison theorem established.
Bayesian model infers strengths from noisy tennis match outcomes.
problem Ranking tennis players from match outcomes.
method Bayesian approach to infer unobserved strengths and mapping function.
result Bayesian approach robust to different model specifications.
On Kahler manifolds with Ricci curvature lower bound, assuming the real analyticity of the metric, we establish a sharp relative volume comparison theorem for small balls. The model spaces being compared to are complex space forms, i.e, Kahler manifolds with constant holomorphic sectional curvature. Moreover, we give a…
Novel CNN-based gaze scanpath comparison distinguishes experts from novices in dental radiograph interpretation.
problem Distinguishing expertise in dental radiograph interpretation based on gaze behavior.
method Convolutional neural networks (CNN) process scene information at the fixation level, using image patches as input to compare gaze scanpaths.
result 93% accuracy in distinguishing experts from novices using image patch features.