A new framework evaluates LLMs by considering judge reliability.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New methods correct bias in LLM-as-a-Judge evaluations, but reliability depends on judge quality and model calibration.
No-knowledge alarms detect misaligned LLM judges without trusting them.
VLM judges rank well but score poorly; task difficulty and annotation quality affect interval width.
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing
Researchers use LLMs to judge other LLMs, but this study provides a new geometric perspective to understand when it works.
RACER optimizes LLM-as-judge accuracy with dynamic reasoning selection.
Optimizes budgeted evaluations of LLMs by allocating queries to judges efficiently.
MultiwayPAM clusters LLM-as-a-Judge scores to reveal evaluator bias.
New method improves LLM judge accuracy by accounting for dependencies in aggregated binary labels.
Study identifies and measures biases in legal case data.
CJE calibrates cheap LLM judges against an oracle, achieving high accuracy at a fraction of the cost.
ValueBlindBench tests LLM-generated investment rationales for validity before returns are known.
CARE improves LLM aggregation by accounting for shared confounders.
LLM evaluation suffers from systematic biases and lacks reliable positive judgments.
Corrects bias in LLM-as-a-judge evaluations using adaptive calibration.
Unified framework to bridge human and LLM judgments.
To make AI systems broadly useful for challenging real-world tasks, we need them to learn complex human goals and preferences. One approach to specifying complex goals asks humans to judge during training which agent behaviors are safe and useful, but this approach can fail if the task is too complicated for a human to…
Paper improves uncertainty estimation in LLM-as-a-judge systems.
New models can't beat existing ones, so debiasing methods only slightly reduce needed labels.
We study a bad arm existing checking problem in which a player's task is to judge whether a positive arm exists or not among given K arms by drawing as small number of arms as possible. Here, an arm is positive if its expected loss suffered by drawing the arm is at least a given threshold. This problem is a formalizati…
This paper proposes a method to use LLMs as auxiliary evaluators in place of human judges.
This paper revisits the problem of analyzing multiple ratings given by different judges. Different from previous work that focuses on distilling the true labels from noisy crowdsourcing ratings, we emphasize gaining diagnostic insights into our in-house well-trained judges. We generalize the well-known DawidSkene model…
When dealing with subjective, noisy, or otherwise nebulous features, the "wisdom of crowds" suggests that one may benefit from multiple judgments of the same feature on the same object. We give theoretically-motivated `feature multi-selection' algorithms that choose, among a large set of candidate features, not only wh…
Detects unusual inputs to neural networks to prevent flawed predictions.
From doctors diagnosing patients to judges setting bail, experts often base their decisions on experience and intuition rather than on statistical models. While understandable, relying on intuition over models has often been found to result in inferior outcomes. Here we present a new method, select-regress-and-round, f…
We give several new criteria to judge whether a simple convex polytope in a Euclidean space is combinatorially equivalent to a product of simplices. These criteria are mixtures of combinatorial, geometrical and topological conditions that are inspired by the ideas from toric topology.
PDO optimizes LLM prompts without labels, improving performance.
We consider an axisymmetric closed hypersurface evolving by its mean curvature with driving force under singular initial hypersurface. We study this problem by level set method. We give some criteria to judge whether the interface evolution is fattening or non-fattening.
The paper proposes count echo state networks for forecasting graduate student enrollments.
Unsupervised scheme ranks sentences in text documents based on semantic importance.
We propose a novel method for fact-checking on knowledge graphs based on debate dynamics. The underlying idea is to frame the task of triple classification as a debate game between two reinforcement learning agents which extract arguments -- paths in the knowledge graph -- with the goal to justify the fact being true (…
We propose a novel method for automatic reasoning on knowledge graphs based on debate dynamics. The main idea is to frame the task of triple classification as a debate game between two reinforcement learning agents which extract arguments -- paths in the knowledge graph -- with the goal to promote the fact being true (…
We introduce advocacy learning, a novel supervised training scheme for attention-based classification problems. Advocacy learning relies on a framework consisting of two connected networks: 1) Advocates (one for each class), each of which outputs an argument in the form of an attention map over the input, and 2) a …
Traditionally, the vision community has devised algorithms to estimate the distance between an original image and images that have been subject to perturbations. Inspiration was usually taken from the human visual perceptual system and how the system processes different perturbations in order to replicate to what exten…
We conduct a large-scale, systematic study to evaluate the existing evaluation methods for natural language generation in the context of generating online product reviews. We compare human-based evaluators with a variety of automated evaluation procedures, including discriminative evaluators that measure how well machi…
The paper introduces Relative Bias to quantify LLM bias systematically.
New framework assesses LLM security risks in BFSI.
Unified model combines scores and rankings for grant panel review.
Online symptom checkers have significant potential to improve patient care, however their reliability and accuracy remain variable. We hypothesised that an artificial intelligence (AI) powered triage and diagnostic system would compare favourably with human doctors with respect to triage and diagnostic accuracy. We per…
Algorithm identifies best arm with biased proxy and selective ground truth audits.
In this paper, we present a data science automation system called Prediction Factory. The system uses several key automation algorithms to enable data scientists to rapidly develop predictive models and share them with domain experts. To assess the system's impact, we implemented 3 different interfaces for creating pre…
In this paper we investigate the moduli space of parabolic Higgs bundles over a punctured Riemann surface with varying weights at the punctures. We show that the harmonic metric depends analytically on the weights and the stable Higgs bundle. This gives a Higgs bundle generalisation of a theorem of McOwen on the existe…
Feature Learning aims to extract relevant information contained in data sets in an automated fashion. It is driving force behind the current deep learning trend, a set of methods that have had widespread empirical success. What is lacking is a theoretical understanding of different feature learning schemes. This work p…
Characterizes periods of abelian differentials on surfaces.
New method improves speech synthesis quality.
Determines conditions for abelian differentials with specific singularities.
In this paper, the author considers the numerical computation of CVA for large systems by Mote Carlo methods. He introduces two types of stochastic mesh methods for the computations of CVA. In the first method, stochastic mesh method is used to obtain the future value of the derivative contracts. In the second method, …