No-knowledge alarms detect misaligned LLM judges without trusting them.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
RACER optimizes LLM-as-judge accuracy with dynamic reasoning selection.
R-AutoEval+ improves model evaluation efficiency and reliability using adaptive synthetic data.
Large language models correlate in errors, even with different architectures and providers.
Language model benchmarks often misrepresent true understanding, revealing vulnerabilities in evaluation methods.
UCFE benchmarks LLMs in financial tasks with human feedback.
Develops methods to correct bias in AI feedback for more accurate alignment.
FinReflectKG - EvalBench benchmarks financial KG extraction from SEC 10-K filings.
New models can't beat existing ones, so debiasing methods only slightly reduce needed labels.
Top-H decoding improves text generation by balancing creativity and coherence.