VLM judges rank well but score poorly; task difficulty and annotation quality affect interval width.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Martingale Doppelgänger-Eval benchmarks VLMs on candlestick evidence vs. trend extrapolation
Study identifies key parameters and input dimensions making LLMs and VLMs brittle.
PyFi uses adversarial agents to train VLMs on financial image understanding.
VERA-V uses variational inference to discover vulnerabilities in multimodal vision-language models.
The study examines how much data is needed for generative and vision-language models to make reliable predictions.
Demon aligns diffusion models without retraining or backpropagation.
LangDA improves domain adaptation for semantic segmentation by learning context-aware scene descriptions.
HyperINF improves influence function estimation for large models with better accuracy and efficiency.
MM-DREX adapts LLM experts for financial trading via dynamic routing.
DoRA improves adaptation efficiency for large models by factoring norms and fusing kernels.