Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

103205308410 · Jun 202019922001200920182026
48 results for Rigorous Evaluation

A spin network is a cubic ribbon graph labeled by representations of SU(2)\mathrm{SU}(2). Spin networks are important in various areas of Mathematics (3-dimensional Quantum Topology), Physics (Angular Momentum, Classical and Quantum Gravity) and Chemistry (Atomic Spectroscopy). The evaluation of a spin network is an integ…

2009-02-18abs ↗pdf ↗

Synthetic experiments are crucial for assessing causal machine learning methods.

problem Current empirical evaluations of causal machine learning methods are insufficient and unreliable.
method Propose principles for conducting rigorous empirical analyses with synthetic data.
result Rigorous synthetic experiments are essential for building trust in causal machine learning methods.

SMEs provide a transparent testbed for RL evaluation.

problem Lack of precise, white-box diagnostics in RL environments.
method Synthetic Monitoring Environments (SMEs) with fully configurable task characteristics and known optimal policies.
result SMEs allow for precise evaluation of RL algorithms, revealing the impact of specific environmental properties.

New rigorous uncertainty bounds for Gaussian Process regression.

problem Need for frequentist uncertainty bounds in applications like learning-based control.
method Introduce new uncertainty bounds that are rigorous and practically useful.
result New bounds are less conservative and more useful for practical applications.

This paper evaluates LLMs for technical market analysis, finding GPT-4 Turbo and FinGPT outperform passive benchmarks.

problem Evaluating LLMs for technical market analysis in financial markets.
method Structured evaluation of five LLMs (GPT-4 Turbo, Claude 3 Opus, Gemini 1.5 Pro, Llama 3 70B, FinGPT) on four tasks: candlestick pattern recognition, directional signal generation, backtesting, and financial report comprehension.
result GPT-4 Turbo and FinGPT outperform passive benchmarks in simulated backtesting, with GPT-4 Turbo achieving the highest annualized return and Sharpe ratio.

We provide faster algorithms for the problem of Gaussian summation, which occurs in many machine learning methods. We develop two new extensions - an O(Dp) Taylor expansion for the Gaussian kernel with rigorous error bounds and a new error control scheme integrating any arbitrary approximation method - within the best …

2012-06-27abs ↗pdf ↗

The paper evaluates and benchmarks electricity price forecasting models.

problem Lack of rigorous evaluation methods and open datasets.
method Literature review, cross-market comparison, open datasets, and python toolbox.
result Best practices for electricity price forecasting are proposed.

The paper constructs semistrict monoidal 2-categories from foam evaluations.

problem Creating examples of semistrict monoidal 2-categories.
method Using a closed foam evaluation formula as input, the paper rigorously constructs semistrict monoidal 2-categories.
result The constructed monoidal 2-categories are semistrict, have duals and adjoints, and carry a spatial duality structure.

Noise titration benchmarks time series forecasting models rigorously.

problem Evaluation of time series forecasting models is often flawed due to lack of interventionist methods.
method Interventionist benchmarking using Gaussian noise titration of dynamical systems.
result Fern model outperforms state-of-the-art models in non-stationary conditions.

A rigorous ML pipeline for binary classification in biomedical studies, focusing on pancreatic cancer.

problem Handling bias in ML models for complex biomedical data.
method Customizable ML analysis pipeline with 9 algorithms, hyperparameter optimization, and thorough evaluation.
result Comparison of ML algorithms to ExSTraCS, highlighting interpretability and bias handling.

Paper proposes a sequential statistical test for comparing imitation learning policies with near-optimal stopping.

problem Challenges in rigorously comparing imitation learning policies due to small sample sizes and potential p-hacking.
method Sequential statistical test that adapts the number of trials based on intermediate results, achieving near-optimal stopping.
result Reduces the number of evaluation trials by up to 32% compared to state-of-the-art baselines, saving significant time and effort.

Machine learning suffers from poor design, data, and evaluation practices.

problem Ad hoc design, poor data hygiene, and lack of statistical rigor in model evaluation.
method Examines the entire machine learning process from design to evaluation, highlighting common pitfalls and providing recommendations.
result Common pitfalls in machine learning research and development are identified and actionable recommendations are provided.

This paper evaluates targeted data poisoning attacks by focusing on the hardest samples, improving evaluation and defense strategies.

problem The effectiveness of targeted data poisoning attacks is often overestimated due to average evaluation methods.
method The paper introduces metrics to identify the hardest and easiest to poison samples based on clean model information.
result The proposed metrics reliably stratify samples by poisoning vulnerability, enabling rigorous worst-case evaluation and proactive defense.

Authors provide a fair comparison of GNNs for graph classification.

problem Lack of reproducibility and rigorousness in experimental procedures for GNNs.
method Controlled and uniform framework with over 47,000 experiments.
result GNNs do not fully exploit structural information on some datasets.

Survey on RL reproducibility using real-world robots.

problem Difficulty in reproducing RL results due to algorithm variance, environment stochasticity, and hyper-parameters.
method Investigates issues leading to irreproducible RL research and standardizes evaluation approach.
result Shows how to manage and standardize evaluation of RL algorithms for unbiased comparison.

Develops a validated trading framework for market microstructure signals.

problem Overfitting and lookahead bias in algorithmic trading.
method Interpretable hypothesis-driven signal generation, reinforcement learning, strict out-of-sample testing.
result Modest annualized returns with strong downside protection and market-neutral characteristics.

A hybrid impurity measure balances theoretical soundness and computational efficiency.

problem Developing a robust impurity measure for decision trees.
method Integrates Tsallis entropy with an exponential polarization component.
result Simple parametric measures outperform ITC, but ITC variants are competitive with strong theoretical guarantees.

Quantum method improves CVaR evaluation under correlated fields.

problem Accurately evaluating CVaR in high-dimensional, correlated material uncertainty.
method Quantum-enhanced inference framework using stabilized IQAE.
result Quantum method achieves lower oracle complexity than classical methods.

Graphical lasso models ASR utterance dependencies for consistent WER estimation.

problem Modeling dependent structure among ASR utterances for accurate significance analysis.
method Graphical lasso for dependency modeling, followed by blockwise bootstrap resampling.
result Statistically consistent variance estimator of WER under mild conditions.

Study evaluates deep learning methods for dermatology, finding they perform poorly under non-ideal conditions.

problem Lack of robustness of deep learning methods in dermatology under real-world conditions.
method Simulated non-ideal conditions on user-submitted dermatology images.
result Deep learning methods show significant drop in accuracy and prediction changes under non-ideal conditions.

FINCH dataset enables financial Text-to-SQL tasks, improving model evaluation.

problem Lack of a large-scale financial dataset for Text-to-SQL tasks.
method Curated financial dataset (FINCH), benchmarking reasoning and language models.
result Proposed FINCH Score for more accurate financial model evaluation.

Interpretable framework evaluates structure learning methods for causal discovery from observational data.

problem Evaluation of structure learning methods under assumption violations in causal discovery.
method Six-dimensional evaluation metric (DOS) tailored for causal discovery.
result Amortized causal discovery delivers results with high proximity to the optimal solution.

This article evaluates explanations without ground truth in IML.

problem Lack of ground truth for evaluating explanations in IML.
method Rigorously defined problem, reviewed existing efforts, summarized three aspects of explanation, designed a unified evaluation framework.
result Unified evaluation framework for different scenarios in practice.

New measures quantify diversity of latent representations using metric space magnitude.

problem Evaluating the diversity of latent representations in machine learning models.
method Developed magnitude-based measures for latent representations, stable under data perturbations.
result Demonstrated superior performance across various domains and tasks.

The paper develops a new model to evaluate policies in complex temporal/spatial experiments.

problem Evaluating the impact of policies in experiments with temporal and spatial dependencies.
method Temporal/spatio-temporal Varying Coefficient Decision Process (VCDP) model, decomposing ATE into DE and IE.
result Effective estimation and inference of DE and IE with rigorous statistical analysis.

OASIS optimizes ER evaluation by reducing labelling needs with optimal sampling.

problem Extreme class imbalance in ER leads to high labelling costs.
method OASIS uses a biased instrumental distribution and Bayesian updates to focus on unlabelled items.
result OASIS estimates F-measure, precision, recall converge to true values with significant labelling reductions.

This research improves calorimeter simulations by creating a faster model.

problem Efficiently simulating detailed calorimeter data for high-energy physics.
method Developed a conditional normalizing flow model for superresolution.
result The model successfully reproduces reference distributions.

Develops a method to simulate rare dangerous events in autonomous systems.

problem Rare dangerous events in safety-critical systems are hard to test in real-world settings.
method Combines exploration, exploitation, and optimization techniques for rare-event simulation.
result Provides rigorous guarantees for the performance of the method.

The CLT fails for LLM evaluations with small data, leading to underestimation of uncertainty.

problem Inaccurate uncertainty estimates in LLM evaluations with small datasets.
method Alternative frequentist and Bayesian methods for uncertainty quantification.
result CLT-based methods underestimate uncertainty in small data settings.

New algorithm maintains privacy while improving model performance in selective release.

problem Privacy degradation and slow convergence in DPSGD.
method Differentially Private Selective Release based on Clipped Gradients (DPSR-CG).
result Maintains strict privacy guarantees while achieving exceptional model performance.

This paper considers the challenge of evaluating a set of classifiers, as done in shared task evaluations like the KDD Cup or NIST TREC, without expert labels. While expert labels provide the traditional cornerstone for evaluating statistical learners, limited or expensive access to experts represents a practical bottl…

2012-12-05abs ↗pdf ↗

Adversarial method finds rare catastrophic failures in safety-critical agents.

problem Evaluating safety-critical learning systems for catastrophic failures.
method Adversarial evaluation approach focusing on rare adversarial situations.
result Adversarial evaluation finds catastrophic failures and estimates failure rates faster.

Benchmark assesses LLMs' causal inference skills, revealing significant limitations.

problem Lack of rigorous evaluation of LLMs' causal inference capabilities.
method CausalPitfalls benchmark with structured challenges and grading rubrics.
result Significant limitations in current LLMs' statistical causal inference.

Proposes a method to choose thresholds for LLM evaluation metrics.

problem Ensuring reliable large language models (LLMs) with correct threshold selection.
method Identify risks, stakeholders' risk tolerance, and use ground-truth data to determine thresholds.
result Demonstrates a concrete example with the Faithfulness metric and HaluBench dataset.